Resource transfer data method and apparatus, computer device and storage medium

By combining multiple anomaly detection models with the anomaly degree of both the target feature dimension and the candidate feature dimension, the problem of low accuracy in identifying abnormal resource transfer records in existing technologies is solved, and high-accuracy anomaly detection of resource transfer records is achieved.

CN115271712BActive Publication Date: 2026-05-22TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TENCENT TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2021-04-29
Publication Date
2026-05-22

Smart Images

  • Figure CN115271712B_ABST
    Figure CN115271712B_ABST
Patent Text Reader

Abstract

The application relates to a resource transfer data processing method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a target feature dimension set; acquiring a target resource transfer record set to be identified, the target resource transfer record set comprising a plurality of target resource transfer records; acquiring target resource transfer features of each target resource transfer record on the target feature dimension, to form a target resource transfer feature set corresponding to the target resource transfer record; determining an anomaly detection model set, the anomaly detection model set comprising a plurality of different anomaly detection models; performing anomaly detection on the target resource transfer feature set by using the anomaly detection model, to obtain a model detection result of the target resource transfer record by using the anomaly detection model; and performing statistics on the model detection result of the target resource transfer record, to obtain an anomaly detection result of the target resource transfer record. The method can improve the accuracy of anomaly identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, computer device, and storage medium for transferring data from resources. Background Technology

[0002] With the development of computer technology, techniques for anomaly detection using computers have emerged. This technology allows computer devices to identify input features and then obtain corresponding model detection results.

[0003] In traditional technologies, experts typically summarize rules based on historical abnormal resource transfer records. Computer devices can then use these rules to identify unknown resource transfer records to determine whether they are abnormal. However, this method is limited by expert experience and has low accuracy. Summary of the Invention

[0004] Therefore, it is necessary to provide a resource transfer data processing method, apparatus, computer equipment, and storage medium that can improve the accuracy of identifying abnormal resource transfer records, in order to address the aforementioned technical problems.

[0005] A resource transfer data processing method includes: acquiring a target feature dimension set; the target feature dimension set is selected from a candidate feature dimension set based on the dimensional anomaly of the candidate feature dimension; the dimensional anomaly is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set; acquiring a target resource transfer record set to be identified, the target resource transfer record set including multiple target resource transfer records, acquiring the target resource transfer features of each target resource transfer record on the target feature dimension, forming a target resource transfer feature set corresponding to the target resource transfer record; determining an anomaly detection model set, the anomaly detection model set including multiple different anomaly detection models; performing anomaly detection on the target resource transfer feature set using the anomaly detection models to obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension; and statistically analyzing the model detection results of the target resource transfer record to obtain the anomaly detection result of the target resource transfer record.

[0006] A resource transfer data processing apparatus, comprising: a target feature dimension acquisition module, configured to acquire a target feature dimension set; the target feature dimension set is selected from a candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the dimensional anomaly is determined based on the feature distribution of a first resource transfer feature of the candidate feature dimension in the corresponding historical feature set; and a resource transfer feature selection module, configured to acquire a set of target resource transfer records to be identified, the target resource transfer record set including multiple target resource transfer records, and acquire the target resource transfer features of each target resource transfer record on the target feature dimension to form the target resource transfer record corresponding to the target resource transfer record. The system includes: a resource transfer feature set; a detection model determination module for determining an anomaly detection model set, which includes multiple different anomaly detection models; an anomaly detection module for performing anomaly detection on the target resource transfer feature set using the anomaly detection models to obtain the model detection results of the anomaly detection models on the target resource transfer records; wherein at least one anomaly detection model obtains the model detection results based on the distribution results of the target resource transfer features in the target feature set corresponding to the feature dimension; and a detection result statistics module for statistically analyzing the model detection results of the target resource transfer records to obtain the anomaly detection results of the target resource transfer records.

[0007] A computer device includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: acquiring a target feature dimension set; the target feature dimension set is selected from a candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the dimensional anomaly is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set; acquiring a target resource transfer record set to be identified, the target resource transfer record set including multiple target resource transfer records, acquiring the target resource transfer features of each target resource transfer record on the target feature dimension, forming a target resource transfer feature set corresponding to the target resource transfer record; determining an anomaly detection model set, the anomaly detection model set including multiple different anomaly detection models; performing anomaly detection on the target resource transfer feature set using the anomaly detection models to obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension; and statistically analyzing the model detection results of the target resource transfer record to obtain the anomaly detection result of the target resource transfer record.

[0008] A computer-readable storage medium storing a computer program, which, when executed by a processor, performs the following steps: obtaining a target feature dimension set; the target feature dimension set is selected from a candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the dimensional anomaly is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set; obtaining a target resource transfer record set to be identified, the target resource transfer record set including multiple target resource transfer records, obtaining the target resource transfer features of each target resource transfer record on the target feature dimension, forming a target resource transfer feature set corresponding to the target resource transfer record; determining an anomaly detection model set, the anomaly detection model set including multiple different anomaly detection models; performing anomaly detection on the target resource transfer feature set using the anomaly detection models to obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension; and statistically analyzing the model detection results of the target resource transfer record to obtain the anomaly detection result of the target resource transfer record.

[0009] The aforementioned resource transfer data processing method, apparatus, computer equipment, and storage medium, on the one hand, employ multiple different anomaly detection models for anomaly detection and comprehensively statistically analyze the model detection results corresponding to these anomaly detection models to obtain the anomaly detection result of the target resource transfer record. This allows for the determination of the anomaly detection result of the target resource transfer record based on multiple different anomaly detection strategies, effectively improving the accuracy of resource transfer record identification. On the other hand, during anomaly detection, the target resource transfer features of the target resource transfer record in the target feature dimension set are obtained to form a target resource transfer feature set for anomaly detection. The target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly degree of the candidate feature dimension. The dimensional anomaly degree is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set. Therefore, the distribution result can well reflect the anomaly of the resource transfer features. Since at least one of the multiple anomaly detection models obtains the model detection result based on the distribution result of the target resource transfer features in the target feature set corresponding to the feature dimension, the anomaly of the resource transfer features is fully considered during the anomaly detection process, further improving the accuracy of resource transfer record identification. Attached Figure Description

[0010] Figure 1 This is an application environment diagram of the resource transfer data processing method in some embodiments;

[0011] Figure 2 This is a flowchart illustrating the resource transfer data processing method in some embodiments;

[0012] Figure 3 This is a schematic diagram illustrating the historical feature set obtained in some embodiments;

[0013] Figure 4 This is a schematic diagram illustrating the target resource transfer feature set obtained in some embodiments;

[0014] Figure 5 This is a schematic diagram illustrating the division of resource transfer features within a feature set in some embodiments;

[0015] Figure 6 This is a schematic diagram of clustering results in some embodiments;

[0016] Figure 7 This is a flowchart illustrating the resource transfer data processing method in some specific embodiments;

[0017] Figure 8 This is a schematic diagram illustrating anomaly detection in some embodiments;

[0018] Figure 9 This is a structural block diagram of the resource transfer data processing apparatus in some embodiments;

[0019] Figure 10 This is a diagram showing the internal structure of a computer device in some embodiments. Detailed Implementation

[0020] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0021] As provided in the resource transfer data processing method of this application, the data involved in the operation of this method, such as the target resource transfer record set, the anomaly detection results of each target resource transfer record, the anomaly detection model set, etc., can all be stored on the blockchain. The blockchain generates different query codes for different data and returns them to the computer device. The computer device can query the corresponding data from the blockchain based on the query code. For example, based on the query code of the target resource transfer record, the target resource transfer record and the corresponding anomaly detection results can be queried from the blockchain.

[0022] In some embodiments, the resource transfer data processing method, apparatus, computer equipment, and storage medium provided in this application can be implemented using artificial intelligence technology. Wherein:

[0023] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0024] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0025] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Based on the classification of learning methods, machine learning includes at least one of supervised learning, unsupervised learning, or reinforcement learning.

[0026] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0027] The solutions provided in this application involve technologies such as machine learning in artificial intelligence, and are specifically illustrated through the following embodiments:

[0028] The resource transfer data processing method provided in this application embodiment can be applied to, for example, Figure 1The application environment shown includes a server 102, a first terminal 104, and a second terminal 106. The first terminal 104 can include multiple terminals, meaning at least two, such as terminals 104A and 104B. The server 102 can deploy multiple different anomaly detection models. The first terminal 104 can send resource transfer requests to the server, which can respond to each request and perform resource transfers. After the resource transfer is completed, multiple target resource transfer records to be identified are generated. The server 102 can obtain the target resource transfer features of each record in the target feature dimension, forming a target resource transfer feature set corresponding to each record. Anomalies are detected on the target resource transfer feature set using the anomaly detection model, yielding the model detection results for the target resource transfer records. The model detection results for the target resource transfer records are then statistically analyzed to obtain the anomaly detection results for the target resource transfer records.

[0029] Server 102 can send the anomaly detection results of multiple target resource transfer records to the second terminal 106 at preset time intervals. Alternatively, it can send the obtained anomaly detection results to the second terminal after receiving a request from the second terminal. Or, when the obtained anomaly detection results indicate that the target resource transfer record is abnormal, the server can send the abnormal target resource transfer record to the terminal and send an alarm message to the terminal.

[0030] The first terminal 104 may have an application for resource transfer installed. The second terminal 106 may be implemented using the same device as one of the first terminals 104.

[0031] The server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The first terminal 104 and the second terminal 106 can be smartphones, tablets, laptops, desktop computers, smart speakers, smartwatches, etc., but are not limited to these. The terminals and the server can be directly or indirectly connected via wired or wireless communication, and this application does not impose any restrictions on this connection.

[0032] The resource transfer data processing method or apparatus provided in this application embodiment includes a blockchain composed of multiple servers, where the servers are nodes on the blockchain.

[0033] In some embodiments, such as Figure 2 As shown, a resource transfer data processing method is provided, which can be applied to... Figure 1Taking the server in the example, the following steps are included:

[0034] Step 202: Obtain the target feature dimension set; the target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the dimensional anomaly is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set.

[0035] In this context, "resources" refers to items that can be circulated online, including at least one of virtual and physical items. Virtual items can specifically include, but are not limited to, account values ​​for various accounts, funds, stocks, bonds, virtual avatar products, virtual recharge cards, and game equipment. Physical items can be any item with a physical form that can be owned by a user, specifically including but not limited to electronic products, toys, handicrafts, or autographed photos. Resource transfer can involve moving resources from one storage medium to another, transferring the resource value corresponding to a resource from one account to another, or transferring the resource value corresponding to a resource to another account. The storage medium can be a computer device with resource storage capabilities, such as a bank's server or storage device. A resource transfer record refers to the data record generated during a resource transfer process; for example, a resource transfer record could be transaction data generated during an online shopping transaction. Resource transfer records prior to the current moment can be called historical resource transfer records. Multiple historical resource transfer records constitute a historical resource transfer record set.

[0036] Resource transfer characteristics refer to the values ​​of various fields in a resource transfer record. Different resource transfer characteristics correspond to different fields. Feature dimensions are divided based on fields; one field can correspond to one feature dimension, or multiple fields of the same type can correspond to one feature dimension. For example, feature dimensions can include at least one of the following: time dimension, frequency dimension, user dimension, or amount dimension. Further, under the time dimension, a resource transfer characteristic can be the time of the resource transfer operation; under the frequency dimension, a resource transfer characteristic can be at least one of the following: the frequency of the resource transfer operation or the frequency of resource transfers within a specified time period; under the user dimension, a resource transfer characteristic can be at least one of the following: the receiving user to whom the resource transfer operation is directed, the number of receiving users, or the characteristics of the receiving users (in a transaction scenario, the receiving user can also be called the counterparty); under the amount dimension, a resource transfer characteristic can be at least one of the following: the value of the transferred resources or the value of resource transfers within a specified time period, etc.

[0037] Each candidate feature dimension in the candidate feature dimension set can have its corresponding dimensional anomaly score obtained based on the historical resource transfer record set. The dimensional anomaly score of a candidate feature dimension reflects the importance of that feature dimension in the anomaly detection process. There is a positive correlation between the dimensional anomaly score and the importance of the candidate feature dimension; that is, the more important the candidate feature dimension, the higher its dimensional anomaly score. This importance is reflected in the fact that during anomaly detection, the more important the candidate feature dimension, the more relevant the resource transfer features under that candidate feature dimension are to anomalies.

[0038] Candidate feature dimensions can be specific numerical values, such as X (where X is a real number greater than 0), or they can be levels obtained by dividing based on the size of the numerical values, such as first-level, second-level, third-level, N-level, etc. The larger the numerical value, the higher the level.

[0039] Dimensional anomaly is determined based on the feature distribution of resource transfer features of a candidate feature dimension within the corresponding historical feature set. The historical feature set corresponding to a candidate feature dimension refers to the set of resource transfer features on that candidate feature dimension in all or part of the resource transfer record set. The feature distribution of the first resource transfer feature of the candidate feature dimension within the corresponding historical feature set reflects the occurrence of each resource transfer feature under that feature dimension in the historical resource transfer record set, including the number of resource transfer features in that feature dimension and the frequency of occurrence of each resource transfer feature.

[0040] like Figure 3 The diagram shown illustrates how historical feature sets are obtained in some embodiments. (Reference) Figure 3 The historical resource transfer record set includes N historical resource transfer records, namely record 1, record 2, ..., record N. Each resource transfer record includes M candidate feature dimensions. Taking candidate feature dimension 1 as an example, the resource transfer value of record 1 in candidate feature dimension 1 is X1, the resource transfer value of record 2 in candidate feature dimension 1 is X2, ..., the resource transfer value of record N in candidate feature dimension 1 is XN. Then the historical feature set includes N resource transfer features X1, X2, ..., XN. These resource transfer features can be different or some resource transfer features can be the same, for example, X1=X2=XN.

[0041] The target feature dimension set is selected from the candidate feature dimension set based on the dimensionality anomaly of the candidate feature dimensions. In some embodiments, the candidate feature dimensions in the candidate feature dimension set can be sorted according to their dimensionality anomaly, and a preset number of candidate feature dimensions can be selected based on the sorting result. These selected candidate feature dimensions constitute the target feature dimension set. For example, the candidate feature dimensions in the candidate feature dimension set can be sorted in descending order according to their dimensionality anomaly, and a preset number of candidate feature dimensions at the top of the sorted list can be selected as the target feature dimensions. In other embodiments, candidate feature dimensions with dimensionality anomalies greater than a preset value can be selected from the candidate feature dimension set as target feature dimensions. The preset value can be set as needed.

[0042] The process of selecting the target feature dimension from the candidate feature dimension set can be regarded as a dimensionality reduction process. The server can use algorithms such as the Copled Biased Random Walk (CBRW) algorithm and the principal component analysis (PCA) algorithm to reduce the feature dimension.

[0043] Specifically, the server can predetermine the target feature dimension set based on the historical resource transfer record set and save it locally. When it is necessary to perform anomaly identification on the resource transfer record to be identified, the server can obtain the target feature dimension set from the local machine. Alternatively, the server can obtain the target feature dimension set from other computer devices, which can predetermine and save the target feature dimension set based on the historical resource transfer record set.

[0044] Step 204: Obtain the set of target resource transfer records to be identified. The set of target resource transfer records includes multiple target resource transfer records. Obtain the target resource transfer features of each target resource transfer record in the target feature dimension to form the target resource transfer feature set corresponding to the target resource transfer record.

[0045] The target resource transfer record set includes a collection of multiple target resource transfer records. "Multiple" means at least two. A target resource transfer record refers to a resource transfer record that needs to be identified. A target resource transfer record can be a resource transfer record from the historical resource transfer record set, or a resource transfer record received from the first terminal at the current moment.

[0046] Specifically, the server can retrieve multiple resource transfer records to be identified from a database storing historical resource transfer records, thus obtaining a target resource transfer record set; or the server can retrieve multiple resource transfer records sent by the first terminal within the current time period, thus obtaining a target resource transfer record set. After obtaining the target resource transfer record set, for each target resource transfer record, the server obtains the target resource transfer features of that target resource transfer record in each target feature dimension, forming a target resource transfer feature set corresponding to that target resource transfer record.

[0047] like Figure 4 The diagram shown illustrates, in one embodiment, how the server obtains the target resource transfer feature set corresponding to the target resource transfer record. (Reference) Figure 4 A target resource transfer record includes four resource transfer features, each corresponding to a different feature dimension. Feature dimension 2 and feature dimension 4 within the dashed box are target feature dimensions. The target resource transfer feature of this target resource transfer record in feature dimension 2 is X2, and the target resource transfer feature in feature dimension 2 is X4. Therefore, the resulting target resource transfer feature set includes two resource transfer features: X2 and X4.

[0048] Step 206: Determine the set of anomaly detection models, which includes multiple different anomaly detection models.

[0049] The anomaly detection model is a network model capable of detecting anomalies according to a specific anomaly detection strategy. Further, the anomaly detection model can be a machine learning model, such as a supervised learning model, a semi-supervised learning model, or an unsupervised learning model. Specifically, the anomaly detection model can be Random Forest, Isolation Forest (iForest), HBOS (Histogram-based Outlier Score), COPOD (Copula-Based Outlier Detection), Auto Encoder, CBLOF (Cluster-Based Local Outlier Factor), etc. In some embodiments, the anomaly detection model can also be referred to as a weak learner.

[0050] Anomaly detection strategies refer to the methods used to detect anomalies. Specifically, anomaly detection strategies can be categorized or clustered based on resource transfer characteristics. Different anomaly detection models employ different anomaly detection strategies.

[0051] Specifically, each anomaly detection model can be a functional module configured in the server. The server calls each anomaly detection model to perform anomaly detection on the resource transfer features in the target resource transfer feature set, thereby obtaining the model detection result of the target resource transfer record corresponding to the target resource transfer feature set. Alternatively, each anomaly detection model can be a functional module configured in different terminals or servers. The server sends detection requests to these terminals or servers respectively, triggering them to perform anomaly detection on the resource transfer features in the target resource transfer feature set based on the corresponding anomaly detection model, thereby obtaining the model detection result of the target resource transfer record corresponding to the target resource transfer feature set.

[0052] Step 208: Perform anomaly detection on the target resource transfer feature set using an anomaly detection model to obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein, at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer features in the target feature set corresponding to the feature dimension.

[0053] Among them, the model detection result is the detection result of the anomaly detection model on the category to which the resource transfer record belongs. The model detection result of the anomaly detection model on the resource transfer record can be characterized by the category identifier of the category.

[0054] In addition, the model detection results can also carry the model identifiers of each anomaly detection model. Based on this, the server can determine which anomaly detection model output each detection result. Since the anomaly detection strategies used by each anomaly detection model are different, the model detection results obtained after anomaly detection may be the same or different.

[0055] Specifically, the server can use each anomaly detection model in the anomaly detection model set to perform anomaly detection on the target resource transfer feature set, and obtain the model detection result of each anomaly detection model for the target resource transfer record.

[0056] At least one anomaly detection model obtains its detection result based on the distribution of target resource transfer features within the target feature set corresponding to its feature dimension. The distribution result refers to the classification of each target resource transfer feature within its corresponding target feature set. Specifically, the distribution result can be represented by classification labels, such as "abnormal" or "normal," or by probability values ​​corresponding to each category, such as 80% or 90%. Alternatively, the distribution result can also be information such as the distribution density of classification probabilities. The target feature set is a feature set consisting of all or some resource transfer features of the target resource transfer record set under a certain feature dimension. The resource transfer features in the target feature set can correspond to a specific feature distribution state, which can be used to classify the resource transfer features and obtain the distribution result. The feature distribution state can be represented by at least one form, such as a dendrogram, histogram, or normal distribution plot.

[0057] In some embodiments, at least one anomaly detection model obtains its detection result based on the distribution of target resource transfer features in the target feature set corresponding to its feature dimension. Specifically, this can be implemented as follows: all anomaly detection models in the anomaly detection model set obtain their detection results based on the distribution of target resource transfer features in the target feature set corresponding to its feature dimension.

[0058] It is understandable that different anomaly detection models can use different methods to classify target resource transfer features, thereby determining the distribution of each target resource transfer feature in its respective target feature set.

[0059] In some embodiments, at least one anomaly detection model can determine the target feature set corresponding to a certain target feature dimension, analyze and classify the distribution state of the target feature set corresponding to the target feature dimension according to the corresponding anomaly detection strategy, so as to obtain the distribution result corresponding to the target feature dimension; then, the distribution results under each target feature dimension are integrated to obtain the distribution result of each resource transfer feature in the target feature set corresponding to its respective feature dimension. Alternatively, at least one anomaly detection model can also perform an overall analysis of the resource transfer features under each target feature dimension, for example, determining the overall distribution state of the resource transfer features under all target feature dimensions, and classifying them based on this distribution state to obtain the overall distribution result corresponding to all target feature dimensions.

[0060] The distribution results of anomaly detection models can characterize the classification results of various target resource transfer features. Based on this, the distribution results can be analyzed to determine the model detection results for target resource transfer records. Furthermore, the distribution results, represented by classification labels or probability values, can be converted into numerical forms, and the converted results can be used as the corresponding model detection results. For example, the probability value can be compared with a probability threshold; a higher probability value results in a detection result of 1, and a lower probability value results in a detection result of 0. The probability threshold can be determined based on the specific scenario.

[0061] Step 210: Statistically analyze the model detection results of the target resource transfer records to obtain the anomaly detection results of the target resource transfer records.

[0062] The anomaly detection result can be the result of whether the target resource transfer record is an abnormal resource transfer record. Therefore, the anomaly detection result can be that the target resource transfer record is an abnormal resource transfer record, a normal resource transfer record, or a suspicious resource transfer record. Statistical analysis of the model detection results can refer to performing statistical calculations on the number of model detection results, and using the result of these calculations as the anomaly detection result. Specifically, the anomaly detection result of the target resource transfer record can be determined as an abnormal resource transfer record when the number of model detection results meets a certain quantity condition. This quantity condition can be greater than a set quantity threshold, which can be determined based on the actual situation. The process of obtaining anomaly detection results based on the model detection results corresponding to each anomaly detection model can be considered as an ensemble learning process of each anomaly detection model. When the anomaly detection model is called a weak learner, the set of anomaly detection models obtained by the ensemble learning of these weak learners can be called a strong learner.

[0063] In some embodiments, statistical analysis of model detection results can be achieved using at least one of the following methods: bagging, boosting, stacking, etc.

[0064] In the above-mentioned resource transfer data processing method, on the one hand, since multiple different anomaly detection models are used for anomaly detection, and the model detection results corresponding to these anomaly detection models are comprehensively statistically analyzed to obtain the anomaly detection results of the target resource transfer record, the anomaly detection results of the target resource transfer record can be determined based on multiple different anomaly detection strategies, which effectively improves the accuracy of resource transfer records. On the other hand, since, during anomaly detection, the target resource transfer features of the target resource transfer record in the target feature dimension set are obtained to form the target resource transfer feature set for anomaly detection, the target feature dimension set is selected from the candidate feature dimension set according to the dimensional anomaly degree of the candidate feature dimension. The dimensional anomaly degree is determined according to the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set. Therefore, the distribution result can well reflect the anomaly of the resource transfer features. Since at least one of the multiple anomaly detection models is based on the distribution result of the target resource transfer features in the target feature set corresponding to the feature dimension to obtain the model detection result, the anomaly of the resource transfer features is fully considered in the anomaly detection process, further improving the accuracy of resource transfer record identification.

[0065] In some embodiments, the step of obtaining the dimensionality anomaly of the candidate feature dimension includes: obtaining the first feature distribution value of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set, and obtaining the representative feature distribution value corresponding to the historical feature set; obtaining the feature anomaly corresponding to the first resource transfer feature based on the difference between the first feature distribution value and the representative feature distribution value; and obtaining the dimensionality anomaly corresponding to the candidate feature dimension based on the feature anomaly corresponding to the first resource transfer feature.

[0066] The first resource transfer feature refers to any resource transfer feature under the candidate feature dimension. The first feature distribution value of the first resource transfer feature in the corresponding historical feature set is positively correlated with the number of times the first resource transfer feature appears in the historical feature set. For example, the first resource transfer feature can be the probability of the first resource transfer feature appearing in the historical feature set. For example, if the historical feature set includes 100,000 monetary values, of which 10,000 monetary values ​​are 50, then the first distribution value of the resource transfer feature is 1 / 10.

[0067] The representative feature distribution value corresponding to the historical feature set is positively correlated with the frequency of occurrence of the most frequent resource transfer feature in the historical feature set. This means that if a certain resource transfer feature appears most frequently in the historical feature set, then from a statistical perspective, that resource transfer feature can be used to represent the set. Specifically, the representative feature distribution value can be the probability of occurrence of the most frequent resource transfer feature, i.e.:

[0068]

[0069] in, It refers to probability. To represent the characteristic distribution value, , This is the historical feature set. For example, suppose the historical feature set includes 100,000 monetary values, among which the most frequently occurring monetary value is 60, which appears 50,000 times (i.e., there are 50,000 monetary values ​​of 60 in the historical feature set). Then the representative feature distribution value corresponding to this historical feature set is 1 / 2.

[0070] Specifically, after obtaining the first feature distribution value of the first resource transfer feature in the corresponding historical feature set and the representative feature distribution value of the historical feature set, the server further obtains the difference between the first feature distribution value and the representative feature distribution value to obtain the feature anomaly score. This feature anomaly score reflects the degree of anomalousness of the first resource transfer feature in the historical feature set; a larger feature anomaly score indicates a greater degree of anomalousness of the first resource transfer feature in the historical feature set. The server further obtains the dimension anomaly score corresponding to the candidate feature dimension based on the feature anomaly score corresponding to the first resource transfer feature.

[0071] In some embodiments, the difference between the first characteristic distribution value and the representative characteristic distribution value can specifically be the absolute difference between the first characteristic distribution value and the representative characteristic distribution value. In other embodiments, the difference between the first characteristic distribution value and the representative characteristic distribution value can be calculated using the following formula:

[0072] First characteristic of resource transfer The difference between the first feature distribution value and the representative feature value in the corresponding historical feature set =

[0073] In some embodiments, the server obtains the dimensional anomaly of the candidate feature dimension based on the feature anomaly of the first resource transfer feature. This can be done by statistically analyzing the feature anomalies of each resource transfer feature under the candidate feature dimension. Specifically, this can be done by summing, averaging, or calculating the median of the feature anomalies of each resource transfer feature.

[0074] In the above embodiments, the feature anomaly degree corresponding to the first resource transfer feature is obtained based on the difference between the first feature distribution value and the representative feature distribution value, and the dimensional anomaly degree corresponding to the candidate feature dimension is obtained based on the feature anomaly degree corresponding to the first resource transfer feature. The overall distribution of features in the historical feature set can be taken into account to obtain an accurate dimensional anomaly degree.

[0075] In some embodiments, obtaining the dimensional anomaly of the candidate feature dimension based on the feature anomaly of the first resource transfer feature includes: obtaining the number of times the first resource transfer feature and second resource transfer features of different feature dimensions co-occur in a historical resource transfer record set; obtaining the anomaly propagation weight between the first resource transfer feature and the second resource transfer feature based on the co-occurrence count and the feature anomaly; propagating the feature anomaly of the second resource transfer feature to the first resource transfer feature based on the anomaly propagation weight to obtain the propagation anomaly of the first resource transfer feature; and calculating the propagation anomaly of the first resource transfer feature to obtain the dimensional anomaly of the candidate feature dimension.

[0076] The co-occurrence of a first resource transfer feature and a second resource transfer feature of different dimensions refers to the occurrence of both features in the same historical resource transfer record within the historical resource transfer record set. For example, if a historical resource transfer record includes a transaction amount of 50 yuan and the transaction channel is Alipay, then the resource transfer features "50" and "Alipay" co-occur in that historical resource transfer record. Anomaly propagation weight represents the proportion of anomaly severity of one resource transfer feature propagated to another; a larger anomaly propagation weight indicates more anomaly severity is propagated.

[0077] Specifically, the server can obtain the co-occurrence probability of the first resource transfer feature and the second resource transfer feature with different feature dimensions in the historical resource transfer record set according to the proportion of the number of times they co-occur in the total number of historical resource transfer records. Furthermore, it can obtain the feature co-occurrence degree between the first resource transfer feature and the second resource transfer feature based on the co-occurrence probability, as shown in the following formula:

[0078]

[0079] in, This refers to the second resource transfer characteristic. and the characteristics of the first resource transfer Feature co-occurrence, This refers to the second resource transfer characteristic. and the characteristics of the first resource transfer The co-occurrence probability, This refers to the characteristics of the first resource transfer. The probability of its occurrence.

[0080] The server can further leverage the second resource transfer feature. The anomaly propagation weight is obtained by considering the co-occurrence and anomaly of features between the first resource transfer feature and the first resource transfer feature. Specifically, the server can obtain the anomaly propagation weight between the two features based on the product of the co-occurrence and anomaly of features.

[0081] In some embodiments, the anomaly propagation weight can be calculated with reference to the following formula, wherein, This refers to the anomaly degree of the features:

[0082]

[0083] As mentioned earlier, if there is a correlation between two resource transfer features, the feature anomaly can be passed from one resource transfer feature to the other. Based on this, after obtaining the anomaly propagation weight, the server can further propagate the feature anomaly of the second resource transfer feature to the first resource transfer feature based on the anomaly propagation weight, thereby obtaining the propagation anomaly of the first resource transfer feature. Finally, the propagation anomaly of all first resource transfer features under the candidate feature dimension is counted to obtain the dimension anomaly of the candidate feature dimension.

[0084] In the above embodiments, the anomaly propagation weight between the first resource transfer feature and the second resource transfer feature with different feature dimensions is obtained based on the number of co-occurrences of the first resource transfer feature and the feature anomaly degree of the first resource transfer feature in the historical resource transfer record set. Then, feature anomaly degree propagation is performed based on the anomaly propagation weight, and finally the propagation anomaly degree of the first resource transfer feature is obtained. These propagation anomalies are counted to obtain the dimensional anomaly degree of the feature dimension in which the first resource transfer feature is located. This not only takes into account the overall feature distribution under the feature dimension, but also takes into account the correlation influence between different feature dimensions, so the dimensional anomaly degree is more accurate and can better reflect the importance of the feature dimension in anomaly detection.

[0085] In some embodiments, the transfer of the feature anomaly degree of the second resource transfer feature to the first resource transfer feature based on the anomaly propagation weight to obtain the propagation anomaly degree of the first resource transfer feature includes: taking the resource transfer features of each historical resource transfer record in the historical resource transfer record set as nodes, connecting resource transfer features with co-occurrence relationships to obtain a feature connection graph, wherein the nodes of the second resource transfer feature and the nodes of the first resource transfer feature in the feature connection graph have connection edges; in the feature connection graph, the feature anomaly degree of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge, and the feature anomaly degree of the first resource transfer feature when the iteration stopping condition is met is taken as the propagation anomaly degree of the first resource transfer feature.

[0086] Specifically, the resource transfer features of each historical resource transfer record in the historical resource transfer record set are used as nodes, and resource transfer features that have co-occurrence relationships are connected to obtain a feature connection graph:

[0087]

[0088] Here, V is composed of resource transfer features. Then, the server can iteratively update the feature anomaly degree of the nodes of the first resource transfer feature in the feature connectivity graph using a random walk approach, based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weights corresponding to the connecting edges. Let... Let be the probability distribution of the random walk at time t, then we have:

[0089]

[0090] in, This is the matrix transpose, i.e.:

[0091]

[0092] In the process of iterative updates, to ensure convergence, let:

[0093]

[0094] in, It can be configured as needed, for example... Values ​​can be taken in the range [0.85, 0.95]. When the iteration stopping condition is met, This will converge to a static probability distribution, i.e.:

[0095]

[0096] Based on the above formula, the final static probability distribution can be derived. It is unrelated to its initialization. The stopping condition for iteration is obtained from two iterations. The absolute difference between them does not exceed a preset threshold, or the number of iterations reaches a preset number of iterations. The preset threshold and the preset number of iterations can be set as needed. For example, the preset threshold can be 0.001 and the preset number of iterations can be 100.

[0097] obtained through iteration Then, the propagation anomaly degree of each resource transfer feature can be obtained, as follows:

[0098]

[0099]

[0100] In a specific implementation, after obtaining the transmission anomaly degree, the server can sum the transmission anomalies of each resource transfer feature under this feature dimension to obtain the dimension anomaly degree of this feature dimension.

[0101] In the above embodiments, by constructing a feature connection graph, the feature anomaly degree of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge. This can take into account the correlation between feature dimensions to the greatest extent, improve the accuracy of the propagation anomaly degree as much as possible, and thus improve the accuracy of the feature dimension anomaly degree.

[0102] In some embodiments, an anomaly detection model is used to detect anomalies in the target resource transfer feature set to obtain the model detection result of the anomaly detection model on the target resource transfer record. This includes: obtaining resource transfer features corresponding to feature dimensions in the target resource transfer feature set corresponding to the target resource transfer record to obtain target feature sets corresponding to each feature dimension; obtaining the distribution partitioning method of the target feature set by the anomaly detection model; partitioning the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension; and determining the model detection result of the anomaly detection model on the target resource transfer record based on the distribution result obtained by the anomaly detection model.

[0103] The distribution partitioning method involves dividing resource transfer features in the target feature set into corresponding feature intervals. This partitioning can be based on a threshold or on distribution intervals. Threshold-based partitioning involves comparing the resource transfer feature with a feature partitioning threshold. If the resource transfer feature is less than the threshold, it is partitioned into feature interval A; if it is greater than or equal to the threshold, it is partitioned into another feature interval B. Distribution interval-based partitioning involves determining multiple feature partitioning intervals, each corresponding to a range of feature values ​​for the resource transfer feature. The feature value of the resource transfer feature is compared to this range to assign it to the appropriate feature partitioning interval. The feature value of the resource transfer feature can be a specific numerical value, such as the time value of the resource transfer operation, the value of the transferred resource, or the number of resource transfer operations. In some embodiments, the distribution partitioning method can be determined based on the anomaly detection strategy of each anomaly detection model.

[0104] In some embodiments, target resource transfer features corresponding to each target feature dimension are selected from the target resource transfer feature set corresponding to the target resource transfer record. The target resource transfer features under a target feature dimension constitute a target feature set. The resource transfer features in the target feature set are divided based on the distribution partitioning method so that each resource transfer feature is divided into the corresponding feature interval. Each feature interval can correspond to a distribution result, thereby obtaining the distribution result of the target resource transfer feature in the target feature set of the feature dimension.

[0105] In some embodiments, the server obtains the distribution results of the target resource transfer features in the target feature set of the corresponding feature dimension, performs statistics on the distribution results corresponding to each target feature dimension, obtains the distribution results of the target resource transfer records under the corresponding anomaly detection model, transforms the distribution results of the target resource transfer records under the corresponding anomaly detection model, and uses the transformed results as the model detection results of the anomaly detection model for the target resource transfer records.

[0106] The above embodiments divide the target resource transfer features into the target feature set corresponding to the feature dimensions, which can determine the distribution results under each feature dimension. Even without the distance between the target resource transfer features, the distribution results corresponding to the target resource transfer features can be accurately determined. This can quickly determine the distribution results of the target resource transfer features, thereby effectively improving the efficiency of resource transfer record identification.

[0107] In some embodiments, the distribution partitioning method corresponding to the feature set includes a threshold-based partitioning method. Obtaining the distribution partitioning method of the target feature set by the anomaly detection model, and partitioning the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method, to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension includes: obtaining a feature distribution structure tree, which includes multiple child nodes; taking the initial node of the feature distribution structure tree as the current child node corresponding to the target resource transfer feature set, obtaining the current feature dimension corresponding to the current child node, and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension; determining the distribution result of the target resource transfer features in the current feature set based on the current feature partitioning threshold and the resource transfer features of the target resource transfer feature set in the current feature dimension; determining the next child node corresponding to the target resource transfer feature set based on the distribution result, taking the next node as the updated current child node, and returning to the steps of obtaining the current feature dimension corresponding to the current child node and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension, until the child nodes corresponding to the target resource transfer feature set are updated.

[0108] The feature distribution structure tree is a branching tree constructed based on resource transfer features, and can be at least one of an isolated tree or a random tree. Each resource transfer feature can serve as a node in the feature distribution structure tree. The feature distribution structure tree can be used as an anomaly detection model.

[0109] The number of feature distribution structure trees can be at least one. When there are multiple feature distribution structure trees, all or some of them can be used together as an anomaly detection model. These feature distribution structure trees within the anomaly detection model can partition resource transfer features in parallel, assigning the target resource transfer features corresponding to the target resource transfer record to the corresponding child nodes. Furthermore, the set of target resource transfer features corresponding to a specific target resource transfer record can be input into each feature distribution structure tree, and each structure tree outputs the distribution result for that target resource transfer record. The distribution results of each structure tree are then integrated to obtain the overall distribution result corresponding to the feature distribution structure tree, which serves as the distribution result for the corresponding anomaly detection model. Similarly, the set of resource transfer features corresponding to other target resource transfer records can be input into each feature distribution structure tree to obtain the corresponding distribution results.

[0110] The current feature segmentation threshold can be predetermined or determined based on the feature values ​​of the resource transfer features of the target resource transfer record under the current feature dimension. For example, at least one of the mean, median, or variance of the feature values ​​of the resource transfer features of the target resource transfer record under the current feature dimension can be used as the current feature segmentation threshold.

[0111] The node where child nodes have been updated can be a leaf node, meaning there is no next node, where the child node corresponding to the resource transfer feature is a leaf node. Furthermore, the distribution result corresponding to the leaf node containing the target resource transfer feature can be determined as the distribution result within the target feature set of the feature dimension containing the target resource transfer feature.

[0112] In some embodiments, the process of determining the distribution of target resource transfer features in the current target feature set based on the current feature partitioning threshold and the resource transfer features of the target resource transfer record in the current feature dimension can be as follows: The target resource transfer features of the target resource transfer record in the current feature dimension are compared with the current feature partitioning threshold; when the target resource transfer features of the target resource transfer record in the current feature dimension are less than the current feature partitioning threshold, the target resource transfer features of the target resource transfer record in the current feature dimension are partitioned to the first node; when the target resource transfer features of the target resource transfer record in the current feature dimension are greater than or equal to the current feature partitioning threshold, the target resource transfer features of the target resource transfer record in the current feature dimension are partitioned to the second node. Then, taking the first node as an example, the next child node of the first node is used as the updated current child node, and the above process is repeated. The second node is similar and will not be described in detail here.

[0113] Specifically, taking the target resource transfer characteristics—the frequency of resource transfer operations, the characteristics of the receiving user, and the value of the transferred resources—as an example, each target resource transfer characteristic is treated as a feature dimension. In the first-level partitioning process, the frequency of resource transfer operations is compared with a frequency threshold to divide the resource transfer characteristics into the corresponding first and second nodes. In the second-level partitioning process, taking one side as an example, the characteristics of the receiving user are compared with the user's feature attributes to divide the resource transfer characteristics in the first node into the third and fourth nodes. In the third-level partitioning process, taking one side as an example, the transferred resource value is compared with a resource value threshold to divide the resource transfer characteristics in the third node into the fifth and sixth nodes. At this point, the child nodes corresponding to the resource transfer characteristics are leaf nodes, and the node update is considered complete.

[0114] In the above embodiments, the current feature dimension corresponding to the current child node is gradually obtained, and then the target resource transfer record is divided into target resource transfer features in the current feature dimension based on the current feature division threshold. Each division of resource transfer features can be considered as the completion of a level division. The feature dimensions of each level are interconnected and progressive, so accurate distribution results can be obtained, and thus accurate user identification results can be obtained.

[0115] In some embodiments, determining the model detection result of the anomaly detection model for the target resource transfer record based on the distribution results obtained by the anomaly detection model includes: determining the child nodes corresponding to the target resource transfer feature set based on the distribution results; counting the number of child nodes corresponding to the target resource transfer feature set to obtain the path length of the target resource transfer feature set in the feature distribution structure tree; determining the first anomaly detection value corresponding to the target resource transfer record based on the path length, wherein the first anomaly detection value is negatively correlated with the path length; and determining the model detection result of the target resource transfer record based on the first anomaly detection value.

[0116] The child nodes corresponding to the target resource transfer feature set can be any nodes from the root node to the leaf node of the target resource transfer feature set. The number of these child nodes can be determined as the path length of the target resource transfer record in the feature distribution structure tree. The first anomaly detection value is a detection value that can assess whether the target resource transfer record is an anomalous resource transfer record.

[0117] In some embodiments, the process of determining the first anomaly detection value corresponding to the target resource transfer record based on the path length can be as follows: When there is only one feature distribution structure tree, the path length corresponding to the target resource transfer record determined by the feature distribution structure tree is determined. An exponential function is constructed with the path length as the exponent and a preset constant value as the base. The path length corresponding to the target resource transfer record is substituted into the exponential function, and the resulting function value is the first anomaly detection value. When there are multiple feature distribution structure trees, the expected value of the path length corresponding to the target resource transfer record is determined based on the path length. An exponential function is constructed with the expected value of the path length as the exponent and a preset constant value as the base. The expected value of the path length corresponding to the target resource transfer record is substituted into the exponential function, and the resulting function value is the first anomaly detection value.

[0118] The anomaly detection value is obtained based on the path length of the child node corresponding to the target resource transfer record in the feature distribution structure tree. The shorter the path length, that is, the fewer times it needs to be segmented, the further the resource transfer feature corresponding to the target resource transfer record is from the normal data point, and the larger the anomaly detection value. Furthermore, to ensure that the anomaly detection value is negatively correlated with the path length, the exponent in the exponential function can be negative.

[0119] In some embodiments, the process of determining the model detection result of the anomaly detection model for the target resource transfer record based on the first anomaly detection value can be as follows: comparing the first anomaly detection value with a first anomaly detection value threshold; when the first anomaly detection value is greater than the first anomaly detection value threshold, the model detection result of the target resource transfer record is determined as an anomaly resource transfer record. The first anomaly detection value threshold can be a predetermined fixed value, or it can be calculated based on the feature values ​​of resource transfer features in the target feature set.

[0120] Specifically, the process of determining the first anomaly detection value corresponding to the target resource transfer record based on the path length can be as follows:

[0121] Construct a feature distribution structure tree based on n training samples, and determine the average path length of the feature distribution structure tree using the following formula:

[0122]

[0123] in, It is the harmonic number, and its value can be estimated as .

[0124] For the target resource transfer feature x of the target resource transfer record, the first anomaly detection value corresponding to the target resource transfer record is determined by the following formula:

[0125]

[0126] in, Let x be the expected path length of a sample x in a set of feature distribution structure trees.

[0127] The anomaly detection value determination method in the above embodiments can reliably detect anomalies even without determining the distance or density between resource transfer features, based on the node partitioning of the feature distribution structure tree. Compared with distance and density calculation, it greatly reduces computational consumption and has the advantages of near-linear complexity and low memory consumption.

[0128] In some embodiments, the construction process of the feature distribution structure tree can involve building multiple feature distribution structure trees based on multiple training samples. The training samples may not have corresponding labels, thus achieving the construction of the feature distribution structure trees in an unsupervised manner. The specific process of constructing the feature distribution structure trees is described below:

[0129] Given \(n\) sample data \(X = \{x_1, x_2, \ldots, x_n\}\), and these \(n\) sample data are resource transfer features in \(d\) dimensions. Randomly select a resource transfer feature \(q\) and its splitting value \(p\), and recursively split the data set \(X\), that is, based on the splitting value \(p\), divide the sample data corresponding to the current child node into two or more nodes until any of the following conditions is met: 1. The tree reaches the restricted height; 2. There is only one sample on the node; 3. All features of the samples on the node are the same.

[0130] Furthermore, assume that \(T\) is a node of the feature distribution structure tree. \(T\) can be a leaf node or an internal node with child nodes \((T_l, T_r)\).

[0131] In some embodiments, the process of gradually splitting the resource transfer features within a certain feature set can be as Figure 5 shown. For each split, determine the resource transfer feature \(q\) and the splitting value \(p\) in the feature dimension, divide the corresponding resource transfer feature \(q\) into the corresponding interval based on the splitting value \(p\), and further split the resource transfer features within the interval. The splitting line corresponding to \(p\) is as Figure 3 shown by the dashed line in. If a certain resource transfer feature \(q < p\), then divide this resource transfer feature into \(T_l\); if a certain resource transfer feature \(q\geq p\), then divide this resource transfer feature into \(T_r\). The process of successive splitting can be as Figure 3 shown. For the sake of saving space, Figure 3 only shows the splitting process on one side in.

[0132] Furthermore, for the target resource transfer feature \(x\) of the target resource transfer record, determine the path length of the sample point \(x\), that is, determine the number of edges passed from the root node to the leaf node of the feature distribution structure tree. The path length can be determined by means of binary search. Output the feature distribution structure tree set based on the constructed feature distribution structure tree, which is the feature distribution structure forest. Then, the first anomaly detection value corresponding to the target resource transfer record can be determined based on the feature distribution structure forest.

[0133] In some embodiments, the distribution partitioning method corresponding to the feature set includes a partitioning method based on distribution intervals. Obtaining the distribution partitioning method of the target feature set by the anomaly detection model, and partitioning the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method, to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension includes: obtaining the feature partitioning interval set corresponding to the target feature set in the anomaly detection model, the feature partitioning interval set including multiple feature partitioning intervals; obtaining the number of resource transfer features in each feature partitioning interval of the target feature set; determining the distribution density corresponding to the feature partitioning interval based on the number of features, and using the distribution density as the distribution result of the target resource transfer features in the target feature set of the corresponding feature dimension.

[0134] Each feature partition corresponds to a range of feature values ​​for resource transfer characteristics. The width of the feature value ranges for each feature partition can be the same or different. Distribution density may include probability density, etc.

[0135] In some embodiments, a target feature set corresponding to each target feature dimension can be determined, and the distribution density corresponding to the feature division interval can be determined based on the resource transfer features in the target feature set, thereby obtaining each distribution result. After obtaining the distribution results corresponding to each target feature dimension, these distribution results can be statistically analyzed, and the overall distribution result of the resource transfer features of the target resource transfer record can be obtained based on the statistical results.

[0136] The above embodiments divide the resource transfer features in the target feature set based on multiple feature division intervals, determine the corresponding distribution density based on the number of features of the resource transfer features in each feature division interval, and then obtain the distribution result. Even without label information, the resource transfer features of the target resource transfer record can be accurately divided based on the distribution density, and thus an accurate distribution result can be obtained.

[0137] In some embodiments, determining the model detection result of the anomaly detection model for the target resource transfer record based on the distribution results obtained by the anomaly detection model includes: determining the feature anomaly detection value corresponding to the target resource transfer feature based on the distribution density, wherein the feature anomaly detection value is negatively correlated with the distribution density; statistically analyzing the feature anomaly detection values ​​corresponding to each target resource transfer feature in the target resource transfer feature set to obtain the second anomaly detection value corresponding to the target resource transfer record; and determining the model detection result of the anomaly detection model for the target resource transfer record based on the second anomaly detection value.

[0138] Among them, the feature anomaly detection value is a detection value that can assess whether the target resource transfer record is an abnormal resource transfer record. Furthermore, the feature anomaly detection value can be obtained by performing specific statistical operations on the feature anomaly detection value corresponding to the resource transfer feature determined by the distribution density, such as performing reciprocal processing or exponential calculation. Specifically, to ensure that the feature anomaly detection value is negatively correlated with the distribution density, the reciprocal of the distribution density can be determined as the feature anomaly detection value.

[0139] The second anomaly detection value is a detection value that can assess whether the target resource transfer record is an abnormal resource transfer record.

[0140] Furthermore, after determining the feature anomaly detection value corresponding to each resource transfer feature in the target resource transfer feature set corresponding to the target resource transfer record, the feature anomaly detection value corresponding to each target feature dimension can be determined. The feature anomaly detection values ​​under each target feature dimension are statistically analyzed to obtain the second anomaly detection value corresponding to the target resource transfer record.

[0141] Specifically, for the target resource transfer feature set p corresponding to the target resource transfer record, the feature anomaly detection value can be represented as a probability density. When the target feature has d dimensions, the second anomaly detection value corresponding to the target resource transfer record can be determined by the following formula:

[0142]

[0143] In some embodiments, the process of determining the model detection result of the anomaly detection model for the target resource transfer record based on the second anomaly detection value can be as follows: comparing the second anomaly detection value with a second anomaly detection value threshold; when the second anomaly detection value is greater than the second anomaly detection value threshold, the model detection result of the target resource transfer record is determined as an anomaly resource transfer record. The second anomaly detection value threshold can be a predetermined fixed value, or it can be obtained by calculating the feature values ​​of resource transfer features in the target feature set.

[0144] The above embodiments obtain anomaly detection values ​​based on distribution density, and then obtain the distribution results corresponding to the anomaly detection model based on the anomaly detection values. Even without label information, the resource transfer characteristics of the target resource transfer record can be accurately divided based on distribution density, the resource transfer characteristics with low distribution density can be determined, and thus an accurate distribution result can be obtained.

[0145] In some embodiments, anomaly detection is performed on the target resource transfer feature set using an anomaly detection model to obtain the model detection result of the target resource transfer record. This includes: obtaining a reference cluster, which is obtained by clustering resource transfer records based on resource transfer features, and the number of resource transfer records in the reference cluster is greater than the record number threshold corresponding to the normal record cluster; determining the distance between the target resource transfer record and the reference cluster based on the target resource transfer feature set; determining the record anomaly degree corresponding to the target resource transfer record based on the distance, where the record anomaly degree is positively correlated with the distance; and determining the model detection result of the target resource transfer record based on the record anomaly degree.

[0146] The reference cluster serves as a reference for determining the model detection results of target resource transfer records. A normal record cluster refers to a cluster where resource transfer records have a higher probability of being normal compared to records in other clusters. The record quantity threshold corresponding to the normal record cluster is determined based on the total number of target resource transfer records in the target resource transfer record set. This threshold ensures that the number of target resource transfer records in the normal record cluster constitutes an absolute majority of the total number; this absolute majority is a configurable parameter. Its value ranges from 0.5 to 1, and is generally taken as 0.9. The number of target resource transfer records in the reference cluster accounts for the vast majority of the total number, so the reference cluster can also be called a large cluster, and other clusters outside the large cluster are called small clusters.

[0147] In some embodiments, parameter clusters can be generated through the following steps: First, resource transfer records are clustered based on resource transfer characteristics to obtain multiple clusters. The number of target resource transfer records in each cluster is counted, and clusters with a number greater than a record count threshold are identified as reference clusters. In a specific embodiment, the clustering method can employ the k-means clustering algorithm. The resource transfer records used in the clustering process can be historical resource transfer records or target resource transfer records.

[0148] like Figure 6 The image shown is a clustering result diagram obtained by clustering resource transfer records based on resource transfer features under the target feature dimension in a specific embodiment. (Refer to...) Figure 7Clustering yields four clusters: C1, C2, C3, and C4. C2 and C4 are large clusters, while C1 and C3 are small clusters. For a target resource transfer record in a set of target resource transfer records, its distance to each cluster can be calculated. Understandably, the closer the distance to the center of C2 or C4 (i.e., the cluster center in k-means), the more normal the target resource transfer record is, and vice versa. When the shortest distance among the calculated distances is the distance to C2 or C4, the target resource transfer record is considered a normal resource transfer record.

[0149] Based on this, the server can determine the feature vector of a target resource transfer record according to the target resource transfer feature set, calculate the distance between the feature vector and each reference cluster, and determine the record anomaly degree corresponding to the target resource transfer record based on the minimum distance value. The record anomaly degree is positively correlated with the distance, that is, the greater the distance, the greater the record anomaly degree. In specific implementation, the distance in the embodiments of this application can be Euclidean distance.

[0150] Furthermore, the server can determine the model detection result of the target resource transfer record based on the record anomaly degree. Specifically, the server can use the record anomaly degree as the model detection result of the target resource transfer record; alternatively, the server can determine whether the target resource transfer record is normal or abnormal based on the relationship between the record anomaly degree and a preset anomaly degree threshold. When the record anomaly degree is greater than the preset anomaly degree threshold, the obtained model detection result is abnormal; otherwise, the obtained model detection result is normal. In other embodiments, the server can also statistically analyze the record anomaly degree of each target resource transfer feature in the target resource transfer record set, and determine the model detection result of the target resource transfer record with the largest preset proportional distance as abnormal.

[0151] In the above embodiments, by obtaining a reference cluster, the distance between the target resource transfer record and the reference cluster is determined based on the target resource transfer feature set. The record anomaly degree corresponding to the target resource transfer record is determined based on the distance. The record anomaly degree is positively correlated with the distance. The model detection result of the target resource transfer record is determined based on the record anomaly degree. Anomaly detection can be performed by combining the overall feature classification of the resource transfer record. The obtained model detection result can reflect whether the resource transfer record is abnormal as a whole.

[0152] In some embodiments, anomaly detection is performed on the target resource transfer feature set using an anomaly detection model to obtain the model detection result of the target resource transfer record. This includes: first, calculating the K-nearest distance of the target resource transfer record in the resource transfer record set; second, calculating the reachability distance of the target resource transfer record based on the K-nearest distance; third, calculating the local reachability density based on the reachability distance; and finally, calculating the local anomaly factor based on the local reachability density. The calculated local anomaly factor is then determined as the model detection result of the target resource transfer record. The resource transfer record set may include historical resource transfer records or target resource transfer records.

[0153] Among the points closest to data point p, the distance between the k-th nearest point and p is called the K-nearest neighbor distance of p, denoted as k-distance(p). The definition of reachable distance is related to the K-nearest neighbor distance. Given the parameter k, the reachable distance reach_dist(p, o) from data point p to data point o is the maximum value of the K-nearest neighbor distance of o and the direct distance between data point p and o. That is:

[0154]

[0155] Local reachability density is defined based on reachability distance. For a data point p, those data points whose distance to p is less than or equal to k-distance(p) are called its k-nearest-neighbor, denoted as . Local reachability density of data point p It is the reciprocal of its average reachability distance to neighboring data points, i.e.:

[0156]

[0157] According to the definition of local reachability density, if a data point is relatively distant from other points, then its local reachability density is obviously low. However, the anomaly of a data point is not measured by its absolute local density, but by its relative density with its neighboring data points, thus allowing for uneven data distribution and varying densities. The local anomaly factor is defined using local relative density. The local relative density (local anomaly factor) of data point p is equal to the average local reachability density of p's neighbors. Locally accessible density of data point p The ratio, that is:

[0158]

[0159] In some embodiments, the model detection results of the target resource transfer record are statistically analyzed to obtain the anomaly detection results of the target resource transfer record, including: determining the number of abnormal results in each model detection result of the target resource transfer record where the model detection result is abnormal; when the number of abnormal results exceeds the anomaly number threshold, the target resource transfer record is determined to be an abnormal resource transfer record.

[0160] Here, "abnormal" in model detection refers to the model detecting that the target resource transfer record is an abnormal resource transfer record. The threshold for the number of abnormal records can be determined based on the actual situation. It can be a pre-set fixed threshold, or it can be determined based on the number of target resource transfer records. For example, the number of target resource transfer records can be multiplied by a fixed coefficient, and the product can be used as the threshold for the number of abnormal records.

[0161] In some embodiments, when the number of abnormal results exceeds an abnormality threshold, the server can determine that the abnormality detection model exceeding the abnormality threshold identifies the target resource transfer record as an abnormal resource transfer record.

[0162] In some embodiments, when there are multiple target resource transfer records, if the number of abnormal results exceeds the abnormal number threshold, all target resource transfer records can be identified as abnormal resource transfer records, or the target resource transfer records corresponding to the number of abnormal results exceeding the abnormal number threshold can be identified as abnormal resource transfer records.

[0163] In the above embodiments, when the model detection results exceeding the abnormality threshold are determined to be abnormal, the target resource transfer record is identified as an abnormal resource transfer record. The results of multiple abnormality detection models are integrated to obtain the abnormality detection result. Compared with obtaining the abnormality detection result through a single abnormality detection model, the result obtained in this way has higher accuracy.

[0164] In some embodiments, the model detection result can be represented by the probability value of the target resource transfer record being an abnormal resource transfer record. Further, the model detection results of each anomaly detection model in the anomaly detection model set for the target resource transfer record are statistically analyzed to obtain the anomaly detection result of the target resource transfer record. This includes: determining the model probability of each anomaly detection model for the target resource transfer record being an abnormal resource transfer record based on the model detection results of each anomaly detection model; obtaining comparison information of the model probability of each anomaly detection model relative to the corresponding probability threshold, and converting the model probability of each anomaly detection model into a voting result of the target resource transfer record being an abnormal resource transfer record based on the comparison information; and statistically analyzing the voting results corresponding to each anomaly detection model to obtain the anomaly detection result of the target resource transfer record.

[0165] The probability threshold can be determined based on the actual situation. It can be a pre-set fixed threshold or determined based on the model probabilities corresponding to each anomaly detection model. For example, the average model probabilities of each anomaly detection model can be used as the probability threshold. Further, a rough range of probability thresholds can be determined, and each probability threshold within this range can be used as a candidate probability threshold. This candidate probability thresholds are then iterated through one by one. Based on the selected candidate probability threshold, the anomaly detection result for the user is determined. The anomaly detection result is compared with the corresponding user's tag in the tag database. If the two match, for example, both indicate that the corresponding user is an abnormal resource transfer record, then the selected candidate probability threshold is considered appropriate and is used as the probability threshold for the anomaly detection model. If the two do not match, then the selected candidate probability threshold is considered inappropriate, and the process iterates through the next candidate probability threshold until an appropriate candidate probability threshold is selected. By iterating through multiple candidate probability thresholds, a suitable probability threshold can be determined, and an accurate and reliable anomaly detection model can be obtained based on the selected probability threshold.

[0166] In some embodiments, statistical analysis of the voting results can involve determining the number of votes that identify an abnormal resource transfer record. When the number of votes exceeds a voting threshold, the anomaly detection result of the target resource transfer record is determined as an abnormal resource transfer record. When the number of votes is less than or equal to the voting threshold, the anomaly detection result of the target resource transfer record is determined as a normal resource transfer record. The voting threshold can be determined based on actual circumstances; it can be a pre-set fixed threshold or determined based on the number of anomaly detection models. For example, if the total number of anomaly detection models is used as the voting threshold, the server will only determine the target resource transfer record as an abnormal resource transfer record if the voting results corresponding to all anomaly detection models are abnormal resource transfer records.

[0167] The above embodiments determine the voting results corresponding to each anomaly detection model, and statistically analyze these voting results to integrate the voting information of multiple anomaly detection models to obtain accurate anomaly detection results.

[0168] like Figure 7 The diagram shown is a flowchart illustrating the resource processing method provided in this application embodiment, in a specific instance. (Reference) Figure 7The server first accesses the current resource transfer business scenario, and initially selects the feature dimensions corresponding to the resource transfer records from the database corresponding to the business scenario to obtain a set of candidate feature dimensions. Then, it sorts each candidate feature dimension in the set according to the dimension anomaly degree, and selects a preset number of target feature dimensions based on the sorting results. When anomaly identification is required, the server obtains the resource transfer features of the target resource transfer record to be identified on each target feature dimension to obtain the target resource transfer feature set of the target resource transfer record. Then, it performs anomaly detection on the target resource transfer feature set based on three different anomaly detection models. Finally, the model detection results output by the three models are statistically analyzed to fuse the detection results of the three models to obtain the anomaly identification result of the target resource transfer record.

[0169] This application also provides an application scenario in which the resource transfer data processing method described above is applied. In this application scenario, the resource transfer data processing method provided in the embodiments of this application can be used to identify fraudulent transactions. The transaction data record generated by each fraudulent transaction is the resource transfer record in the embodiments of this application. By identifying anomalies in this transaction data record, it can be determined whether the transaction is a fraudulent transaction. Fraudulent transactions generally involve sellers providing purchase fees to help designated online store sellers buy goods to increase sales and credit ratings, and filling in fake positive reviews. In this way, online stores can obtain better search rankings. For example, when searching "by sales volume" on the platform, the store with high sales volume (even if fake) will be more easily found by buyers.

[0170] Specifically, the application of this resource transfer data processing method in this application scenario is as follows:

[0171] (a) Determine the set of target feature dimensions.

[0172] 1. The server obtains the first feature distribution value of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set, and obtains the representative feature distribution value corresponding to the historical feature set.

[0173] 2. The server obtains the feature anomaly degree corresponding to the first resource transfer feature based on the difference between the first feature distribution value and the representative feature distribution value.

[0174] Specifically, the number of times the first resource transfer feature and the second resource transfer feature of different feature dimensions are co-occurred in the historical resource transfer record set is obtained. Based on the number of co-occurrences and the feature anomaly degree, the anomaly transmission weight between the first resource transfer feature and the second resource transfer feature is obtained. Based on the anomaly transmission weight, the feature anomaly degree of the second resource transfer feature is transmitted to the first resource transfer feature to obtain the transmission anomaly degree of the first resource transfer feature. The transmission anomaly degree of the first resource transfer feature is calculated to obtain the dimension anomaly degree of the candidate feature dimension.

[0175] When transferring the feature anomaly degree of the second resource transfer feature to the first resource transfer feature based on the anomaly propagation weight to obtain the propagation anomaly degree of the first resource transfer feature, the specific implementation is as follows: taking the resource transfer features of each historical resource transfer record in the historical resource transfer record set as nodes, connecting resource transfer features with co-occurrence relationships to obtain a feature connection graph, wherein the nodes of the second resource transfer feature and the nodes of the first resource transfer feature in the feature connection graph have connection edges; in the feature connection graph, the feature anomaly degree of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge, and the feature anomaly degree of the first resource transfer feature when the iteration stopping condition is met is taken as the propagation anomaly degree of the first resource transfer feature.

[0176] 3. The server obtains the dimension anomaly degree corresponding to the candidate feature dimension based on the feature anomaly degree corresponding to the first resource transfer feature.

[0177] 4. The server selects the target feature dimension from the candidate feature dimension set based on the dimensional anomaly of the candidate feature dimension, and forms the target feature dimension set.

[0178] (ii) Anomaly detection.

[0179] refer to Figure 8 The anomaly detection process includes: First, the server selects resource transfer features of the target feature dimension from the target resource transfer records (i.e., feature selection) to obtain a target resource transfer feature set. Then, it performs anomaly detection based on three anomaly detection models in the anomaly detection model set, obtaining the detection results of each model. Finally, the detection results of the three models are statistically analyzed and fused to obtain the anomaly detection result. The specific steps are as follows:

[0180] 1. The server obtains a set of target resource transfer records to be identified. The set of target resource transfer records includes multiple target resource transfer records. The server obtains the target resource transfer features of each target resource transfer record in the target feature dimension and forms a target resource transfer feature set corresponding to the target resource transfer record.

[0181] 2. The server determines a set of anomaly detection models, which includes three different anomaly detection models. These models are used to perform anomaly detection on the target resource transfer feature set, yielding the model detection results for the target resource transfer records. Specifically:

[0182] 1) Anomaly detection is performed using anomaly detection model 1. The server obtains a feature distribution structure tree, which includes multiple child nodes. The initial node of the feature distribution structure tree is used as the current child node corresponding to the target resource transfer feature set. The current feature dimension corresponding to the current child node is obtained, and the current feature partitioning threshold of the current feature set corresponding to the current feature dimension is obtained. Based on the current feature partitioning threshold and the resource transfer features of the target resource transfer feature set in the current feature dimension, the distribution result of the target resource transfer features in the current feature set is determined. Based on the distribution result, the next child node corresponding to the target resource transfer feature set is determined, and the next node is used as the updated current child node. The steps of obtaining the current feature dimension corresponding to the current child node and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension are returned until all child nodes corresponding to the target resource transfer feature set are updated. The number of child nodes corresponding to the target resource transfer feature set is counted to obtain the path length of the target resource transfer feature set in the feature distribution structure tree. Based on the path length, the first anomaly detection value corresponding to the target resource transfer record is determined. The first anomaly detection value is negatively correlated with the path length. Based on the first anomaly detection value, the model detection result of the target resource transfer record is determined.

[0183] 2) Anomaly detection is performed using anomaly detection model 2. The server obtains the feature partitioning interval set corresponding to the target feature set in the anomaly detection model. The feature partitioning interval set includes multiple feature partitioning intervals. The number of resource transfer features in the target feature set in each feature partitioning interval is obtained. Based on the number of features, the distribution density corresponding to the feature partitioning interval is determined. The distribution density is used as the distribution result of the target resource transfer feature in the target feature set of its respective feature dimension. Based on the distribution density, the feature anomaly detection value corresponding to the target resource transfer feature is determined. The feature anomaly detection value is negatively correlated with the distribution density. The feature anomaly detection values ​​corresponding to each target resource transfer feature in the target resource transfer feature set are statistically analyzed to obtain the second anomaly detection value corresponding to the target resource transfer record. Based on the second anomaly detection value, the model detection result of the anomaly detection model for the target resource transfer record is determined.

[0184] 3) Anomaly detection is performed using anomaly detection model 2. The server obtains a reference cluster, which is obtained by clustering resource transfer records based on resource transfer features. The number of resource transfer records in the reference cluster is greater than the threshold number of records corresponding to the normal record cluster. The distance between the target resource transfer record and the reference cluster is determined based on the target resource transfer feature set. The record anomaly degree corresponding to the target resource transfer record is determined based on the distance. The record anomaly degree is positively correlated with the distance. The model detection result of the target resource transfer record is determined based on the record anomaly degree.

[0185] When integrating the detection results of the three models, the server determines the number of abnormal results in each model detection result of the target resource transfer record. When the number of abnormal results exceeds the abnormal number threshold, the target resource transfer record is determined to be an abnormal resource transfer record.

[0186] The method provided in this application can accurately identify fraudulent transaction data and further identify the users corresponding to such data as abnormal users. When fraudulent transactions are identified, the server can send a notification to the second terminal. The server can also restrict the transactions of abnormal users' accounts to prevent them from engaging in fraudulent transactions again within a preset time period.

[0187] It should be understood that, although Figure 2 and Figure 8 The steps in the flowchart are shown sequentially as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order in which these steps are executed, and they can be performed in other orders. Figure 2 and Figure 8 At least some of the steps in the process may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but may be executed at different times. The execution order of these steps or stages is not necessarily sequential, but may be executed in turn or alternately with other steps or at least some of the steps or stages in other steps.

[0188] In some embodiments, such as Figure 9 As shown, a resource transfer data processing device 900 is provided. This device can be a software module, a hardware module, or a combination of both as part of a computer device. Specifically, the device includes:

[0189] The target feature dimension acquisition module 902 is used to acquire the target feature dimension set; the target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the dimensional anomaly is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set;

[0190] The resource transfer feature selection module 904 is used to obtain a set of target resource transfer records to be identified. The set of target resource transfer records includes multiple target resource transfer records. The module obtains the target resource transfer features of each target resource transfer record in the target feature dimension and forms a set of target resource transfer features corresponding to the target resource transfer record.

[0191] The detection model determination module 906 is used to determine the anomaly detection model set, which includes multiple different anomaly detection models;

[0192] Anomaly detection module 908 is used to perform anomaly detection on the target resource transfer feature set through anomaly detection model, and obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein, at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension;

[0193] The detection result statistics module 910 is used to statistically analyze the model detection results of the target resource transfer records and obtain the anomaly detection results of the target resource transfer records.

[0194] The aforementioned resource transfer data processing device, on the one hand, employs multiple different anomaly detection models for anomaly detection and comprehensively statistically analyzes the model detection results corresponding to these anomaly detection models to obtain the anomaly detection results of the target resource transfer record. This allows for the determination of the anomaly detection results of the target resource transfer record based on multiple different anomaly detection strategies, effectively improving the accuracy of resource transfer record identification. On the other hand, during anomaly detection, the device acquires the target resource transfer features of the target resource transfer record in the target feature dimension set to form a target resource transfer feature set for anomaly detection. The target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly degree of the candidate feature dimension. The dimensional anomaly degree is determined based on the feature distribution of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set. Therefore, the distribution result can well reflect the anomaly of the resource transfer features. Since at least one of the multiple anomaly detection models obtains the model detection result based on the distribution result of the target resource transfer features in the target feature set corresponding to the feature dimension, the anomaly of the resource transfer features is fully considered during the anomaly detection process, further improving the accuracy of resource transfer record identification.

[0195] In some embodiments, the above apparatus further includes: a dimension anomaly acquisition module, configured to acquire a first feature distribution value of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set, and acquire a representative feature distribution value corresponding to the historical feature set; obtain the feature anomaly corresponding to the first resource transfer feature based on the difference between the first feature distribution value and the representative feature distribution value; and obtain the dimension anomaly corresponding to the candidate feature dimension based on the feature anomaly corresponding to the first resource transfer feature.

[0196] In some embodiments, the dimension anomaly acquisition module is further configured to acquire the number of times the first resource transfer feature and the second resource transfer feature with different feature dimensions co-occur in the historical resource transfer record set; obtain the anomaly transmission weight between the first resource transfer feature and the second resource transfer feature based on the co-occurrence number and feature anomaly; transmit the feature anomaly of the second resource transfer feature to the first resource transfer feature based on the anomaly transmission weight to obtain the transmission anomaly of the first resource transfer feature; and calculate the transmission anomaly of the first resource transfer feature to obtain the dimension anomaly of the candidate feature dimension.

[0197] In some embodiments, the dimension anomaly acquisition module is further configured to use the resource transfer features of each historical resource transfer record in the historical resource transfer record set as nodes, connect resource transfer features that have a co-occurrence relationship to obtain a feature connection graph, wherein the nodes of the second resource transfer feature in the feature connection graph have connection edges with the nodes of the first resource transfer feature; in the feature connection graph, the feature anomaly of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge, and the feature anomaly of the first resource transfer feature when the iteration stopping condition is met is taken as the propagation anomaly of the first resource transfer feature.

[0198] In some embodiments, the anomaly detection module is further configured to: obtain resource transfer features corresponding to feature dimensions from the target resource transfer feature set corresponding to the target resource transfer record; obtain the distribution partitioning method of the target feature set by the anomaly detection model; partition the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method; obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension; and determine the model detection result of the anomaly detection model for the target resource transfer record based on the distribution result obtained by the anomaly detection model.

[0199] In some embodiments, the distribution partitioning method corresponding to the feature set includes a threshold-based partitioning method. The anomaly detection module is further configured to obtain a feature distribution structure tree, which includes multiple child nodes; take the initial node of the feature distribution structure tree as the current child node corresponding to the target resource transfer feature set; obtain the current feature dimension corresponding to the current child node; obtain the current feature partitioning threshold of the current feature set corresponding to the current feature dimension; determine the distribution result of the target resource transfer feature in the current feature set based on the current feature partitioning threshold and the resource transfer feature of the target resource transfer feature set in the current feature dimension; determine the next child node corresponding to the target resource transfer feature set based on the distribution result; take the next node as the updated current child node; return to the steps of obtaining the current feature dimension corresponding to the current child node and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension, until the child nodes corresponding to the target resource transfer feature set are updated.

[0200] In some embodiments, the anomaly detection module is further configured to: determine the child nodes corresponding to the target resource transfer feature set based on the distribution results; count the number of child nodes corresponding to the target resource transfer feature set to obtain the path length of the target resource transfer feature set in the feature distribution structure tree; determine the first anomaly detection value corresponding to the target resource transfer record based on the path length, wherein the first anomaly detection value is negatively correlated with the path length; and determine the model detection result of the target resource transfer record based on the first anomaly detection value.

[0201] In some embodiments, the distribution partitioning method corresponding to the feature set includes a partitioning method based on distribution intervals. The anomaly detection module is further used to obtain the feature partitioning interval set corresponding to the target feature set in the anomaly detection model. The feature partitioning interval set includes multiple feature partitioning intervals. The module obtains the number of resource transfer features in each feature partitioning interval in the target feature set. Based on the number of features, the module determines the distribution density corresponding to the feature partitioning interval and uses the distribution density as the distribution result of the target resource transfer features in the target feature set of the feature dimension.

[0202] In some embodiments, the anomaly detection module is further configured to determine the feature anomaly detection value corresponding to the target resource transfer feature based on the distribution density, wherein the feature anomaly detection value is negatively correlated with the distribution density; to statistically analyze the feature anomaly detection values ​​corresponding to each target resource transfer feature in the target resource transfer feature set to obtain the second anomaly detection value corresponding to the target resource transfer record; and to determine the model detection result of the anomaly detection model for the target resource transfer record based on the second anomaly detection value.

[0203] In some embodiments, the anomaly detection module is further configured to obtain a reference cluster, which is obtained by clustering resource transfer records based on resource transfer features, and the number of resource transfer records in the reference cluster is greater than the record number threshold corresponding to the normal record cluster; determine the distance between the target resource transfer record and the reference cluster based on the target resource transfer feature set; determine the record anomaly degree corresponding to the target resource transfer record based on the distance, and the record anomaly degree is positively correlated with the distance; and determine the model detection result of the target resource transfer record based on the record anomaly degree.

[0204] In some embodiments, the detection result statistics module is further used to determine the number of abnormal results in each model detection result of the target resource transfer record where the model detection result is abnormal; when the number of abnormal results exceeds the abnormal number threshold, the target resource transfer record is determined to be an abnormal resource transfer record.

[0205] Specific limitations regarding the resource transfer data processing device can be found in the limitations of the resource transfer data processing method described above, and will not be repeated here. Each module in the aforementioned resource transfer data processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0206] In some embodiments, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown, the computer device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The database stores resource transfer data processing data. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a resource transfer data processing method.

[0207] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0208] In some embodiments, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.

[0209] In some embodiments, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above method embodiments.

[0210] In some embodiments, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and executes the computer instructions, causing the computer device to perform the steps in the above-described method embodiments.

[0211] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0212] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0213] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for processing resource transfer data, characterized in that, The method for identifying fraudulent transactions includes: Obtain a target feature dimension set; the target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions; the step of obtaining the dimensional anomaly of the candidate feature dimensions includes: obtaining a co-occurrence probability based on the proportion of the co-occurrence frequency of the first resource transfer feature of the candidate feature dimension and the second resource transfer feature of different feature dimensions in the total number of historical resource transfer records; obtaining the feature co-occurrence degree between the first resource transfer feature and the second resource transfer feature based on the co-occurrence probability; obtaining the first feature distribution value of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set, and obtaining the representative feature distribution value corresponding to the historical feature set; obtaining the feature anomaly degree corresponding to the first resource transfer feature based on the difference between the first feature distribution value and the representative feature distribution value; obtaining the anomaly propagation weight based on the feature co-occurrence degree and the feature anomaly degree; and transferring the historical resource transfer records... Resource transfer features of each historical resource transfer record in the record set are used as nodes. Resource transfer features with co-occurrence relationships are connected to obtain a feature connection graph. In the feature connection graph, the nodes of the second resource transfer feature are connected to the nodes of the first resource transfer feature. In the feature connection graph, the feature anomaly degree of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge. The feature anomaly degree of the first resource transfer feature when the iteration stops is taken as the propagation anomaly degree of the first resource transfer feature. The propagation anomaly degree of the first resource transfer feature is calculated to obtain the dimension anomaly degree of the candidate feature dimension. The feature anomaly degree is determined according to the feature distribution of the first resource transfer feature in the corresponding historical feature set. Different resource transfer features correspond to different fields. One field or multiple fields of the same type correspond to one feature dimension. Obtain a set of target resource transfer records to be identified, the set of target resource transfer records includes multiple target resource transfer records, obtain the target resource transfer features of each target resource transfer record on the target feature dimension, and form a target resource transfer feature set corresponding to the target resource transfer record; Determine an anomaly detection model set, which includes multiple different anomaly detection models; The anomaly detection model is used to perform anomaly detection on the target resource transfer feature set to obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein, at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension; The model detection results of the target resource transfer records are statistically analyzed to obtain the anomaly detection results of the target resource transfer records.

2. The method according to claim 1, characterized in that, The step of performing anomaly detection on the target resource transfer feature set using the anomaly detection model to obtain the model detection result of the anomaly detection model on the target resource transfer record includes: Obtain the resource transfer features corresponding to the feature dimensions from the target resource transfer feature set corresponding to the target resource transfer record, and obtain the target feature set corresponding to each feature dimension respectively; Obtain the distribution partitioning method of the target feature set by the anomaly detection model, and partition the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension; The model detection result of the anomaly detection model for the target resource transfer record is determined based on the distribution results obtained from the anomaly detection model.

3. The method according to claim 2, characterized in that, The distribution partitioning method corresponding to the feature set includes a threshold-based partitioning method. The method of obtaining the distribution partitioning method of the target feature set by the anomaly detection model, and partitioning the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method, to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension includes: Obtain the feature distribution structure tree, which includes multiple child nodes; The initial node of the feature distribution structure tree is taken as the current child node corresponding to the target resource transfer feature set. The current feature dimension corresponding to the current child node is obtained, and the current feature partitioning threshold of the current feature set corresponding to the current feature dimension is obtained. Based on the current feature segmentation threshold and the resource transfer features of the target resource transfer feature set in the current feature dimension, determine the distribution result of the target resource transfer features in the current feature set; Based on the distribution results, determine the next child node corresponding to the target resource transfer feature set, take the next node as the updated current child node, return to the steps of obtaining the current feature dimension corresponding to the current child node, and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension, until the child nodes corresponding to the target resource transfer feature set are updated.

4. The method according to claim 3, characterized in that, The determination of the model detection result of the anomaly detection model for the target resource transfer record based on the distribution result obtained by the anomaly detection model includes: Based on the distribution results, the child nodes corresponding to the target resource transfer feature set are determined; the number of child nodes corresponding to the target resource transfer feature set is counted to obtain the path length of the target resource transfer feature set in the feature distribution structure tree; A first anomaly detection value corresponding to the target resource transfer record is determined based on the path length, and the first anomaly detection value is negatively correlated with the path length. The model detection result of the target resource transfer record is determined based on the first anomaly detection value.

5. The method according to claim 2, characterized in that, The distribution partitioning method corresponding to the feature set includes a partitioning method based on distribution intervals. The method of obtaining the distribution partitioning method of the target feature set by the anomaly detection model, and partitioning the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method, to obtain the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension includes: Obtain the feature partitioning interval set corresponding to the target feature set in the anomaly detection model, wherein the feature partitioning interval set includes multiple feature partitioning intervals; Obtain the number of resource transfer features in each feature partitioning interval of the target feature set; Based on the number of features, the distribution density corresponding to the feature division interval is determined, and the distribution density is used as the distribution result of the target resource transfer feature in the target feature set of the feature dimension.

6. The method according to claim 5, characterized in that, The determination of the model detection result of the anomaly detection model for the target resource transfer record based on the distribution result obtained by the anomaly detection model includes: Based on the distribution density, the feature anomaly detection value corresponding to the target resource transfer feature is determined, and the feature anomaly detection value is negatively correlated with the distribution density; The anomaly detection values ​​corresponding to each target resource transfer feature in the target resource transfer feature set are statistically analyzed to obtain the second anomaly detection value corresponding to the target resource transfer record; The anomaly detection model's detection result for the target resource transfer record is determined based on the second anomaly detection value.

7. The method according to claim 1, characterized in that, The step of performing anomaly detection on the target resource transfer feature set using the anomaly detection model to obtain the model detection results of the anomaly detection model for the target resource transfer records includes: Obtain a reference cluster, which is obtained by clustering resource transfer records based on resource transfer characteristics. The number of resource transfer records in the reference cluster is greater than the record number threshold corresponding to the normal record cluster. The distance between the target resource transfer record and the reference cluster is determined based on the target resource transfer feature set; The record anomaly degree corresponding to the target resource transfer record is determined based on the distance, and the record anomaly degree is positively correlated with the distance; The model detection result of the target resource transfer record is determined based on the record anomaly degree.

8. The method according to any one of claims 1 to 7, characterized in that, The statistical analysis of the model detection results of the target resource transfer records to obtain the anomaly detection results of the target resource transfer records includes: Determine the number of abnormal results where the model detection result is abnormal among the various model detection results of the target resource transfer record; When the number of abnormal results exceeds the abnormal number threshold, the target resource transfer record is determined to be an abnormal resource transfer record.

9. A resource transfer data processing device, characterized in that, The device for identifying fraudulent transactions includes: The target feature dimension acquisition module is used to acquire a target feature dimension set; the target feature dimension set is selected from the candidate feature dimension set based on the dimensional anomaly of the candidate feature dimensions. The dimensional anomaly acquisition module is used to obtain a co-occurrence probability based on the proportion of the co-occurrence frequency of the first resource transfer feature of the candidate feature dimension and the second resource transfer feature of different feature dimensions in the total number of historical resource transfer records; to obtain the feature co-occurrence degree between the first resource transfer feature and the second resource transfer feature based on the co-occurrence probability; to obtain the first feature distribution value of the first resource transfer feature of the candidate feature dimension in the corresponding historical feature set, and to obtain the representative feature distribution value corresponding to the historical feature set; to obtain the feature anomaly degree corresponding to the first resource transfer feature based on the difference between the first feature distribution value and the representative feature distribution value; to obtain the anomaly propagation weight based on the feature co-occurrence degree and the feature anomaly degree; and to use the resource transfer features of each historical resource transfer record in the historical resource transfer record set as nodes, and to identify nodes with co-occurrence... Resource transfer features of a relationship are connected to obtain a feature connection graph, wherein nodes of the second resource transfer feature and nodes of the first resource transfer feature are connected by edges in the feature connection graph. In the feature connection graph, the feature anomaly degree of the nodes of the first resource transfer feature is iteratively updated based on the feature anomaly degree of the second resource transfer feature and the anomaly propagation weight corresponding to the connection edge. The feature anomaly degree of the first resource transfer feature when the iteration stops is taken as the propagation anomaly degree of the first resource transfer feature. The propagation anomaly degree of the first resource transfer feature is calculated to obtain the dimensional anomaly degree of the candidate feature dimension. The feature anomaly degree is determined based on the feature distribution of the first resource transfer feature in the corresponding historical feature set. Different resource transfer features correspond to different fields; one or more fields of the same type correspond to one feature dimension. The resource transfer feature selection module is used to obtain a set of target resource transfer records to be identified, the set of target resource transfer records includes multiple target resource transfer records, and obtains the target resource transfer features of each target resource transfer record on the target feature dimension to form a target resource transfer feature set corresponding to the target resource transfer record; The detection model determination module is used to determine an anomaly detection model set, which includes multiple different anomaly detection models; An anomaly detection module is used to perform anomaly detection on the target resource transfer feature set using the anomaly detection model, and obtain the model detection result of the anomaly detection model on the target resource transfer record; wherein, at least one anomaly detection model obtains the model detection result based on the distribution result of the target resource transfer feature in the target feature set corresponding to the feature dimension; The detection result statistics module is used to statistically analyze the model detection results of the target resource transfer records to obtain the anomaly detection results of the target resource transfer records.

10. The resource transfer data processing apparatus according to claim 9, characterized in that, The anomaly detection module is further configured to obtain resource transfer features corresponding to feature dimensions from the target resource transfer feature set corresponding to the target resource transfer record, thereby obtaining target feature sets corresponding to each feature dimension; obtain the distribution partitioning method of the anomaly detection model for the target feature set, and partition the target resource transfer features in the target feature set of the corresponding feature dimension based on the distribution partitioning method, thereby obtaining the distribution result of the target resource transfer features in the target feature set corresponding to the corresponding feature dimension; and determine the model detection result of the anomaly detection model for the target resource transfer record based on the distribution result obtained by the anomaly detection model.

11. The apparatus according to claim 10, characterized in that, The distribution partitioning method corresponding to the feature set includes a threshold-based partitioning method. The anomaly detection module is further used to obtain a feature distribution structure tree, which includes multiple child nodes; taking the initial node of the feature distribution structure tree as the current child node corresponding to the target resource transfer feature set, obtaining the current feature dimension corresponding to the current child node, and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension; determining the distribution result of the target resource transfer feature in the current feature set based on the current feature partitioning threshold and the resource transfer feature of the target resource transfer feature set in the current feature dimension; determining the next child node corresponding to the target resource transfer feature set based on the distribution result, taking the next node as the updated current child node, and returning to the steps of obtaining the current feature dimension corresponding to the current child node and obtaining the current feature partitioning threshold of the current feature set corresponding to the current feature dimension, until the child nodes corresponding to the target resource transfer feature set are updated.

12. The apparatus according to claim 11, characterized in that, The anomaly detection module is further configured to: determine the child nodes corresponding to the target resource transfer feature set based on the distribution results; count the number of child nodes corresponding to the target resource transfer feature set to obtain the path length of the target resource transfer feature set in the feature distribution structure tree; determine the first anomaly detection value corresponding to the target resource transfer record based on the path length, wherein the first anomaly detection value is negatively correlated with the path length; and determine the model detection result of the target resource transfer record based on the first anomaly detection value.

13. The resource transfer data processing apparatus according to claim 10, characterized in that, The distribution partitioning method corresponding to the feature set includes a partitioning method based on distribution intervals. The anomaly detection module is also used to obtain the feature partitioning interval set corresponding to the target feature set in the anomaly detection model. The feature partitioning interval set includes multiple feature partitioning intervals. The module obtains the number of resource transfer features in the target feature set in each feature partitioning interval. Based on the number of features, the module determines the distribution density corresponding to the feature partitioning interval and uses the distribution density as the distribution result of the target resource transfer feature in the target feature set of the feature dimension.

14. The apparatus according to claim 13, characterized in that, The anomaly detection module is further configured to determine the feature anomaly detection value corresponding to the target resource transfer feature based on the distribution density, wherein the feature anomaly detection value is negatively correlated with the distribution density; to statistically analyze the feature anomaly detection values ​​corresponding to each target resource transfer feature in the target resource transfer feature set to obtain the second anomaly detection value corresponding to the target resource transfer record; and to determine the model detection result of the anomaly detection model for the target resource transfer record based on the second anomaly detection value.

15. The apparatus according to claim 9, characterized in that, The anomaly detection module is also used to obtain a reference cluster, which is obtained by clustering resource transfer records based on resource transfer features. The number of resource transfer records in the reference cluster is greater than the record number threshold corresponding to the normal record cluster. The distance between the target resource transfer record and the reference cluster is determined based on the target resource transfer feature set; the record anomaly degree corresponding to the target resource transfer record is determined based on the distance, and the record anomaly degree is positively correlated with the distance; the model detection result of the target resource transfer record is determined based on the record anomaly degree.

16. The apparatus according to any one of claims 9 to 15, characterized in that, The detection result statistics module is also used to determine the number of abnormal results in each model detection result of the target resource transfer record; when the number of abnormal results exceeds the abnormal number threshold, the target resource transfer record is determined to be an abnormal resource transfer record.

17. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 8.

18. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.

19. A computer program product comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 8.