Training sample generation method and device, electronic equipment and storage medium
By determining and fusing the correlation information of three sample sets to generate a multivariate Gaussian distribution model, the problem of inaccurate sample distribution in the existing training sample generation method is solved, and the performance of the fraud detection model is improved.
Patent Information
- Application Number
- CN202410037019.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-08
- Publication Date
- 2025-07-08
AI Technical Summary
In the existing training sample generation method, the interpolation method assumes that the sample features are uniformly distributed and the feature connection is ignored. However, the estimation distribution method is inaccurate in the sample distribution due to insufficient data volume of suspected normal transaction sample objects, which affects the model classification performance.
By determining three initial sample sets: normal transactions, abnormal transactions and suspected normal transactions, the correlation information of each sample set is calculated, the correlation information is fused to generate a multivariate Gaussian distribution model, and multiple samples are performed to generate new samples.
It improves the accuracy of feature distribution estimation of suspected normal transaction sample objects, alleviates the problem of insufficient distribution accuracy caused by insufficient sample data, and improves the classification performance of the model.
Smart Images

Figure CN120278719A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, this application relates to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating training samples. Background Art
[0002] Fraud detection refers to analyzing data and behavior patterns to discover potential fraud and taking corresponding measures to prevent losses. In recent years, in the field of fraud detection, machine learning technology has been widely used. Using machine learning algorithms, we can train an identification model using historical data and use this model to predict potential fraud samples.
[0003] Existing methods for generating training samples generally include the following solutions:
[0004] 1) Interpolation method: Assume that the sample set of the category to be generated is S. Randomly sample two sample vectors A and B from S, and then directly calculate the mean vector of A and B as the new sample;
[0005] 2) Estimated distribution method: Assume that the sample distribution of the category to be generated is a multivariate Gaussian distribution. Estimate the distribution parameters using all samples in the sample set of this category, and then sample from the estimated distribution to generate the required new samples.
[0006] The interpolation method assumes that the features of the samples are uniformly distributed and ignores the relationship between sample features. However, in real life, the sample data in various scenarios more conforms to the multivariate Gaussian distribution. The existing estimated distribution method, although assuming the multivariate Gaussian distribution as the premise of the sample distribution, is more in line with the actual situation. However, due to the insufficient amount of sample object data for suspected normal transactions, the estimated sample distribution may not be accurate enough. Therefore, the newly sampled samples may deviate from the true distribution, and using samples with deviations to train the model will damage the classification performance of the model. Summary of the Invention
[0007] Embodiments of this application provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating training samples, which can solve the above problems of the prior art. The technical solutions are as follows:
[0008] According to the first aspect of the embodiments of this application, a method for generating training samples is provided. The method includes:
[0009] Determine three initial sample sets, which include a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object.
[0010] Based on the transaction characteristics of each sample object in each initial sample set, determine the first association information of each initial sample set. The first association information is used to characterize the association degree of the characteristics of each dimension in the corresponding initial sample set.
[0011] According to the first association information of each initial sample set, determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions. Perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
[0012] According to the second aspect of the embodiments of the present application, there is provided a method for identifying an abnormal transaction object, and the method includes:
[0013] Obtain the transaction characteristics of the object to be identified.
[0014] Input the transaction characteristics of the object to be identified into a pre-trained object recognition model to obtain the recognition result output by the object recognition model.
[0015] The recognition result is used to indicate that the object to be identified is an object with a normal transaction or an object with an abnormal transaction.
[0016] The object recognition model is trained by the first sample set, the second sample set, and the third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. At least one sample in the third sample set is generated by the method provided in the first aspect.
[0017] According to the third aspect of the embodiments of the present application, there is provided a training sample generation device, and the device includes:
[0018] A sample set module, which is used to determine three initial sample sets, including a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object.
[0019] A first association module, which is used to determine the first association information of each initial sample set based on the transaction characteristics of each sample object in each initial sample set. The first association information is used to characterize the association degree of the characteristics of each dimension in the corresponding initial sample set.
[0020] A sampling module, which is used to determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the first association information of each initial sample set, and perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
[0021] As an optional implementation manner, the sampling module includes:
[0022] A fusion sub-module, which is used to fuse the first association information of each initial sample set to obtain second association information. The second association information is used to characterize the association degree of the characteristic values of each dimension in all initial sample sets.
[0023] A distribution determination sub-module, which is used to use the second association information as the association degree between the characteristics of each dimension corresponding to the sample objects with suspected normal transactions, and determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the third sample set and the second association information.
[0024] As an optional implementation manner, the first association information is a covariance matrix, and the covariance matrix includes the covariance between the characteristic values of each dimension and the characteristic values of other dimensions in the corresponding sample set.
[0025] The fusion sub-module is specifically used to perform weighted summation on the covariance matrices of each initial sample set to obtain a target covariance matrix, and the second association information is the target covariance matrix.
[0026] As an optional implementation manner, the fusion sub-module is specifically used to:
[0027] Obtain the weights corresponding to each of the three initial sample sets, where the weights corresponding to the first sample set and the second sample set are not greater than the weights corresponding to the third sample set.
[0028] Based on the weights corresponding to the three initial sample sets respectively, perform weighted fusion on the three initial sample sets to obtain second correlation information.
[0029] As an alternative implementation manner, the distribution determination sub-module is specifically configured to:
[0030] Determine the mean of the eigenvalues of each dimension of the third sample set as the sample mean of the multivariate Gaussian distribution model, and use the target covariance matrix as the covariance matrix of the multivariate Gaussian distribution model to obtain a multivariate Gaussian distribution model describing the distribution of the new samples.
[0031] As an alternative implementation manner, the number of samplings does not exceed the number of samples in any one of the initial sample sets.
[0032] As an alternative implementation manner, it further includes:
[0033] A label setting module, configured to set first training labels for the first sample set, the third sample set, and the new samples; set second training labels for the second sample set;
[0034] Wherein, the first training label is used to represent that the corresponding sample object is a sample object of a normal transaction;
[0035] The second training label is used to represent that the corresponding sample object is a sample object of an abnormal transaction.
[0036] According to the fourth aspect of the embodiments of the present application, an identification device for abnormal transaction objects is provided, including:
[0037] A to-be-identified object acquisition module, configured to acquire the transaction characteristics of the to-be-identified object;
[0038] A model judgment module, configured to input the transaction characteristics of the to-be-identified object into a pre-trained object identification model to obtain an identification result output by the object identification model;
[0039] Wherein, the identification result is used to indicate that the to-be-identified object is an object of a normal transaction or an object of an abnormal transaction;
[0040] The object identification model is trained by a first sample set, a second sample set, and a third sample set, each sample set includes a plurality of initial samples, and the initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects of normal transactions, the transaction characteristics of sample objects of abnormal transactions, and the transaction characteristics of sample objects of suspected normal transactions;
[0041] At least one sample in the third sample set is generated based on the method provided in the first aspect.
[0042] According to the fifth aspect of the embodiments of the present application, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above method.
[0043] According to the sixth aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method are implemented.
[0044] According to the seventh aspect of the embodiments of the present application, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the steps of the above method are implemented.
[0045] The beneficial effects brought by the technical solutions provided by the embodiments of the present application are as follows:
[0046] By determining three initial sample sets, which include a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object. Further, the first association information of each initial sample set is determined. The first association information is used to characterize the degree of association of the characteristic values of each dimension in the corresponding initial sample set. Since the dimensions of the transaction characteristics of all sample objects are the same, calculating the first association information can summarize the information of different initial sample sets for processing on the same scale, and the first association information can reflect to a certain extent the distribution of the characteristic values of the initial sample set. It is realized that due to the small data volume of the third sample set, the present application uses the first association information of each initial sample set to obtain the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions. This distribution can more accurately characterize the transaction characteristics of the sample objects with suspected normal transactions. Then, multiple samplings are performed according to the characteristic distribution to obtain multiple new samples of the third sample set. The embodiments of the present application can effectively alleviate the problem that the sample distribution estimation is not accurate enough due to the insufficient sample data volume in the third sample set. Description of the Drawings
[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description in the embodiments of the present application.
[0048] Figure 1 It is a schematic diagram of the system architecture for implementing the method of generating training samples provided by the embodiments of the present application;
[0049] Figure 2Schematic flowchart of a method for generating training samples provided by an embodiment of the present application;
[0050] Figure 3 Schematic flowchart of a method for generating training samples provided by an embodiment of the present application;
[0051] Figure 4 Schematic flowchart of a method for identifying abnormal transaction objects provided by an embodiment of the present application;
[0052] Figure 5 Schematic diagram of an application scenario for identifying fraudulent merchants provided by an embodiment of the present application;
[0053] Figure 6 Schematic diagram of the structure of a device for generating training samples provided by an embodiment of the present application;
[0054] Figure 7 Schematic flowchart of a device for identifying abnormal transaction objects provided by an embodiment of the present application;
[0055] Figure 8 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] The embodiments of the present application will be described below with reference to the accompanying drawings in the present application. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present application, and do not constitute limitations on the technical solutions of the embodiments of the present application.
[0057] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", and "the" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present application mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements, and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components, and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used herein may include a wireless connection or a wireless coupling. The term "and / or" used herein indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".
[0058] To make the purpose, technical solutions, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0059] First, introduce and explain several terms related to this application:
[0060] Distribution: Also known as "probability distribution". Common distributions include uniform distribution, Gaussian distribution, etc. For example, human height can be approximately considered to conform to a Gaussian distribution (that is, most people are near the average height, and the number of people deviating from the average height is less).
[0061] Gaussian distribution: Also called normal distribution. A univariate Gaussian distribution can be determined by two parameters, one is the mean μ, and the other is the variance δ. For a multivariate Gaussian distribution, it is determined by the mean vector and the covariance matrix.
[0062] Distribution estimation: Or called "distribution estimation", which means calculating the parameters of a distribution based on existing sample data. For example, if it is known that the height of all people in a certain place (the entire sample) conforms to a Gaussian distribution (μ, δ), and if you want to know the values of the mean μ and the variance δ, then you need to count the specific height values of all people and calculate that the mean height is 160 cm and the variance is 36. This mean 160 and variance 36 are the mean μ and variance δ required by the Gaussian distribution. This process of calculating the unknown parameters of a distribution based on actual sample data is called "distribution estimation".
[0063] Sampling: Refers to randomly generating new samples from a certain distribution. For example, it is now known that the height in a certain place follows a Gaussian distribution with a mean of 160 and a variance of 36 If you want to generate 100 new samples that conform to this Gaussian distribution, then you can randomly generate 100 random numbers from this Gaussian distribution These 100 random numbers will also conform to this Gaussian distribution (that is, the mean of these 100 samples will also be close to 160, and the variance will be close to 36).
[0064] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling machines to have the functions of perception, reasoning, and decision-making.
[0065] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0066] Fraud detection refers to the process of analyzing data and behavior patterns to identify potential fraud and taking corresponding measures to prevent losses. In recent years, machine learning technology has been widely used in the field of fraud detection. Using machine learning algorithms, an identification model can be trained using historical data and used to predict potential fraud samples. This method is more efficient and accurate than traditional manual analysis and can be adjusted and optimized at any time to adapt to the changing characteristics of fraud samples.
[0067] The performance of machine learning methods depends heavily on the number of samples. For the problem of detecting fraud merchants on an online shopping platform, there are merchants who have been complained by customers but are confirmed to be legal. The embodiments of this application refer to these merchants as suspicious legal merchants. The characteristics of these merchant samples are between those of normal legal merchants and fraud merchants. If the number of these samples in the training set can be increased, the performance of the fraud detection model trained by the machine learning method can be improved. This method of increasing the number of samples is usually referred to as data augmentation or data generation.
[0068] Existing methods for generating training samples generally include the following solutions:
[0069] 1) Interpolation method: Assume that the sample set of the category to be generated is S. Randomly sample two sample vectors A and B from S, and then directly calculate the mean vector of A and B as the new sample.
[0070] 2) Estimation of distribution method: Assume that the sample distribution of the category to be generated is a multivariate Gaussian distribution. Estimate the parameters of the distribution using all the samples in the sample set of this category, and then sample from the estimated distribution to generate the required new samples.
[0071] The interpolation method assumes that the features of the samples are uniformly distributed and ignores the relationship between sample features. However, in real-life scenarios, the sample data in various scenarios is more in line with the multivariate Gaussian distribution. The existing estimation of distribution method, although assuming the multivariate Gaussian distribution as the premise of the sample distribution, is more in line with the actual situation. However, due to the insufficient amount of sample object data for suspected normal transactions, the estimated sample distribution may not be accurate enough. Therefore, the newly sampled samples may deviate from the true distribution, and using samples with deviations to train the model will damage the classification performance of the model.
[0072] The method, apparatus, electronic device, computer-readable storage medium, and computer program product for generating training samples provided by this application aim to solve the above technical problems in the prior art.
[0073] The technical solutions of the embodiments of this application and the technical effects produced by the technical solutions of this application will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can refer to, draw on, or combine with each other. For the same terms, similar features, and similar implementation steps in different embodiments, they will not be described repeatedly.
[0074] Figure 1 It is a schematic diagram of the system architecture for implementing the method of generating training samples provided by the embodiments of this application. The method of generating training samples provided by the embodiments of this application can be realized based on the coordination between the server and the terminal. The terminal 100 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two. The terminal 100 sends the transaction information of each transaction of the transaction object to the server. The server extracts the transaction characteristics of the transaction object according to the historical transaction information uploaded by the transaction object, and classifies the transaction characteristics of each transaction object according to whether the transaction object is complained and whether it is determined to have bad transaction behavior after the complaint, and obtains which initial sample set. The initial sample set includes the first sample set, the second sample set, and the third sample set. Among them, each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of the sample objects with normal transactions, the transaction characteristics of the sample objects with abnormal transactions, and the transaction characteristics of the sample objects with suspected normal transactions; the transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object; determine the first association information of each initial sample set, and the first association information is used to characterize the degree of association of the characteristic values of each dimension in the corresponding initial sample set; perform weighted fusion on the first association information of each initial sample set to obtain the second association information, and the second association information is used to characterize the degree of association of the characteristic values of each dimension in all initial sample sets; according to the third sample set and the second association information, determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions, and perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
[0075] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal 100 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, and no limitation is made in the embodiments of the present application.
[0076] In some embodiments, the terminal or the server may implement the method for generating training samples provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or a software module in an operating system; it may be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a news APP or an e-commerce APP; it may also be a small program, that is, a program that only needs to be downloaded to a browser environment to run; it may also be a small program that can be embedded in any APP. In short, the above computer program may be any form of application program, module or plug-in.
[0077] In each embodiment of the present application, the collection and processing of relevant data should strictly comply with the requirements of relevant national laws and regulations when applied in practice, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope authorized by laws and regulations and the personal information subject.
[0078] In the embodiments of the present application, a method for generating training samples is provided, as Figure 2 shown, the method includes:
[0079] S101. Determine three initial sample sets, where the three initial sample sets include a first sample set, a second sample set, and a third sample set.
[0080] Each sample set in the embodiments of the present application includes multiple initial samples.
[0081] The initial samples of the first sample set are the transaction characteristics of the sample objects of normal transactions. The sample objects of normal transactions in the embodiments of the present application refer to the sample objects that have not had complaints about abnormal behaviors.
[0082] The sample objects in the embodiments of the present application may refer to merchants or customers.
[0083] The embodiments of the present application do not specifically define bad behaviors. For example, they may include behaviors such as fraud, passing off inferior goods as good ones, and refusing after-sales service.
[0084] The initial samples of the second sample set are the transaction characteristics of the sample objects of abnormal transactions. The sample objects of abnormal transactions in the embodiments of the present application refer to the sample objects that have been complained about and confirmed to have bad behaviors.
[0085] The initial samples of the third sample set are the transaction characteristics of the sample objects of suspected normal transactions. The sample objects of suspected normal transactions in the embodiments of the present application may refer to the sample objects that have been complained about but are considered to have no bad behaviors.
[0086] It should be noted that the essence of the initial samples in the third sample set determined in the embodiments of the present application is still normal, that is, although there are some complaints, it is actually found through detection and verification that the sample object has not had bad behaviors.
[0087] In some embodiments, the sample objects of normal transactions may also refer to the sample objects with a transaction positive review rate greater than the first preset threshold (for example, 99%).
[0088] In some embodiments, the sample objects of abnormal transactions may be the sample objects with a transaction positive review rate less than the second preset threshold (for example, 50%).
[0089] In some embodiments, the sample objects of suspected normal transactions may also refer to the sample objects with a transaction positive review rate between the first preset threshold and the third preset threshold (for example, 90%).
[0090] It should be understood that the characteristic values of each dimension of the sample objects of suspected normal transactions are between the characteristics of the sample objects of normal transactions and the sample objects of abnormal transactions.
[0091] The transaction characteristics of a sample object in the embodiments of the present application include the characteristic values of multiple dimensions of the sample object. The embodiments of the present application do not specifically limit the types of dimensions. For example, they may be the number of transactions in the recent 7 days, the transaction amount in the recent 7 days, the number of transactions in the recent 30 days, the unit price per customer, the repurchase rate, and so on.
[0092] In practical applications, the collection of the first sample set and the second sample set is relatively easy. Therefore, the number of samples in the first sample set and the second sample set is relatively abundant. However, since the third sample set involves manual verification and also needs to be finally verified as normal transactions, the number of samples is relatively small.
[0093] S102. Determine the first association information of each initial sample set, where the first association information is used to characterize the degree of association of the characteristic values of each dimension in the corresponding initial sample set.
[0094] After the embodiments of the present application obtain three initial sample sets, the first association information of each initial sample set is further determined. Since the dimensions of the transaction characteristics of all sample objects in the embodiments of the present application are the same, therefore, the degree of association of the characteristic values of each dimension in each initial sample set can be calculated as the first association information.
[0095] When the number of samples in the initial sample set is relatively sufficient, the first association information can more accurately reflect the distribution characteristics of the transaction characteristics of the corresponding type of sample objects. In practical applications, the number of samples in the first sample set and the second sample set is relatively sufficient, while the number of samples in the third sample set, that is, the sample objects of suspected normal transactions, is small. Therefore, the first association information of the third sample set directly obtained based on the third sample set cannot accurately describe the distribution of the transaction characteristics of the sample objects of suspected normal transactions.
[0096] The embodiments of the present application lay a foundation for obtaining the laws of the characteristic values of each dimension in each initial sample set by determining the first association information of each initial sample set, and also realize the conversion of the originally inconsistent information (i.e., the number of samples) in each initial sample set to a smaller range with a smaller quantity (i.e., the range of the degree of association of the characteristic values), so as to generate training samples more accurately.
[0097] S103. According to the first association information of each of the initial sample sets, determine the characteristic distribution of the transaction characteristics of the sample objects of suspected normal transactions, and perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
[0098] Due to the problem that the characteristic distribution of the transaction characteristics of the sample objects of suspected normal transactions determined only based on the third sample set is inaccurate due to the small data volume of the third sample set, the embodiments of the present application recognize the characteristic that the transaction characteristics of the sample objects of suspected normal transactions are between the transaction characteristics of the sample objects of normal transactions and abnormal transactions, and use the first association information of each initial sample set to obtain the characteristic distribution of the transaction characteristics of the sample objects of suspected normal transactions, thereby improving the accuracy of obtaining the characteristic distribution of the transaction characteristics of the sample objects of suspected normal transactions, and thus the finally obtained new samples are also more accurate.
[0099] The method for generating training samples according to the embodiments of the present application determines three initial sample sets, which include a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object. Further, the first association information of each initial sample set is determined. The first association information is used to characterize the degree of association of the characteristic values of each dimension in the corresponding initial sample set. Since the dimensions of the transaction characteristics of all sample objects are the same, calculating the first association information can summarize the information of different initial sample sets to the same scale for processing, and the first association information can reflect the distribution of the characteristic values of the initial sample set to a certain extent. It is realized that due to the small data volume of the third sample set, the present application uses the first association information of each initial sample set to obtain the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions. This distribution can more accurately characterize the transaction characteristics of the sample objects with suspected normal transactions. Then, multiple samplings are performed according to the characteristic distribution to obtain multiple new samples of the third sample set. The embodiments of the present application can effectively alleviate the problem that the sample distribution estimation is not accurate enough due to the insufficient sample data volume in the third sample set.
[0100] On the basis of the above embodiments, as an alternative embodiment, determining the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the first association information of each of the initial sample sets includes:
[0101] Fusing the first association information of each initial sample set to obtain second association information, where the second association information is used to characterize the degree of association of the characteristic values of each dimension in all initial sample sets;
[0102] Taking the second association information as the degree of association between the characteristics of each dimension corresponding to the sample objects with suspected normal transactions, and determining the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the third sample set and the second association information.
[0103] In the embodiments of the present application, fusing the first association information of each initial sample set may refer to first weighting each first association information with corresponding weights, and then adding all the weighted results and taking the average value to obtain the second association information.
[0104] In an embodiment of the present application, according to the characteristic that the transaction characteristics of the sample objects of the suspected normal transactions are between the transaction characteristics of the sample objects of the normal transactions and the abnormal transactions, it is considered that the distribution of the eigenvalue of each dimension of the first sample set and the second sample set will be helpful for constructing the training samples corresponding to the sample objects of the suspected normal transactions, and the first correlation information contains the distribution of the eigenvalue, so the present application uses the relationship between the eigenvalues of each dimension contained in the first correlation information of the first sample set and the second sample set to estimate a more accurate sample distribution of the third sample set, alleviating the problem that the sample distribution estimation is not accurate enough due to the insufficient sample data volume in the third sample set.
[0105] The embodiment of the present application realizes the problem that since the data volume of the third sample set is small, the characteristic distribution of the transaction characteristics of the sample objects of the suspected normal transactions directly obtained from the third sample set may be inaccurate. Therefore, on the basis of the third sample set, combined with the second correlation information, using the second correlation information to represent the correlation degree information of the eigenvalues of each dimension in all initial sample sets, a more accurate characteristic distribution of the transaction characteristics of the sample objects of the suspected normal transactions is obtained. In this way, the samples obtained by sampling based on this characteristic distribution can have the transaction characteristics of the sample objects of the suspected normal transactions more accurately.
[0106] The embodiment of the present application realizes that since the data volume of the third sample set is small, the present application obtains the second correlation information by weighted fusion of the first correlation information of each initial sample set. The second correlation information is used to represent the correlation degree of the eigenvalues of each dimension in all initial sample sets, and the second correlation information is used to correct the characteristic distribution of the third sample set to obtain the characteristic distribution of the transaction characteristics of the sample objects of the suspected normal transactions. This distribution can more accurately represent the transaction characteristics of the sample objects of the suspected normal transactions. Then, multiple samplings are performed according to the characteristic distribution to obtain multiple new samples of the third sample set. The embodiment of the present application can effectively alleviate the problem that the sample distribution estimation is not accurate enough due to the insufficient sample data volume in the third sample set.
[0107] On the basis of the above embodiments, as an optional embodiment, the first correlation information is a covariance matrix, and the covariance matrix includes the covariance between the eigenvalue of each dimension and the eigenvalues of other dimensions in the corresponding sample set.
[0108] Covariance is used in probability theory and statistics to measure the overall error between two variables. Covariance represents the overall error between two variables, which is different from variance that only represents the error of one variable. If the change trends of two variables are the same, that is, if one is greater than its own expected value and the other is also greater than its own expected value, then the covariance between the two variables is positive. If the change trends of two variables are opposite, that is, one is greater than its own expected value while the other is less than its own expected value, then the covariance between the two variables is negative.
[0109] In the embodiments of the present application, the covariance between the eigenvalue of one dimension and the eigenvalues of other dimensions can reflect the correlation between the eigenvalues of each dimension in a sample set. The eigenvalues of dimensions with high correlation can more accurately reflect the general characteristics of a sample set.
[0110] It should be understood that the calculation formula of covariance can be expressed as:
[0111] COV(X,Y) = E[(X - E(X))(Y - E(Y))];
[0112] Among them, COV represents covariance, E represents expectation, and X and Y represent two different transaction characteristics;
[0113] The following uses a specific example to illustrate the process of calculating the covariance of two transaction characteristics:
[0114] Transaction characteristic X is [1.1, 1.9, 3], and transaction characteristic Y is [5.0, 10.4, 14.6];
[0115] The expectation of transaction characteristic X, E(X) = (1.1 + 1.9 + 3) / 3 = 2;
[0116] The expectation of transaction characteristic Y, E(Y) = (5.0 + 10.4 + 14.6) / 3 = 10;
[0117] The common expectation of transaction characteristic X and transaction characteristic Y is:
[0118] E(XY) = (1.1×5.0 + 1.9×10.4 + 3×14.6) / 3 = 23.02;
[0119] The covariance of transaction characteristic X and transaction characteristic Y is:
[0120] Cov(X,Y) = E(XY) - E(X)E(Y) = 23.02 - 2×10 = 3.02.
[0121] Based on the above embodiments, as an alternative embodiment, the association information of each initial sample set is fused to obtain second association information, including:
[0122] Perform weighted summation on the covariance matrices of each initial sample set to obtain a target covariance matrix, and the second association information is the target covariance matrix.
[0123] The formula for calculating the second association information in the embodiments of this application can be expressed as:
[0124]
[0125] wherein, represents the first association information of the first sample set, represents the first association information of the second sample set, represents the first association information of the third sample set, k1 represents the weight of the first sample set, k2 represents the weight of the second sample set, k3 represents the weight of the third sample set, and d represents the number of dimensions of the transaction feature. In the embodiments of this application, the weights of each sample set are all positive numbers and k1 + k2 + k3 = 1.
[0126] Based on the above embodiments, as an alternative embodiment, the corresponding weights of the first sample set and the second sample set are not greater than the weight corresponding to the third sample set.
[0127] It should be noted that since the embodiments of this application use the first sample set and the second sample set with a large amount of data to correct the feature distribution of the third sample set, the third sample set needs to play a dominant role, and the weight corresponding to the third sample set can be set to not less than the weights corresponding to the first sample set and the second sample set. In some embodiments, the weights corresponding to the first sample set and the second sample set can be 0.25, and the weight corresponding to the third sample set is 0.5.
[0128] Based on the above embodiments, as an alternative embodiment, according to the third sample set and the second association information, determine the feature distribution of the transaction features of the sample objects of suspected normal transactions, including:
[0129] Determine the mean value of the feature values of each dimension of the third sample set to obtain the sample mean of the multivariate Gaussian distribution model, and use the second association information as the covariance matrix of the multivariate Gaussian distribution model to obtain a multivariate Gaussian distribution model describing the distribution of new samples.
[0130] Specifically, in this application, for each dimension, calculate the transaction features of each sample object in the third sample set, and the mean value of the feature values in this dimension, so as to obtain the sample mean of the multivariate Gaussian distribution model. For example, if the n-dimensional transaction feature of sample object i is defined as If there are m sample objects in the third sample set, the sample mean μ m can be expressed as:
[0131] The multivariate Gaussian distribution constructed in the embodiments of this application can be expressed as Sampling a sample from this multivariate Gaussian distribution and marking this sample as a newly added sample in the third sample set.
[0132] It should be noted that due to reasons such as the limited number of existing samples, the distributions estimated according to the existing samples in the embodiments of this application, including the corrected distributions, cannot fully reflect the true distribution of the transaction characteristics of the sample objects of suspected normal transactions. Therefore, if too many new samples are sampled from the estimated distribution, it may introduce incorrect distribution information and contaminate the overall data. Therefore, the number of sampling times in the embodiments of this application is less than the number of initial samples in any one of the initial sample sets.
[0133] Please refer to Figure 3 , which exemplarily shows a schematic flowchart of the method for generating training samples provided by the embodiments of this application. As shown in the figure, it includes:
[0134] Obtaining the transaction characteristics of multiple sample objects, classifying the transaction characteristics of multiple sample objects according to whether the sample objects are complained and whether they are confirmed to have bad transaction behaviors after the complaint, and obtaining a first sample set, a second sample set, and a third sample set. Among them, the first sample set can be expressed as N is the total number of sample objects of normal transactions, and the second sample set can be expressed as H is the total number of sample objects of abnormal transactions; the third sample set can be expressed as: M is the total number of sample objects of suspected normal transactions.
[0135] Obtaining the covariance matrix of the first sample set according to the transaction characteristics of all sample objects in the first sample set Obtaining the covariance matrix of the second sample set according to the transaction characteristics of all sample objects in the second sample set Obtaining the covariance matrix of the third sample set according to the transaction characteristics of all sample objects in the third sample set
[0136] Performing weighted averaging on the three covariance matrices to obtain a covariance matrix Calculating the sample mean according to all samples in S m and constructing a multivariate Gaussian distribution
[0137] From this multivariate Gaussian distribution Take a sample from And mark the sample as a new sample in the third sample set.
[0138] On the basis of the above embodiments, as an optional embodiment, the method for generating training samples in the embodiment of the present application further includes:
[0139] For the first sample set, the third sample set and the newly added samples, a first training label is set, where the first training label is used to indicate that the corresponding sample objects are sample objects of normal transactions;
[0140] A second training label is set for the second sample set, and the second training label is used to characterize that the corresponding sample object is a sample object of abnormal transaction.
[0141] In order to subsequently identify the objects of abnormal transactions through a neural network model, when actually training the model, a first training label will be set for the first sample set, the third sample set and the newly added samples. The first training label is used to characterize the corresponding sample objects as normal transaction objects. A second training label is set for the initial sample set of the second category. The second training label is used to characterize the corresponding sample objects as abnormal transaction objects. In this way, a binary classification model can be trained to distinguish between abnormal transaction objects and normal transaction objects.
[0142] See also Figure 4 , which exemplarily shows a flow chart of a method for identifying abnormal transaction objects provided in an embodiment of the present application, as shown in the figure, including:
[0143] S201, obtaining transaction features of the object to be identified;
[0144] S202, inputting the transaction features of the object to be identified into a pre-trained object recognition model to obtain a recognition result output by the object recognition model;
[0145] The identification result is used to indicate whether the object to be identified is a normal transaction object or an abnormal transaction object;
[0146] The object recognition model is trained by a first sample set, a second sample set and a third sample set, each sample set including a plurality of initial samples, the initial samples in the first sample set, the second sample set and the third sample set are respectively: transaction features of sample objects of normal transactions, transaction features of sample objects of abnormal transactions, and transaction features of sample objects of suspected normal transactions; at least one sample in the third sample set is generated based on the method described in any one of claims 1 to 7.
[0147] See also Figure 5, which exemplarily shows a schematic diagram of the scenario to which the embodiments of the present application are applied for identifying fraudulent merchants. As shown in the figure, after a merchant completes a transaction through the first terminal 401, the first terminal 401 uploads the transaction information and the merchant identifier to the transaction database 402 for storage.
[0148] The first server 403 regularly / irregularly obtains the newly uploaded transaction information of each merchant from the transaction database 402, updates the transaction characteristics of the merchant according to the transaction information, and determines whether to classify the transaction characteristics of the merchant into the first sample set, the second sample set, or the third sample set according to whether the merchant is complained about (which can be recorded in the transaction information) and whether it is determined to have an abnormal transaction after the complaint. Moreover, through the method for generating training samples provided in the above embodiments of the present application, a plurality of new samples of the third sample set are obtained and supplemented into the third sample set.
[0149] The second server 404 regularly or irregularly obtains the first sample set, the second sample set, and the third sample set from the first server, sets a first training label for the first sample set and the third sample set. The first training label is used to represent that the corresponding sample object is a sample object of a normal transaction; a second training label is set for the second sample set. The second training label is used to represent that the corresponding sample object is a sample object of a normal transaction, and a binary classification model is trained (updated) using the three sample sets and the corresponding labels. This binary classification model can distinguish between abnormal transaction objects and normal transaction objects.
[0150] When a shopper enters an e-shop of a certain merchant through the second terminal 405, the second terminal 405 sends a query request to the third server 406. The query request includes the shop identifier of the merchant. The third server 406 requests the recent transaction characteristics of the merchant from the first server 403 according to the shop identifier of the merchant, calls the binary classification model trained by the second server, inputs the requested transaction characteristics into the binary classification model, and obtains a prediction result of whether the merchant is an abnormal transaction object. If it is determined that the merchant is an abnormal transaction object, a prompt message is sent to the shopper through the second terminal, thereby reducing the shopping risk.
[0151] The embodiments of the present application provide a training sample generation device, as Figure 6 shown. The training sample generation device may include: a sample set module 601, a first association module 602, and a sampling module 603, where
[0152] A sample set module 601 is used to determine three initial sample sets, which include a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object.
[0153] A first association module 602 is used to determine the first association information of each initial sample set. The first association information is used to characterize the association degree of the characteristic values of each dimension in the corresponding initial sample set.
[0154] A sampling module 603 is used to determine the characteristic distribution of the transaction characteristics of sample objects with suspected normal transactions according to the first association information of each initial sample set, and perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
[0155] The device in the embodiment of the present application can execute the method provided in the embodiment of the present application, and its implementation principle is similar. The actions performed by each module in the device in the embodiments of the present application correspond to the steps in the methods in the embodiments of the present application. For the detailed function descriptions of each module of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.
[0156] As an optional implementation manner, the sampling module includes:
[0157] A fusion sub-module is used to fuse the first association information of each initial sample set to obtain second association information. The second association information is used to characterize the association degree of the characteristic values of each dimension in all initial sample sets.
[0158] A distribution determination sub-module is used to use the second association information as the association degree between the characteristics of each dimension corresponding to the sample object with suspected normal transactions, and determine the characteristic distribution of the transaction characteristics of the sample object with suspected normal transactions according to the third sample set and the second association information.
[0159] As an optional implementation manner, the first association information is a covariance matrix, and the covariance matrix includes the covariance between the characteristic value of each dimension and the characteristic values of other dimensions in the corresponding sample set.
[0160] The fusion sub-module is specifically used to perform weighted summation on the covariance matrices of each initial sample set to obtain a target covariance matrix, and the second association information is the target covariance matrix.
[0161] As an alternative implementation, the fusion sub-module is specifically configured to:
[0162] Obtain the weights corresponding to each of the three initial sample sets, where the weights corresponding to the first sample set and the second sample set are not greater than the weight corresponding to the third sample set;
[0163] Based on the weights corresponding to each of the three initial sample sets, perform weighted fusion on the three initial sample sets to obtain second correlation information.
[0164] As an alternative implementation, the distribution determination sub-module is specifically configured to:
[0165] Determine the mean of the eigenvalues of each dimension of the third sample set as the sample mean of the multivariate Gaussian distribution model, and use the target covariance matrix as the covariance matrix of the multivariate Gaussian distribution model to obtain a multivariate Gaussian distribution model describing the distribution of the new samples.
[0166] As an alternative implementation, the number of samplings does not exceed the number of samples in any one of the initial sample sets.
[0167] As an alternative implementation, it further includes:
[0168] A label setting module, configured to set first training labels for the first sample set, the third sample set, and the new samples; and set second training labels for the second sample set;
[0169] Wherein, the first training label is used to represent that the corresponding sample object is a sample object of a normal transaction; the second training label is used to represent that the corresponding sample object is a sample object of an abnormal transaction.
[0170] An embodiment of the present application provides an identification device for abnormal transaction objects, as Figure 7 shown. The training sample generation device may include: a to-be-identified object acquisition module 701 and a model judgment module 702, where
[0171] The to-be-identified object acquisition module 701 is configured to acquire the transaction characteristics of the to-be-identified object;
[0172] The model judgment module 702 is configured to input the transaction characteristics of the to-be-identified object into a pre-trained object recognition model to obtain the recognition result output by the object recognition model;
[0173] Wherein, the recognition result is used to indicate that the to-be-identified object is a normal transaction object or an abnormal transaction object;
[0174] The object recognition model of the embodiment of the present application is trained by a first sample set, a second sample set, and a third sample set. Each sample set includes a plurality of initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of the sample objects of normal transactions, the transaction characteristics of the sample objects of abnormal transactions, and the transaction characteristics of the sample objects of suspected normal transactions; at least one sample in the third sample set is generated based on the training sample generation method provided in each of the above embodiments.
[0175] The object recognition model of the embodiment of the present application is trained by supervised learning. Specifically, a first training label is set for the first sample set and the third sample set, and a second training label is set for the second sample set. The first training label is used to represent that the corresponding sample object is a sample object of a normal transaction; the second training label is used to represent that the corresponding sample object is a sample object of an abnormal transaction. Based on the sample set and the corresponding training label, the initial model is trained until convergence, and the object recognition model is obtained.
[0176] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of a method for generating training samples. Compared with the related art, it can be achieved that by determining three initial sample sets, the three initial sample sets include a first sample set, a second sample set, and a third sample set. Among them, each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects of normal transactions, the transaction characteristics of sample objects of abnormal transactions, and the transaction characteristics of sample objects of suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object. Further, the first association information of each initial sample set is determined. The first association information is used to characterize the degree of association of the characteristic values of each dimension in the corresponding initial sample set. Since the dimensions of the transaction characteristics of all sample objects are the same, calculating the first association information can summarize the information of different initial sample sets to the same scale for processing, and the first association information can reflect the distribution of the characteristic values of the initial sample set to a certain extent. It is realized that due to the small data volume of the third sample set, in this application, the first association information of each initial sample set is weighted and fused to obtain the second association information. The second association information is used to characterize the degree of association of the characteristic values of each dimension in all initial sample sets. The second association information is used to correct the characteristic distribution of the third sample set to obtain the characteristic distribution of the transaction characteristics of the sample objects of suspected normal transactions. This distribution can more accurately characterize the transaction characteristics of the sample objects of suspected normal transactions. Then, multiple samplings are performed according to the characteristic distribution to obtain multiple new samples of the third sample set. The embodiment of the present application can effectively alleviate the problem that the sample distribution estimation is inaccurate due to insufficient sample data volume in the third sample set.
[0177] In an alternative embodiment, an electronic device is provided, as Figure 8 shown. Figure 8 The electronic device 4000 shown in the figure includes: a processor 4001 and a memory 4003. Among them, the processor 4001 and the memory 4003 are connected, such as connected through a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004. The transceiver 4004 may be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data, etc. It should be noted that in practical applications, the transceiver 4004 is not limited to one, and the structure of the electronic device 4000 does not constitute a limitation to the embodiment of the present application.
[0178] The processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in connection with the disclosure of this application. The processor 4001 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0179] The bus 4002 may include a path for transmitting information between the above components. The bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. The bus 4002 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 only a thick line is used to represent it here, but it does not mean that there is only one bus or one type of bus.
[0180] The memory 4003 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or it may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited here.
[0181] The memory 4003 is used to store the computer program for implementing the embodiments of the present application, and is controlled by the processor 4001 to execute. The processor 4001 is used to execute the computer program stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0182] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0183] The embodiments of the present application also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.
[0184] The terms "first", "second", "third", "fourth", "1", "2", etc. (if any) in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than the illustrated or described order.
[0185] It should be understood that although the flowchart of the embodiments of the present application indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless otherwise clearly stated in this application, in some implementation scenarios of the embodiments of the present application, the implementation steps in each flowchart can be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage of these sub-steps or stages can also be executed at different times respectively. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present application do not limit this.
[0186] The above are only optional implementation manners of some implementation scenarios of the present application. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present application, other similar implementation means based on the technical idea of the present application also belong to the protection scope of the embodiments of the present application.
Claims
1. A method for generating training samples, characterized in that, Including: Determine three initial sample sets, which include a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects with normal transactions, the transaction characteristics of sample objects with abnormal transactions, and the transaction characteristics of sample objects with suspected normal transactions. The transaction characteristics of a sample object include the characteristic values of multiple dimensions of the sample object. Based on the transaction characteristics of each sample object in each initial sample set, determine the first association information of each initial sample set. The first association information is used to characterize the degree of association of the characteristics of each dimension in the corresponding initial sample set. According to the first association information of each initial sample set, determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions, and perform multiple samplings according to the characteristic distribution to obtain multiple new samples of the third sample set.
2. The method according to claim 1, characterized in that, The step of determining the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the first association information of each initial sample set includes: Fuse the first association information of each initial sample set to obtain second association information. The second association information is used to characterize the degree of association of the characteristic values of each dimension in all initial sample sets. Use the second association information as the degree of association between the characteristics of each dimension corresponding to the sample objects with suspected normal transactions. According to the third sample set and the second association information, determine the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions.
3. The method according to claim 2, wherein The first association information is a covariance matrix, and the covariance matrix includes the covariance between the characteristic value of each dimension and the characteristic values of other dimensions in the corresponding sample set. The step of fusing the association information of each initial sample set to obtain second association information includes: Perform weighted summation on the covariance matrices of each initial sample set to obtain a target covariance matrix. The second association information is the target covariance matrix.
4. The method according to claim 2 or 3, characterized in that, The step of fusing the first association information of each initial sample set to obtain second association information includes: Obtain the weights corresponding to the three initial sample sets respectively, where the weights corresponding to the first sample set and the second sample set are not greater than the weight corresponding to the third sample set. Based on the weights corresponding to the three initial sample sets respectively, perform weighted fusion on the three initial sample sets to obtain second association information.
5. The method according to claim 3, wherein The step of determining the characteristic distribution of the transaction characteristics of the sample objects with suspected normal transactions according to the third sample set and the second association information includes: Determine the mean value of the characteristic values of each dimension of the third sample set as the sample mean of the multivariate Gaussian distribution model, and use the target covariance matrix as the covariance matrix of the multivariate Gaussian distribution model to obtain a multivariate Gaussian distribution model describing the distribution of the new samples.
6. The method according to claim 1, wherein The number of samplings does not exceed the number of samples in any one of the initial sample sets.
7. The method according to claim 1, wherein It also includes: Set first training labels for the first sample set, the third sample set, and the new samples. Set second training labels for the second sample set. Among them, the first training label is used to represent that the corresponding sample object is a sample object of normal transaction; The second training label is used to represent that the corresponding sample object is a sample object of abnormal transaction.
8. A method for identifying an abnormal transaction object, characterized in that, It includes: Obtain the transaction characteristics of the object to be recognized; Input the transaction characteristics of the object to be recognized into a pre-trained object recognition model to obtain the recognition result output by the object recognition model; Among them, the recognition result is used to indicate that the object to be recognized is an object of normal transaction or an object of abnormal transaction; The object recognition model is trained by a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects of normal transactions, the transaction characteristics of sample objects of abnormal transactions, and the transaction characteristics of sample objects suspected of normal transactions; at least one sample in the third sample set is generated based on the method described in any one of claims 1-7.
9. A generating device for training samples, characterized in that It includes: A sample set module for determining three initial sample sets, namely a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects of normal transactions, the transaction characteristics of sample objects of abnormal transactions, and the transaction characteristics of sample objects suspected of normal transactions; the transaction characteristics of a sample object include eigenvalue of multiple dimensions of the sample object; A first association module for determining the first association information of each initial sample set based on the transaction characteristics of each sample object in each initial sample set. The first association information is used to represent the association degree of the characteristics of each dimension in the corresponding initial sample set; A sampling module for determining the feature distribution of the transaction characteristics of the sample objects suspected of normal transactions according to the first association information of each initial sample set, and performing multiple samplings according to the feature distribution to obtain multiple new samples of the third sample set.
10. An identification device for abnormal transaction objects, characterized in that, It includes: An object to be recognized acquisition module for acquiring the transaction characteristics of the object to be recognized; A model judgment module for inputting the transaction characteristics of the object to be recognized into a pre-trained object recognition model to obtain the recognition result output by the object recognition model; Among them, the recognition result is used to indicate that the object to be recognized is an object of normal transaction or an object of abnormal transaction; The object recognition model is trained by a first sample set, a second sample set, and a third sample set. Each sample set includes multiple initial samples. The initial samples in the first sample set, the second sample set, and the third sample set are respectively: the transaction characteristics of sample objects of normal transactions, the transaction characteristics of sample objects of abnormal transactions, and the transaction characteristics of sample objects suspected of normal transactions; At least one sample in the third sample set is generated based on the method described in any one of claims 1-7.
11. An electronic device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the steps of the method for generating training samples according to any one of claims 1-7 or the method for identifying abnormal transaction objects according to claim 8.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for generating training samples according to any one of claims 1-7 or the method for identifying abnormal transaction objects according to claim 8.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method for generating training samples according to any one of claims 1-7 or the method for identifying abnormal transaction objects according to claim 8.