Sample data generation method, apparatus and device, and computer readable storage medium
By classifying and statistically analyzing the transaction data, calculating the generation probability of trading features with feature recommendation functions, and automatically generating sample data, solving the problem of low sample data generation efficiency in the existing technology and improving the model training efficiency.
Patent Information
- Application Number
- CN202510031542.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-06
AI Technical Summary
In the prior art, the method of extracting transaction features based on personal experience leads to low efficiency in sample data generation, thereby reducing model training efficiency.
Classify transaction fields through preset field categories, statistically analyze and generate the first type of transaction characteristics, and use the feature recommendation function to calculate the generation probability of the second type of transaction characteristics based on the influencing factor, and automatically generate target sample data.
The automated process of sample data generation is realized, user operations are simplified, and sample data generation efficiency and model training efficiency are improved.
Smart Images

Figure CN119939251A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and in particular, relates to a method, device, equipment and computer-readable storage medium for generating sample data. Background Art
[0002] With the increasing growth of online payments, more and more financial institutions and payment service providers have begun to collect transaction data, extract transaction features related to users' transaction behaviors from transaction data, generate sample data based on transaction features, and train models based on sample data with the above transaction features to provide powerful assistance for personalized services and precision marketing for users. Among them, transaction features can be basic transaction features including user transaction amount distribution, types of purchased goods, purchase time statistics, number of transactions, maximum / minimum transaction amount, total transaction amount, etc. However, the accuracy of model prediction obtained by training sample data with basic transaction features is low.
[0003] In order to improve the accuracy of model prediction, it is currently usually done by experienced business personnel who extract advanced transaction features such as meta features and user behavior pattern features from transaction data based on their personal experience, and train the model based on sample data with basic transaction features and advanced transaction features.
[0004] However, the method of extracting basic transaction features and advanced transaction features based on personal experience will generate a lot of workload, and as the transaction data changes, it is necessary to manually adjust the basic transaction features and advanced transaction features continuously, which is time-consuming and labor-intensive, resulting in low efficiency in generating sample data, thereby reducing the efficiency of model training. Summary of the invention
[0005] The embodiments of the present application provide a method, apparatus, device, computer-readable storage medium, and computer program product for generating sample data, which can improve the efficiency of generating sample data and thereby improve the efficiency of model training.
[0006] In a first aspect, an embodiment of the present application provides a method for generating sample data, the method comprising:
[0007] Acquire first transaction data, where the first transaction data includes a plurality of transaction fields;
[0008] Based on preset field categories, classify the multiple transaction fields and determine multiple field categories corresponding to the multiple transaction fields;
[0009] For each of the field categories, statistical analysis is performed on the multiple transaction fields to obtain a first type of transaction feature;
[0010] Inputting a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of the transaction field, second attribute information of the sample data, and computing power information of a target electronic device, the target electronic device being an electronic device for model training based on the sample data;
[0011] When the generation probability is greater than a preset probability, generating the second type of transaction feature based on second transaction data, where the second transaction data includes the first transaction data and data generated based on the first type of transaction feature;
[0012] Target sample data is generated based on the transaction field, the first type of transaction features, and the second type of transaction features.
[0013] In a possible implementation, before inputting the first influencing factor for generating the second type of transaction feature into the feature recommendation function, the method further includes:
[0014] Acquire a reference influencing factor corresponding to the second type of transaction feature and a preset generation probability corresponding to the reference influencing factor, wherein the reference influencing factor includes reference attribute information of the transaction field, reference attribute information of the sample data, and reference computing power information of the target electronic device;
[0015] The initial feature recommendation function is fitted based on the reference influencing factor and the corresponding preset generation probability to obtain the feature recommendation function.
[0016] In a possible implementation, for each of the field categories, statistical analysis is performed on the multiple transaction fields to obtain a first type of transaction features, including:
[0017] Determining a feature generation mode based on a second influencing factor for generating the first type of transaction feature, the second influencing factor including first attribute information of the transaction field and second attribute information of the sample data;
[0018] For each of the field categories, statistical analysis is performed on the multiple transaction fields based on the feature generation mode to obtain a first type of transaction features.
[0019] In a possible implementation, after determining the multiple field categories and before generating the second-type transaction feature based on the second transaction data, the method further includes:
[0020] Based on the field category, preprocessing the first transaction data to obtain preprocessed first transaction data; the data preprocessing at least includes classifying date type data in the first transaction data according to a preset rule, and standardizing the first transaction data;
[0021] generating first data based on the first type of transaction features and the preprocessed first transaction data;
[0022] The second transaction data is determined according to the preprocessed first transaction data and the first data.
[0023] In a possible implementation manner, before determining the second transaction data according to the preprocessed first transaction data and the first data, the method further includes:
[0024] In a case where the data type corresponding to the field category is text data, performing text analysis on the transaction field corresponding to the field category to obtain a text vectorization feature;
[0025] Generate second data based on the text vectorization feature and the preprocessed first transaction data;
[0026] The determining the second transaction data according to the preprocessed first transaction data and the first data includes:
[0027] determining the second transaction data according to the preprocessed first transaction data, the first data, and the second data;
[0028] The generating target sample data based on the transaction field, the first type of transaction features and the second type of transaction features includes:
[0029] Generate target sample data based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features.
[0030] In a possible implementation, before generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features, and the text vectorization features, the method further includes:
[0031] Based on the second transaction data, generate an additional feature, where the additional feature is any feature other than the first transaction feature, the second transaction feature, and the text vectorization feature;
[0032] The generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features includes:
[0033] Generate target sample data based on the transaction field, the first category transaction features, the second category transaction features, the text vectorization features and the additional category features.
[0034] In a possible implementation, generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features, the text vectorization features, and the additional type of features includes:
[0035] Determine a first target transaction feature according to the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature;
[0036] Performing feature selection on the first target transaction features based on a feature selection algorithm, and determining a first score for each transaction feature in the first target transaction features;
[0037] Performing feature selection on the first target transaction features based on a distribution verification algorithm to determine a second score for each transaction feature in the first target transaction features;
[0038] Based on the weights corresponding to the first score and the second score respectively, performing a weighted summation on the first score and the second score to obtain a third score;
[0039] Determining the transaction feature corresponding to the third score greater than a preset threshold as a second target transaction feature;
[0040] Generate target sample data based on the transaction field and the second target transaction feature.
[0041] In a second aspect, an embodiment of the present application provides a device for generating sample data, the device comprising:
[0042] A first acquisition module, configured to acquire first transaction data, where the first transaction data includes a plurality of transaction fields;
[0043] A classification module, configured to classify the plurality of transaction fields based on preset field categories, and determine a plurality of field categories corresponding to the plurality of transaction fields;
[0044] A first analysis module, configured to perform statistical analysis on the plurality of transaction fields for each of the field categories to obtain a first type of transaction feature;
[0045] An input module, used to input a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of the transaction field, second attribute information of the sample data, and computing power information of a target electronic device, the target electronic device being an electronic device for model training based on the sample data;
[0046] a first generating module, configured to generate the second type of transaction feature based on second transaction data when the generation probability is greater than a preset probability, the second transaction data including the first transaction data and data generated based on the first type of transaction feature;
[0047] The second generating module is used to generate target sample data based on the transaction field, the first type of transaction characteristics and the second type of transaction characteristics.
[0048] In a third aspect, an embodiment of the present application provides an electronic device, the device comprising: a processor and a memory storing computer program instructions;
[0049] When the processor executes the computer program instructions, the processor implements any possible implementation method of the first aspect described above.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implements a method in any possible implementation method of the first aspect described above.
[0051] In a fifth aspect, an embodiment of the present application provides a computer program product. When instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes a method as any possible implementation method in the first aspect above.
[0052] The embodiment of the present application classifies multiple transaction fields based on preset field categories, determines multiple field categories corresponding to multiple transaction fields in the first transaction data, and statistically analyzes multiple transaction fields for each field category to obtain first-class transaction features, thereby automatically generating first-class transaction features based on preset field categories. By taking the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device as inputs, and calculating the generation probability of the second-class transaction features based on the feature recommendation function, it is possible to automatically evaluate the generation conditions of the second-class transaction features based on the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device. By generating the second-class transaction features based on the second transaction data when the generation probability is greater than the preset probability, and generating the target sample data based on the transaction field, the first-class transaction features, and the second-class transaction features, it is possible to automatically generate the second-class transaction features and the target sample data when the generation conditions of the second-class transaction features are met. In this way, through the embodiments of the present application, it is possible to realize the automated process of generating first-class transaction features, determining whether to generate second-class transaction features, and generating second-class transaction features and sample data, thereby simplifying user operations, improving the efficiency of sample data generation, and thereby improving model training efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0054] Figure 1 It is a flowchart of a first method for generating sample data provided in an embodiment of the present application;
[0055] Figure 2 is a flow chart of a second method for generating sample data provided in an embodiment of the present application;
[0056] Figure 3 is a structural schematic diagram of a sample data generation device provided in an embodiment of the present application;
[0057] Figure 4 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0058] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0059] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0060] In addition, the acquisition, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0061] With the increasing growth of online payments, more and more financial institutions and payment service providers have begun to collect transaction data, extract transaction features related to users' transaction behaviors from the transaction data, generate sample data based on the transaction features, and train models based on sample data with the above transaction features, so as to provide strong support for users' personalized services and precision marketing.
[0062] Among them, transaction features can be basic transaction features including the user's transaction amount distribution, purchased commodity types, purchase time statistics, transaction times, maximum / minimum transaction amounts, total transaction amounts, etc. Basic transaction features can usually only provide surface information and provide a rough description of user behavior. Models trained based on sample data with basic transaction features often lack depth and detail. For complex user transaction behavior patterns, basic transaction features may not provide sufficient explanatory power. Therefore, the prediction accuracy of the model trained based on sample data with basic transaction features is low.
[0063] In order to improve the accuracy of model prediction, currently experienced business personnel usually extract advanced transaction features such as meta features and user behavior pattern features from transaction data based on personal experience, and train models based on sample data with basic transaction features and advanced transaction features. Among them, advanced transaction features can reveal the deep patterns, trends and associations behind the data, which also helps to conduct more fine-grained user analysis in features, such as changes in user behavior patterns and consumption habits. The models generated by these extended features will also be of great help to personalized services and precision marketing.
[0064] However, the method of extracting basic transaction features and advanced transaction features based on personal experience will generate a lot of workload, and as the transaction data changes, it is necessary to manually adjust the basic transaction features and advanced transaction features continuously, which is time-consuming and labor-intensive, resulting in low efficiency in generating sample data, thereby reducing the efficiency of model training.
[0065] Thus, in order to solve the problems of the prior art, the embodiments of the present application provide a method, apparatus, device, computer-readable storage medium and computer program product for generating sample data. The method for generating sample data can be applied to scenarios of model training and model prediction.
[0066] The following first introduces the method for generating sample data provided in the embodiment of the present application.
[0067] Figure 1 FIG. 1 is a flow chart of a method for generating sample data provided by an embodiment of the present application. The method for generating sample data can be executed by an electronic device or server with data processing capability. Figure 1 As shown, the method for generating sample data provided in the embodiment of the present application includes the following steps:
[0068] S110, obtaining first transaction data, where the first transaction data includes a plurality of transaction fields;
[0069] S120. Classify multiple transaction fields based on preset field categories, and determine multiple field categories corresponding to the multiple transaction fields;
[0070] S130, for each field category, performing statistical analysis on multiple transaction fields to obtain a first type of transaction features;
[0071] S140, inputting a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of a transaction field, second attribute information of sample data, and computing power information of a target electronic device, where the target electronic device is an electronic device for model training based on the sample data;
[0072] S150. When the generation probability is greater than a preset probability, generate a second type of transaction feature based on the second transaction data, where the second transaction data includes the first transaction data and data generated based on the first type of transaction feature;
[0073] S160: Generate target sample data based on the transaction field, the first type of transaction features, and the second type of transaction features.
[0074] The embodiment of the present application classifies multiple transaction fields based on preset field categories, determines multiple field categories corresponding to multiple transaction fields in the first transaction data, and statistically analyzes multiple transaction fields for each field category to obtain first-class transaction features, thereby automatically generating first-class transaction features based on preset field categories. By taking the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device as inputs, and calculating the generation probability of the second-class transaction features based on the feature recommendation function, it is possible to automatically evaluate the generation conditions of the second-class transaction features based on the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device. By generating the second-class transaction features based on the second transaction data when the generation probability is greater than the preset probability, and generating the target sample data based on the transaction field, the first-class transaction features, and the second-class transaction features, it is possible to automatically generate the second-class transaction features and the target sample data when the generation conditions of the second-class transaction features are met. In this way, through the embodiments of the present application, it is possible to realize the automated process of generating first-class transaction features, determining whether to generate second-class transaction features, and generating second-class transaction features and sample data, thereby simplifying user operations, improving the efficiency of sample data generation, and thereby improving model training efficiency.
[0075] The specific implementation methods of the above steps are introduced below.
[0076] In some embodiments, in S110, the first transaction data may include information such as transaction amount, transaction time, transaction type, information of both parties to the transaction, etc. The first transaction data may exist in a data table. The data table of the first transaction data may include multiple transaction fields. The multiple transaction fields may include, for example, user identification, transaction identification, commodity name, commodity type, transaction date, transaction time, whether the transaction is paid, payment duration, whether the transaction is in installments, transaction amount, preferential amount, final payment amount, transaction location, transaction notes, merchant identification, merchant name, merchant category, merchant address, merchant business category, merchant status, merchant channel name, institution identification, institution name, etc. In addition, the above transaction fields can be regarded as original transaction features.
[0077] In the embodiment of the present application, the first transaction data can be a separate data table, or can be obtained by merging multiple data tables, which is not limited here. For example, the data table corresponding to the first transaction data can be obtained by merging the user transaction information table and the merchant information table.
[0078] In some embodiments, in S120, the preset field category may be a field category to which a pre-set transaction field belongs. The preset field category may include user transaction information fields, time information fields, amount information fields, commodity merchant information fields, additional information fields, etc.
[0079] Taking the above multiple transaction fields including user ID, transaction ID, commodity name, commodity type, transaction date, transaction time, whether the transaction is paid, payment duration, whether the transaction is in installments, transaction amount, discount amount, final payment amount, transaction location, transaction remarks, merchant ID, merchant name, merchant category, merchant address, merchant business category, merchant status, merchant channel name, organization ID, organization name, etc. as an example, multiple transaction fields are classified to determine multiple field categories corresponding to the multiple transaction fields, as shown below:
[0080] Specifically, user ID, transaction ID, whether the transaction is paid, whether the transaction is in installments, transaction location, etc. can be marked as user transaction information fields, and the data types corresponding to such fields can be marked as strings, numeric data, Boolean data, etc.
[0081] Transaction date, transaction time, payment duration, etc. can be marked as time information fields, and the data type corresponding to such fields can be marked as time format.
[0082] Transaction amount, discount amount, final payment amount, etc. can be marked as amount information fields, and the data type corresponding to such fields can be marked as amount value format.
[0083] Product name, business category, address information, merchant name, merchant type, organization name, etc. can be marked as product and merchant information fields, and the data type corresponding to such fields can be marked as text format;
[0084] Transaction notes can be marked as additional information fields, and the data type corresponding to this type of field can be marked as text format.
[0085] In some embodiments, in S130, the first type of transaction features may be the basic transaction features described above. The first type of transaction features corresponding to different transaction fields may be as follows:
[0086] User transaction behavior characteristics:
[0087] 1) The number of successful and failed transactions for each user, and the ratio of successful and failed transactions;
[0088] 2) The number, maximum value, and entropy of transactions completed by each user at different merchant types, cities, provinces, acquiring institutions, etc.;
[0089] 3) The number and maximum value of transactions completed by each user in different months, dates, days of the week, and whether they are working days;
[0090] 4) Statistics on the number of transactions completed by each user at the same merchant;
[0091] Amount characteristics:
[0092] 1) Statistical information such as the maximum, minimum, average, and total transaction amount of each user;
[0093] 2) Statistics of the maximum, minimum, average, and total amount of each user under different conditions such as month, date, day of the week, weekend, etc.;
[0094] 3) Statistical information such as the maximum, minimum, average, and total transaction amounts of each user after time decay weighting;
[0095] 4) Various statistical characteristics of the amount extracted after grouping by other categorical variables;
[0096] 5) Dimensionality reduction features of each user’s transaction amount sequence (by time);
[0097] Time characteristics:
[0098] 1) The time of each user’s first and last transaction and the time difference between them;
[0099] 2) The ratio of weekends to weekdays for each user’s transactions;
[0100] 3) Statistical characteristics of the number of transactions and the amount of each user in the past few months;
[0101] 4) The average, maximum, and minimum time differences before and after each user's transaction;
[0102] 5) Transaction statistics of each user during holidays.
[0103] After determining the first type of transaction features, the first type of transaction features can be regarded as transaction fields of the first transaction data, and first data can be generated based on the first type of transaction features and the first transaction data, and then second transaction data can be generated based on the first transaction data and the first data.
[0104] As an example, if the first transaction data exists in the form of a data table, after determining the first type of transaction features, the first type of transaction features can be added to the original transaction fields in the data table, and then the first type of transaction features can be used as the calculation target. The first data corresponding to the first type of transaction features can be calculated based on the first transaction data, and the first data can be added to the corresponding position of the data table to obtain the second transaction data.
[0105] Based on this, in order to ensure that the generated first-category transaction features have reasonable types and quantities, thereby ensuring both the accuracy and efficiency of model training, in some embodiments, the above S130 may specifically include:
[0106] Determining a feature generation mode based on a second influencing factor for generating a first type of transaction feature, the second influencing factor including first attribute information of a transaction field and second attribute information of sample data;
[0107] For each field category, multiple transaction fields are statistically analyzed based on the feature generation pattern to obtain the first category of transaction features.
[0108] Here, the feature generation mode may include a fast mode, a general mode, and a large number of modes. The number and types of first-category transaction features corresponding to different feature generation modes may be different. Among them, the fast mode can generate first-category transaction features of the number of transaction fields × 10; the general mode can generate first-category transaction features of the number of transaction fields × 50; the large number of modes can generate first-category transaction features of the number of transaction fields × 100. In addition, the types of first-category transaction features corresponding to the general mode may be more than those in the fast mode, and the types of first-category transaction features corresponding to the large number of modes may be more than those in the general mode. For example, based on the general mode, first-category transaction features such as aggregation features (i.e., groupby combined with statistical functions), time difference statistical features, time interval features (such as statistical features of the number of transactions / transaction amounts according to time) can be generated to extract interaction features between different times, while the fast mode may not generate the above-mentioned first-category transaction features.
[0109] Since the number and types of first-category transaction features corresponding to different feature generation modes may be different, the time required to generate the first-category transaction features using different feature generation modes is generally different.
[0110] In addition, the second influencing factor may be an influencing factor for determining a feature generation mode. The second influencing factor may include first attribute information of a transaction field and second attribute information of sample data. The first attribute information may include the number of transaction fields and field categories. The second attribute information may include the number of sample data, expected generation dimensions, and expected generation duration. The expected generation dimensions of the sample data may include the number of first-class transaction features that are expected to be generated. The expected generation duration may include the duration for generating the first-class transaction features and the duration for generating sample data based on the first-class transaction features.
[0111] In this way, by determining the feature generation pattern based on the second influencing factor, and for each field category, performing statistical analysis on multiple transaction fields based on the feature generation pattern, the first category of transaction features are obtained, so that the generated first category of transaction features can have reasonable types and quantities, thereby ensuring the accuracy and efficiency of model training at the same time.
[0112] In some embodiments, in S140, the second type of transaction features may be the advanced transaction features mentioned above. The second type of transaction features may include at least one of meta features, time series features, energy features, frequency features, transaction network features, user behavior pattern features, and the like.
[0113] The first influencing factor may be an influencing factor for determining whether to generate the second type of transaction feature. The first influencing factor may include the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device. Among them, the first attribute information may include the number of transaction fields and the field category. The second attribute information may include the number of sample data, the expected generation dimension, and the expected generation duration. The expected generation dimension of the sample data may include the number of the second type of transaction features expected to be generated. The expected generation duration may include the duration of generating the second type of transaction feature and the duration of generating the sample data based on the second type of transaction feature. In addition, the target electronic device may be an electronic device for model training based on the sample data, and the computing power information of the target electronic device may include the category information and the number of the target electronic device. Among them, the category information of the target electronic device may be associated with the processor performance, parallel processing capability, storage performance, memory bandwidth and capacity of the target electronic device.
[0114] In addition, the feature recommendation function may be a pre-fitted functional relationship for determining whether to generate the second type of transaction feature. The feature recommendation functions corresponding to different second type of transaction features may be different. For example, the meta feature may correspond to feature recommendation function A; the time series feature may correspond to feature recommendation function B. The feature recommendation function may take the first influencing factor as an independent variable and the generation probability of the second type of transaction feature as an independent variable. By inputting the first influencing factor into the feature recommendation function, the generation probability of the second type of transaction feature may be calculated.
[0115] As an example, the feature generation interface can display multiple second-category transaction features. The user can check at least one second-category transaction feature in the feature generation interface, input the specific values corresponding to each first influencing factor, and click the "OK" control, so that the electronic device where the feature generation interface is located, or the server corresponding to the feature generation interface, determines the feature recommendation function corresponding to each second-category transaction feature, and inputs the specific values corresponding to each first influencing factor into the feature recommendation function, and calculates the generation probability of the second-category transaction feature.
[0116] Based on this, in order to subsequently determine the generation probability of the second type of transaction features based on the feature recommendation function, in some embodiments, before the above S140, the method may further include:
[0117] Obtaining a reference influencing factor corresponding to the second type of transaction feature and a preset generation probability corresponding to the reference influencing factor, the reference influencing factor including reference attribute information of the transaction field, reference attribute information of the sample data, and reference computing power information of the target electronic device;
[0118] The initial feature recommendation function is fitted based on the reference influencing factors and their corresponding preset generation probabilities to obtain a feature recommendation function.
[0119] Here, the preset generation probability may include 0 and 1. The reference attribute information may include the number of transaction fields and the field category. The reference attribute information may include the reference number of sample data, the reference generation dimension, and the reference generation duration. The reference generation dimension of the sample data may include the number of second-category transaction features generated. The reference generation duration may include the duration for generating the second-category transaction features and the duration for generating the target sample data based on the second-category transaction features. The target sample data may be sample data having the second-category transaction features. In addition, the target electronic device may be an electronic device for model training based on the sample data, and the reference computing power information of the target electronic device may include the category information and quantity of the target electronic device. Among them, the category information of the target electronic device may be associated with the processor performance, parallel processing capability, storage performance, memory bandwidth and capacity of the target electronic device.
[0120] In addition, the initial feature recommendation function may be a linear function. The initial feature recommendation functions corresponding to different second-category transaction features may be different. For each second-category transaction feature, the feature recommendation function may be obtained by fitting the initial feature recommendation function based on the reference influencing factor and its corresponding preset generation probability.
[0121] The embodiment of the present application obtains a feature recommendation function by fitting an initial feature recommendation function based on a reference influencing factor and its corresponding preset generation probability, and can subsequently determine the generation probability of the second type of transaction feature based on the feature recommendation function.
[0122] In some embodiments, in S150, the preset probability may be a pre-set probability threshold for determining whether to generate the second type of transaction feature. The preset probability may be, for example, 0.5. If the generation probability is greater than the preset probability, the second type of transaction feature may be generated based on the transaction data. If the generation probability is less than or equal to the preset probability, the second type of transaction feature may not be generated.
[0123] Different second-category transaction features can be generated in different ways. For example, if the second-category transaction feature is a meta feature, the target model can be trained using the transaction data sample and its corresponding label to obtain the predicted value corresponding to the transaction data sample, and then the predicted value is used as a new transaction field in the transaction data sample, and then multiple transaction fields in the transaction data sample are statistically analyzed to obtain statistical features such as the maximum value, minimum value, and average value, i.e., the meta feature. Among them, if the target model is a user loyalty evaluation model, the transaction data sample can be the user's credit card historical transaction data, the label can be a pre-calibrated user loyalty, and the predicted value can be the predicted user loyalty.
[0124] In addition, if the second type of transaction features are transaction network features, the transaction network features can be determined based on graph network models or community detection. If the second type of transaction features are user behavior pattern features, behavior sequence analysis such as Markov model or hidden Markov model can be used to generate user behavior sequence data features, or cluster analysis methods such as K-means or hierarchical clustering can be used to classify user transactions and annotate user behavior model features.
[0125] Based on this, in order to improve the accuracy of the second type of transaction features, in some embodiments, after determining the multiple field categories and before generating the second type of transaction features based on the second transaction data, the method may further include:
[0126] Based on the field category, the first transaction data is preprocessed to obtain the preprocessed first transaction data; the data preprocessing at least includes classifying the date type data in the first transaction data according to a preset rule, and performing standardization processing on the first transaction data;
[0127] Generate first data based on the first type of transaction features and the preprocessed first transaction data;
[0128] The second transaction data is determined according to the preprocessed first transaction data and the first data.
[0129] Here, classifying the date type data in the first transaction data according to the preset rules may include classifying the date type data according to months, quarters, working days, holidays, etc., so as to determine which quarter a certain date belongs to, which day of the week it belongs to, whether it is a holiday, etc. In addition, standardizing the first transaction data may include performing unique hot encoding on the data belonging to enumeration features or classification features in the transaction data. For example, encoding "yes" in "whether the transaction is in installments" as [0,1], and encoding "no" in "whether the transaction is in installments" as [1,0].
[0130] As an example, data preprocessing may include at least one of the following processes:
[0131] Remove duplicate records: Delete duplicate transaction records by transaction ID;
[0132] Handling missing values: Fill, delete or interpolate missing data based on the classification of transaction fields;
[0133] Correct errors: Correct obvious errors in the data, such as incorrect transaction amounts and time formats, based on the classification of transaction fields. When it is impossible to obtain correct data, decide whether to delete outliers or perform special processing (such as replacing them with medians). At the same time, some data can be discarded.
[0134] Data formatting: Unify data formats according to the classification of transaction fields, such as date and time format, currency unit, etc.;
[0135] Data encoding: Convert data (such as user ID, merchant category) into numerical data, and convert the time data in the time information field into meaningful data such as the day of the week, quarter, whether it is a weekend, whether it is a holiday, etc.
[0136] Data normalization: Normalization or standardization, using one-hot encoding for enumeration or categorical features.
[0137] The specific process of generating the first data based on the first type of transaction features and the preprocessed first transaction data, and determining the second transaction data based on the preprocessed first transaction data and the first data may be as follows:
[0138] If the first transaction data exists in the form of a data table, after preprocessing the first transaction data, the data table may be updated based on the preprocessed first transaction data. After determining the first type of transaction features, the first type of transaction features may be added to the updated data table, and then the first type of transaction features may be used as a calculation target, and the first data corresponding to the first type of transaction features may be calculated based on the preprocessed first transaction data, and the first data may be added to the corresponding position of the updated data table to obtain the second transaction data.
[0139] The embodiment of the present application performs data preprocessing on the first transaction data based on the field category to obtain the preprocessed first transaction data, and determines the second transaction data based on the preprocessed first transaction data, so as to improve the accuracy of the second transaction data and further improve the accuracy of the second type of transaction features.
[0140] In some embodiments, in S160, the second type of transaction feature can be one of the intermediate and advanced transaction features such as meta features, time series features, energy features, frequency features, transaction network features, user behavior pattern features, etc., or multiple, which are not limited here. After determining the second type of transaction features, the second type of transaction features can be regarded as transaction fields of transaction data. By assigning values to multiple transaction fields, first type of transaction features, and second type of transaction features based on transaction data, target sample data can be generated.
[0141] Based on this, in order to further improve the accuracy of model training, in some embodiments, before determining the second transaction data according to the preprocessed first transaction data and the first data, the method may further include:
[0142] When the data type corresponding to the field category is text data, text analysis is performed on the transaction field corresponding to the field category to obtain text vectorization features;
[0143] Second data is generated based on the text vectorization feature and the preprocessed first transaction data.
[0144] Based on this, the above-mentioned determination of the second transaction data based on the preprocessed first transaction data and the first data may specifically include:
[0145] The second transaction data is determined according to the preprocessed first transaction data, the first data, and the second data.
[0146] Based on this, the above S160 may specifically include:
[0147] Generate target sample data based on transaction fields, first-category transaction features, second-category transaction features, and text vectorization features.
[0148] Here, the text vectorization feature may include at least one of an interaction feature, a TfidfVectorizer feature, and a Word2Vec feature. The TfidfVectorizer feature may be, for example, a feature obtained by performing dimensionality reduction processing on a sequence of merchant identification, merchant name, merchant type, channel name, organization type, etc. corresponding to each user based on a Tfidf matrix.
[0149] The method of generating Word2Vec features may be, for example: first extract the word embedding features of transaction fields such as merchant ID, merchant name, merchant type, name channel, transfer remarks, etc., and then calculate the maximum value, minimum value, average value, standard deviation and other statistical features corresponding to each of the above transaction fields according to the user ID to obtain the Word2Vec features.
[0150] After determining the text vectorization feature, the text vectorization feature can be regarded as a transaction field of the first transaction data, and the text vectorization feature is used as a calculation target to calculate the second data corresponding to the text vectorization feature based on the preprocessed first transaction data.
[0151] As an example, if the first transaction data exists in the form of a data table, after preprocessing the first transaction data, the data table can be updated based on the preprocessed first transaction data. After determining the first type of transaction features, the first type of transaction features can be first added to the updated data table, and then the first type of transaction features are used as the calculation target, and the first data corresponding to the first type of transaction features are calculated based on the preprocessed first transaction data, and the first data is added to the corresponding position of the updated data table. Similarly, after determining the text vectorization features, the text vectorization features can be first added to the updated data table, and then the text vectorization features are used as the calculation target, and the second data corresponding to the text vectorization features are calculated based on the preprocessed first transaction data, and the second data is added to the corresponding position of the updated data table to obtain the second transaction data including the first transaction data, the first data and the second data.
[0152] Afterwards, by assigning values to multiple transaction fields, first-category transaction features, second-category transaction features, and text vectorization features based on the transaction data, target sample data can be generated.
[0153] The embodiments of the present application can further improve the accuracy of model training by generating target sample data based on transaction fields, first-category transaction features, second-category transaction features, and text vectorization features, and performing model training based on the target sample data.
[0154] In addition, since feature extraction through empirical analysis often ignores the multi-dimensionality of data, some unnoticed dimensions (such as text data) may be important features, resulting in incomplete features and failure to ensure the universality of the model. Therefore, the embodiment of the present application performs text analysis on the transaction field corresponding to the field category when the data type corresponding to the field category is text data to obtain text vectorization features, generates target sample data based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features, and performs model training based on the target sample data, thereby ensuring the universality of model training.
[0155] Based on this, in order to further improve the accuracy of model training, in some embodiments, before generating target sample data based on the transaction fields, the first type of transaction features, the second type of transaction features and the text vectorization features, the method may further include:
[0156] Based on the second transaction data, an additional feature is generated, where the additional feature is any feature other than the first-category transaction feature, the second-category transaction feature, and the text vectorization feature.
[0157] Based on this, the target sample data is generated based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features, which may specifically include:
[0158] Generate target sample data based on transaction fields, first-category transaction features, second-category transaction features, text vectorization features, and additional-category features.
[0159] Here, the additional class features may include, for example, at least one of principal component analysis (PDA) features and linear discriminant analysis (LDA) features. By performing linear discriminant analysis on the transaction data and its corresponding labels, a feature subset with classification information may be determined to obtain LDA features.
[0160] After the additional class features are determined, the additional class features can be regarded as transaction fields of the transaction data. By assigning values to multiple transaction fields, first-class transaction features, second-class transaction features, text vectorization features, and additional class features based on the transaction data, target sample data can be generated.
[0161] The embodiments of the present application can further improve the accuracy of model training by generating target sample data based on transaction fields, first-category transaction features, second-category transaction features, text vectorization features, and additional-category features, and performing model training based on the target sample data.
[0162] Based on the above feature generation process, a large number of features can be generated. If the model is trained directly based on the sample data with the above large number of features, it will affect the accuracy and efficiency of model training. Therefore, in order to improve the accuracy and efficiency of model training, in some embodiments, after generating the first type of transaction features, the second type of transaction features, the text vectorization features and the additional type of features, the above features can be selected to determine the available second target transaction features, and then the target sample data can be generated based on the transaction field and the second target transaction features.
[0163] Based on this, the target sample data is generated based on the transaction field, the first type of transaction features, the second type of transaction features, the text vectorization features and the additional type features, which may specifically include:
[0164] Determine a first target transaction feature according to the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature;
[0165] Performing feature selection on the first target transaction feature based on a feature selection algorithm, and determining a first score for each transaction feature in the first target transaction feature;
[0166] Performing feature selection on the first target transaction feature based on the distribution verification algorithm to determine a second score for each transaction feature in the first target transaction feature;
[0167] Based on the weights corresponding to the first score and the second score respectively, performing a weighted summation on the first score and the second score to obtain a third score;
[0168] Determine the transaction feature corresponding to the third score greater than the preset threshold as the second target transaction feature;
[0169] Generate target sample data based on the transaction field and the second target transaction feature.
[0170] Here, the first target transaction feature may be a transaction feature obtained by concatenating the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature. The first target transaction feature may be recorded as a feature matrix X. The feature selection algorithm may include at least one of a Boruta algorithm, a lightgbm algorithm, an xgboost algorithm, and a catboost algorithm. The distribution validation algorithm may be an Adversarial validation algorithm. Among them, the Boruta algorithm may be used to select available features in the first target transaction feature.
[0171] The way to select available features using Boruta algorithm can be as follows:
[0172] 1) Randomly sort the first features in the feature matrix X;
[0173] 2) Concatenate the randomly sorted second feature with the first feature to obtain a new feature matrix Y;
[0174] 3) Based on the Boruta algorithm, the feature matrix Y is used as input to train a feature selection model that can output feature importance;
[0175] 4) Calculate the first score of the feature in the feature matrix Y based on the feature selection model;
[0176] 5) determining a target score with the largest value among the first scores corresponding to the plurality of second features;
[0177] 6) If the first score corresponding to the first feature is greater than the target score, the first feature is marked as an important feature; if the first score corresponding to the first feature is less than the target score, the first feature is marked as an unimportant feature and deleted from the first target transaction feature.
[0178] In addition, the lightgbm algorithm, the xgboost algorithm, and the catboost algorithm can be used to select important features in the first target transaction feature. The adversarial validation algorithm can be used to remove features with inconsistent distribution in the first target transaction feature.
[0179] As an example, the first score may include a first sub-score and a second sub-score. Based on the Boruta algorithm, the first sub-score of each transaction feature in the first target transaction feature may be determined, and the first sub-score may be used to select available features in the first target transaction feature; based on the lightgbm algorithm, the xgboost algorithm, or the catboost algorithm, the second sub-score of each transaction feature in the first target transaction feature may be determined, and the second sub-score may be used to select important features in the first target transaction feature; based on the Adversarial validation algorithm, the second score of each transaction feature in the first target transaction feature may be determined, and the second score may be used to delete features with inconsistent distribution in the first target transaction feature. Afterwards, by weighted summing the first sub-score, the second sub-score, and the second score based on the weights corresponding to the first sub-score, the second sub-score, and the second score, the third score of each transaction feature in the first target transaction feature may be obtained. Among them, the weights of the first sub-score and the second sub-score may be positive weights, and the weight of the second score may be negative weights.
[0180] The embodiment of the present application can reduce redundant features in the target sample data and retain the best features by determining the transaction feature corresponding to the third score greater than a preset threshold as the second target transaction feature, and generating target sample data based on the transaction field and the second target transaction feature, thereby improving the accuracy and efficiency of model training.
[0181] In addition, the above sample data can also be used for model prediction to improve the accuracy and efficiency of model prediction.
[0182] In order to better describe the entire solution, some specific examples are given based on the above embodiments.
[0183] For example, Figure 2 As shown, a method for generating sample data provided by an embodiment of the present application may include the following steps:
[0184] S21. Acquire first transaction data, where the first transaction data includes a plurality of transaction fields;
[0185] S22, classifying multiple transaction fields to obtain multiple field categories;
[0186] S23. Preprocess the first transaction data based on the field category to obtain preprocessed first transaction data.
[0187] S24, determining a feature generation mode based on a second influencing factor for generating the first type of transaction feature;
[0188] S25. For each field category, statistical analysis is performed on multiple transaction fields based on the feature generation mode to obtain a first category of transaction features;
[0189] S26. Generate first data based on the first type of transaction features and the preprocessed first transaction data;
[0190] S27, determine whether the data type corresponding to the field category is text data, if so, execute S28, if not, execute S211;
[0191] S28. Perform text analysis on the transaction fields corresponding to the field category to obtain text vectorization features;
[0192] S29, generating second data based on the text vectorization feature and the preprocessed first transaction data;
[0193] S210, determining second transaction data according to the preprocessed first transaction data, the first data, and the second data;
[0194] S211, inputting the first influencing factor for generating the second type of transaction feature into the feature recommendation function to determine the generation probability of the second type of transaction feature;
[0195] S212: if the generation probability is greater than the preset probability, generate a second type of transaction feature based on the second transaction data;
[0196] S213, generating additional class features based on the second transaction data;
[0197] S214, determining a first target transaction feature according to the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature;
[0198] S215, performing feature selection on the first target transaction feature based on the feature selection algorithm and the distribution verification algorithm to obtain a second target transaction feature;
[0199] S216: Generate target sample data based on the transaction field and the second target transaction feature.
[0200] Therefore, the embodiment of the present application can automatically determine which types of transaction features to generate (including at most the first type of transaction features, the second type of transaction features, text vectorization features and additional type features), ensuring the dimension and diversity of transaction features and meeting the training requirements of various models. In addition, by automatically cross-validating and screening a large number of generated transaction features, the most helpful features for model training can be selected to avoid excessive time consumption during model training and prediction.
[0201] Based on the sample data generation method provided in the above embodiment, the present application also provides a specific implementation of a sample data generation device. Please refer to the following embodiment.
[0202] like Figure 3 As shown, the sample data generation device 300 provided in the embodiment of the present application includes the following modules:
[0203] A first acquisition module 310, configured to acquire first transaction data, where the first transaction data includes a plurality of transaction fields;
[0204] A classification module 320, for classifying the plurality of transaction fields based on preset field categories, and determining a plurality of field categories corresponding to the plurality of transaction fields;
[0205] A first analysis module 330, for performing statistical analysis on multiple transaction fields for each field category to obtain a first type of transaction features;
[0206] An input module 340 is used to input a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of a transaction field, second attribute information of sample data, and computing power information of a target electronic device, where the target electronic device is an electronic device for model training based on the sample data;
[0207] A first generating module 350, configured to generate a second type of transaction feature based on the second transaction data when the generation probability is greater than a preset probability, the second transaction data including the first transaction data and data generated based on the first type of transaction feature;
[0208] The second generating module 360 is used to generate target sample data based on the transaction field, the first type of transaction features and the second type of transaction features.
[0209] The sample data generation device 300 is described in detail below, as shown below:
[0210] In some embodiments, the sample data generating device 300 may further include:
[0211] A second acquisition module is used to obtain a reference influence factor corresponding to the second type of transaction feature and a preset generation probability corresponding to the reference influence factor before inputting the first influence factor for generating the second type of transaction feature into the feature recommendation function, wherein the reference influence factor includes reference attribute information of the transaction field, reference attribute information of the sample data, and reference computing power information of the target electronic device;
[0212] The fitting module is used to fit the initial feature recommendation function based on the reference influencing factor and its corresponding preset generation probability to obtain the feature recommendation function.
[0213] In some embodiments, the first analysis module 330 may specifically include:
[0214] A first determination submodule, configured to determine a feature generation mode based on a second influencing factor for generating a first type of transaction feature, wherein the second influencing factor includes first attribute information of a transaction field and second attribute information of sample data;
[0215] The analysis submodule is used to perform statistical analysis on multiple transaction fields based on the feature generation mode for each field category to obtain the first category of transaction features.
[0216] In some embodiments, the sample data generating device 300 may further include:
[0217] a processing module, configured to, after determining the plurality of field categories and before generating the second type of transaction features based on the second transaction data, perform data preprocessing on the first transaction data based on the field categories to obtain preprocessed first transaction data; the data preprocessing at least includes classifying date type data in the first transaction data according to a preset rule and performing standardization processing on the first transaction data;
[0218] A third generating module, configured to generate first data based on the first type of transaction features and the preprocessed first transaction data;
[0219] The determination module is used to determine the second transaction data according to the preprocessed first transaction data and the first data.
[0220] In some embodiments, the sample data generating device 300 may further include:
[0221] A second analysis module is used for, before determining the second transaction data according to the preprocessed first transaction data and the first data, performing text analysis on the transaction field corresponding to the field category when the data type corresponding to the field category is text data, to obtain a text vectorization feature;
[0222] The fourth generating module is used to generate the second data based on the text vectorization feature and the preprocessed first transaction data.
[0223] Based on this, the determination module may specifically include:
[0224] The second determining submodule is used to determine the second transaction data according to the preprocessed first transaction data, the first data and the second data.
[0225] Based on this, the second generating module 360 may specifically include:
[0226] The first generation submodule is used to generate target sample data based on transaction fields, first-category transaction features, second-category transaction features, and text vectorization features.
[0227] In some embodiments, the second generating module 360 may further include:
[0228] The second generating submodule is used to generate additional class features based on the second transaction data before generating target sample data based on the transaction fields, the first class transaction features, the second class transaction features and the text vectorization features. The additional class features are any features other than the first class transaction features, the second class transaction features and the text vectorization features.
[0229] Based on this, the first generation submodule may specifically include:
[0230] A generating unit is used to generate target sample data based on transaction fields, first-category transaction features, second-category transaction features, text vectorization features, and additional-category features.
[0231] In some embodiments, the generating unit may be specifically configured to:
[0232] Determine a first target transaction feature according to the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature;
[0233] Performing feature selection on the first target transaction feature based on a feature selection algorithm, and determining a first score for each transaction feature in the first target transaction feature;
[0234] Performing feature selection on the first target transaction feature based on the distribution verification algorithm to determine a second score for each transaction feature in the first target transaction feature;
[0235] Based on the weights corresponding to the first score and the second score respectively, performing a weighted summation on the first score and the second score to obtain a third score;
[0236] Determine the transaction feature corresponding to the third score greater than the preset threshold as the second target transaction feature;
[0237] Generate target sample data based on the transaction field and the second target transaction feature.
[0238] The embodiment of the present application classifies multiple transaction fields based on preset field categories, determines multiple field categories corresponding to multiple transaction fields in the first transaction data, and statistically analyzes multiple transaction fields for each field category to obtain first-class transaction features, thereby automatically generating first-class transaction features based on preset field categories. By taking the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device as inputs, and calculating the generation probability of the second-class transaction features based on the feature recommendation function, it is possible to automatically evaluate the generation conditions of the second-class transaction features based on the first attribute information of the transaction field, the second attribute information of the sample data, and the computing power information of the target electronic device. By generating the second-class transaction features based on the second transaction data when the generation probability is greater than the preset probability, and generating the target sample data based on the transaction field, the first-class transaction features, and the second-class transaction features, it is possible to automatically generate the second-class transaction features and the target sample data when the generation conditions of the second-class transaction features are met. In this way, through the embodiments of the present application, it is possible to realize the automated process of generating first-class transaction features, determining whether to generate second-class transaction features, and generating second-class transaction features and sample data, thereby simplifying user operations, improving the efficiency of sample data generation, and thereby improving model training efficiency.
[0239] Based on the method for generating sample data provided in the above embodiment, the embodiment of the present application also provides a specific implementation of the electronic device. Figure 4 A schematic diagram of an electronic device 400 provided in an embodiment of the present application is shown.
[0240] The electronic device 400 may include a processor 410 and a memory 420 storing computer program instructions.
[0241] Specifically, the processor 410 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0242] The memory 420 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 420 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. Where appropriate, the memory 420 may include a removable or non-removable (or fixed) medium. Where appropriate, the memory 420 may be inside or outside the electronic device 400. In a particular embodiment, the memory 420 is a non-volatile solid-state memory.
[0243] The memory may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to the first aspect of the present application.
[0244] The processor 410 implements any one of the methods for generating sample data in the above embodiments by reading and executing computer program instructions stored in the memory 420 .
[0245] In one example, the electronic device 400 may further include a communication interface 430 and a bus 440. Figure 4 As shown, the processor 410, the memory 420, and the communication interface 430 are connected via a bus 440 and communicate with each other.
[0246] The communication interface 430 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0247] Bus 440 includes hardware, software or both, and the parts of electronic equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 440 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.
[0248] Exemplarily, the electronic device 400 may be a mobile phone, a tablet computer, a laptop computer, a PDA, an in-vehicle electronic device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA).
[0249] The electronic device can execute the method for generating sample data in the embodiment of the present application, thereby realizing the combination Figures 1 to 3 The invention describes a method and apparatus for generating sample data.
[0250] In addition, in combination with the sample data generation method in the above embodiments, the present application embodiment may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by a processor, any of the sample data generation methods in the above embodiments is implemented.
[0251] In combination with the sample data generation method in the above embodiments, the present application embodiment may provide a computer program product for implementation. When the instructions in the computer program product are executed by a processor of an electronic device, any of the sample data generation methods in the above embodiments is implemented.
[0252] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0253] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0254] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0255] The above reference is according to the method of the embodiment of the present application, the flow chart of the device (system) and the computer program product and / or the block diagram described various aspects of the present application.It should be understood that each square box in the flow chart and / or the block diagram and the combination of each square box in the flow chart and / or the block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the realization of the function / action specified in one or more square boxes of the flow chart and / or the block diagram.Such a processor can be but is not limited to a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit.It can also be understood that each square box in the block diagram and / or the flow chart and the combination of the square boxes in the block diagram and / or the flow chart can also be realized by the dedicated hardware that performs the specified function or action, or can be realized by the combination of dedicated hardware and computer instructions.
[0256] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. A method for generating sample data, characterized in that: include: Acquire first transaction data, where the first transaction data includes a plurality of transaction fields; Based on preset field categories, classify the multiple transaction fields and determine multiple field categories corresponding to the multiple transaction fields; For each of the field categories, statistical analysis is performed on the multiple transaction fields to obtain a first type of transaction feature; Inputting a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of the transaction field, second attribute information of the sample data, and computing power information of a target electronic device, the target electronic device being an electronic device for model training based on the sample data; When the generation probability is greater than a preset probability, generating the second type of transaction feature based on second transaction data, where the second transaction data includes the first transaction data and data generated based on the first type of transaction feature; Target sample data is generated based on the transaction field, the first type of transaction features, and the second type of transaction features.
2. The method according to claim 1, characterized in that Before inputting the first influencing factor for generating the second type of transaction feature into the feature recommendation function, the method further includes: Acquire a reference influencing factor corresponding to the second type of transaction feature and a preset generation probability corresponding to the reference influencing factor, wherein the reference influencing factor includes reference attribute information of the transaction field, reference attribute information of the sample data, and reference computing power information of the target electronic device; The initial feature recommendation function is fitted based on the reference influencing factor and the corresponding preset generation probability to obtain the feature recommendation function.
3. The method according to claim 1, characterized in that For each of the field categories, statistical analysis is performed on the multiple transaction fields to obtain a first type of transaction features, including: Determining a feature generation mode based on a second influencing factor for generating the first type of transaction feature, the second influencing factor including first attribute information of the transaction field and second attribute information of the sample data; For each of the field categories, statistical analysis is performed on the multiple transaction fields based on the feature generation mode to obtain a first type of transaction features.
4. The method according to claim 1, characterized in that After determining the plurality of field categories and before generating the second-type transaction features based on the second transaction data, the method further includes: Based on the field category, preprocessing the first transaction data to obtain preprocessed first transaction data; the data preprocessing at least includes classifying date type data in the first transaction data according to a preset rule, and standardizing the first transaction data; generating first data based on the first type of transaction features and the preprocessed first transaction data; The second transaction data is determined according to the preprocessed first transaction data and the first data.
5. The method according to claim 4, characterized in that Before determining the second transaction data according to the preprocessed first transaction data and the first data, the method further includes: In a case where the data type corresponding to the field category is text data, performing text analysis on the transaction field corresponding to the field category to obtain a text vectorization feature; Generate second data based on the text vectorization feature and the preprocessed first transaction data; The determining the second transaction data according to the preprocessed first transaction data and the first data includes: determining the second transaction data according to the preprocessed first transaction data, the first data, and the second data; The generating target sample data based on the transaction field, the first type of transaction features and the second type of transaction features includes: Generate target sample data based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features.
6. The method according to claim 5, characterized in that Before generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features, and the text vectorization features, the method further includes: Based on the second transaction data, generate an additional feature, where the additional feature is any feature other than the first transaction feature, the second transaction feature, and the text vectorization feature; The generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features and the text vectorization features includes: Generate target sample data based on the transaction field, the first category transaction features, the second category transaction features, the text vectorization features and the additional category features.
7. The method according to claim 6, characterized in that The generating target sample data based on the transaction field, the first type of transaction features, the second type of transaction features, the text vectorization features and the additional type features includes: Determine a first target transaction feature according to the first type of transaction feature, the second type of transaction feature, the text vectorization feature, and the additional type of feature; Performing feature selection on the first target transaction features based on a feature selection algorithm, and determining a first score for each transaction feature in the first target transaction features; Performing feature selection on the first target transaction features based on a distribution verification algorithm to determine a second score for each transaction feature in the first target transaction features; Based on the weights corresponding to the first score and the second score respectively, performing a weighted summation on the first score and the second score to obtain a third score; Determining the transaction feature corresponding to the third score greater than a preset threshold as a second target transaction feature; Generate target sample data based on the transaction field and the second target transaction feature.
8. A device for generating sample data, characterized in that: The device comprises: A first acquisition module, configured to acquire first transaction data, where the first transaction data includes a plurality of transaction fields; A classification module, configured to classify the plurality of transaction fields based on preset field categories, and determine a plurality of field categories corresponding to the plurality of transaction fields; A first analysis module, configured to perform statistical analysis on the plurality of transaction fields for each of the field categories to obtain a first type of transaction feature; An input module, used to input a first influencing factor for generating a second type of transaction feature into a feature recommendation function to determine a generation probability of the second type of transaction feature; the first influencing factor includes first attribute information of the transaction field, second attribute information of the sample data, and computing power information of a target electronic device, the target electronic device being an electronic device for model training based on the sample data; a first generating module, configured to generate the second type of transaction feature based on second transaction data when the generation probability is greater than a preset probability, the second transaction data including the first transaction data and data generated based on the first type of transaction feature; The second generating module is used to generate target sample data based on the transaction field, the first type of transaction characteristics and the second type of transaction characteristics.
9. An electronic device, characterized in that: The electronic device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the method for generating sample data according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method for generating sample data according to any one of claims 1 to 7 is implemented.
11. A computer program product, characterized in that When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the method for generating sample data as described in any one of claims 1 to 7.