Feature Derivation Method, Apparatus, Non-Volatile Storage Medium, and Electronic Device

By constructing and screening initial derivative features, and using feature verification instructions and fitness functions, the problem of extracting data relationships from massive data in data mining is solved, efficient and automated feature derivatives are achieved, and manpower investment is reduced.

CN114661750BActive Publication Date: 2025-06-03INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210278066.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-21
Publication Date
2025-06-03
Estimated Expiration
2042-03-21

AI Technical Summary

Technical Problem

During the data mining process, it is difficult for the prior art to extract rich and effective data relationships from massive data, and the feature derivation process lacks randomness, easily loses derivative features with high correlation, and requires a lot of manpower.

Method used

By constructing initial derivative features based on multiple basic features, and generating feature verification instructions that correspond to the initial derivative features one by one, the target derived features are filtered out based on the execution results of the target database, and the feature cross-variation is performed using methods such as fitness functions and genetic algorithms to screen out derivative features with data mining significance.

Benefits of technology

It realizes the extraction of rich and effective data relationships from massive data, reduces manpower investment, improves the efficiency and effect of feature derivation, and can dig out meaningful data derivation features that are difficult for human resources to detect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114661750B_ABST
    Figure CN114661750B_ABST
Patent Text Reader

Abstract

The present invention discloses a feature derivation method, apparatus, non-volatile storage medium and electronic device, which are applied to the field of financial technology. Among them, the method includes: constructing initial derivative features based on a plurality of basic features, where the basic features are used to describe the data stored in the target database, and each initial derivative feature includes at least two basic features; generating first feature verification instructions corresponding one-to-one to the initial derivative features; and screening out first target derivative features from the initial derivative features based on the first execution result of the target database for the first feature verification instructions. The present invention solves the technical problem of how to extract rich and effective data relationships from a large amount of data during data mining.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fintech, and in particular, to a feature derivation method, device, non-volatile storage medium, and electronic device. Background Art

[0002] With the explosive growth of Internet products, the volume and scale of data storage have also shown a geometric growth trend. By monitoring and comparing new incoming data and historical data, providing effective reference for decision-making has become a key topic in machine learning, and how to extract effective relationships from massive data has become the focus and difficulty in the data mining process. In the process of learning massive historical data, financial Internet enterprises will formulate a large number of basic features, and complete the scoring process of relevant scorecards by learning such feature indicators, including corporate credit scoring, loan business scoring, etc. Currently, each enterprise will combine its own business characteristics and perform corresponding feature derivation based on the existing basic features.

[0003] In related technologies, most of the real-time data stored in enterprises is basically source-attached data (raw data, basically without any processing), and not much processing is done on the data. In order to effectively improve the learning efficiency and learning effect, relevant feature derivation is performed based on the original basic features to generate a series of derivative features that more prominently reflect user portrait features. Such derivation mostly adopts the methods of manual derivation or templated derivation, and determines the derivation effect according to the IV value and WOE value of the derived features, etc. The basic features are combined with the derivative features with better prediction characteristics to jointly construct the input parameters of the model and perform mathematical modeling to complete the data mining process. Obviously, due to the fixed thinking of people, and the features derived by templated derivation are inevitably affected by the preset template, the entire derivation process lacks randomness, and relevant derivative features with high relevance but not easily noticed by staff will be lost to varying degrees. At the same time, feature derivation also requires a large amount of manpower to formulate and consume manpower to manually verify the derived features.

[0004] To address the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] Embodiments of the present invention provide a feature derivation method, device, non-volatile storage medium, and electronic device to at least solve the technical problem of how to extract rich and effective data relationships from massive data during data mining.

[0006] According to one aspect of an embodiment of the present invention, a feature derivation method is provided, including: constructing initial derived features based on a plurality of basic features, where the basic features are used to describe data stored in a target database, and each of the initial derived features includes at least two of the basic features; generating first feature verification instructions corresponding to the initial derived features one by one; and screening out first target derived features from the initial derived features based on a first execution result of the target database for the first feature verification instructions.

[0007] Optionally, the screening out first target derived features from the initial derived features based on the first execution result of the target database for the first feature verification instructions includes: dividing the initial derived features into a first set and a second set; determining a first fitness corresponding to each of the initial derived features based on the first execution result and a fitness function, where the fitness function is used to calculate the fitness of each of the initial derived features according to the first execution result corresponding to each of the initial derived features and the set to which each of the initial derived features belongs; and screening out the first target derived features from the initial derived features based on the first fitness.

[0008] Optionally, the screening out the first target derived features from the initial derived features based on the first fitness includes: randomly generating a fitness threshold within a predetermined numerical range; comparing the magnitudes of the first fitness and the fitness threshold, and determining the initial derived features corresponding to the first fitness greater than the fitness threshold as the first target derived features.

[0009] Optionally, the method further includes: obtaining a second target derived feature and a third target derived feature included in the first target derived feature, where the second target derived feature belongs to the first set and the third target derived feature belongs to the second set; exchanging the basic features included in the second target derived feature and the third target derived feature respectively to obtain a plurality of cross-genetic features; and screening out target genetic features from the plurality of cross-genetic features.

[0010] Optionally, screening out target genetic features from the multiple cross-genetic features includes: dividing the multiple cross-genetic features into a third set and a fourth set; generating second feature verification instructions corresponding to the multiple cross-genetic features one by one; determining a second fitness corresponding to each cross-genetic feature based on a second execution result of the second feature verification instruction for the target database and the fitness function, where the fitness function is used to calculate the fitness of each cross-genetic feature according to the second execution result corresponding to each cross-genetic feature and the set to which each cross-genetic feature belongs; and screening out the target genetic features from the cross-genetic features based on the second fitness.

[0011] Optionally, screening out target genetic features from the cross-genetic features includes: obtaining a mutation factor, where the mutation factor is used to indicate a replacement ratio for replacing basic features included in the cross-genetic features; replacing the basic features in each cross-genetic feature according to the mutation factor to obtain mutated genetic features; and screening out the target genetic features from the mutated genetic features.

[0012] Optionally, before screening out first target derivative features from the initial derivative features based on a first execution result of the first feature verification instruction for the target database, the method further includes: sending the first feature verification instruction to the target database, where the first feature verification instruction includes database query statements corresponding to the initial derivative features one by one; and receiving the first execution result fed back by the target database, where the first execution result is used to indicate whether the target database can successfully execute the database query statements.

[0013] According to another aspect of the embodiments of the present invention, there is also provided a feature derivation device, including: a construction module, configured to construct initial derivative features based on multiple basic features, where the basic features are used to describe data stored in a target database, and each initial derivative feature includes at least two of the basic features; a generation module, configured to generate first feature verification instructions corresponding to the initial derivative features one by one; and a screening module, configured to screen out first target derivative features from the initial derivative features based on a first execution result of the first feature verification instruction for the target database.

[0014] According to still another aspect of the embodiments of the present invention, there is also provided a non-volatile storage medium, where the non-volatile storage medium includes a stored program, and when the program runs, it controls a device where the non-volatile storage medium is located to execute the feature derivation method described in any one of the above.

[0015] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including a processor for running a program, wherein when the program runs, it executes the feature derivation method described in any one of the above.

[0016] In the embodiments of the present invention, based on multiple basic features, initial derivative features are constructed, where the basic features are used to describe the data stored in the target database, and each initial derivative feature includes at least two basic features; first feature verification instructions corresponding one-to-one to the initial derivative features are generated; based on the first execution results of the target database for the first feature verification instructions, the first target derivative features are screened out from the initial derivative features, achieving the purpose of constructing derivative features with data mining significance for the data stored in a specific database, thereby realizing the technical effect of mining meaningful data derivative features that are difficult to detect by humans, and further solving the technical problem of how to extract rich and effective data relationships from massive data during data mining. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0018] Figure 1 A hardware structure block diagram of a computer terminal for implementing the feature derivation method is shown;

[0019] Figure 2 It is a flowchart of the feature derivation method provided by the embodiments of the present invention;

[0020] Figure 3 It is a flowchart of feature derivation using the genetic method according to an alternative embodiment of the present invention;

[0021] Figure 4 It is a structure block diagram of the feature derivation device provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0024] First, some nouns or terms that appear during the description of the embodiments of the present application are applicable to the following explanations:

[0025] The basic features, which can also be called the existing features or original features of the data. During the data mining process, the basic features of each data can reflect information in one aspect of the data. For example, in the financial field, the basic features can include various types such as "the last 6 months, the last 12 months, account amount, overdue days, maximum overdue days in history, account balance, valid period of collateral, collateral evaluation date, collateral audit date, collateral evaluation amount", etc.

[0026] The derivative features refer to the new features obtained after combining and arranging the basic features or feature learning. For example, "account amount in the last 6 months" as a derivative feature is obtained by combining the two basic features of "the last 6 months" and "account amount". Conducting data mining based on the derivative features can reflect multi-dimensional information of the data.

[0027] The genetic algorithm is a computational model that simulates the natural selection and genetic mechanism of Darwin's theory of evolution in the biological evolution process, and is a method for searching for the optimal solution by simulating the natural evolution process.

[0028] According to an embodiment of the present invention, an embodiment of a method for feature derivation is provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.

[0029] Through the feature derivation method provided by the embodiments of the present invention, valuable derivative features can be mined based on financial big data and the basic features of the data. Providing the mined derivative features to financial analysts such as risk control administrators or other financial product developers can help them better grasp the user portraits and behavioral characteristics of the population corresponding to the financial big data, and thus better utilize the data to complete relevant work.

[0030] The method embodiment provided by the first embodiment of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the feature derivation method is shown. As Figure 1 shown, the computer terminal 10 may include one or more (shown as 102a, 102b,..., 102n in the figure) processors (the processor may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), and a memory 104 for storing data. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only illustrative and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0031] It should be noted that the above-mentioned one or more processors and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any combination thereof. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0032] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the feature derivation method in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the feature derivation method of the above application program. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0033] The display can be, for example, a touch-screen liquid crystal display (LCD), and the liquid crystal display enables a user to interact with the user interface of the computer terminal 10.

[0034] Figure 2 is a schematic flowchart of the feature derivation method provided according to the embodiments of the present invention, as Figure 2 shown, the method includes the following steps:

[0035] Step S202, based on a plurality of basic features, construct initial derived features, where the basic features are used to describe the data stored in the target database, and each initial derived feature includes at least two basic features.

[0036] It should be noted that in the case where the target database is a database storing big data in the financial field, each basic feature can reflect the information of one dimension of the data in the database. For example, for the financial field, the basic features can include any one of multiple types such as "last 6 months, last 12 months, account amount, overdue days, maximum historical overdue days, account balance, valid period of collateral, collateral evaluation date, collateral audit date, collateral evaluation amount", etc.

[0037] The initial derived features are the unvalidated derived features constructed according to the basic features in this embodiment. For example, "account amount in the last 6 months" as a derived feature is obtained by combining the two basic features of "last 6 months" and "account amount". Since it is not validated, the initial derived features may be unreasonable derived features. For example, it may construct illogical derived features such as "last 6 months last 12 months", and such initial derived features are useless for data analysts and will be eliminated through subsequent processing.

[0038] Step S204, generate first feature verification instructions corresponding to the initial derived features one by one.

[0039] It should be noted that the first feature verification instruction should correspond one by one to the initial derived features to verify whether each initial derived feature is valid. Among them, the first feature verification instruction can be an instruction that the target database can read and understand, or it can be pseudocode that can be understood by the target database after compilation. Among the initial derived features, the invalid initial derived features can be considered as illogical and unreasonable derived features. For example, the basic features in the invalid initial derived features can include information on two dimensions of data, but there is no data including the information on these two dimensions in the target database. Then this initial derived feature is meaningless for the data mining work of the target database. On the contrary, the valid initial derived features can reflect the multi-dimensional information of some data in the target database.

[0040] Step S206, based on the first execution result of the target database for the first feature verification instruction, screen out the first target derived features from the initial derived features.

[0041] In this step, the target database can try to execute the first feature verification instruction. For example, search in the target database to determine whether there is data in the database that can simultaneously reflect the multi-dimensional information required by the initial derived feature corresponding to the first feature verification instruction. If such data exists in the target database, the first execution result is valid or verified. Further, the initial derived features corresponding to the first feature verification instruction with a valid first execution result can be screened out as the first target derived features, that is, this derived feature is a meaningful derived feature and can be retained for subsequent use.

[0042] As an optional embodiment, before screening out the first target derived features from the initial derived features, the first execution result can be obtained in the following way: send the first feature verification instruction to the target database, where the first feature verification instruction includes a database query statement corresponding one by one to the initial derived features; receive the first execution result fed back by the target database, where the first execution result is used to indicate whether the target database can successfully execute the database query statement.

[0043] Optionally, the database query statement can be an execution SQL (Structured Query Language, abbreviated as SQL) statement generated corresponding to the description of the basic features in the derived features. For example, when the basic feature is "in the past 3 months", the corresponding SQL statement can be: ETL_WRK_DT >= ADD_MONTHS(date_sub(current_date, 1), -3); when the basic feature is "transaction number is 00508", the corresponding SQL statement can be: tr_cde = '00508'.

[0044] Further, for a derived feature including multiple basic features, the connection method of the SQL statement executed by the target database can use "and" to splice feature combinations. For example, when the derived feature includes two basic features, the target database can execute a judgment by testing the execution feedback result of'select count(1) from table where the SQL statement corresponding to the first basic feature and the SQL statement corresponding to the second basic feature'. If the feedback is a non-zero number, it is determined that the execution feedback result is a successful execution, and the current test result is set to pass. Otherwise, it can be fed back that the execution feedback result is a failed execution and the test fails. During the above test process, the SQL statement can find in the target database whether there is data that conforms to both the first basic feature and the second basic feature at the same time. If such data exists, the execution feedback result is returned as a successful execution, which also means that the derived feature is meaningful for data mining work and is worth retaining.

[0045] Through the above steps, based on multiple basic features, initial derived features are constructed, where the basic features are used to describe the data stored in the target database, and each initial derived feature includes at least two basic features; first feature verification instructions corresponding to the initial derived features are generated; based on the first execution result of the target database for the first feature verification instructions, the first target derived features are screened out from the initial derived features, achieving the purpose of constructing derived features with data mining significance for the data stored in a specific database, thus realizing the technical effect of mining meaningful data derived features that are difficult to detect by humans, and further solving the technical problem of how to extract rich and effective data relationships from a large amount of data during data mining.

[0046] As an optional embodiment, screening out the first target derived features from the initial derived features based on the first execution result of the target database for the first feature verification instructions includes the following process: dividing the initial derived features into a first set and a second set; determining the first fitness corresponding to each initial derived feature based on the first execution result and the fitness function, where the fitness function is used to calculate the fitness of the initial derived feature according to the first execution result corresponding to each initial derived feature and the set to which each initial derived feature belongs; screening out the first target derived features from the initial derived features based on the first fitness.

[0047] In this alternative embodiment, multiple initial derivative features can be randomly assigned and evenly distributed into a first set and a second set. In a genetic algorithm, the first set and the second set can be referred to as the initial population. Since the first execution results of each initial derivative feature have been obtained from the execution feedback results of the target database, the first fitness of each initial derivative feature can be calculated based on a fitness function on this basis. The fitness function can also be called an evaluation function, which is used to reflect the quality of an initial derivative feature in a set. An initial derivative feature with a higher first fitness can be considered a better derivative feature and is more meaningful for subsequent data mining work. Therefore, the first target derivative features can be selected based on the first fitness corresponding to each initial derivative feature.

[0048] Optionally, the fitness function can be constructed using this functional form: where f(x) represents the value of the fitness, i represents the serial number of the initial derivative feature, n represents the total number of all initial derivative features, k is a constant (for example, it can be set to 0.1), μ is a hyperparameter (for example, it can be set to 0.25), represents the proportion of the SQL statements corresponding to the initial derivative features that are successfully executed in the target database among all initial derivative features, represents the execution passing rate of the SQL statements corresponding to all initial derivative features in the target database.

[0049] As an alternative embodiment, the first target derivative features can be selected from the initial derivative features in the following manner: randomly generate a fitness threshold within a predetermined numerical range; compare the size of the first fitness with the fitness threshold, and determine the initial derivative features corresponding to the first fitness greater than the fitness threshold as the first target derivative features.

[0050] In this alternative embodiment, a method for screening the first target derivative features is provided, that is, the initial derivative features corresponding to the first fitness not greater than the fitness threshold are "eliminated", and the initial derivative features that "survive" are determined as the first target derivative features. The above process can be considered as a round of adaptive screening of the derivative features, screening out the inferior derivative features to avoid interfering with the work of subsequent data miners, and retaining the derivative features with more data mining value.

[0051] As an alternative embodiment, more derivative features can be generated through the following steps: obtain the second target derivative features and the third target derivative features included in the first target derivative features, where the second target derivative features belong to the first set and the third target derivative features belong to the second set; exchange the basic features included in the second target derivative features and the third target derivative features respectively to obtain multiple cross-genetic features; screen out the target genetic features from the multiple cross-genetic features.

[0052] This optional embodiment can be analogous to a genetic algorithm. By performing basic feature exchanges between derivative features, derivative feature types that do not exist in the initial derivative features can be generated, greatly expanding the scope of materials for derivative features provided to data miners. For example, the second target derivative feature may include basic feature A and basic feature B, and the third target derivative feature may include basic feature C and basic feature D. By performing a basic feature exchange between the second target derivative feature and the third target derivative feature, basic feature B and basic feature D can be exchanged to obtain two cross-genetic features. The first cross-genetic feature includes basic feature A and basic feature D, and the second cross-genetic feature includes basic feature C and basic feature B. It should be noted that the cross-genetic feature is also a type of derivative feature, and each cross-genetic feature includes at least two basic features. Further, by repeatedly performing the above steps on the first target derivative features in the first set and the second set, many cross-genetic features can be obtained. At this time, only by examining the advantages and disadvantages of each of the many cross-genetic features, the target genetic features can be screened out as meaningful derivative features provided to data miners.

[0053] Optionally, screening out the target genetic features from multiple cross-genetic features can also be done in a way that depends on fitness. As an optional embodiment, to screen out the target genetic features from multiple cross-genetic features, the following steps can be taken: divide the multiple cross-genetic features into a third set and a fourth set; generate second feature verification instructions corresponding one-to-one to the multiple cross-genetic features; based on the second execution results of the second feature verification instructions for the target database and the fitness function, determine the second fitness corresponding to each cross-genetic feature, where the fitness function is used to calculate the fitness of the cross-genetic feature according to the second execution result corresponding to each cross-genetic feature and the set to which each cross-genetic feature belongs; based on the second fitness, screen out the target genetic features from the cross-genetic features.

[0054] Optionally, the above screening process can be the same as the process of screening out the first target derivative features from the initial derivative features, so the same fitness function can also be used. The second feature verification instructions can be execution SQL (Structured Query Language, abbreviated as SQL) statements generated corresponding to the descriptions of the basic features in the cross-genetic features. The target database tries to execute this SQL statement to obtain a feedback result, that is, whether the cross-genetic feature corresponding to the second feature verification instruction is meaningful for the data in the target database. Then, further determine the advantages and disadvantages of each cross-genetic feature according to the fitness function, and screen out the target genetic features based on the advantages and disadvantages.

[0055] As an alternative embodiment, screening for target genetic features from cross-genetic features may further include the following steps: obtaining a mutation factor, where the mutation factor is used to indicate the replacement ratio for replacing the basic features included in the cross-genetic features; according to the mutation factor, replacing the basic features in each cross-genetic feature to obtain mutated genetic features; and screening for target genetic features from the mutated genetic features.

[0056] In this alternative embodiment, by mutating the cross-genetic features to a certain extent, some derivative features that are not generated during the entire feature derivation process can be created, further expanding the range of feature types included in all derivative features. Specifically, the mutation factor can be used to indicate what proportion of the basic features in each cross-genetic feature are to be replaced, or it can also indicate the probability of replacing each basic feature in the cross-genetic feature. For example, if the mutation factor is 0.7, then 70% of the basic features in a cross-genetic feature can be replaced, or it can also be in a way that each basic feature in the cross-genetic feature has a 70% probability of being replaced. Optionally, when mutating the cross-genetic features, the basic features to be replaced are not randomly replaced by any basic features. The newly inserted basic features can have a certain association with the replaced basic features. For example, the newly inserted basic features and the replaced basic features can be of the same type of basic features but with different values. For example, "the last 3 months", "the last 6 months", and "the last 12 months" can be defined as a group of basic features of a time value type. If a basic feature in a cross-genetic feature needs to be replaced during the mutation process and this basic feature comes from this group of basic features, the newly inserted group of basic features will also only come from this group of basic features and not from other groups of basic features. The above actions can, to a certain extent, maintain the stability of the entire cross-genetic feature and prevent some meaningless mutations that would otherwise reduce the generation efficiency of derivative features.

[0057] Figure 3 is a flowchart of feature derivation using a genetic method according to an alternative embodiment of the present invention, as Figure 3 shown, and the process may include the following steps:

[0058] Step S1, collect basic features and put them into the feature library.

[0059] Step S2: In a single-granularity random inflation manner, perform single-granularity feature random inflation combination processing based on existing basic features. For example, the basic features include the most recent 6 months, the most recent 12 months, account amount, overdue days, maximum historical overdue days, account balance, valid period of collateral, collateral evaluation date, collateral audit date, collateral evaluation amount, etc. Using the most recent 6 months as the single-particle expansion benchmark, perform random inflation combination to generate derivatives such as "account amount in the most recent 6 months", "account balance in the most recent 6 months", "collateral evaluation amount within the most recent 6 months", "collateral audit date within the most recent 6 months", etc. The derivatives generated by this part of the combination are the initial derivatives. Among them, there will be some illogical derivatives, such as "the most recent 6 months and the most recent 12 months". Such features will be screened and eliminated in the subsequent genetic algorithm. Although they are not deleted here, it ensures obtaining the largest feature sample space.

[0060] Step S3: Initialize the gene population. Set corresponding assertions for all existing population samples through the execution feedback of SQL statements by the target database. If an SQL error occurs during execution, set the assertion of this population sample to False, indicating execution failure; if the execution passes, set the assertion to True, indicating execution success. Among them, the initialized gene population can be the first set and the second set described in the above embodiments, that is, divide the initial derivatives into two populations, and the initial derivatives in the population can also be called population samples.

[0061] Step S4: Calculate the fitness of each sample in the population. The fitness function is as follows:

[0062]

[0063] Among them, f(x) represents the value of fitness, i represents the serial number of the initial derivative, n represents the total number of all initial derivatives, k is a constant (for example, it can be set to 0.1), μ is a hyperparameter (for example, it can be set to 0.25), represents the proportion of the SQL statements corresponding to the initial derivatives that are successfully executed in the target database among all initial derivatives, and represents the execution passing rate of the SQL statements corresponding to all initial derivatives in the target database.

[0064] Step S5: By screening, cross - inheritance, and mutation of the samples in the population, the genes of the population samples with higher fitness in the population are retained, and the genetic samples with lower fitness are eliminated. Here, the population samples are the derived features, and the genes are the basic features in each derived feature. Since feature derivation has high requirements for multi - dimensional features, a relatively high crossover rate and mutation rate can be adopted. For example, the crossover rate can be set to 0.9, and the mutation rate can be set to 0.7. Setting the crossover rate to 0.9 means that 90% of the basic features between two cross - inherited features are exchanged. Through cross - inheritance and mutation, it can be ensured that the genetic process of feature derivation generates as many feature dimensions as possible, ensuring the diversity of feature derivation dimensions.

[0065] Step S6: After steps S1 to S5, one round of genetic - algorithm - based feature derivation is completed. Then, it can jump to step S4 for the next round of feature derivation. After multiple rounds of derivation, when the total number of existing features in the derived features is basically unchanged, the entire process of feature derivation using the genetic algorithm is completed, and multi - dimensional derived features that can be output to data miners are obtained.

[0066] According to an embodiment of the present invention, there is also provided a feature derivation device for implementing the above - mentioned feature derivation method. Figure 4 It is a structural block diagram of the feature derivation device provided according to an embodiment of the present invention, as Figure 4 shown. The feature derivation device includes: a construction module 42, a generation module 44, and a screening module 46. The following is an explanation of the feature derivation device.

[0067] The construction module 42 is used to construct initial derived features based on multiple basic features. Here, the basic features are used to describe the data stored in the target database, and each initial derived feature includes at least two basic features.

[0068] The generation module 44 is connected to the above - mentioned construction module 42 and is used to generate first feature verification instructions corresponding one - to - one to the initial derived features.

[0069] The screening module 46 is connected to the above - mentioned generation module 44 and is used to screen out first - target derived features from the initial derived features based on the first execution result of the target database for the first feature verification instructions.

[0070] It should be noted here that the above - mentioned construction module 42, generation module 44, and screening module 46 correspond to steps S202 to S206 in the embodiment. The examples and application scenarios implemented by the multiple modules and the corresponding steps are the same, but are not limited to the content disclosed in the above - mentioned embodiment. It should be noted that the above - mentioned modules, as part of the device, can run in the computer terminal 10 provided in the embodiment.

[0071] Embodiments of the present invention may provide a computer device. Optionally, in this embodiment, the above computer device may be at least one of multiple network devices in a computer network. The computer device includes a memory and a processor.

[0072] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the feature derivation method and device in the embodiments of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, to implement the above-mentioned feature derivation method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely set relative to the processor, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.

[0073] The processor can call the information and application programs stored in the memory through a transmission device to execute the following steps: based on multiple basic features, construct initial derived features, where the basic features are used to describe the data stored in the target database, and each initial derived feature includes at least two basic features; generate first feature verification instructions corresponding one-to-one to the initial derived features; based on the first execution result of the target database for the first feature verification instructions, screen out the first target derived features from the initial derived features.

[0074] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing the relevant hardware of the terminal device through a program, and this program can be stored in a non-volatile storage medium. The storage medium may include: flash drive, read-only memory (ROM), random access memory (RAM), magnetic disk, optical disk, etc.

[0075] Embodiments of the present invention also provide a non-volatile storage medium. Optionally, in this embodiment, the above non-volatile storage medium can be used to save the program code executed by the feature derivation method provided in the above embodiment.

[0076] Optionally, in this embodiment, the above non-volatile storage medium may be in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.

[0077] Optionally, in this embodiment, the non-volatile storage medium is configured to store program code for performing the following steps: constructing initial derivative features based on a plurality of basic features, where the basic features are used to describe the data stored in the target database, and each initial derivative feature includes at least two basic features; generating first feature verification instructions corresponding one-to-one to the initial derivative features; and screening out first target derivative features from the initial derivative features based on the first execution results of the target database for the first feature verification instructions.

[0078] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0079] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0080] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of units can be a logical function division, and in actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in an electrical or other form.

[0081] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0082] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0083] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0084] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A feature derivation method, characterized in that, it includes: Based on multiple basic features, construct initial derived features, where the basic features are used to describe the data stored in the target database, each basic feature is used to reflect the information of one dimension of the data in the target database, and each initial derived feature includes at least two of the basic features; Generate first feature verification instructions corresponding one-to-one to the initial derived features, where the first feature verification instructions are used to verify whether each initial derived feature is valid; Based on the first execution result of the target database for the first feature verification instructions, screen out the first target derived features from the initial derived features; Wherein, before screening out the first target derived features from the initial derived features based on the first execution result of the target database for the first feature verification instructions, the method further includes: sending the first feature verification instructions to the target database, where the first feature verification instructions include database query statements corresponding one-to-one to the initial derived features; receiving the first execution result feedback by the target database, where the first execution result is used to indicate whether the target database can successfully execute the database query statements; Wherein, screening out the first target derived features from the initial derived features based on the first execution result of the target database for the first feature verification instructions includes: dividing the initial derived features into a first set and a second set; based on the first execution result and a fitness function, determine the first fitness corresponding to each initial derived feature, where the fitness function is used to calculate the fitness of each initial derived feature according to the first execution result corresponding to each initial derived feature and the set to which each initial derived feature belongs; based on the first fitness, screen out the first target derived features from the initial derived features; Wherein, the method further includes: obtaining a second target derived feature and a third target derived feature included in the first target derived feature, where the second target derived feature belongs to the first set and the third target derived feature belongs to the second set; exchanging the basic features included in the second target derived feature and the third target derived feature respectively to obtain a plurality of cross-genetic features; screening out target genetic features from the plurality of cross-genetic features; Among them, the fitness function adopts the functional form: It is constructed, where f(x) represents the value of fitness, i represents the serial number of the initial derivative features, n represents the number of all initial derivative features, k is a constant, and μ is a hyperparameter. It represents the proportion of the SQL statements corresponding to the initial derivative features that are successfully executed in the target database among all the initial derivative features. It represents the execution pass rate of the SQL statements corresponding to all the initial derivative features in the target database.

2. The method according to claim 1, characterized in that, screening out the first target derived features from the initial derived features based on the first fitness includes: Randomly generate a fitness threshold within a predetermined numerical range; Compare the magnitude of the first fitness with the fitness threshold, and determine the initial derived feature corresponding to the first fitness greater than the fitness threshold as the first target derived feature.

3. The method according to claim 1, characterized in that, screening out target genetic features from the plurality of cross-genetic features includes: Dividing the plurality of cross-genetic features into a third set and a fourth set; Generate second feature verification instructions corresponding one by one to the multiple cross-genetic features; Based on the second execution result of the second feature verification instruction for the target database and the fitness function, determine the second fitness corresponding to each cross-genetic feature, where the fitness function is used to calculate the fitness of each cross-genetic feature according to the second execution result corresponding to each cross-genetic feature and the set to which each cross-genetic feature belongs; Based on the second fitness, screen out the target genetic features from the cross-genetic features.

4. The method according to claim 1, wherein, The screening out the target genetic features from the cross-genetic features includes: Obtain a mutation factor, where the mutation factor is used to indicate the replacement ratio for replacing the basic features included in the cross-genetic features; According to the mutation factor, replace the basic features in each cross-genetic feature to obtain mutated genetic features; Screen out the target genetic features from the mutated genetic features.

5. A feature derivation device, wherein, comprising: A construction module, configured to construct initial derived features based on multiple basic features, where the basic features are used to describe the data stored in the target database, each basic feature is used to reflect the information of one dimension of the data in the target database, and each initial derived feature includes at least two of the basic features; A generation module, configured to generate first feature verification instructions corresponding one by one to the initial derived features, where the first feature verification instructions are used to verify whether each initial derived feature is valid; A screening module, configured to screen out first target derived features from the initial derived features based on the first execution result of the first feature verification instruction for the target database; wherein, before screening out the first target derived features from the initial derived features based on the first execution result of the first feature verification instruction for the target database, the device is further configured to: send the first feature verification instruction to the target database, where the first feature verification instruction includes database query statements corresponding one by one to the initial derived features; receive the first execution result feedback by the target database, where the first execution result is used to indicate whether the target database can successfully execute the database query statements; wherein, the screening module is further configured to divide the initial derived features into a first set and a second set; based on the first execution result and the fitness function, determine the first fitness corresponding to each initial derived feature, where the fitness function is used to calculate the fitness of each initial derived feature according to the first execution result corresponding to each initial derived feature and the set to which each initial derived feature belongs; based on the first fitness, screen out the first target derived features from the initial derived features; Wherein, the device is further configured to obtain a second target derivative feature and a third target derivative feature included in the first target derivative feature, wherein the second target derivative feature belongs to the first set, and the third target derivative feature belongs to the second set; exchange the basic features included in the second target derivative feature and the third target derivative feature respectively to obtain a plurality of cross-genetic features; and screen out target genetic features from the plurality of cross-genetic features; Among them, the fitness function adopts the functional form: It is constructed, where f(x) represents the value of fitness, i represents the serial number of the initial derivative feature, n represents the number of all initial derivative features, k is a constant, and μ is a hyperparameter. It represents the proportion of the SQL statements corresponding to the initial derivative features that are successfully executed in the target database among all the initial derivative features. It represents the execution pass rate of the SQL statements corresponding to all the initial derivative features in the target database.

6. A non-volatile storage medium, characterized in that, the non-volatile storage medium includes a stored program, wherein when the program runs, it controls the device where the non-volatile storage medium is located to execute the feature derivation method according to any one of claims 1 to 4.

7. An electronic device, characterized in that, it includes a processor, and the processor is used to run a program, wherein when the program runs, it executes the feature derivation method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Feature processing method and system for machine learning

    CN110781978A

  • Feature derivation method and device

    CN113297185A