Feature preprocessing method and device, storage medium and electronic equipment

By mapping the original feature distribution to a uniform distribution method, the stability and prediction ability problems caused by data distribution differences in model training are solved, and a more efficient neural network training effect is achieved.

CN120011724AInactive Publication Date: 2025-05-16ANT ZHIXIN HANGZHOU INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510505926.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has the problem of data distribution differences in the processing of numerical features in model training, which affects the stability and prediction ability of the model.

Method used

By obtaining the sorted original feature distribution, the ordered quantile intervals connected by the head and tail are determined, and the original feature value is mapped to the target feature value, so that the feature distribution is mapped to the approximate uniform distribution.

Benefits of technology

It improves the stability and generalization capabilities of the neural network, solves the problems of outliers and information loss, and achieves better model training effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011724A_ABST
    Figure CN120011724A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a feature preprocessing method and device, a storage medium and electronic equipment, and the method comprises the steps: firstly obtaining sorted original feature distribution, determining a plurality of ordered end-to-end quantile intervals according to the original feature distribution and the number of the quantile intervals, then, for each original feature value in the plurality of original feature values, obtaining a plurality of original feature values; according to the original feature value, a starting quantile and a terminating quantile of a target quantile interval in which the original feature value falls, the quantity of the quantiles, and a sorting sequence number of the starting quantile or the terminating quantile in a plurality of quantiles corresponding to the plurality of quantile intervals, determining the number of the quantiles in the target quantile interval; the original feature values are mapped into corresponding target feature values, and the multiple target feature values obtained through mapping are used for training a neural network model, so that the original feature distribution can be mapped to approximate uniform distribution, and the stability and generalization ability of the neural network can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to computer technology, and in particular to a method, device, storage medium and electronic equipment for feature preprocessing. Background Art

[0002] In actual application scenarios, the original data used for some model training include numerical features, and these original data often come from different scenarios or calibers. The data distribution of different columns in these data may vary greatly. For example, some columns are normally distributed, while some columns are skewedly distributed. This will seriously affect the stability and predictive ability of the trained model. Summary of the invention

[0003] The purpose of the embodiments of this specification is to provide a method, device, storage medium and electronic device for feature preprocessing.

[0004] The embodiment of this specification provides a method for feature preprocessing, which can map the original feature distribution to an approximate uniform distribution, thereby helping to improve the stability and generalization ability of the neural network. The method includes: Acquire a sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values; According to the original feature distribution and the number of quantile intervals, a plurality of quantile intervals are determined in an ordered manner and connected end to end, wherein the number of original feature values ​​falling into each quantile interval is the same; For each of the multiple original feature values, the original feature value is mapped to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the mapped multiple target feature values ​​are used to train the neural network model.

[0005] Furthermore, determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals includes: According to the original feature distribution and the number of quantile intervals, a plurality of original feature value sets that do not intersect with each other are determined from the original feature distribution in a corresponding order, wherein each original feature value set includes the same number of original feature data; According to the multiple original feature value sets, multiple quantile intervals that are ordered and connected end to end are determined.

[0006] Furthermore, determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals includes: According to the original feature distribution and the number of quantile intervals, a plurality of ordered quantile points are determined from the original feature distribution, wherein the number of original feature values ​​between every two adjacent quantile points is the same; According to the multiple quantile points, a plurality of ordered end-to-end connected quantile intervals are determined.

[0007] Further, for each original feature value of the multiple original feature values, mapping the original feature value to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting sequence number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals includes: For each of the multiple original feature values, the original feature value is mapped to the corresponding target feature value according to the ratio of a first difference between the original feature value and the starting quantile of the target quantile interval in which the original feature value falls and a second difference between the ending quantile of the target quantile interval and the starting quantile, the number of quantile intervals, and the sorting order of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals.

[0008] Furthermore, the method further comprises: The number of quantile intervals is determined according to the data volume of the original feature distribution.

[0009] Furthermore, determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals includes: determining whether the original feature distribution is uniform with respect to the number of quantile intervals; If so, a plurality of ordered end-to-end connected quantile intervals are determined according to the original feature distribution and the number of quantile intervals.

[0010] Furthermore, the step of determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals further includes: If the original feature distribution is not uniform relative to the number of quantile intervals, adjusting the original feature distribution so that the adjusted original feature distribution is uniform relative to the number of quantile intervals; According to the adjusted original feature distribution and the number of quantile intervals, a plurality of quantile intervals connected end to end in order are determined.

[0011] Further, the adjusting the original feature distribution includes: The original feature distribution is adjusted by removing at least one original feature value in the original feature distribution.

[0012] The embodiment of this specification also provides a device for feature preprocessing, including: An acquisition module, used to acquire the sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values; A determination module, configured to determine a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​falling into each quantile interval is the same; A mapping module is used to map each original feature value among the multiple original feature values ​​to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting sequence number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the multiple target feature values ​​obtained by mapping are used to train the neural network model.

[0013] The embodiments of the present specification also provide a storage medium, wherein the storage medium stores a computer program, and the computer program is suitable for being loaded by a processor and executing the steps of the above method.

[0014] An embodiment of the present specification also provides an electronic device, comprising: a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the above method.

[0015] The embodiments of the present specification also provide a computer program product having at least one instruction stored thereon, wherein the at least one instruction implements the steps of the above method when executed by a processor.

[0016] According to the scheme of the embodiments of the present specification, by first obtaining the sorted original feature distribution, and determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals, and then for each original feature value among the plurality of original feature values, according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantiles, and the sorting sequence number of the starting quantile or the ending quantile in the plurality of quantiles corresponding to the plurality of quantile intervals, the original feature value is mapped to the corresponding target feature value, and feature data for training the neural network model can be obtained, thereby enabling the original feature distribution to be mapped to an approximate uniform distribution, which helps to improve the stability and generalization ability of the neural network. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1A flowchart of a method for feature preprocessing provided in an embodiment of this specification; Figure 2 A schematic diagram of a structure of a device for feature preprocessing provided in an embodiment of this specification; Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of this specification more clear, the technical solutions of this specification will be clearly and completely described below in combination with the specific embodiments of this specification and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of this specification, not all of them. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0019] See also Figure 1 , is a flow chart of a method for feature preprocessing provided in an embodiment of this specification. In an embodiment of this specification, the method for feature preprocessing is applied to a device for feature preprocessing (hereinafter referred to as "preprocessing device") or an electronic device equipped with a preprocessing device described in an embodiment of this specification. Figure 1 The process shown is described in detail, and the method for feature preprocessing may specifically include the following steps: S102, obtaining a sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values.

[0020] In some embodiments, a plurality of original feature values ​​are first obtained, and then the plurality of original feature values ​​are sorted from small to large to obtain the sorted original feature distribution. For example, 1000 original feature values ​​are obtained, and the 1000 original feature values ​​are sorted from small to large to obtain the sorted original feature distribution, which is the sorting result for the 1000 original feature values. In some embodiments, the sorted original feature distribution provided by other devices is obtained. In some embodiments, the plurality of original feature values ​​are actual feature data generated in the target scene, and the plurality of original feature values ​​may correspond to a plurality of data sources and / or a plurality of sub-scenes in the target scene. In some embodiments, the plurality of original feature values ​​correspond to a plurality of distributions, for example, a plurality of original feature values ​​include data in different columns, wherein the data in some columns are normally distributed, and the data in another part of the columns are skewed distributed.

[0021] S104, determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​falling into each quantile interval is the same.

[0022] In some embodiments, the number of quantile intervals is preset, for example, the number of quantile intervals is preset to 10, then according to the original feature distribution and the preset number of quantile intervals, 10 quantile intervals corresponding to the original feature distribution and connected to the first order are determined. In some embodiments, ordered means that there is a size order between multiple quantile intervals, for example, the values ​​in the first quantile interval are all smaller than the values ​​in the second quantile interval, the values ​​in the second quantile interval are all smaller than the values ​​in the third quantile interval, and so on, and each quantile interval is ordered. In some embodiments, the first quantile interval is continuous, that is, the next quantile interval starts from the end point of the previous quantile interval, but the next quantile interval does not include the end point of the previous quantile interval (that is, it is ensured that an original feature value will only fall into one quantile interval). For example, the first quantile interval is [0,100] (that is, the minimum original feature value of the first quantile interval is 0 and the maximum original feature value is 100), and the second quantile interval is (100,1000] (that is, the second quantile interval starts from the end point 100 of the first quantile interval, but does not include 100). It should be noted that in some embodiments, the next quantile interval can also be expressed as including the previous quantile interval. The form of the end point of a partition, for example, the first quantile interval is expressed as [0,100], and the second quantile interval is expressed as [100,1000]. In this case, the value 100 can be regarded as falling into the first quantile interval, that is, as falling into the smaller quantile interval. In some embodiments, the number of original feature values ​​falling into each quantile interval can be determined based on the total number of original feature values ​​and the number of quantile intervals. For example, the total number of original feature values ​​is 1000, and the number of quantile intervals is 10, then there are 100 original feature values ​​in each quantile interval. It should be noted that since the number of original feature values ​​falling into each quantile interval is the same, the sample size between each quantile interval is uniform.

[0023] S106, for each original feature value among the multiple original feature values, map the original feature value to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting sequence number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the mapped multiple target feature values ​​are used to train the neural network model.

[0024] In some embodiments, the starting quantile of the target quantile interval is also the starting point of the target quantile interval, and the ending quantile of the target quantile interval is also the end point of the target quantile interval. For example, the first quantile interval is [0, 100], and the original characteristic value 10 falls into the first quantile interval. Then the first quantile interval is also the target quantile interval where the characteristic value 10 falls. The starting quantile of the target quantile interval is 0, and the ending quantile is 100. In some embodiments, the ending quantile of each quantile interval is used as the quantile corresponding to the quantile interval; in some embodiments, the starting quantile of the first quantile interval is used as the first quantile; as an example, the sorting number of the starting quantile of the first quantile interval is recorded as 0, the sorting number of the ending quantile of the first quantile interval is recorded as 1, and the sorting number of the ending quantile of the second quantile interval is recorded as 2, and so on. The sorting number of the ending quantile of the nth quantile interval is recorded as n. In some embodiments, the sorting sequence number of the starting quantile or the ending quantile among the multiple quantiles corresponding to the multiple quantile intervals, that is, the order of the target quantile interval among the multiple quantile intervals, for example, the sorting sequence number is 2, which indicates that the target quantile interval is the second quantile interval. In some embodiments, the target feature value mapped to the original feature value falls into the interval [0,1], that is, through mapping, each original feature value is mapped to a value within the interval. Since the number of original feature values ​​falling into each quantile interval is the same, the sample size of each quantile interval is uniform, but since the data distribution of each quantile interval may not be uniform, when there are enough quantiles, the target distribution (that is, the distribution of each target feature value) is approximately uniform. In some embodiments, considering that the more quantiles there are, the closer the target distribution is to a uniform distribution, and further considering that the more quantiles there are, the greater the amount of calculation, the appropriate number of quantile intervals can be set based on the data magnitude of the original feature values ​​or based on experience; in some embodiments, based on the actual situation of the target distribution, the number of quantile intervals can be adjusted and feature preprocessing can be re-performed. For example, if the initial number of quantile intervals is 10, but the target distribution is not ideal, the number of quantile intervals can be increased to make the target distribution more similar to a uniform distribution, or a balance between the target distribution effect and the amount of calculation can be achieved by adjusting the number of quantile intervals.

[0025] In some embodiments, each target feature value obtained after preprocessing is used to train a neural network model. For example, for a risk control scenario, the original feature values ​​in the scenario can be preprocessed based on the solution of the embodiment of this specification so that each target feature value obtained after preprocessing can be used to train the risk control model. This specification does not limit the specific application scenario of the target feature value. For example, it can also be used in the training of other models such as credit models where the training data is numerical values.

[0026] According to the scheme of the embodiments of the present specification, the sorted original feature distribution is first obtained, and a plurality of ordered end-to-end connected quantile intervals are determined according to the original feature distribution and the number of quantile intervals. Then, for each original feature value among the plurality of original feature values, the original feature value is mapped to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantiles, and the sorting sequence number of the starting quantile or the ending quantile in the plurality of quantiles corresponding to the plurality of quantile intervals. The mapped plurality of target feature values ​​are used to train the neural network model, thereby enabling the original feature distribution to be mapped to an approximate uniform distribution, which helps to improve the stability and generalization ability of the neural network. In addition, the data dimension remains unchanged before and after preprocessing in the scheme of the embodiments of the present specification, so the unchanged data brings about an additional increase in computing cost.

[0027] The present application has found that the existing technology may be sensitive to outliers and have information loss problems in the standardization of input data in the training of neural network models. For example, the existing MinMaxScaler (a data preprocessing solution) solution scales the eigenvalues ​​to a given range through linear transformation. However, since MinMaxScaler uses the minimum and maximum values ​​of the original data for scaling, if there are outliers in the data, the scaling range of the entire data set will be affected, making the outliers less obvious, which may have a negative impact on certain machine learning algorithms. For another example, the existing Quantile Transformation (QT) solution can achieve data conversion by mapping the quantiles of the original data to the quantiles of a new uniform distribution. Although this method is robust to outliers, it also has some problems that may cause information loss. The values ​​within a quantile interval will lose relativity after the transformation. The scheme of the embodiments of the present specification can be regarded as a quantile linear encoding scheme, which can map the original data distribution to an approximate uniform distribution. Since the mapping process is based on the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting sequence of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, the preprocessing scheme of the present specification is not sensitive to outliers. When processing data containing outliers, it affects the distribution of a single quantile interval at most. Therefore, the outlier sensitivity problem of the MinMaxScaler scheme can be solved. In addition, the preprocessing scheme of the present specification realizes a one-to-one mapping between the original distribution and the target distribution, and maintains the relative relationship between the values. Therefore, compared with the existing quantile transformation scheme, this scheme does not have the problem of information loss. Based on business data, the applicant compared the solution of the embodiments of this specification with various existing preprocessing solutions (including the above-mentioned MinMaxScaler and QT solutions). Based on the comparison results, the solution of this specification ranks best in evaluation indicators (such as model loss value, AUC (Area Under the Curve) value, the difference in cumulative distribution between good and bad samples, etc.).

[0028] In some embodiments, the step S104 includes: according to the original feature distribution and the number of quantile intervals, determining multiple sets of original feature values ​​that do not intersect each other from the original feature distribution in a corresponding order, wherein each set of original feature values ​​includes the same number of original feature data; according to the multiple sets of original feature values, determining multiple quantile intervals that are connected end to end in an ordered manner. In some embodiments, the corresponding order is to sort the original feature values ​​in the original feature distribution from small to large. As an example, there are 1000 original feature values ​​in the original feature distribution obtained by sorting from small to large, and the number of quantile intervals is 10, then one quantile interval should contain 100 (1000 / 10=100) original feature values, then the 1st to 100th original feature values ​​can be taken from the original feature distribution according to the sorting to form the first original feature value set, and the 101st to 200th original feature values ​​can be taken to form the second original feature value set, and so on until the last 100 original feature values ​​are taken to form the tenth original feature value set, then each set contains 100 original feature data, and each set is ordered and has no intersection. In some embodiments, the maximum value in each set is also the end point of the corresponding quantile interval, and is the starting point of the next quantile interval (if any) (where the starting point of the first quantile interval is the minimum value in the first set).

[0029] In some embodiments, the step S104 includes: determining a plurality of ordered quantiles from the original feature distribution according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​between every two adjacent quantiles is the same; and determining a plurality of ordered end-to-end quantile intervals according to the plurality of quantiles. In some embodiments, according to the original feature distribution and the number of quantile intervals, the sorting position of the original feature value corresponding to each quantile in the original feature distribution can be determined. For example, there are 1000 original feature values ​​in the original feature distribution obtained by sorting from small to large, and the number of quantile intervals is 10, then one quantile interval should contain 100 (1000 / 10=100) original feature values, thereby determining that the first quantile is the original feature value ranked 100th in the original feature distribution (the original feature value ranked first is the starting point of the first quantile interval and can also be regarded as the 0th quantile), the second quantile is the original feature value ranked 200th in the original feature distribution, and so on, until all quantiles are determined. In some embodiments, based on the determined multiple quantiles, multiple quantile intervals that are ordered and connected end to end can be determined. Continuing with the above example, if the smallest value in the ranking is 0, the first quantile (that is, the original feature value ranked at the 100th place) is 1000, then the first quantile interval is [0, 1000], and if the second quantile (that is, the original feature value ranked at the 200th place) is 15000, then based on the first quantile and the second quantile, the second quantile interval can be determined to be (1000, 15000], and so on, 10 quantile intervals that are ordered and connected end to end can be determined.

[0030] In some embodiments, step S106 includes: for each of the multiple original feature values, according to the ratio of the first difference between the original feature value and the starting quantile of the target quantile interval in which the original feature value falls and the second difference between the ending quantile of the target quantile interval and the starting quantile, the number of quantile intervals, and the sorting sequence number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, the original feature value is mapped to the corresponding target feature value. As an example, the target feature value to which the original feature value is mapped can be calculated based on the following formula: Among them, x represents the original characteristic value, and the i-th quantile interval that x falls into is recorded as [b i , b i+1 ], represents the target feature value after x mapping, n is the number of quantile intervals, based on the above formula, the original feature distribution in [b i , b i+1] interval, and linearly maps it to [i / n, (i + 1) / n].

[0031] In some embodiments, it further includes: determining the number of quantile intervals according to the amount of data of the original feature distribution. In some embodiments, a mapping relationship between the amount of data of different magnitudes and the number of quantile intervals is pre-established, and the number of quantile intervals can be determined based on the mapping relationship and the amount of data of the original feature distribution, wherein the larger the data magnitude, the larger the number of quantile intervals, for example, the number of quantile intervals corresponding to the amount of data below 1 million is 100, and the number of quantile intervals corresponding to the amount of data above 1 million is 1000.

[0032] In some embodiments, step S104 includes: determining whether the original feature distribution is uniform relative to the number of quantile intervals; if so, determining a plurality of ordered end-to-end connected quantile intervals based on the original feature distribution and the number of quantile intervals. In some embodiments, determining whether the original feature distribution is uniform relative to the number of quantile intervals means determining whether the number of quantile intervals can satisfy the requirement that the number of original feature values ​​in each quantile interval is the same. If so, each quantile interval can be directly determined. For example, 1000 original feature data need to be divided into 10 ordered quantile intervals, which can be satisfied. Therefore, each quantile interval can be directly determined based on the 1000 original feature data.

[0033] In some embodiments, the step S104 further includes: if the original feature distribution is not uniform relative to the number of quantile intervals, adjusting the original feature distribution so that the adjusted original feature distribution is uniform relative to the number of quantile intervals; determining a plurality of ordered end-to-end quantile intervals according to the adjusted original feature distribution and the number of quantile intervals. In some embodiments, if the number of original feature values ​​in each quantile interval cannot be the same based on the current original feature distribution and the number of quantile intervals, the original feature distribution can be adjusted accordingly so that the adjusted original feature distribution is uniform relative to the number of quantile intervals. In some embodiments, the original feature number of each quantile interval after adjustment is the integer value of the ratio between the total amount of adjusted original feature values ​​and the number of quantile intervals. The implementation method of determining a plurality of ordered end-to-end quantile intervals according to the adjusted original feature distribution and the number of quantile intervals is the same or similar to the implementation method of determining a plurality of ordered end-to-end quantile intervals according to the original feature distribution and the number of quantile intervals, and will not be repeated here. It should be noted that any implementation method for adjusting the original feature distribution so that it is uniform relative to the number of quantile intervals should be included in the protection scope of this specification.

[0034] In some embodiments, the adjusting the original feature distribution includes: adjusting the original feature distribution by removing at least one original feature value in the original feature distribution. In some embodiments, the original feature distribution can be adjusted by removing the smallest original feature value in the original feature distribution so that the number of original feature values ​​in each quantile interval is the same.

[0035] Figure 2 A schematic diagram of a structure of a device for feature preprocessing provided in an embodiment of this specification, the device for feature preprocessing (hereinafter referred to as "preprocessing device 1") can be implemented as all or part of an electronic device through software, hardware or a combination of both. According to some embodiments, the preprocessing device 1 includes an acquisition module 11, a determination module 12 and a mapping module 13.

[0036] An acquisition module 11 is used to acquire the sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values; A determination module 12 is used to determine a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​falling into each quantile interval is the same; A mapping module 13 is used to map each original feature value among the multiple original feature values ​​to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the multiple target feature values ​​obtained by mapping are used to train the neural network model.

[0037] In some embodiments, the determination module 12 is used to: determine, according to the original feature distribution and the number of quantile intervals, from the original feature distribution in a corresponding order, a plurality of original feature value sets that do not overlap with each other, wherein each original feature value set includes the same number of original feature data; and determine, according to the plurality of original feature value sets, a plurality of quantile intervals that are connected end to end in an ordered manner.

[0038] In some embodiments, the determination module 12 is used to: determine a plurality of ordered quantile points from the original feature distribution according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​between every two adjacent quantile points is the same; and determine a plurality of ordered end-to-end connected quantile intervals according to the plurality of quantile points.

[0039] In some embodiments, the mapping module 13 is used to: for each of the multiple original feature values, map the original feature value to a corresponding target feature value based on the ratio of a first difference between the original feature value and the starting quantile of the target quantile interval in which the original feature value falls and a second difference between the ending quantile of the target quantile interval and the starting quantile, the number of quantile intervals, and a sorting number of the starting quantile or the ending quantile among the multiple quantiles corresponding to the multiple quantile intervals.

[0040] In some embodiments, the preprocessing device 1 is further used to determine the number of quantile intervals according to the data volume of the original feature distribution.

[0041] In some embodiments, the determination module 12 is used to: determine whether the original feature distribution is uniform relative to the number of quantile intervals; if so, determine a plurality of ordered end-to-end connected quantile intervals based on the original feature distribution and the number of quantile intervals.

[0042] In some embodiments, the determination module 12 is also used to: if the original feature distribution is not uniform relative to the number of quantile intervals, adjust the original feature distribution so that the adjusted original feature distribution is uniform relative to the number of quantile intervals; and determine a plurality of ordered end-to-end connected quantile intervals based on the adjusted original feature distribution and the number of quantile intervals.

[0043] In some embodiments, the adjusting the original feature distribution includes: adjusting the original feature distribution by removing at least one original feature value in the original feature distribution.

[0044] The above device embodiments correspond to the method embodiments. For specific descriptions, please refer to the description of the method embodiments, which will not be repeated here. The device embodiments are obtained based on the corresponding method embodiments and have the same technical effects as the corresponding method embodiments. For specific descriptions, please refer to the corresponding method embodiments.

[0045] The embodiment of the present specification also provides a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor to execute the method of the embodiment of the present specification.

[0046] The embodiments of the present specification also provide a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor to execute the method of the embodiments of the present specification.

[0047] The embodiments of this specification also provide Figure 3 The structural diagram of the electronic device shown in FIG. Figure 3At the hardware level, the electronic device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory, and may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above method.

[0048] Of course, in addition to software implementation, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is to say, the executor of the following processing flow is not limited to each logic unit, but can also be hardware or logic devices.

[0049] In the 1990s, it was very clear whether the improvement of a technology was hardware improvement (for example, improvement of the circuit structure of diodes, transistors, switches, etc.) or software improvement (improvement of the method flow). However, with the development of technology, many improvements of the method flow today can be regarded as direct improvements of the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that the improvement of a method flow cannot be implemented with hardware entity modules. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to ask chip manufacturers to design and make dedicated integrated circuit chips. Moreover, nowadays, instead of manually making integrated circuit chips, this kind of programming is mostly implemented by "logic compiler" software, which is similar to the software compiler used when developing and writing programs, and the original code before compilation must also be written in a specific programming language, which is called hardware description language (HDL). There is not only one kind of HDL, but many kinds, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also know that it is only necessary to program the method flow slightly in the above-mentioned hardware description languages ​​and program it into the integrated circuit, and then it is easy to obtain the hardware circuit that implements the logic method flow.

[0050] The controller may be implemented in any suitable manner, for example, the controller may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (e.g., software or firmware) executable by the (micro)processor, a logic gate, a switch, an application-specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, examples of which include but are not limited to the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320, and the memory controller may also be implemented as part of the control logic of the memory. It is also known to those skilled in the art that in addition to implementing the controller in a purely computer-readable program code manner, the controller may be implemented in the form of a logic gate, a switch, an application-specific integrated circuit, a programmable logic controller, and an embedded microcontroller by logically programming the method processing steps. Therefore, such a controller may be considered as a hardware component, and the means for implementing various functions included therein may also be considered as a structure within the hardware component. Or even, the means for implementing various functions may be considered as both a software module for implementing the method and a structure within the hardware component.

[0051] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.

[0052] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0053] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems, or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0054] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0055] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0056] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational processing steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The processing steps of the functions specified in one or more boxes.

[0057] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0058] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0059] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0060] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0061] It should be understood by those skilled in the art that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0063] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0064] The above description is only an embodiment of the present specification and is not intended to limit the present specification. For those skilled in the art, the present specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present specification shall be included in the scope of the claims of the present specification.

Claims

1. A method for feature preprocessing, comprising: Acquire a sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values; According to the original feature distribution and the number of quantile intervals, a plurality of quantile intervals are determined in an ordered manner and connected end to end, wherein the number of original feature values ​​falling into each quantile interval is the same; For each of the multiple original feature values, the original feature value is mapped to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the mapped multiple target feature values ​​are used to train the neural network model.

2. The method according to claim 1, wherein determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals comprises: According to the original feature distribution and the number of quantile intervals, a plurality of original feature value sets that do not intersect with each other are determined from the original feature distribution in a corresponding order, wherein each original feature value set includes the same number of original feature data; According to the multiple original feature value sets, multiple quantile intervals that are ordered and connected end to end are determined.

3. The method according to claim 1, wherein determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals comprises: According to the original feature distribution and the number of quantile intervals, a plurality of ordered quantile points are determined from the original feature distribution, wherein the number of original feature values ​​between every two adjacent quantile points is the same; According to the multiple quantile points, a plurality of ordered end-to-end connected quantile intervals are determined.

4. The method according to claim 1, wherein for each original feature value in the plurality of original feature values, according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantile intervals, and the sorting sequence number of the starting quantile or the ending quantile in the plurality of quantiles corresponding to the plurality of quantile intervals, the original feature value is mapped to the corresponding target feature value, comprising: For each of the multiple original feature values, the original feature value is mapped to the corresponding target feature value according to the ratio of a first difference between the original feature value and the starting quantile of the target quantile interval in which the original feature value falls and a second difference between the ending quantile of the target quantile interval and the starting quantile, the number of quantile intervals, and the sorting order of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals.

5. The method according to claim 1, further comprising: The number of quantile intervals is determined according to the data volume of the original feature distribution.

6. The method according to claim 1, wherein determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals comprises: determining whether the original feature distribution is uniform with respect to the number of quantile intervals; If so, a plurality of ordered end-to-end connected quantile intervals are determined according to the original feature distribution and the number of quantile intervals.

7. The method according to claim 6, wherein determining a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals further comprises: If the original feature distribution is not uniform relative to the number of quantile intervals, adjusting the original feature distribution so that the adjusted original feature distribution is uniform relative to the number of quantile intervals; According to the adjusted original feature distribution and the number of quantile intervals, a plurality of quantile intervals connected end to end in order are determined.

8. The method according to claim 7, wherein adjusting the original feature distribution comprises: The original feature distribution is adjusted by removing at least one original feature value in the original feature distribution.

9. A device for feature preprocessing, comprising: An acquisition module, used to acquire the sorted original feature distribution, wherein the original feature distribution includes a plurality of original feature values; A determination module, configured to determine a plurality of ordered end-to-end connected quantile intervals according to the original feature distribution and the number of quantile intervals, wherein the number of original feature values ​​falling into each quantile interval is the same; A mapping module is used to map each original feature value among the multiple original feature values ​​to a corresponding target feature value according to the original feature value, the starting quantile and the ending quantile of the target quantile interval in which the original feature value falls, the number of quantiles, and the sorting sequence number of the starting quantile or the ending quantile in the multiple quantiles corresponding to the multiple quantile intervals, wherein the multiple target feature values ​​obtained by mapping are used to train the neural network model.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. An electronic device, characterized in that: include: A processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the steps of the method as claimed in any one of claims 1 to 8.

12. A computer program product having at least one instruction stored thereon, characterized in that: When the at least one instruction is executed by the processor, the steps of the method described in any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data processing method and apparatus

    CN106599899A