A data processing method, apparatus, device, and storage medium
The data processing method standardizes preprocessing for multiple AI algorithms by storing transformed feature data and scripts, addressing inefficiencies in data processing and enabling rapid generation of tailored input data.
Patent Information
- Application Number
- CN201911082031.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-11-07
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2039-11-07
AI Technical Summary
The lack of a general data preprocessing scheme in the prior art leads to different requirements for feature data input by different AI algorithms, resulting in slow data preprocessing operation speed and low efficiency.
A data processing method is provided, by performing data processing operations in the setting process on the original feature data, obtaining feature data to be converted, and storing it according to the set data structure. At the same time, scripts for different conversion operations are stored for at least two algorithms, so as to quickly obtain input feature data that meets the personalized requirements of each algorithm.
It realizes the standardization of the data preprocessing process, improves the speed and efficiency of data preprocessing, meets the personalized requirements of different algorithms, and shortens the algorithm development cycle.
Smart Images

Figure CN112784149B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technologies, and in particular, to a data processing method, apparatus, device, and storage medium. Background Art
[0002] Today, in the face of information overload and a plethora of complex product content, people have diverse needs for various reasons. As a platform tool, a recommendation system analyzes users' historical behaviors and uses corresponding algorithms to recommend products that meet their needs. Among them, the currently popular AI (Artificial Intelligence) algorithms in recommendation systems, such as LR (Logistic Regression), GBDT (Gradient Boosting Decision Tree), FM (Factorization Machine), or NN (Neural Network), etc., have strict requirements for their respective input feature data. Before applying these AI algorithms, corresponding data processing, that is, data preprocessing, needs to be performed on the original feature data. For example, a computer cannot directly recognize text, and feature data in text form needs to be converted into feature data in digital vector form. Different AI algorithms require different formats of input feature data. Common data formats include LIBSVM format, DMATRIX format, or TENSOR format. Therefore, before inputting the feature data into the algorithm, the original feature data needs to be format-converted. In particular, for LR, FM, and NN, normalization processing of the original feature data is also required to obtain the input feature data, otherwise, the weight gradients corresponding to some features will disappear or explode. In some business scenarios, although GBDT and XGBOOST do not require complex preprocessing of the original feature data, conventional data processing operations such as removing invalid data, oversampling / undersampling, etc. are still necessary.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] Because different AI algorithms have different requirements for the input of feature data (or feature vectors), data preprocessing has become an indispensable part of the entire business system that uses AI algorithms. However, there is currently a lack of a general data preprocessing solution, resulting in slow data preprocessing operations and low efficiency. Summary of the Invention
[0005] Embodiments of the present invention provide a data processing method, apparatus, device and storage medium to quickly obtain input feature data that meets the personalized requirements of various AI algorithms.
[0006] In a first aspect, an embodiment of the present invention provides a data processing method, the method comprising:
[0007] Perform data processing operations of a set process on the original feature data according to the business scenario to obtain feature data to be converted;
[0008] The feature data to be converted is stored according to a set data structure, and a script for performing different conversion operations on the feature data to be converted is stored for each of at least two algorithms, so as to obtain input feature data required by the corresponding algorithm by executing the script;
[0009] The data processing operation of the set flow is a preprocessing operation that is required to be performed on the feature data of at least two algorithms.
[0010] In a second aspect, an embodiment of the present invention further provides a data processing device, the device comprising:
[0011] A processing module is used to perform a data processing operation of a set process on the original feature data according to the business scenario to obtain the feature data to be converted;
[0012] A storage module, used to store the feature data to be converted according to a set data structure, and to store scripts for performing different conversion operations on the feature data to be converted for each of at least two algorithms, so as to obtain input feature data required by the corresponding algorithm by executing the scripts;
[0013] The data processing operation of the set flow is a preprocessing operation that is required to be performed on the feature data of at least two algorithms.
[0014] In a third aspect, an embodiment of the present invention further provides a device, the device comprising:
[0015] one or more processors;
[0016] A memory for storing one or more programs;
[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the steps of the data processing method provided by any embodiment of the present invention.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of a data processing method as provided in any embodiment of the present invention.
[0019] The embodiments in the above invention have the following advantages or beneficial effects:
[0020] By performing data processing operations on the original feature data according to the business scenario to obtain the feature data to be converted; storing the feature data to be converted according to the set data structure, and respectively storing scripts for performing different conversion operations on the feature data to be converted for each of at least two algorithms, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts;
[0021] Among them, the data processing operation of the set process is a technical means for the preprocessing operations commonly required by the feature data of at least two algorithms, which realizes the purpose of standardizing the data preprocessing process, solves the technical problems of slow speed and low efficiency of data preprocessing operations, and achieves the beneficial effect of improving the speed and efficiency of data preprocessing. Through the data processing method provided by the embodiments of the present invention, the input feature data that meets the personalized requirements of each algorithm can be quickly obtained. Description of the Drawings
[0022] Figure 1 is a flowchart of a data processing method provided by Embodiment 1 of the present invention;
[0023] Figure 2 is a schematic diagram of the storage structure of a single feature data provided by Embodiment 1 of the present invention;
[0024] Figure 3 is a flowchart of a data processing method provided by Embodiment 2 of the present invention;
[0025] Figure 4 is provided by Embodiment 2 of the present invention and Figure 3 is a flowchart of another data processing method corresponding thereto;
[0026] Figure 5 is a schematic diagram of the structure of a data processing device provided by Embodiment 3 of the present invention;
[0027] Figure 6 is a schematic diagram of the structure of a device provided by Embodiment 4 of the present invention. Detailed Description of the Invention
[0028] The present invention will be further described in detail below with reference to the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. In addition, it should be noted that for the sake of description, only parts related to the present invention are shown in the drawings, rather than all the structures.
[0029] Embodiment 1
[0030] Figure 1 This is a flowchart of a data processing method provided in Embodiment 1 of the present invention. This method can be executed by a data processing device, which can be implemented in software and / or hardware, and is usually integrated into a terminal, such as a computer. As Figure 1 shown, the method specifically includes the following steps:
[0031] Step 110: Perform data processing operations on the original feature data according to the business scenario to obtain the to-be-converted feature data.
[0032] Among them, the original feature data refers to the feature data that has not undergone any processing. For example, the original feature data representing the age feature may be 20, 60, or 80, etc.; the original feature data representing the color feature may be feature data in text form such as red, green, or blue; the original feature data representing the gender feature may be male or female. The data processing operations of the set process refer to the data processing operations that the original feature data usually undergoes in the process of realizing its data value. For example, it can be understood that a computer cannot recognize text-form data. Therefore, if one wants to perform certain learning or analysis on feature data through a computer, first, the text-form feature data needs to be converted into digital-form feature data that can be recognized by the computer, that is, perform text feature digitization operations on the text-form feature data to convert the text-form feature data into digital-form feature data that the computer can recognize. Therefore, if the original feature data includes text-form feature data, the data processing operations of the set process usually include text feature digitization operations. On the other hand, before inputting the feature data into various algorithm models, usually, filtering operations are also performed on the original feature data to eliminate the invalid values in the original feature data and avoid interference of the invalid values on the performance of the algorithm model. The invalid values can be specifically determined according to the physical meaning of the feature. For example, for the age feature of a person, if there is age feature data with a value of 300 years old, it can be determined that 300 is the invalid value in the age feature data because the lifespan of humans generally does not reach 300 years.
[0033] Exemplarily, the data processing operations of the set process include at least one of the following operations: text feature digitization operations, filtering operations, data type classification operations, and performing label encoding operations on the character-type feature data among them, data type classification operations, and determining the outlier feature data of the numerical type among them, and performing strong assignment operations on the outlier feature data, and performing assignment operations on the missing feature values.
[0034] Among them, the data type classification operation specifically classifies the feature data into numerical feature data and character feature data. The label encoding operation on the character-type feature data specifically digitizes the feature data in character form so that the computer can recognize it. Commonly used label encoding operations include one-hot encoding, stringindexer encoding, and various embedding methods. The outlier feature data specifically refers to data that differs significantly from most of the feature data. For example, for the age feature data, there are 10 in total, which are 20, 25, 26, 30, 29, 22, 24, 23, 27, 65. Then the value 65 can be considered as the outlier feature data under this feature data. Outlier feature data is usually deleted as an invalid value. However, in some cases where it is impossible to determine whether the outlier feature data is a valid value or an invalid value, it is usually chosen to strongly assign a value to the outlier feature data in a certain way to weaken the negative impact brought by the outlier feature data. Commonly used methods for determining outlier feature data include: the 3Sigma principle method of normal distribution, the box plot determination method, and the standard score method Z-score. Commonly used methods for strongly assigning values to outlier feature data include: the median or mean method, the difference method, and the similar sample method.
[0035] Step 120: Store the to-be-converted feature data according to the set data structure, and store scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms respectively, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts, where the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms.
[0036] Among them, the conversion operation includes at least one of the following: normalization operation, regularization operation, exponentiation operation, and logarithm operation. The feature data to be converted is the feature data obtained by performing one or several basic processing operations on the original feature data. For example, a filtering operation is performed on the original feature data, and the feature data obtained after filtering is determined as the feature data to be converted. Or, a text feature digitization operation and a filtering operation are performed on the original feature data, and the obtained feature data is determined as the feature data to be converted. The data processing operation of the setting process is a preprocessing operation that the feature data of at least two algorithms commonly need to perform. For example, for LR, GBDT, and NN in the AI algorithm, before inputting the feature data into the algorithm, a text feature digitization operation and a filtering operation need to be performed on the original feature data. Then, through the data processing method provided by the embodiments of the present invention, a text feature digitization operation and a filtering operation are pre-performed on the original feature data to obtain the feature data to be converted, and the feature data to be converted is stored. At the same time, scripts for personalized conversion operations that need to be performed based on the feature data to be converted before inputting the feature data into each algorithm are respectively stored. For example, simply, before inputting the feature data into the algorithm LR, if a normalization operation needs to be performed on the feature data to be converted; before inputting the feature data into the algorithm GBDT, if an exponentiation operation needs to be performed on the feature data to be converted; before inputting the feature data into the algorithm NN, if a regularization operation needs to be performed on the feature data to be converted; then when storing the feature data to be converted, scripts for performing a normalization operation, an exponentiation operation, and a regularization operation on the feature data to be converted are respectively stored. When it is necessary to obtain the feature data input to the algorithm LR, by executing the pre-stored script for performing a normalization operation on the pre-stored feature data to be converted, the feature data that meets the input requirements of the algorithm LR can be quickly obtained. When it is necessary to obtain the feature data input to the algorithm GBDT, by executing the pre-stored script for performing a logarithm operation on the pre-stored feature data to be converted, the feature data that meets the input requirements of the algorithm GBDT can be quickly obtained. When it is necessary to obtain the feature data input to the algorithm NN, by executing the pre-stored script for performing a regularization operation on the pre-stored feature data to be converted, the feature data that meets the input requirements of the algorithm NN can be quickly obtained. It should be noted that the above examples are used to explain the technical solutions of the embodiments of the present invention, rather than limiting the technical solutions of the embodiments of the present invention.
[0037] By storing the to-be-converted feature data obtained after basic data processing, targeted conversion operations can be performed on the to-be-converted feature data according to the personalized requirements of different algorithms, meeting the different requirements of different algorithms for input feature data. Moreover, the purpose of obtaining the input feature data in real time according to personalized requirements can be achieved without the need to pre-generate a set of input feature data that meets specific requirements for each algorithm, thus achieving the purpose of saving storage space. At the same time, algorithm developers can flexibly and quickly select pre-stored scripts for performing different conversion operations on the to-be-converted feature data according to development needs, and obtain the personalized input feature data required by the corresponding algorithm by executing the scripts.
[0038] Further, storing the to-be-converted feature data according to a set data structure and respectively storing scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms includes:
[0039] Storing the to-be-converted feature data based on a data structure of the Map type;
[0040] Storing the scripts based on a data structure of the ArrayList.
[0041] Specifically, the to-be-converted feature data is stored in a data structure of the Map type. In the data structure of the Map type, the key value key corresponds to the feature name of each feature element. For example, the feature name of the age feature element is age, and the feature name of the gender feature element is gender, etc. The numerical value value represents a class named transformer. The class named transformer contains two attribute values. One is the above-mentioned to-be-converted feature data, and the other is a data structure of the ArrayList. Each element in the data structure of the ArrayList is an executable script excute. An executable script excute contains a method for performing a conversion operation on the to-be-converted feature data and the corresponding input parameters of the method. Since the length of the ArrayList can change, the ArrayList can contain multiple executable scripts excute and can execute multiple feature conversion scripts. Therefore, algorithm developers can conveniently select different conversion methods based on the structural flexibility of the ArrayList to obtain various personalized feature data. For details, please refer to Figure 2Schematic diagram of a storage structure for a single feature data. In the data structure Map, the corresponding key value key represents the name name of the current feature, such as feature names like age, gender, etc., and the corresponding numerical value value represents a class named transformer. The class 210 named transformer includes two attribute values, namely the feature data to be transformed of the current feature and an executable script Array List for performing transformation operations on the feature data to be transformed of the current feature <excute>, the executable script excute 220 further includes: 1. A specific feature transformation method; 2. The feature data to be transformed for executing the feature transformation method is the feature data to be transformed stored in the attribute of the transformer class; 3. The parameters required for the feature transformation method. Each time specific feature data is read, a transformation operation is triggered. Algorithm developers can freely and flexibly select one of multiple transformation scripts to perform transformation operations with different purposes on the feature data to be transformed, so as to obtain feature data that meets different requirements. Under the Map type data structure, for each feature (such as age, gender, category of purchased goods, consumption level, or positive review rate, etc.), not only the feature data to be transformed (i.e., the data that can be recognized by the algorithm for this feature) is stored, but also as many representation forms as possible (i.e., different feature transformation methods) are stored. These representation forms are stored in the data structure of each feature in the form of executable scripts before being selected, and they only occupy very little storage space. Only when a certain representation form is selected, the feature data to be transformed of the current feature is transformed according to the executable script corresponding to this representation form, and the required feature representation form is obtained and output.
[0042] In the data processing method provided by the embodiment of the present invention, by storing the feature data to be transformed (i.e., the data that has undergone basic processing such as text feature digitization operation and / or filtering operation and can be recognized by a computer), and at the same time storing at least two scripts for performing personalized transformation operations on the feature data to be transformed, the technical means is realized that enables algorithm developers to freely and flexibly select one of multiple scripts to perform transformation operations with different purposes on the feature data to be transformed, improves the data preprocessing speed, shortens the algorithm development cycle, and at the same time has good scalability, that is, algorithm developers can supplement new transformation scripts at any time. It should be noted that in the data processing method provided in this embodiment, the feature data to be transformed obtained after the original feature data undergoes some basic data processing operations required by multiple algorithms is stored, rather than the feature data obtained after the transformation operation is performed. This is because the transformation operations required by different algorithms are different. If the corresponding transformation operation is performed on the original feature data according to the transformation requirements of each algorithm, and the transformed data obtained is stored, not only will storage space be wasted, but also some basic data processing operations required by multiple algorithms will be repeatedly executed, and it is impossible to exhaust them, and the purpose of meeting the requirements of various algorithms cannot be achieved.
[0043] The technical solution of this embodiment obtains the feature data to be converted by performing data processing operations required by multiple algorithms on the original feature data in advance, stores the feature data to be converted, and simultaneously stores personalized conversion scripts required by multiple algorithms, so as to quickly obtain input feature data that meets the personalized requirements of each algorithm based on the feature data to be converted, achieving the purpose of standardizing the data preprocessing process, solving the technical problems of slow speed and low efficiency of data preprocessing operations, and obtaining the beneficial effect of improving the speed and efficiency of data preprocessing. Through the data processing method provided by the embodiments of the present invention, input feature data that meets the personalized requirements of each algorithm can be quickly obtained.
[0044] Embodiment 2
[0045] Figure 3 The flowchart of a data processing method provided by Embodiment 2 of the present invention. Based on the above embodiment, this embodiment further illustrates the data processing method by taking the business scenario as an example of an e-commerce product push scenario. The explanations of the same or corresponding terms as those in the above embodiment will not be repeated here.
[0046] See Figure 3 , the data processing method provided by this embodiment specifically includes the following steps:
[0047] Step 310: Perform a filtering operation on the original feature data to obtain first feature data.
[0048] Specifically, performing a filtering operation on the original feature data includes:
[0049] Count the missing values of the target feature data included in the original feature data. If the percentage of the number of missing values in the total number of original feature data exceeds a set threshold, the target feature data is deleted. Among them, the missing values include at least one of the following: null value, invalid value, and NULL value, and the invalid value of the feature data is determined according to the physical meaning of the feature.
[0050] In the e-commerce product push scenario, the numerical value 0 is regarded as an invalid value. Therefore, when there is only one non-zero unique value of the feature data of a certain feature, it means that the effective information entropy corresponding to this feature is 0 and has no meaning. Therefore, this feature is deleted. The reason why the numerical value 0 is regarded as an invalid value is that in most algorithms, the final mathematical model can be approximately expressed as f(wTx), where x represents the feature vector and w represents the weight of the current feature. If x is 0, then after multiplying by the weight, it is still 0 and will not affect the output. Therefore, the numerical value 0 can be regarded as an invalid value.
[0051] Step 320: Perform a data type classification operation on the first feature data and a label encoding operation on the character-type feature data therein to obtain second feature data.
[0052] Specifically, uniformly set the data type of the numerical feature data to the float type and the data type of the character-type feature data to the String type. Common label encoding operations include one-hot encoding, stringindexer encoding, and various embedding encodings.
[0053] Step 330: Calculate the mean and variance of the numerical feature data in the first feature data, and determine the mean and variance as the third feature data.
[0054] Among them, for continuous numerical feature data, its mean and variance are basic numerical statistical results, which are convenient for algorithm development engineers to understand the correctness of the feature data.
[0055] Step 340: Determine the outlier feature data in the numerical feature data in the first feature data, and perform a strong assignment operation on the outlier feature data to obtain fourth feature data.
[0056] Determine the outlier feature data in the numerical feature data based on the box plot, and perform a strong assignment operation on the outlier feature data using the upper limit value and / or lower limit value of the box plot. Specifically, if the value of the outlier feature data is greater than the upper limit value of the box plot, then strongly assign the value of the outlier feature data to the upper limit value of the box plot; if the value of the outlier feature data is less than the lower limit value of the box plot, then strongly assign the value of the outlier feature data to the lower limit value of the box plot.
[0057] Step 350: Determine the second feature data, the third feature data, and the fourth feature data as the feature data to be converted, store the feature data to be converted, and store scripts for performing different conversion operations on the feature data to be converted for each of at least two algorithms respectively, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts; wherein, the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms.
[0058] Furthermore, for the fourth feature data, according to different business scenarios, algorithm development engineers can determine a reasonable normalization method based on business experience and store the script corresponding to the determined normalization method, so as to implement the normalization operation on the feature data to be converted by executing the script. Usually, the 2-norm calculation method can be used to normalize the numerical feature data.
[0059] See Figure 4 The flow schematic diagram of another data processing method corresponding to Figure 3 is as follows:
[0060] Step 410: Count the missing values, non-zero unique values, and the total number of the original feature data for each feature in the original feature data, and eliminate the features whose missing value ratio reaches the threshold. Figure 4 Taking the threshold as 80% as an example, eliminate the features with only one non-zero unique value.
[0061] Step 420: Standardize the data types of each feature in the original feature data, uniformly set the data type of the numerical feature data to the float type, uniformly set the data type of the character feature data to the String type, and perform label encoding operations on the character feature data among them.
[0062] Step 430: Without considering the feature missing values, calculate the mean and variance of each feature data.
[0063] Step 440: Based on the box plot, determine the outlier feature data in the numerical feature data, and strongly assign the outlier feature data to the upper limit or the lower limit of the box plot. Take the upper limit of the box plot as the maximum value of the feature data, and take the lower limit of the box plot as the minimum value of the feature data.
[0064] Step 450: Draw a histogram of the values taken by each numerical feature data, and determine the normalization method according to the value distribution of each numerical feature data. For the character feature data, merge its labels according to the Pearson Chi2 test method.
[0065] Step 460: Prepare one or more feature transformation schemes and corresponding feature transformation parameters for the non-missing values of each feature. For the missing values of each feature, fill 0.0 for the numerical features and fill 0 for the character features.
[0066] Step 470: Store the to-be-transformed feature data obtained above (including the feature data remaining after the filtering operation, the mean and variance of each feature, the feature data obtained after performing label encoding operations on the character feature data, the numerically feature data after strong assignment, etc.) and the feature transformation schemes prepared for the non-missing values of each feature.
[0067] The technical solution of this embodiment takes the e-commerce product push scenario as an example and gives a specific example of the data processing operation for the set process. The data processing operation of the set process given in this embodiment is only used to explain the present invention and does not limit the technical solution of the present invention. The data processing operation of the set process needs to be adaptively adjusted according to different business scenarios. By pre-executing the data processing operations commonly required by multiple algorithms on the original feature data, the to-be-converted feature data is obtained, and the to-be-converted feature data is stored. At the same time, personalized conversion scripts required by multiple algorithms are stored, so as to quickly obtain the input feature data that meets the personalized requirements of each algorithm based on the to-be-converted feature data through the personalized conversion scripts, achieving the purpose of standardizing the data preprocessing process, solving the technical problems of slow speed and low efficiency of data preprocessing operations, and obtaining the beneficial effect of improving the speed and efficiency of data preprocessing.
[0068] The following is an embodiment of the data processing device provided by the embodiment of the present invention. This device and the data processing method of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the data processing device, reference may be made to the embodiment of the data processing method.
[0069] Embodiment III
[0070] Figure 5 FIG. 10 is a schematic structural diagram of a data processing device provided by Embodiment III of the present invention. The device specifically includes: a processing module 510 and a storage module 520;
[0071] Among them, the processing module 510 is used to perform the data processing operation of the set process on the original feature data according to the business scenario to obtain the to-be-converted feature data; the storage module 520 is used to store the to-be-converted feature data according to the set data structure, and respectively store scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts; the data processing operation of the set process is a preprocessing operation commonly required by the feature data of at least two algorithms.
[0072] In an embodiment of the present invention, the data processing operation of the set process includes at least one of the following operations: text feature digitization operation, filtering operation, data type classification operation, and label encoding operation on character-type feature data among them, data type classification operation, and determining outlier feature data of numerical type among them, and performing a strong assignment operation on the outlier feature data, and performing an assignment operation on missing feature values.
[0073] In an embodiment of the present invention, when the data processing operation of the set process includes a filtering operation, the processing module 510 includes:
[0074] A statistical unit for counting the missing values of the target feature data included in the original feature data;
[0075] A deletion unit for deleting the target feature data if the percentage of the number of the missing values in the total number of the original feature data exceeds a set threshold;
[0076] Wherein, the missing values include at least one of the following: null values and invalid values, and the invalid values of the feature data are determined according to the physical meaning of the feature.
[0077] In an embodiment of the present invention, when the data processing operations of the set process include data type classification operations, determining outlier feature data of numerical types therein, and performing a strong assignment operation on the outlier feature data, the processing module 510 is specifically configured to:
[0078] Determine outlier feature data in the numerical feature data based on a box plot, and perform a strong assignment operation on the outlier feature data by using the upper limit value and / or the lower limit value of the box plot.
[0079] In an embodiment of the present invention, the business scenario includes an e-commerce product push scenario.
[0080] In an embodiment of the present invention, the storage module 520 is specifically configured to:
[0081] Store the to-be-converted feature data based on a data structure of the Map type;
[0082] Store the script based on a data structure of an Array List.
[0083] In an embodiment of the present invention, the conversion operations include at least one of the following: normalization operation, regularization operation, exponentiation operation, and logarithm operation.
[0084] The technical solution of this embodiment obtains the to-be-converted feature data by pre-executing data processing operations commonly required by multiple algorithms on the original feature data, stores the to-be-converted feature data, and stores personalized conversion scripts required by multiple algorithms at the same time, so as to quickly obtain input feature data that meets the personalized requirements of each algorithm based on the personalized conversion scripts and the to-be-converted feature data, realizes the purpose of standardizing the data preprocessing process, solves the technical problems of slow speed and low efficiency of data preprocessing operations, and achieves the beneficial effect of improving the speed and efficiency of data preprocessing. Through the data processing method provided by the embodiment of the present invention, input feature data that meets the personalized requirements of each algorithm can be quickly obtained.
[0085] The data processing device provided by the embodiments of the present invention can execute the data processing method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the data processing method.
[0086] Embodiment 4
[0087] Figure 6 It is a schematic structural diagram of a device provided by Embodiment 4 of the present invention. Figure 6 It shows a block diagram of an exemplary device 12 suitable for implementing the embodiments of the present invention. Figure 6 The shown device 12 is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present invention.
[0088] As Figure 6 shown, the device 12 is presented in the form of a general-purpose computing device. The components of the device 12 may include but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0089] The bus 18 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0090] The device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the device 12, including volatile and non-volatile media, removable and non-removable media.
[0091] The system memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be used to read and write non-removable, non-volatile magnetic media ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6 Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical medium) can be provided. In these cases, each drive can be connected to the bus 18 through one or more data medium interfaces. The system memory 28 may include at least one program product having a set of program modules (such as a processing module 410 and a storage module 420 in a data processing device), and these program modules are configured to execute the functions of various embodiments of the present invention.
[0092] A program / utility 40 having a set of program modules 42 (such as a processing module 410 and a storage module 420 in a data processing device) can be stored in, for example, the system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. The implementation of a network environment may be included in each or some combination of these examples. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present invention.
[0093] The device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the device 12, and / or communicate with any device that enables the device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 22. In addition, the device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0094] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the steps of a data processing method provided by an embodiment of the present invention. The method includes:
[0095] Performing a data processing operation on the original feature data according to a set process in a service scenario to obtain to-be-converted feature data;
[0096] Store the to-be-converted feature data according to the set data structure, and respectively store scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts;
[0097] Wherein, the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms.
[0098] Of course, those skilled in the art can understand that the processor can also implement the technical solutions of the data processing methods provided in any embodiment of the present invention.
[0099] Embodiment Five
[0100] Embodiment Five of the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the data processing method provided in any embodiment of the present invention. The method includes:
[0101] Perform the data processing operation of the set process on the original feature data according to the service scenario to obtain the to-be-converted feature data;
[0102] Store the to-be-converted feature data according to the set data structure, and respectively store scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts;
[0103] Wherein, the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms.
[0104] The computer storage medium of the embodiments of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0105] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0106] The program code contained on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0107] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0108] Those of ordinary skill in the art should understand that the various modules or steps of the present invention described above may be implemented using a general-purpose computing device. They may be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they may be implemented using program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they may be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them may be fabricated into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0109] Note that the above is only a preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.< / excute>
Claims
1. A data processing method, characterized in that, Including: Performing a data processing operation of a set process on the original feature data according to the business scenario to obtain the feature data to be converted; Storing the feature data to be converted according to the set data structure, and respectively storing scripts for performing different conversion operations on the feature data to be converted for each of at least two algorithms, so as to obtain the input feature data required by the corresponding algorithm by executing the scripts; the algorithms are artificial intelligence AI algorithms; Wherein, the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms; the input feature data is data that meets the algorithm input requirements; The performing a data processing operation of a set process on the original feature data according to the business scenario to obtain the feature data to be converted includes: Performing a filtering operation on the original feature data to obtain first feature data; Performing a data type classification operation on the first feature data and performing a label encoding operation on the character-type feature data therein to obtain second feature data; Calculating the average value and variance of the numerical feature data in the first feature data, and determining the average value and variance as third feature data; Determining the outlier feature data in the numerical feature data in the first feature data, and performing a strong assignment operation on the outlier feature data to obtain fourth feature data; Determining the second feature data, the third feature data, and the fourth feature data as the feature data to be converted.
2. The method according to claim 1, wherein The performing a filtering operation on the original feature data includes: Counting the missing values of the target feature data included in the original feature data; If the percentage of the number of the missing values in the total number of the original feature data exceeds a set threshold, deleting the target feature data; Wherein, the missing values include at least one of the following: null values and invalid values, and the invalid values of the feature data are determined according to the physical meaning of the feature.
3. The method according to claim 1, wherein The performing a strong assignment operation on the outlier feature data includes: Determining the outlier feature data in the numerical feature data based on the box plot, and performing a strong assignment operation on the outlier feature data by using the upper limit value and / or the lower limit value of the box plot.
4. The method according to any one of claims 1-3, characterized in that, The business scenario includes an e-commerce product push scenario.
5. The method according to any one of claims 1 to 3, characterized in that, The storing the feature data to be converted according to the set data structure, and respectively storing scripts for performing different conversion operations on the feature data to be converted for each of at least two algorithms includes: Storing the feature data to be converted based on a Map-type data structure; Storing the scripts based on an Array List data structure.
6. The method according to any one of claims 1 to 3, characterized in that, The conversion operations include at least one of the following: normalization operation, regularization operation, exponentiation operation, and logarithm operation.
7. A data processing device, characterized in that, Including: A processing module, configured to perform a data processing operation of a set process on the original feature data according to the business scenario to obtain the feature data to be converted; A storage module, configured to store the to-be-converted feature data according to a set data structure, and respectively store scripts for performing different conversion operations on the to-be-converted feature data for each of at least two algorithms, so as to obtain input feature data required by the corresponding algorithm by executing the scripts; the algorithms are artificial intelligence (AI) algorithms; Wherein, the data processing operation of the set process is a preprocessing operation commonly required for the feature data of at least two algorithms; the input feature data is data that meets the input requirements of the algorithms; The processing module is specifically configured to perform a filtering operation on the original feature data to obtain first feature data; Perform a data type classification operation on the first feature data and a label encoding operation on the character-type feature data therein to obtain second feature data; Calculate the average value and variance of the numerical feature data in the first feature data, and determine the average value and variance as third feature data; Determine the outlier feature data in the numerical feature data in the first feature data, and perform a strong assignment operation on the outlier feature data to obtain fourth feature data; Determine the second feature data, the third feature data, and the fourth feature data as the to-be-converted feature data.
8. A device, characterized in that, The device includes: One or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data processing method steps as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the data processing method steps as described in any one of claims 1-6.
Citation Information
Patent Citations
Data processing platform and data processing method
CN109343833A
Air traffic control flight path big data-based airlines operation situation law analysis method
CN110335507A