Data processing method and device, electronic equipment and computer readable storage medium
By aggregating the object data feature and training the integration tree model, the object features are automatically screened, and the problem of time-consuming and low accuracy in manual screening in the prior art is solved, and more efficient and accurate object behavior prediction is achieved.
Patent Information
- Application Number
- CN202410046863.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the object feature screening process relies on manual experience, which is time-consuming and difficult, resulting in low accuracy in object behavior prediction and lack of effective ways to improve.
By obtaining the object data set, performing feature aggregation and training the integration tree model, extracting the candidate aggregated feature subset, using the trained prediction model to determine the feature importance indicator, and automatically filtering the target feature subset for classification processing.
It improves the accuracy and automation of object behavior prediction, reduces feature screening time, avoids missed and missed selection, and improves the efficiency of feature engineering.
Smart Images

Figure CN120296572A_ABST
Abstract
Description
Technical Field
[0001] This application relates to data processing technologies, and in particular, to a data processing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] Currently, in the process of information recommendation of Internet applications, it is necessary to aggregate and filter a large number of object features. The aggregated features are used to characterize the historical data of the object, and the features corresponding to the user's historical data can be used to predict the future behavior of the object, providing a decision-making basis for subsequent operations. In related technologies, the screening process of object features is carried out based on manual experience, which has a high dependence on professionals. When the number of object features is huge, it is very time-consuming and difficult to complete manually, and it is easy to miss or misselect. The screening basis is too simple and violent, resulting in low accuracy of the predicted results.
[0003] In related technologies, there is no good way to improve the accuracy of object behavior prediction processing. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can efficiently screen object features and improve the accuracy of object behavior prediction.
[0005] The technical solution of the embodiments of this application is implemented as follows:
[0006] Embodiments of this application provide a data processing method, and the method includes:
[0007] Obtain an object data set and an initialized ensemble tree model, where the object data set includes historical data and actual type labels respectively corresponding to multiple objects, and each object corresponds to at least one piece of historical data;
[0008] Perform feature aggregation processing on each piece of historical data to obtain at least one aggregated feature respectively corresponding to each object, and combine each aggregated feature into an aggregated feature set;
[0009] Extract a candidate aggregated feature subset from the aggregated feature set;
[0010] Based on the candidate aggregated feature subset, call the initialized ensemble tree model for training processing to obtain a trained ensemble tree model and training process data, where the training process data is data generated during the classification process of the ensemble tree model for the candidate aggregated feature subset;
[0011] Train the initialized prediction model based on the training process data to obtain a trained prediction model, where the trained prediction model is used to determine the importance index of each of the aggregated features;
[0012] Call the trained prediction model based on the aggregated feature set for prediction processing to obtain the importance index of each of the aggregated features;
[0013] Determine a target aggregated feature subset based on the importance index of each of the aggregated features, where the target aggregated feature subset is used to perform classification processing on the multiple objects by calling the trained ensemble tree model.
[0014] An embodiment of the present application provides a data processing device, including:
[0015] A data acquisition module, configured to obtain an object data set and an initialized ensemble tree model, where the object data set includes historical data and actual type labels respectively corresponding to multiple objects, and each object corresponds to at least one piece of historical data;
[0016] A feature processing module, configured to perform feature aggregation processing on each piece of historical data to obtain at least one aggregated feature respectively corresponding to each object, and combine each of the aggregated features into an aggregated feature set; extract a candidate aggregated feature subset from the aggregated feature set;
[0017] A training processing module, configured to call the initialized ensemble tree model based on the candidate aggregated feature subset for training processing to obtain a trained ensemble tree model and training process data, where the training process data is data generated during the classification processing of the candidate aggregated feature subset by the ensemble tree model; train the initialized prediction model based on the training process data to obtain a trained prediction model, where the trained prediction model is used to determine the importance index of each of the aggregated features;
[0018] A prediction processing module, configured to call the trained prediction model based on the aggregated feature set for prediction processing to obtain the importance index of each of the aggregated features; determine a target aggregated feature subset based on the importance index of each of the aggregated features, where the target aggregated feature subset is used to perform classification processing on the multiple objects by calling the trained ensemble tree model.
[0019] An embodiment of the present application provides an electronic device for data processing, where the electronic device for data processing includes:
[0020] A memory, configured to store computer-executable instructions;
[0021] A processor, when executing computer-executable instructions or computer programs stored in the memory, implements the data processing method provided by an embodiment of the present application.
[0022] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the data processing method provided by an embodiment of the present application.
[0023] The embodiments of the present application have the following beneficial effects:
[0024] Aggregate historical data to obtain an aggregate feature set. Compared with the related art that relies on manual experience to screen high-complexity tasks, which is prone to missed selection or misselection, the generation and screening efficiency of object features is accelerated. Obtain a candidate aggregate feature subset from the aggregate feature set to train an initialized ensemble tree model, use the training process data to train a prediction model, and predict the trained prediction model based on the aggregate feature set, so as to obtain the importance index of each aggregate feature, and then determine the target aggregate feature subset according to the importance index, so that the trained ensemble tree model performs classification processing. During the training process of the ensemble tree model, the object behavior can be predicted more accurately, the classification accuracy of the ensemble tree model in the prediction task is maximized, the degree of automation is higher, and the prediction accuracy is higher. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is a schematic diagram of an application mode of the data processing method provided by an embodiment of the present application;
[0026] Figure 2 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application;
[0027] Figures 3A to 3E is a schematic flowchart of the data processing method provided by an embodiment of the present application;
[0028] Figure 3F is a schematic diagram of the structure of a prediction model provided by an embodiment of the present application;
[0029] Figure 3G is a schematic flowchart of the data processing method provided by an embodiment of the present application;
[0030] Figure 4 is an optional flowchart of the data processing method provided by an embodiment of the present application;
[0031] Figure 5 is a schematic diagram of the content input by the data processing method provided by an embodiment of the present application;
[0032] Figure 6A is a schematic flowchart of the process based on the training process of the ensemble tree model provided by an embodiment of the present application;
[0033] Figures 6B to 6C It is a schematic diagram of the principle provided by the embodiments of the present application;
[0034] Figure 6D It is a schematic diagram of the structure of the prediction model provided by the embodiments of the present application. Specific embodiments
[0035] In order to make the objectives, technical solutions, and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limitations on the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0036] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0037] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0038] It should be noted that the collection and processing of relevant data in the present application (for example: the online duration data of users, the social data of users, payment records) should strictly comply with the requirements of relevant national laws and regulations during actual application, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing behaviors within the scope of authorization of laws and regulations and the personal information subject.
[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other relevant parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0040] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0041] Before further elaborating on the embodiments of the present application, the nouns and terms involved in the embodiments of the present application are described. The nouns and terms involved in the embodiments of the present application are applicable to the following explanations.
[0042] 1) Extreme Gradient Boosting (XGBoost): It is a gradient boosting framework. Gradient boosting corrects the residuals of all existing weak learners by adding new weak learners. Finally, multiple learners are added together for the final prediction, and the accuracy is higher than that of a single learner. The extreme gradient boosting model introduces a regularization term in the loss function to control and reduce overfitting during the training process. It not only calculates the pseudo-residuals using the first derivative but also calculates the second derivative to approximately and quickly prune and construct new base learners. It is mainly used to solve classification and regression problems and is a powerful machine learning algorithm.
[0043] 2) Binary logistic regression function: Binary classification, also known as logistic regression, is a classification method that divides a set of samples into two different categories. Binary logistic regression analysis is applicable to studying data where the dependent variable is a binary variable, that is, a variable with only two possible outcomes. For example, the dependent variable is expressed in forms such as "yes" or "no", "agree" or "disagree", "occur" or "not occur".
[0044] 3) Manual analysis: Different mathematical methods are used to aggregate the different historical behavior information of users at different times to generate a large number of user aggregation features with different meanings. According to experience, the user aggregation features highly correlated with the historical behavior of the target user are manually selected for machine learning model analysis. For example: Aggregate the historical behavior information of game users, manually experience the game, and manually select the user aggregation features highly correlated with the historical behavior of the target user based on game experience.
[0045] 4) Automated Feature Engineering (Autofe): Existing automated feature engineering algorithms can automatically screen the features in the input table to obtain a set of excellent feature sets to replace the original features.
[0046] 5) Receiver Operating Characteristic Curve (ROC): It is used to show the prediction accuracy of the X-axis against the Y-axis. The receiver operating characteristic curve reflects the relationship between sensitivity and specificity. The abscissa X-axis is 1 - specificity, also known as the false positive rate (false alarm rate). The closer the X-axis is to zero, the higher the accuracy; the ordinate Y-axis is called sensitivity, also known as the true positive rate (sensitivity). The larger the Y-axis, the better the accuracy.
[0047] 6) Area Under Curve (AUC): The Receiver Operating Characteristic (ROC) curve divides the graph into two parts. The area under the curve is called AUC, which is used to represent the prediction accuracy. The higher the AUC value, that is, the larger the area under the curve, the higher the prediction accuracy. The closer the curve is to the upper left corner (the smaller X and the larger Y), the higher the prediction accuracy.
[0048] The embodiments of the present application provide a data processing method, apparatus, electronic device, and computer-readable storage medium, which can efficiently screen object features and improve the accuracy of object behavior prediction.
[0049] The following describes the exemplary applications of the electronic device provided by the embodiments of the present application. The electronic device provided by the embodiments of the present application can be implemented as various types of user terminals such as terminal devices, such as laptop computers, tablet computers, desktop computers, set-top boxes, smart TVs, mobile devices (e.g., mobile phones, portable music players, personal digital assistants, dedicated messaging devices, portable game devices), vehicle-mounted terminals, Virtual Reality (VR) devices, Augmented Reality (AR) devices, etc., or can be implemented as a server. Below, the exemplary applications when the electronic device is implemented as a server will be described.
[0050] Reference Figure 1 , Figure 1 is a schematic diagram of the application mode of the data processing method provided by the embodiments of the present application; for example, Figure 1 involves a server 200, a network 300, and a terminal device 400. The terminal device 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.
[0051] In some embodiments, the user is a game user, a game application program is installed in the terminal device 400, the server 200 is a server of a game platform, and object historical data is stored in the database 500.
[0052] For example, the user starts an application in the terminal device 400 to perform game-related operations. The terminal device 400 sends a user operation instruction to the server 200 via the network 300 to generate behavior data in the application. The server 200 obtains object historical data from the database 500 and invokes the data processing method provided in the embodiments of the present application to predict object behavior. The server 200 sends the recommendation information corresponding to the result of the object behavior prediction to the terminal device 400 via the network 300 to implement corresponding information recommendation for the user. The information recommendation behavior of the server 200 for the terminal device 400 can be for the game application or presented in other applications on the terminal device 400, such as an instant messaging client. Through the data processing method provided in the embodiments of the present application, it is possible to quickly and efficiently aggregate and filter object features, more accurately predict object behavior, maximize the classification accuracy of the integrated tree model in the prediction task, improve the accuracy of the prediction behavior, and save the time for feature processing.
[0053] In some embodiments, the data processing method of the embodiments of the present application can also be applied to the following application scenarios: for users using shopping software, obtain information such as the frequency of user purchases and the duration of online browsing of goods using the shopping software, and predict future product recommendation information for users by processing relevant user data; for users using video software, obtain information such as the login frequency and video viewing duration of users in the past period, and predict video recommendation information for users by processing relevant data of users using video software, and so on.
[0054] The embodiments of the present application can be implemented through database technology. A database, in short, can be regarded as a place for storing electronic files in an electronic filing cabinet. Users can perform operations such as adding, querying, updating, and deleting data in the files. The so-called "database" is a data set stored together in a certain way, shared by multiple users, having as little redundancy as possible, and independent of application programs.
[0055] A database management system (DBMS) is a computer software system designed to manage databases and generally has basic functions such as storage, retrieval, security, backup, etc. Database management systems can be classified according to the database models they support, such as relational, XML (Extensible Markup Language); or according to the types of computers they support, such as server clusters, mobile phones; or according to the query languages they use, such as Structured Query Language (SQL), XQuery; or according to the key performance metrics, such as maximum scale, highest running speed; or other classification methods. Regardless of the classification method used, some DBMSs can span categories, for example, supporting multiple query languages simultaneously.
[0056] Embodiments of this application can also be implemented through cloud technology. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the highly developed and applied Internet industry, as well as the promotion of demands such as search services, social networks, mobile commerce, and open collaboration, in the future, each item may have its own hash code identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data all require a powerful system back-end support, which can only be achieved through cloud computing.
[0057] Embodiments of this application can also be implemented through artificial intelligence. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.
[0058] In some embodiments, the server 200 may be an integration of multiple servers. For example, a game server, a feature engineering server, and a recommendation server. Among them, the game server is used to store the historical data of users, perform feature aggregation based on the historical data of users, and the obtained aggregated feature set is used for the training of the integrated tree model. The feature engineering server is used to store the data during the training process of the integrated tree model, train a prediction model based on the data during the training process of the integrated tree model, call the trained prediction model based on the aggregated feature set to perform prediction processing to obtain the importance index of the aggregated features, and infer the corresponding recommendation information based on the target aggregated feature subset determined according to the importance index. The recommendation server is used to send the recommendation information to the user's terminal device.
[0059] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The electronic device may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal device and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.
[0060] See Figure 2 , Figure 2 is a schematic structural diagram of the electronic device provided by the embodiments of the present application. The electronic device may be Figure 1 the server 200 in Figure 2 The server 200 shown in includes: at least one processor 410, a memory 450, and at least one network interface 420. Each component in the server 200 is coupled together through a bus system 440. It can be understood that the bus system 440 is used to implement the connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 440.
[0061] The processor 410 may be an integrated circuit chip with signal processing capabilities. For example: a general-purpose processor, a digital signal processor (DSP, Digital Signal Processor), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.
[0062] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard disk drives, optical disk drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 410.
[0063] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0064] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0065] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0066] A network communication module 452, used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB), etc.;
[0067] In some embodiments, the device provided in the embodiments of the present application can be implemented in software. Figure 2 The data processing device 455 stored in the memory 450 is shown, which can be software in the form of programs and plug-ins, including the following software modules: data acquisition module 4551, feature processing module 4552, training processing module 4553 and prediction processing module 4554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. Figure 2 For the sake of convenience, all the above modules are shown at once, but it should not be considered that the data processing device 455 excludes the implementation that can only include the feature processing module 4552, the training processing module 4553 and the prediction processing module 4554. The functions of each module will be explained below.
[0068] In some embodiments, a terminal or a server may implement the data processing method provided in the embodiments of the present application by running a computer program. For example, the computer program may be a native program or a software module in an operating system; it may be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run, such as a game APP; it may also be a small program, that is, a program that only needs to be downloaded to a browser environment to run; it may also be a small program that can be embedded in any APP. In short, the above computer program may be any form of application program, module or plug-in.
[0069] The exemplary applications and implementations of the server provided in the embodiments of the present application will be combined to illustrate the data processing method provided in the embodiments of the present application.
[0070] Next, the data processing method provided in the embodiments of the present application will be described. As described above, the electronic device for implementing the data processing method in the embodiments of the present application may be a terminal or a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.
[0071] See Figure 3A , Figure 3A is a schematic flowchart of the data processing method provided in the embodiments of the present application, which will be described in combination with Figure 3A the steps shown.
[0072] In step 301, an object data set and an initialized integrated tree model are obtained.
[0073] Exemplarily, the object data set includes historical data and actual type labels respectively corresponding to multiple objects, and each object corresponds to at least one piece of historical data. The object is a user or the user's account. The integrated tree model is XGBoost, a gradient boosting framework for prediction. The initialized integrated tree model at least has the functions of splitting and growing trees for features, predicting sample scores according to sample features, and adding up the scores of each tree to obtain the predicted value of the sample.
[0074] In the process of information recommendation, the prediction of future behaviors needs to refer to the historical data of multiple objects, and each object corresponds to at least one piece of historical data. For an object, the multiple historical data of the object include different types, and the historical data of each type corresponds to different basic features. The historical data may be historical behaviors extracted from hourly historical behavior logs. For example, in a game application program, the multiple historical data corresponding to the object are data of different types of behaviors such as the online duration, payment behavior, or social situation of the game user at different time periods in the past.
[0075] The prediction of future behavior needs to refer to the actual type labels of multiple objects, that is, for the target task to be predicted, the real object type labels are extracted from the user's performance logs. For example, the churn label of game users, which represents whether the user has not logged in to the game for a week. If the user meets the condition of not logging in to the game, the user is regarded as a churn user and the label of the user is set to 1. Otherwise, the label of the user is set to 0.
[0076] In step 302, feature aggregation processing is performed on each historical data to obtain at least one aggregated feature corresponding to each object, and each of the aggregated features is combined into an aggregated feature set.
[0077] Exemplarily, based on the object data set, feature extraction and aggregation processing are performed, and different aggregation methods are used to perform aggregation operations on the features to obtain an aggregated feature set composed of the aggregated features of multiple objects.
[0078] In some embodiments, refer to Figure 3B , Figure 3B which is a schematic flowchart of the data processing method provided by the embodiment of the present application. Figure 3A The step 302 shown can be implemented by Figure 3B steps 3021 to 3023 below, which will be specifically described.
[0079] In step 3021, feature extraction processing is performed on each historical data to obtain historical record features.
[0080] Exemplarily, features are extracted from the objects in each object data set to obtain the corresponding historical record features of each historical data. The historical record features are obtained by performing feature extraction processing on the historical data, and the feature extraction processing can be implemented through the feature extraction layer of the neural network model. The way of feature extraction processing can be: each attribute parameter in the historical data is used as a feature value of one dimension, converted into the form of a feature vector, and dimensionality reduction processing is performed on the feature vector to obtain historical record features.
[0081] In step 3022, based on the pre-configured aggregation type, aggregation processing is performed on each historical record feature to obtain at least one aggregated feature of each object.
[0082] Performing aggregation processing on the historical record features according to the pre-configured aggregation type can construct at least one corresponding aggregated feature for each object.
[0083] Exemplarily, the pre-configured aggregation type includes at least one of the following:
[0084] Type 1, aggregate the historical record features within the pre-configured duration.
[0085] Aggregate the historical record features within a preconfigured duration. It can be for a certain historical record feature, performing a multi-day aggregation operation. The multi-day aggregation is to aggregate the information of the historical record feature in the recent few days. The preconfigured duration is set according to the actual requirements of the application scenario.
[0086] Type 2: Aggregate the historical record features of the preconfigured number of times closest to the current moment.
[0087] It can also be to aggregate the historical record features of the preconfigured number of times closest to the current moment. For example, for a certain historical record feature, perform a time-series aggregation. For example, for the users of a game application, the time-series aggregation is to aggregate the recent login information of the users. It can also be to aggregate the recent consumption information of the users, etc. The preconfigured aggregation types of the historical record features can also include: the mathematical operations used in the aggregation process, such as max, min, sum, avg, max-min, or the length of the period to be aggregated, such as 1 day, 3 days, 7 days, 14 days, 21 days, etc.
[0088] In step 3023, combine at least one aggregation feature of each object to obtain an aggregation feature set.
[0089] Exemplarily, the combination process is to add each aggregation feature to the corresponding set to form an aggregation feature set. Based on the aggregation operations performed on each object in the previous step, a large number of aggregation features with different meanings can be constructed. The constructed aggregation features reflect the construction information of the features, and this construction information includes the preconfigured aggregation type and the user-based feature based on which the construction is performed, etc. Combine these object aggregation features to obtain a set of aggregation features.
[0090] In the embodiments of the present application, by performing feature extraction processing on each historical data and aggregating corresponding features for each object based on the preconfigured aggregation type, the efficiency of generating corresponding aggregation features based on object behavior is improved, the generation of a diverse set of aggregation features is realized, and a search space is provided for the subsequent screening of features.
[0091] Continue to refer to Figure 3A , in step 303, extract a candidate aggregation feature subset from the aggregation feature set.
[0092] Exemplarily, the number of aggregation features in the candidate aggregation feature subset is much smaller than that of the complete aggregation feature set. The extraction process can be achieved in the following way: randomly select aggregation features from the aggregation feature set, and the number of selected aggregation features is the target feature scale to obtain an initialized aggregation feature subset. The target feature scale can be set according to the actual requirements of the application scenario.
[0093] In some embodiments, refer toFigure 3C , Figure 3C is a schematic flowchart of the data processing method provided by an embodiment of the present application. Figure 3A The step 303 shown can be implemented by Figure 3C steps 3031 to 3032 below will be specifically described.
[0094] In step 3031, obtain the preconfigured number of features of the candidate aggregated feature subset.
[0095] Exemplarily, the preconfigured number of features is also the number of features to be selected during the training process. The number of features to be selected is preconfigured according to the target task, and the target feature scale is the expected number of aggregated features to be generated.
[0096] In step 3032, select the preconfigured number of aggregated features from the aggregated feature set to obtain a candidate aggregated feature subset.
[0097] Based on the obtained preconfigured number of features, select aggregated features from the aggregated feature set. The selected aggregated features can be stored in an aggregated feature table with a size of the target feature scale. Refer to Figure 5 the aggregated feature table 503 with the TS scale in. In the table, the id is the corresponding unique identification number, f1 and f2 can represent two basic features, label represents the user's label. For example, when a game user logs in every day within a week, the label corresponding to the user's churn label value is 1. The date in the table represents the statistical date of the label. For example, the statistics of this game user start from this date.
[0098] The representation of the aggregated feature can be in the form of a quadruple, which characterizes the aggregation information of the corresponding feature, that is, it can reflect the preconfigured aggregation type of the above features. For example, multi-day aggregation is to aggregate the corresponding information of the recent few days. For example: (f1, multi-day aggregation, sum, 2) represents the sum of the f1 features in the recent two days, (f2, multi-day aggregation, max, 3) represents the maximum value of the f2 features in the recent three days; for example, time-series aggregation is to aggregate the corresponding information of the recent several times. For example: (f2, time-series aggregation, avg, 2) represents the average value of the f2 features in the recent two times. For the aggregated features of each object, an aggregated feature table marked with labels can be constructed, and from the aggregated feature table with labels, filter out the preconfigured number of aggregated features. Refer to Figure 5 the aggregated feature table 503 with the TS scale in to obtain a candidate aggregated feature subset.
[0099] Continue to refer to Figure 3A , in step 304, based on the candidate aggregated feature subset, call the initialized ensemble tree model for training processing to obtain the trained ensemble tree model and training process data.
[0100] Exemplarily, the training process data is the data generated during the classification process of the integrated tree model for the candidate aggregated feature subset.
[0101] Screen and obtain a candidate aggregated feature subset from the aggregated feature set, perform training processing on the initialized integrated tree model based on the candidate aggregated subset, use the aggregated features in the aggregated feature subsets of some objects as the input of the integrated tree model, use the actual type label as the target output of the integrated tree model, and then minimize the value of the logistic regression loss function for binary classification to reduce the difference between the output value of the integrated tree model and the true label value, so as to implement the training process of the integrated tree model.
[0102] In some embodiments, refer to Figure 3D , Figure 3D is a schematic flowchart of the data processing method provided by the embodiments of the present application. Figure 3A The step 304 shown can be implemented by Figure 3D steps 3041 to 3044 below, and will be specifically described below.
[0103] In step 3041, classify each candidate aggregated feature by calling the initialized integrated tree model to obtain the predicted type of each candidate aggregated feature.
[0104] Exemplarily, the candidate aggregated feature is the aggregated feature in the candidate aggregated feature subset.
[0105] Screen and obtain a candidate aggregated feature subset from the aggregated feature set, use the aggregated features in the candidate aggregated feature subset as the input data of the integrated tree model. After receiving the input data, the integrated tree model classifies the aggregated features. The classification process can obtain a predicted type for each candidate aggregated feature, and this predicted type is used to compare with the true type value. The classification process can be implemented in the following way: use the embedding layer of the neural network model to process features of different categories, map the high-dimensional feature data to a low-dimensional space, construct an embedding layer for each category feature, splice the embedding layers, train the neural network model, and use the output of the trained embedding layer as the category feature for replacement and merge semantically similar items.
[0106] In step 3042, determine the first loss function of the initialized integrated tree model based on the difference between the predicted type and the true type of each candidate aggregated feature.
[0107] The integrated tree model receives the input data to classify the aggregated features, obtains the corresponding predicted type for each candidate aggregated feature, compares it with the true type value to obtain the difference between the two type values, and determines the first loss function of the initialized integrated tree model through the difference between the predicted type and the true type. The first loss function can be any one of cross-entropy loss, logarithmic loss, or information entropy loss.
[0108] Exemplarily, the first loss function of the integrated tree model is used to reduce the difference between the output value of the integrated tree model and the true type value. In the embodiments of the present application, taking the first loss function as the logistic regression loss function for minimizing binary classification as an example, the first loss function can specifically refer to the following formula (1):
[0109] P(y|x,θ)=h θ (x) y (1 - h θ (x)) 1-y (1)
[0110] In the above formula (1), the value of y corresponds to the actual binary classification label value. Therefore, the value of y can only be 1 or 0. h θ (x) y is the posterior probability when the value of y is 1, and 1 - h θ (x) y is the posterior probability when the value of y is 0. The two probabilities are combined to obtain the general form in the above formula (1).
[0111] In step 3043, based on the first loss function, parameter update processing is performed on the initialized integrated tree model to obtain the trained integrated tree model.
[0112] Based on the fact that the first loss function can characterize the difference between the output value of the integrated tree model and the true type value, based on the difference characterized by the first loss function, in each round of the training process of the integrated tree model, a search for the target aggregated feature subset is performed to obtain a set of the best-performing aggregated feature sets recommended so far during the update process, which is used to replace the initialized or the previous-round aggregated feature subset to obtain the trained integrated tree model. The aggregated feature subset with the highest importance index in the current round is the target aggregated feature subset. The target aggregated feature subset can be obtained in the following way: perform a descending sort on the importance index of each aggregated feature to obtain a descending sorted list, and select the aggregated feature combination with the pre-configured number of features at the head as the target aggregated feature subset, or select the aggregated features with the pre-configured number of features whose importance index is greater than the index threshold and combine them as the target aggregated feature subset.
[0113] Continuously execute the above process for the update process of the target aggregated feature subset. If the training process meets the pre-configured conditions, stop executing the above steps. The pre-configured conditions can be that the training duration of the ensemble tree model reaches the pre-configured duration, or the number of training times reaches the pre-configured number. Use the current target aggregation subset of the last round when the execution stops as the best candidate aggregated feature subset.
[0114] In the aggregated feature subset, except for some aggregated features used to train the ensemble tree model, input the remaining aggregated features into the trained ensemble tree model for evaluating the trained ensemble tree model. The trained ensemble tree model can predict the prediction types of the remaining aggregated features. By comparing the output value of the trained ensemble tree model with the corresponding true types, the prediction effect of the ensemble tree model can be evaluated. For example, in the scenario of predicting whether game users will churn, the aggregated features in the remaining user aggregated feature subset can be input into the trained ensemble tree model, and the output of the ensemble tree model can be compared with the true churn labels of the remaining users to predict the probability of the remaining users churning.
[0115] The index for evaluating the prediction results of the ensemble tree model can use the Receiver Operating Characteristic curve (ROC) for auxiliary judgment. By calculating the area enclosed by the Receiver Operating Characteristic curve (ROC) and the coordinate axes, that is, the size of the Area Under the Curve (AUC), judge the evaluation index of the prediction results. The larger the value of the Area Under the Curve (AUC), the higher the prediction effect.
[0116] In step 3044, use at least one of the following parameters in the training process as the training process data: candidate aggregated features participating in the training process, aggregated feature pairs corresponding to the candidate aggregated features, the number of times the candidate aggregated features are processed, the processing order of the candidate aggregated features, and the difference between the prediction type and the true type.
[0117] Training the ensemble tree model based on the candidate feature subset can obtain the data of the training process, including candidate aggregated features participating in the training process, aggregated feature pairs corresponding to the candidate aggregated features during the training process, and the processing order of the candidate aggregated features. The processing order is the sequence information of selecting aggregated features once during the splitting process of the ensemble tree model. For example, if the ensemble tree model sequentially selects aggregated features agg2, agg1, agg8 for splitting, then use the set of feature sequence information (agg2, agg1, agg8) as the processing order information of the candidate aggregated features in the ensemble tree model.
[0118] During the training process of the integrated tree model, a tree is continuously grown through feature splitting. Each round, a tree is learned to fit the residual between the predicted values and the actual values of the previous round of tree model, that is, the difference between the predicted type and the true type. The integrated tree model will explore the contribution degree of each aggregated feature according to the number of times each aggregated feature is used during the training process, so as to score the importance of the aggregated feature. The more times a candidate aggregated feature is used for splitting, the more important the candidate aggregated feature is. The number of times the candidate aggregated feature is processed is the importance information of each aggregated feature calculated by the integrated tree model.
[0119] In the embodiments of the present application, the initialized integrated tree model is called according to the candidate aggregated feature for classification to obtain a corresponding prediction model, and the integrated tree model is updated based on the difference between the true value and the predicted value, which ensures the source of training information of the prediction model, improves the accuracy of determining the difference, and determines the loss function based on the difference, which further improves the update quality of the integrated tree model.
[0120] Continue to refer to Figure 3A In step 305, the initialized prediction model is trained based on the training process data to obtain a trained prediction model.
[0121] Exemplarily, the trained prediction model is used to determine the importance index of each aggregated feature.
[0122] Based on the training process data, the initialized prediction model can be trained. Different training process data are used for training different prediction models. For example, candidate aggregated features are used to train the prediction model of candidate aggregated features, candidate aggregated feature pairs are used to train the prediction model of candidate aggregated feature pairs, and the processed order of candidate aggregated features is used to train the prediction model of candidate aggregated feature sequences.
[0123] In some embodiments, refer to Figure 3E Figure 3E is a schematic flowchart of the data processing method provided by the embodiments of the present application. Figure 3A The step 305 shown can be implemented by Figure 3E steps 3051 to 3053 below, which will be specifically described.
[0124] In step 3051, the initialized prediction model is called based on the training process data for prediction processing to obtain various types of prediction metrics.
[0125] Exemplarily, the weighted sum result of the prediction metrics of each type is the importance index.
[0126] In the embodiments of the present application, the initialized prediction model includes: a first prediction model, a second prediction model, and a third prediction model.
[0127] In the embodiments of the present application, the types of prediction indicators include: the number of times the prediction of candidate aggregated features is processed, the first prediction effect indicator of the ensemble tree model for the pair of aggregated features, and the second prediction effect indicator of the ensemble tree model for the candidate aggregated feature sequence.
[0128] Exemplarily, a first prediction model is used to determine the number of times the prediction is processed. A second prediction model is used to predict the first prediction effect indicator based on the pair of aggregated features, where the first prediction effect indicator characterizes the prediction difference between the predicted type and the true type of the pair of aggregated features.
[0129] For the training process data of the above-mentioned training ensemble tree model, the corresponding prediction models are trained according to different training process data. The first prediction model is used to determine the number of times the candidate aggregated feature is processed. During the splitting process of the ensemble tree model, the number of times the candidate aggregated feature is processed is the number of times each aggregated feature is used during the splitting process. The number of times used corresponds to the contribution degree of the candidate aggregated feature, which characterizes the importance of the candidate aggregated feature. That is, the first prediction model is used to determine the importance information of the candidate aggregated feature calculated by the ensemble tree model.
[0130] The second prediction model is used to predict the pair of candidate aggregated features to obtain the first prediction effect indicator of the ensemble tree model for the pair of aggregated features, which is used to characterize the difference between the predicted type and the true type of the pair of aggregated features. That is, the second prediction model is used to determine the importance information of the pair of candidate aggregated features calculated by the ensemble tree model.
[0131] A third prediction model is used to predict the second prediction effect indicator based on the feature sequence of the candidate aggregated feature, where the second prediction effect indicator characterizes the prediction difference between the predicted type and the true type of the feature sequence.
[0132] The third prediction model is used to predict the pair of candidate aggregated features to obtain the second prediction effect indicator of the ensemble tree model for the aggregated feature sequence, which is used to characterize the difference between the predicted type and the true type of the aggregated feature sequence. That is, the third prediction model is used to determine the importance information of the candidate aggregated feature sequence calculated by the ensemble tree model.
[0133] In the embodiments of the present application, the structures of the first prediction model, the second prediction model, and the third prediction model are the same; the first prediction model includes: an embedding layer, a fully connected layer, and a prediction layer.
[0134] The embedding layer is used to convert the aggregated feature into an embedded feature vector.
[0135] The fully connected layer is used to perform a linear transformation on the embedded feature vector; the linear transformation is to perform global convolution on the feature vector input to the fully connected layer, and according to the number of input samples and the dimension of each sample, output the corresponding samples after expanding the dimension.
[0136] A prediction layer, which is used to perform mapping processing on the result of linear transformation processing to obtain a prediction result. The mapping processing is to map the feature space samples marked by the previous layer calculation to the sample label space, that is, to integrate the features, reduce the influence of the feature position on the classification result, and improve the robustness of the entire network.
[0137] See Figure 3F , Figure 3F is a schematic structural diagram of the prediction model provided by the embodiment of the present application. For example, taking the structure of the first prediction model 311 as an example for illustration, the first prediction model 311 includes: an embedding layer 3111, a fully connected layer 3112, and a prediction layer 3113. The prediction model is trained according to the data in the training process of the integrated tree model, and the determined number of times the prediction is processed is output. The second prediction model 312 and the third prediction model 313 also call the training data in the training process of the integrated tree model for prediction. The structures of the above three prediction models are the same.
[0138] The embedding layers of the above three prediction models all adopt a component-based feature representation method, that is, quickly learn the embedding representation of each aggregated feature quadruple element, and then introduce an attention mechanism to respectively determine the embedding representation methods corresponding to the basic features and pre-configured aggregation types in the quadruple, and quickly assemble the embedding representations of the elements to obtain the embedding representation of the feature aggregation feature. For example, the quadruple of agg8 is represented as (f8, multi-day aggregation, sum, 10). The embedding representation of the basic feature is used as q, and the embedding representations of the aggregation type, aggregation operation, and aggregation duration are used as k and v. For the assembled embedding representation, the corresponding component-based feature representation is obtained as agg8 = (f8, t8, m8, l8), where f8 represents the basic feature corresponding to the component feature, t8 represents the type of aggregating the feature, m8 represents the aggregation operation performed on the feature, and l8 represents the aggregation duration of the feature. That is, the training process data passes through the embedding layer to convert the aggregated features into embedding representations, and then refines the high-order representations corresponding to the input training process data through multiple linear layers. Finally, the prediction layer performs mapping processing on the result of the linear transformation processing and outputs the final prediction result, which represents the type obtained by the classification processing.
[0139] In the embodiments of the present application, when using the training process data of the integrated tree model to call the prediction model for prediction, various types of prediction metrics can be obtained. The diverse metrics provide a basis for updating the integrated tree model in the next round, and make full use of the contribution and cooperation mode of each feature in the feature subset by the integrated tree model to be clarified. The first prediction model can be used to determine the number of times a prediction is processed. The second prediction model can be based on the aggregated features to predict the metrics of the first prediction effect. The third prediction model can predict the second prediction effect metrics based on the feature sequence. The embodiments of the present application effectively utilize this information to construct a better feature subset from multiple perspectives and improve the feature screening effect.
[0140] In some embodiments, in addition to the structures exemplified above, the prediction model can also be implemented using neural network models of other structures for classification processing.
[0141] In step 3052, based on the difference between each prediction metric and the corresponding actual metric, the second loss function of the initialized prediction model is determined.
[0142] Based on step 3051, the training process data includes: candidate aggregated features participating in the training process, aggregated feature pairs corresponding to the candidate aggregated features, the number of times the candidate aggregated features are processed, the processing order of the candidate aggregated features, the difference between the prediction type and the true type; wherein, the candidate aggregated features are the aggregated features in the candidate feature subset.
[0143] In some embodiments, based on the above training process data, step 3052 can be implemented in the following manner: Based on the predicted number of times each candidate aggregated feature is processed and the actual number of times it is processed, a first sub-function is determined.
[0144] Based on the first prediction effect metric and the corresponding first actual difference, a second sub-function is determined, where the first actual difference is the actual difference between the predicted type and the true type of the aggregated feature pair.
[0145] Based on the second prediction effect metric and the corresponding second actual difference, a third sub-function is determined, where the second actual difference is the actual difference between the predicted type and the true type of the feature sequence.
[0146] The sum of the first sub-function, the second sub-function, and the third sub-function is used as the second loss function.
[0147] The three prediction models can respectively determine the difference between the prediction effect metric of the prediction model and the actual type, and respectively determine the corresponding losses to sum up to obtain the total loss, effectively utilizing the details of the application of features in the model, providing help and a favorable basis for the generation of subsequent candidate feature subsets.
[0148] In step 3053, parameter update processing is performed on the initialized prediction model based on the second loss function to obtain the trained prediction model.
[0149] The principle of parameter update processing utilizes the process of backpropagation. In the backpropagation process, the loss function is needed to obtain the gradients of the output layer and hidden layer parameters, and then these gradients are used to update the parameters, so that the calculation result of the loss function reaches a smaller value. Based on gradient descent, error backpropagation is carried out. By calculating the gradient of the loss function and feeding this gradient back to the optimization function to update the weights to minimize the loss function, that is, calculating the error between the predicted value and the true value of the neural network, propagating the error forward to the previous layer, and considering the weight coefficients of the previous neuron at the same time. Since the nodes in the latter layer are connected to multiple nodes in the previous layer, the errors are summed, and the weights are updated based on the error sum and the influence of the weights on the whole.
[0150] In the embodiments of the present application, the loss is determined separately for each model, and the loss values are added to obtain the total loss, which can more accurately determine the gain of each tree generated by each split, facilitating the selection of the node with the largest gain, simplifying the complexity, and accelerating the speed of node splitting. The optimal tree structure is obtained by minimizing the loss function, which helps to optimize the ensemble tree model by selecting the feature with the highest importance index.
[0151] Continue to refer to Figure 3A , in step 306, based on the aggregated feature set, the trained prediction model is called for prediction processing to obtain the importance index of each aggregated feature.
[0152] Exemplarily, the trained prediction model can obtain the importance indexes of different dimensions of each aggregated feature, and the weighted sum result of the importance indexes of each dimension is used as the total importance index of the aggregated feature.
[0153] In some embodiments, referring to Figure 3G , Figure 3G is a schematic flowchart of the data processing method provided by the embodiments of the present application. Figure 3A The step 306 shown can be implemented by Figure 3G steps 3061 to 3064, which will be specifically described below.
[0154] In step 3061, based on the training process data, the first prediction model is called for prediction processing to obtain the predicted number of processed times.
[0155] The three prediction models are trained according to the training process data. Based on the first prediction model, the importance of the candidate aggregated features can be determined, and the importance indexes of the candidate aggregated features are obtained. The training process and principle of the first prediction model have been described above and will not be elaborated here.
[0156] In step 3062, the second prediction model is called based on the training process data for prediction processing to obtain the first prediction effect index.
[0157] Based on the second prediction model, the importance of the candidate aggregation feature pairs can be determined, and the importance index of the candidate aggregation feature pairs is obtained. The training process and principle of the second prediction model have been described above and will not be elaborated here.
[0158] In step 3063, the third prediction model is called based on the training process data for prediction processing to obtain the second prediction effect index.
[0159] Based on the third prediction model, the importance of the candidate aggregation feature sequences can be determined, and the importance index of the aggregation feature sequences is obtained. The training process and principle of the third prediction model have been described above and will not be elaborated here.
[0160] In step 3064, weighted summation processing is performed on each prediction index of each aggregation feature to obtain the importance index of each aggregation feature.
[0161] Using the prediction indexes of each aggregation feature in the above three prediction models respectively for weighted summation to obtain the importance index of each aggregation feature, according to the determined importance information, the corresponding candidate aggregation features, aggregation feature pairs and aggregation feature sequences with high importance scores can be screened out.
[0162] In the embodiment of the present application, based on the first prediction model, the number of times a feature is processed can be predicted. Based on the second and third prediction models, the first and second prediction effect indexes can be obtained, reflecting the corresponding importance indexes from the three perspectives of aggregation features, feature pairs and feature sequences, providing a basis for subsequent feature screening. Selecting features with high importance indexes can help improve the screening results.
[0163] Continue to refer to Figure 3A , in step 307, a target aggregation feature subset is determined based on the importance index of each aggregation feature.
[0164] Exemplarily, the target aggregation feature subset is used to perform classification processing on multiple objects by calling the trained ensemble tree model. For example, classification processing is performed according to the behaviors of game users to obtain users who have the intention to purchase a certain type of product in the game application, and relevant recommendation information of this type of product is sent to these users, or users who have not logged in to the game application recently but have the intention to return are classified, and relevant information for recalling the game is sent to users with a greater intention to return.
[0165] In some embodiments, Figure 3AThe shown step 307 can determine the target aggregated feature subset in any of the following ways:
[0166] Way 1: Perform a descending order sorting process on the importance indicators of each aggregated feature to obtain a descending order list, select the aggregated features with the pre-configured number of features at the head of the descending order list, and combine the selected aggregated features into the target aggregated feature subset.
[0167] Way 2: Select the aggregated features with the pre-configured number of features whose importance indicators are greater than the indicator threshold, and combine the selected aggregated features into the target aggregated feature subset.
[0168] In some embodiments, after performing step 307, in response to not meeting the pre-configured conditions, use the current target aggregated feature subset as the candidate aggregated feature subset, use the current trained ensemble tree model as the initialized ensemble tree model, and transfer to step 304.
[0169] Exemplarily, the pre-configured conditions include at least one of the following:
[0170] Condition 1: The training duration of the ensemble tree model reaches the pre-configured duration;
[0171] Condition 2: The number of training times of the ensemble tree model reaches the pre-configured number of times.
[0172] The pre-configured conditions of the ensemble tree model are determined according to the remaining aggregated features used to evaluate the prediction results. The verification set is composed of the aggregated features excluding those used for training the ensemble tree model. For example, during the process of evaluating the trained ensemble tree model based on the verification set, observe the loss value of the verification curve, and select the number of times corresponding to the model with the minimum loss value as the pre-configured number of times for training the ensemble tree model.
[0173] Exemplarily, if the pre-configured conditions are not met, loop through steps 304 to 307 until the pre-configured conditions are reached to obtain the best target aggregated feature subset.
[0174] In the embodiments of the present application, by aggregating and screening the historical data of the object, the efficiency of processing the object features is accelerated, and the time for processing the object features is reduced. Using the training process data to train the prediction model and making predictions on the trained prediction model based on the aggregated feature set, the importance indicators of each aggregated feature can be obtained, and then the target aggregated feature subset can be determined according to the importance indicators, so that the trained ensemble tree model can perform classification processing, can more accurately predict the object behavior, effectively utilize the contribution and cooperation method of the aggregated features, construct a better feature subset from multiple perspectives, improve the effect of feature screening, and save the execution cycle of the feature engineering work.
[0175] Next, an exemplary application of the data processing method according to the embodiments of the present application in an actual application scenario will be described.
[0176] In the process of information recommendation of Internet applications, it is necessary to predict the future behavior of users. Taking game information recommendation as an example, for example, whether game users will make a return visit, whether they will churn, and whether they will perform paid behavior, etc. The prediction results can be used as a strong basis for decision-making in game operation and market information recommendation. For example: In the activity of targeted placement for game return users, it is necessary to predict the future return willingness of users and recall users with a greater return willingness; in the anti-churn placement activity, it is necessary to predict users with a high churn risk, and then intervene in the users indicated by the prediction results to extend the overall usage period. How to use the historical data of in-game users to quickly and accurately predict the future behavior of users plays an important role in the information recommendation process.
[0177] The link of predicting the future behavior of users includes: generating and screening key object features from the historical behavior information of users, and using a machine learning model to predict the future behavior of users based on the key object features. The machine learning model can make more accurate predictions with the help of high-quality user features. For example: In the task of predicting the churn risk of users, from a large amount of historical game behavior data of users, user aggregation features highly correlated with user churn events are refined (such as: the total online duration of the user in the past month, the average number of assists per day of the user in the past week). User aggregation features can reflect the user's recent stickiness to the game, help estimate the user's churn event, and are regarded as high-quality user aggregation features in the churn prediction scenario. Taking these high-quality user aggregation features as input, the model can more accurately distinguish between churn users and non-churn users, so that in the actual application process of the model, the churn probability of users can be more reasonably estimated and potential churn users can be more accurately identified. The generation and screening of user key features will greatly affect the prediction accuracy of predicting future behavior based on user historical behavior.
[0178] Predicting the future behavior of users relies on generating corresponding object features from the historical data of users. However, the historical data of game users is usually very large, and there are various ways to generate object features. A large number of user aggregation features with different meanings can be generated by aggregating different historical behavior information of users at different time periods using different mathematical methods. The feature scale can be as high as tens of thousands, resulting in extremely difficult screening and extremely low efficiency of manual screening.
[0179] In existing behavior prediction technologies, there are two types of solutions: manual analysis solutions and automated feature engineering solutions. In the manual analysis-based solutions, the degree of intelligence is low, the dependence on professionals is high, and the execution is relatively inefficient. In addition, there is usually a large amount of user historical data, and there are various ways to generate user aggregation features, which can generate user aggregation features in the tens of thousands. There are too many features to be screened, and the complexity of the screening task is too high. It is very difficult for professionals to process such complex tasks efficiently and accurately, and there may be omissions or mis-selections. In the automated feature engineering-based solutions, compared with the manual analysis solutions, they are more intelligent and efficient, but there are certain deficiencies in the existing algorithms. In the automated feature engineering method based on reinforcement learning, the evaluation process of the feature subset is regarded as a black box, and only optimization analysis is carried out based on the performance feedback of the feature subset, missing a lot of key information. In the automated feature engineering methods based on feature importance and label correlation, the feature screening basis is too simple and violent, and it is not applicable to the case where the number of candidate features is particularly large.
[0180] Based on the above problems existing in the prior art, the embodiments of the present application provide a data processing method. By obtaining an object data set and an initialized integrated tree model, performing feature aggregation processing on each historical data to obtain a set of aggregated features, extracting candidate subsets from the set of aggregated features, calling the integrated tree model based on the candidate aggregated subsets for training processing, and training the initialized prediction model based on the training process data to obtain a trained prediction model. By calling the trained model through the aggregated feature subset for prediction processing to obtain importance indicators, and then determining the target aggregated feature subset according to the importance indicators for calling the trained integrated tree model to perform classification processing for multiple objects, it is possible to efficiently generate and screen object features, with a higher degree of automation and improved accuracy of object behavior prediction.
[0181] Specifically, refer to Figure 4 , Figure 4 which is an optional process schematic diagram of the data processing method provided by the embodiments of the present application. Taking the server as the execution entity, it will be described in detail in combination with Figure 4 the steps shown.
[0182] In step 401, an algorithm input is provided according to the target task (the algorithm corresponds to the integrated tree model above).
[0183] Exemplarily, according to the target task, the input of the integrated tree model is obtained. The input content includes the following three parts. Specifically, refer to Figure 5 , Figure 5 which is a schematic diagram of the content input of the data processing method provided by the embodiments of the present application.
[0184] Table 501 is a user basic feature table, denoted as TF, which represents the basic features of the user's historical behavior extracted from the user's hourly historical behavior log. For example, the basic features of the user's games, payments, social interactions, etc. at different times in the past three months.
[0185] Table 502 is a basic label table, denoted as TL. According to the target task, the corresponding real user churn label is extracted from the user performance log, that is, whether the user has not logged in for a week. If this condition is met, the user is considered to have churned, and the user label is 1. Otherwise, the user label is 0.
[0186] Table 503 is an aggregated feature table of the target feature scale, and the target feature scale is denoted as TS. Set the number of high-quality user features expected to be generated, where id is the user's unique identification number, f1 and f2 represent two basic features, and label represents the user churn label. The date in the user basic feature table represents the date when the user's behavior occurred, and the date in the basic label table represents the statistical date of the label (starting from this date, if the user has not logged in for a week, it is considered churned). For example, agg1, agg2, and agg3 represent three aggregated features, where agg is the abbreviation of aggregation. agg1=(f1, multi-day aggregation, sum, 2) indicates that the acquisition method of the aggregated feature agg1 is: sum the f1 features of the user in the past 2 days.
[0187] In step 402, construct the aggregated feature search space.
[0188] Among them, the aggregated feature search space corresponds to the above-mentioned aggregated feature set.
[0189] By adopting different aggregation methods for the basic features in the user basic feature table (Table 501), a large number of user aggregated features with different meanings can be generated for constructing the aggregated feature search space. The constructed aggregated features can reflect the construction information of the corresponding features. The construction information indicates which basic feature in the user basic feature table the aggregated feature is based on, and also indicates the aggregation type adopted for the basic feature (for example: multi-day aggregation, which aggregates the information of the user in recent days; time-series aggregation, which aggregates the login information of the user in recent times), as well as the mathematical operation adopted for the basic feature during the aggregation process (for example: max, min, sum, avg, max-min), and the length of the aggregation period (for example: 1 day, 3 days, 7 days, 14 days, 21 days).
[0190] In step 403, initialize the aggregated feature subset.
[0191] For example, the aggregated feature subset corresponds to the above-mentioned candidate aggregated feature subset. Randomly select the number of aggregated features equal to the target feature scale from the aggregated feature search space to obtain the initialized aggregated feature subset.
[0192] In step 404, the tree model is trained and the trained tree model is evaluated.
[0193] Exemplarily, the tree model is trained using the initialized aggregated feature subset, and the prediction performance of the trained tree model is evaluated with reference to Figure 6A , Figure 6A FIG. is a schematic flow chart of the training process based on the integrated tree model provided by the embodiments of the present application. Figure 6A Steps 601 to 605 shown in
[0194] In step 601, the tree model is trained and the trained tree model is evaluated.
[0195] Exemplarily, with reference to Figure 6B , Figure 6B FIG. is a schematic principle diagram provided by the embodiments of the present application. After the input aggregated feature set, an integrated tree model 610 including aggregated features agg2, agg1, and agg8 is constructed, and the corresponding importance score value is 0.6.
[0196] In step 602, the tree model information is extracted.
[0197] Exemplarily, with reference to Figure 6C , Figure 6C FIG. is a schematic principle diagram provided by the embodiments of the present application. Information is extracted from the tree model to form a data set 620. The data set 620 includes: further obtaining the feature importance score, feature pair information, and feature sequence information according to the extracted information. After obtaining the corresponding information, an estimation model is trained based on the extracted information.
[0198] In step 603, the estimation model is trained.
[0199] Exemplarily, the estimation model can be composed of sub-estimation models with different functions.
[0200] Exemplarily, with reference to Figure 6D , Figure 6DIt is a schematic structural diagram of the prediction model provided by the embodiments of the present application. Based on the obtained tree model information, the prediction model 630 is trained. The prediction model 630 includes: a feature importance prediction model 631, a feature pair importance prediction model 632, and a feature sequence importance prediction model 633. The important features, feature pairs, and feature sequences with high importance scores are screened out by using the constructed information of the ensemble tree model to call the prediction model. The feature pair importance prediction model 632 includes a component attention mechanism 6321 (component atten in the figure), a feature extraction layer 6322, and a linear transformation layer 6323. The feature sequence importance prediction model 633 includes: a component attention mechanism 6331 (component atten in the figure), a feature extraction layer 6332, and a linear transformation layer 6333.
[0201] Taking the feature importance prediction model 631 as an example for illustration, the feature importance prediction model 631 evaluates the importance of the feature agg8. The feature importance model outputs the corresponding importance score through a component attention mechanism 6311 (component atten in the figure), a feature extraction layer 6312, and a linear transformation layer 6313.
[0202] Exemplarily, in step 604, random sampling is performed on the aggregated features, feature pairs, and feature sequences, and the prediction model is used to screen out important features, feature pairs, and feature sequences to construct a new subset of aggregated features. The prediction model screens from the search space of the aggregated features, and the search space is on the scale of tens of thousands. The scale of the search space of the aggregated features is obtained by cumulative multiplication of the number of quadruples of the aggregated features.
[0203] Exemplarily, in step 605, TS aggregated features are obtained.
[0204] TS corresponds to the target scale of Table 503. There are a total of N basic features in the user basic feature table (Table 501), T types of aggregation types, M types of aggregation operations, and L types of aggregation durations. Then, the search space of the aggregated features contains N×T×M×L user aggregated features, that is, the size of the search space of the aggregated features is N×T×M×L.
[0205] Exemplarily, after the screening in step 604, in step 605, the target feature scale (TS) of aggregated features can be obtained. Each aggregated feature is represented by the above quadruple, characterizing the aggregated construction information of the corresponding feature. For example: (f1, multi-day aggregation, sum, 2) represents the sum of the f1 features in the past two days, (f2, multi-day aggregation, max, 3) represents the maximum value of the f2 features in the past three days, and (f2, time-series aggregation, avg, 2) represents the average value of the f2 features in the past two times.
[0206] Exemplarily, aggregate features are generated for each user in the basic label table (Table 501) to construct an aggregate feature table with labels. Some users are sampled from this table for tree model training, and the remaining are used for tree model performance testing.
[0207] Exemplarily, the training process of the integrated tree model is specifically as follows: Some users are sampled from the basic label table, and the aggregate features of their aggregate feature subsets are used as the input of the integrated tree model, and their true churn labels are used as the target output of the integrated tree model. Then, by minimizing the value of the logistic regression loss function for binary classification, the difference between the model output value and the true label value is reduced, thereby training the integrated tree model.
[0208] Exemplarily, the aggregate features of the aggregate feature subsets of the remaining users are input into the trained integrated tree model. The trained integrated tree model can then predict the probability of the remaining users churning. By comparing the output value of the integrated tree model with the true churn labels of the remaining users, the prediction effect of the integrated tree model can be evaluated.
[0209] Exemplarily, the metrics for evaluating the prediction results can be the following: Area Under the Curve (AUC), the area enclosed by the Receiver Operating Characteristic Curve (ROC) and the coordinate axes. The larger this value, the higher the prediction effect.
[0210] In step 405, the best feature subset is updated.
[0211] Exemplarily, the best feature subset corresponds to the above-mentioned target feature subset. Update the aggregate feature subset (corresponding to the above-mentioned target feature subset) with the best performance on the integrated tree model searched so far, denoted as Fset*. Exemplarily, each round of the update process will recommend a set of aggregate feature sets for integrated tree model training and performance testing. Fset* represents the aggregate feature set that maximizes the prediction performance of the tree model among all evaluated aggregate feature sets.
[0212] Exemplarily, the best aggregate feature subset is obtained specifically by executing steps 405 to 408. Step 409 is executed after step 405, or step 406 is executed. If the duration of executing steps 405 to 408 has reached the pre-configured search duration, step 409 is executed. If the pre-configured search duration has not been reached, step 406 is executed after step 405.
[0213] In step 406, the tree model construction information is extracted.
[0214] Exemplarily, refer to Figure 6C, extract the data during the training process from the trained ensemble tree model to form a data set 620. The training process data is the data generated during the classification process of the candidate aggregated feature subset of the ensemble tree model, including the importance information of each aggregated feature calculated by the ensemble tree model, the information of the aggregated feature pairs that appear in the ensemble tree model, and the usage order information of the aggregated features.
[0215] For example, during the training process, the ensemble tree model will continuously grow a tree through feature splitting. Each round, a tree is learned to fit the residual between the predicted value and the actual value of the previous round of tree model. The ensemble tree model will explore the contribution degree of each feature according to the number of times each aggregated feature is used during the training process, so as to score the importance of the aggregated features. The more times a feature is used for splitting, the more important it is. The number of times the candidate aggregated feature is processed is the importance information of each aggregated feature calculated by the ensemble tree model.
[0216] For example, during the training process, the ensemble tree model will sequentially select appropriate aggregated features as nodes for splitting. The processing order of the candidate aggregated features is the usage order information of the aggregated features in the ensemble tree model.
[0217] For example, referring to Figure 6B As shown in the ensemble tree model, if the aggregated features agg2, agg1, and agg8 are sequentially selected for splitting, then the set of feature sequence information (agg2, agg1, agg8) is used as the usage order information of the aggregated features in the ensemble tree model.
[0218] In step 407, train a prediction model to estimate the importance of features, feature pairs, and feature sequences.
[0219] For example, as Figure 6D shown in the prediction model, use the information extracted from the ensemble tree model to call the prediction model to screen out features, feature pairs, and feature sequences with high importance scores. The three prediction models respectively implement three mapping relationships. The feature importance prediction model (corresponding to the first prediction model above) is used to implement the mapping from the aggregated feature to the importance of the aggregated feature provided by the ensemble tree model to determine the predicted number of times of processing; the feature pair importance prediction model (corresponding to the second prediction model above) is used to implement the mapping from the aggregated feature pair to the performance of the ensemble tree model where the aggregated feature pair is located. The performance of the ensemble tree model is the predicted effect value (corresponding to the first prediction effect index above) obtained by testing the tree model under the aggregated feature in step 404, which represents the prediction difference between the predicted type and the true type of the aggregated feature pair; the feature sequence importance prediction model (corresponding to the third prediction model above) is used to implement the mapping from the aggregated feature sequence to the performance of the ensemble tree model where the aggregated feature sequence is located, and predict the second prediction effect index, which represents the prediction difference between the predicted type and the true type of the feature sequence.
[0220] Exemplarily, the construction information of the integrated tree model extracted in step 406 is used to train the three prediction models in step 407 respectively to enhance the prediction accuracy. The above three prediction models have the same structure, including: an embedding layer, a fully connected layer, and a prediction layer. First, through the embedding layer, the aggregated features are converted into embedding representations, then the input high-order representations are refined through multiple linear layers, and finally the prediction layer performs a mapping process on the results of the linear transformation process to obtain the prediction results.
[0221] Exemplarily, the embedding layer of the prediction model in step 407 adopts a component-based feature representation method. Specifically, according to step 402, each aggregated feature can be represented as a quadruple: (basic feature, aggregation type, aggregation operation, aggregation duration). For example, the quadruple of agg8 is represented as (f8, multi-day aggregation, sum, 10), indicating the sum of the feature f8 in the recent ten days. The embodiments of the present application quickly learn the embedding representations of these constituent elements, and then introduce an attention mechanism, taking the embedding representation of the basic feature as q, and the embedding representations of the aggregation type, aggregation operation, and aggregation duration as k and v, so as to quickly assemble the embedding representation of a specific aggregated feature. Exemplarily, the corresponding component-based feature representation obtained is agg8 = (f8, t8, m8, l8).
[0222] In step 408, the aggregated features are screened to construct a new feature subset.
[0223] Exemplarily, based on the prediction model, the aggregated features are screened (as shown in step 604 in Figure 6A ), random sampling is performed on the aggregated features, feature pairs, and feature sequences from the aggregated feature search space, and then the three prediction models are used to screen out the features, feature pairs, and feature sequences with high importance scores for constructing a new aggregated feature subset Fset with the scale of the target feature scale. The new aggregated feature subset Fset is the best aggregated feature subset Fset* of the current round.
[0224] Exemplarily, based on the newly constructed aggregated feature subset, the integrated tree model is continuously trained and evaluated, that is, for the newly constructed aggregated feature subset, steps 404 to 408 are repeatedly executed until the expected search time is reached.
[0225] In step 409, the best aggregated feature subset is output to help the model effectively solve the target prediction task.
[0226] Exemplarily, in response to reaching the expected search time, the best aggregated feature subset Fset* of the current round is output, and the machine learning model is allowed to perform analysis and prediction based on the aggregated features in Fset*, so as to effectively solve the target task.
[0227] Embodiments of the present application can be applied to the field of data processing, extract features and aggregate features from user historical behavior information, train and evaluate a model based on the aggregated feature subset, and obtain predictions of object behavior. In practical application scenarios, embodiments of the present application can efficiently and quickly generate and screen object features and improve prediction accuracy.
[0228] The data processing method provided by the embodiments of the present application has the following beneficial effects:
[0229] In the embodiments of the present application, the process of generating and training object features is automated feature engineering with a high degree of intelligence. It does not rely on manual experience for feature screening, avoiding misselection and omission that are likely to occur when dealing with complex and large-scale features.
[0230] In the solutions of related technologies, dealing with the same amount of object features highly depends on manual experience. In the existing automated feature engineering methods, the application details of each feature in the model are ignored, regarding the application details of each feature in the model as a black box, and failing to effectively utilize the contributions and cooperation methods of each feature in the model to provide help and a favorable basis for the generation of subsequent feature subsets. The solution provided by the embodiments of the present application can make up for the deficiencies, fully utilize the contributions and cooperation methods of each feature in the aggregated feature subset by using the integrated tree model, and effectively use this information to construct a better feature subset from multiple perspectives, improving the effect of feature screening and reducing the execution cycle of feature engineering work. The solution of the present application has strong generalization ability and can be applied to predictions in the information recommendation scenario of Internet application programs to perform information recommendation based on the prediction results.
[0231] The feature representation method of the prior art assigns a vector to each aggregated feature as its embedding representation, and then learns the embedding representations of tens of thousands of aggregated features, which is extremely inefficient to execute. Even if hundreds of completely different aggregated features are analyzed in each round, hundreds of iterations are required to learn the embedding representations of tens of thousands of aggregated features. The component-based feature representation method proposed by the embodiments of the present application can utilize the importance information of a small number of aggregated features to quickly learn the embedding representations of the constituent elements, and then use the embedding representations of these constituent elements to quickly estimate the embedding representations of more aggregated features in the search space, thereby effectively improving the screening effect of excellent features and making the feature representation method more efficient.
[0232] Next, continue to describe the exemplary structure of the software module implementation of the data processing device 455 provided by the embodiments of the present application. In some embodiments, such as Figure 2As shown, the software modules stored in the data processing device 455 of the memory 450 may include: a data acquisition module 4551, configured to obtain an object data set and an initialized integrated tree model, where the object data set includes historical data and actual type labels respectively corresponding to a plurality of objects, and each of the objects corresponds to at least one piece of historical data; a feature processing module 4552, configured to perform feature aggregation processing on each piece of historical data to obtain at least one aggregated feature respectively corresponding to each of the objects, and combine each of the aggregated features into an aggregated feature set; extract a candidate aggregated feature subset from the aggregated feature set; a training processing module 4553, configured to perform training processing based on the candidate aggregated feature subset by invoking the initialized integrated tree model to obtain a trained integrated tree model and training process data, where the training process data is data generated during the classification processing of the candidate aggregated feature subset by the integrated tree model; perform training processing on the initialized prediction model based on the training process data to obtain a trained prediction model, where the trained prediction model is used to determine the importance index of each aggregated feature; a prediction processing module 4554, configured to perform prediction processing by invoking the trained prediction model based on the aggregated feature set to obtain the importance index of each aggregated feature; determine a target aggregated feature subset based on the importance index of each aggregated feature, where the target aggregated feature subset is used to perform classification processing on a plurality of objects by invoking the trained integrated tree model.
[0233] In some embodiments, the feature processing module 4552 is further configured to perform feature extraction processing on each piece of historical data to obtain historical record features; perform aggregation processing on each historical record feature based on a pre-configured aggregation type to obtain at least one aggregated feature of each object; combine at least one aggregated feature of each object to obtain an aggregated feature set.
[0234] In some embodiments, the pre-configured aggregation type includes at least one of the following: aggregate historical record features within a pre-configured aggregation duration; aggregate historical record features of a pre-configured number of times closest to the current moment.
[0235] In some embodiments, the feature processing module 4552 is further configured to obtain a pre-configured feature number of the candidate aggregated feature subset; select an aggregated feature of the pre-configured feature number from the aggregated feature set to obtain a candidate aggregated feature subset.
[0236] In some embodiments, the training processing module 4553 is further configured to perform classification processing on each candidate aggregated feature by invoking an initialized ensemble tree model, to obtain the predicted type of each candidate aggregated feature, where the candidate aggregated feature is an aggregated feature in a candidate aggregated feature subset; determine a first loss function of the initialized ensemble tree model based on the difference between the predicted type and the true type of each candidate aggregated feature; perform parameter update processing on the initialized ensemble tree model based on the first loss function, to obtain a trained ensemble tree model; use at least one of the following parameters in the training process as training process data: the candidate aggregated features participating in the training process, the aggregated feature pairs corresponding to the candidate aggregated features, the number of times the candidate aggregated features are processed, the processing order of the candidate aggregated features, and the difference between the predicted type and the true type.
[0237] In some embodiments, the training processing module 4553 is further configured to perform prediction processing on the training process data by invoking an initialized prediction model, to obtain multiple types of prediction metrics, where the weighted sum result of each prediction metric is an importance metric; determine a second loss function of the initialized prediction model based on the difference between each prediction metric and the corresponding actual metric; perform parameter update processing on the initialized prediction model based on the second loss function, to obtain a trained prediction model.
[0238] In some embodiments, the initialized prediction model includes: a first prediction model, a second prediction model, and a third prediction model; the types of prediction metrics include: the predicted number of times the candidate aggregated feature is processed, a first prediction effect metric of the ensemble tree model for the aggregated feature pair, and a second prediction effect metric of the ensemble tree model for the candidate aggregated feature sequence; the first prediction model is configured to determine the predicted number of times of processing; the second prediction model is configured to predict the first prediction effect metric based on the aggregated feature pair, where the first prediction effect metric represents the prediction difference between the predicted type and the true type of the aggregated feature pair; the third prediction model is configured to predict the second prediction effect metric based on the feature sequence of the candidate aggregated feature, where the second prediction effect metric represents the prediction difference between the predicted type and the true type of the feature sequence.
[0239] In some embodiments, the training process data includes: candidate aggregated features participating in the training process, aggregated feature pairs corresponding to the candidate aggregated features, the number of times the candidate aggregated features are processed, the processing order of the candidate aggregated features, and the difference between the predicted type and the true type; wherein, the candidate aggregated features are aggregated features in the candidate aggregated feature subset; the training processing module 4553 is further configured to determine a first sub-function based on the predicted number of times the candidate aggregated features are processed and the actual number of times they are processed; determine a second sub-function based on the first prediction effect metric and the corresponding first actual difference, where the first actual difference is the actual difference between the predicted type and the true type of the aggregated feature pair; determine a third sub-function based on the second prediction effect metric and the corresponding second actual difference, where the second actual difference is the actual difference between the predicted type and the true type of the feature sequence; and use the sum of the first sub-function, the second sub-function, and the third sub-function as the second loss function.
[0240] In some embodiments, the structures of the first prediction model, the second prediction model, and the third prediction model are the same; the first prediction model includes: an embedding layer, a fully connected layer, and a prediction layer; the embedding layer is configured to convert the aggregated features into embedding feature vectors; the fully connected layer is configured to perform linear transformation processing on the embedding feature vectors; and the prediction layer is configured to perform mapping processing on the result of the linear transformation processing to obtain a prediction result.
[0241] In some embodiments, the trained prediction model includes: the first prediction model, the second prediction model, and the third prediction model; the types of prediction metrics include: the predicted number of times the candidate aggregated features are processed, the first prediction effect metric of the ensemble tree model for the aggregated feature pair, and the second prediction effect metric of the ensemble tree model for the candidate aggregated feature sequence; the prediction processing module 4554 is further configured to call the first prediction model for prediction processing based on the training process data to obtain the predicted number of times the candidate aggregated features are processed; call the second prediction model for prediction processing based on the training process data to obtain the first prediction effect metric; call the third prediction model for prediction processing based on the training process data to obtain the second prediction effect metric; and perform weighted summation processing on each prediction metric of each aggregated feature to obtain the importance metric of each aggregated feature.
[0242] In some embodiments, the prediction processing module 4554 is further configured to determine the target aggregated feature subset in any of the following ways: perform descending order sorting on the importance metrics of each aggregated feature to obtain a descending order list, select the pre-configured number of aggregated features at the head of the descending order list, and combine the selected aggregated features into the target aggregated feature subset; select the pre-configured number of aggregated features with importance metrics greater than the metric threshold, and combine the selected aggregated features into the target aggregated feature subset.
[0243] In some embodiments, the prediction processing module 4554 is further configured to, in response to not meeting the preconfigured conditions, use the current target aggregated feature subset as the candidate aggregated feature subset, and use the current trained ensemble tree model as the initialized ensemble tree model; and proceed to the step of training by invoking the initialized ensemble tree model based on the candidate aggregated feature subset to obtain the trained ensemble tree model and the training process data.
[0244] In some embodiments, the preconfigured conditions include at least one of the following: the training duration of the ensemble tree model reaches the preconfigured duration; the number of training times of the ensemble tree model reaches the preconfigured number of times.
[0245] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions, and the computer program or computer-executable instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer program or computer-executable instructions from the computer-readable storage medium, and the processor executes the computer program or computer-executable instructions, so that the electronic device executes the data processing method described above in the embodiments of the present application.
[0246] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, the processor will be caused to execute the data processing method provided in the embodiments of the present application, for example: Figure 3A the data processing method shown.
[0247] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; or it may be various devices including one or any combination of the above memories.
[0248] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0249] By way of example, the computer-executable instructions may or may not correspond to a file in a file system, and may be stored as part of a file that holds other programs or data, such as: in one or more scripts stored in a HyperText Markup Language (HTML) document, stored in a single file dedicated to the program under discussion, or, stored in multiple cooperating files (e.g., files that store one or more modules, subroutines, or portions of code).
[0250] By way of example, the executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or, on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0251] In summary, through the embodiments of the present application, by aggregating and filtering the historical data of an object, the efficiency of processing the object features is accelerated, and the time for processing the object features is reduced. Training the prediction model using the training process data and making predictions on the trained prediction model based on the aggregated feature set can obtain the importance indicators of each aggregated feature, and then determine the target aggregated feature subset according to the importance indicators, so that the trained ensemble tree model performs classification processing, can more accurately predict the object behavior, effectively utilize the contributions and cooperation methods of the aggregated features, construct a better feature subset from multiple perspectives, improve the effect of feature screening, and save the execution cycle of the feature engineering work.
[0252] The above is only the embodiments of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining an object data set and an initialized ensemble tree model, where the object data set includes historical data and actual type labels respectively corresponding to multiple objects, and each object corresponds to at least one piece of historical data; Performing feature aggregation processing on each piece of historical data to obtain at least one aggregated feature respectively corresponding to each object, and combining each aggregated feature into an aggregated feature set; Extracting a candidate aggregated feature subset from the aggregated feature set; Based on the candidate aggregated feature subset, calling the initialized ensemble tree model for training processing to obtain a trained ensemble tree model and training process data, where the training process data is data generated during the classification processing of the candidate aggregated feature subset by the ensemble tree model; Based on the training process data, performing training processing on an initialized prediction model to obtain a trained prediction model, where the trained prediction model is used to determine the importance index of each aggregated feature; Based on the aggregated feature set, calling the trained prediction model for prediction processing to obtain the importance index of each aggregated feature; Determining a target aggregated feature subset based on the importance index of each aggregated feature, where the target aggregated feature subset is used to perform classification processing on the multiple objects by calling the trained ensemble tree model.
2. The method according to claim 1, characterized in that, The performing feature aggregation processing on each piece of historical data to obtain at least one aggregated feature respectively corresponding to each object, and combining each aggregated feature into an aggregated feature set includes: Performing feature extraction processing on each piece of historical data to obtain historical record features; Based on a pre-configured aggregation type, performing aggregation processing on each historical record feature to obtain at least one aggregated feature of each object; Combining at least one aggregated feature of each object to obtain the aggregated feature set.
3. The method according to claim 2, wherein The pre-configured aggregation type includes at least one of the following: Aggregating the historical record features within a pre-configured duration; Aggregating the historical record features of the pre-configured number of times closest to the current moment.
4. The method according to claim 1, wherein The extracting a candidate aggregated feature subset from the aggregated feature set includes: Obtaining the pre-configured feature number of the candidate aggregated feature subset; Selecting the aggregated features of the pre-configured feature number from the aggregated feature set to obtain the candidate aggregated feature subset.
5. The method according to claim 1, wherein The based on the candidate aggregated feature subset, calling the initialized ensemble tree model for training processing to obtain a trained ensemble tree model and training process data includes: Based on each candidate aggregated feature, calling the initialized ensemble tree model for classification processing to obtain the predicted type of each candidate aggregated feature, where the candidate aggregated feature is the aggregated feature in the candidate aggregated feature subset; Based on the difference between the predicted type and the true type of each candidate aggregated feature, determining the first loss function of the initialized ensemble tree model; Based on the first loss function, performing parameter update processing on the initialized ensemble tree model to obtain a trained ensemble tree model; Use at least one of the following parameters in the training process as training process data: the candidate aggregation features participating in the training process, the aggregation feature pairs corresponding to the candidate aggregation features, the number of times the candidate aggregation features are processed, the processing order of the candidate aggregation features, and the difference between the predicted type and the true type.
6. The method according to claim 1, wherein Training the initialized prediction model based on the training process data to obtain a trained prediction model, including: Invoking the initialized prediction model based on the training process data for prediction processing to obtain various types of prediction metrics, where the weighted sum result of each prediction metric is the importance metric; Determining the second loss function of the initialized prediction model based on the difference between each prediction metric and the corresponding actual metric; Performing parameter update processing on the initialized prediction model based on the second loss function to obtain a trained prediction model.
7. The method according to claim 6, wherein The initialized prediction model includes: a first prediction model, a second prediction model, and a third prediction model; The types of the prediction metrics include: the predicted number of times the candidate aggregation features are processed, the first prediction effect metric of the ensemble tree model for the aggregation feature pairs, and the second prediction effect metric of the ensemble tree model for the candidate aggregation feature sequence; The first prediction model is used to determine the predicted number of times of processing; The second prediction model is used to predict the first prediction effect metric based on the aggregation feature pair, where the first prediction effect metric represents the prediction difference between the predicted type and the true type of the aggregation feature pair; The third prediction model is used to predict the second prediction effect metric based on the feature sequence of the candidate aggregation features, where the second prediction effect metric represents the prediction difference between the predicted type and the true type of the feature sequence.
8. The method according to claim 7, wherein The training process data includes: candidate aggregation features participating in the training process, the aggregation feature pairs corresponding to the candidate aggregation features, the number of times the candidate aggregation features are processed, the processing order of the candidate aggregation features, and the difference between the predicted type and the true type; where the candidate aggregation features are the aggregation features in the candidate aggregation feature subset; Determining the second loss function of the initialized prediction model based on the difference between each prediction metric and the corresponding actual metric, including: Determining a first sub-function based on the predicted number of times of processing and the actual number of times of processing of each candidate aggregation feature; Determining a second sub-function based on the first prediction effect metric and the corresponding first actual difference, where the first actual difference is the actual difference between the predicted type and the true type of the aggregation feature pair; Determining a third sub-function based on the second prediction effect metric and the corresponding second actual difference, where the second actual difference is the actual difference between the predicted type and the true type of the feature sequence; Using the sum of the first sub-function, the second sub-function, and the third sub-function as the second loss function.
9. The method according to claim 7, wherein The structures of the first prediction model, the second prediction model, and the third prediction model are the same; the first prediction model includes: an embedding layer, a fully connected layer, and a prediction layer; The embedding layer is used to convert the aggregated features into embedding feature vectors; The fully connected layer is used to perform linear transformation processing on the embedding feature vectors; The prediction layer is used to perform mapping processing on the result of the linear transformation processing to obtain a prediction result.
10. The method according to claim 1, characterized in that, The trained prediction model includes: a first prediction model, a second prediction model, and a third prediction model; the types of the prediction metrics include: the predicted processed times of the candidate aggregated features, the first prediction effect metric of the ensemble tree model for the aggregated feature pairs, and the second prediction effect metric of the ensemble tree model for the candidate aggregated feature sequences; Performing prediction processing by invoking the trained prediction model based on the aggregated feature set to obtain the importance metric of each of the aggregated features, including: Invoking the first prediction model for prediction processing based on the training process data to obtain the predicted processed times; Invoking the second prediction model for prediction processing based on the training process data to obtain the first prediction effect metric; Invoking the third prediction model for prediction processing based on the training process data to obtain the second prediction effect metric; Performing weighted summation processing on each of the prediction metrics of each of the aggregated features to obtain the importance metric of each of the aggregated features.
11. The method according to any one of claims 1 to 10, characterized in that, Determining a target aggregated feature subset based on the importance metric of each of the aggregated features, including: Determining the target aggregated feature subset by any one of the following methods: Performing descending order sorting on the importance metrics of each of the aggregated features to obtain a descending order list, selecting the aggregated features with a preconfigured number of features at the head of the descending order list, and combining the selected aggregated features into a target aggregated feature subset; Selecting the aggregated features with a preconfigured number of features whose importance metrics are greater than the metric threshold, and combining the selected aggregated features into a target aggregated feature subset.
12. The method according to any one of claims 1 to 10, characterized in that, After determining the target aggregated feature subset based on the importance metric of each of the aggregated features, the method further includes: In response to not meeting the preconfigured conditions, taking the current target aggregated feature subset as a candidate aggregated feature subset, and taking the current trained ensemble tree model as an initialized ensemble tree model; Transferring to the step of invoking the initialized ensemble tree model based on the candidate aggregated feature subset for training processing to obtain a trained ensemble tree model and training process data.
13. The method according to claim 12, wherein The preconfigured conditions include at least one of the following: The training duration of the ensemble tree model reaches a preconfigured duration; The number of training times of the ensemble tree model reaches a preconfigured number of times.
14. A data processing device, characterized in that, The apparatus includes: A data acquisition module, configured to acquire an object data set and an initialized ensemble tree model, where the object data set includes historical data and actual type labels respectively corresponding to multiple objects, and each object corresponds to at least one historical data; A feature processing module, configured to perform feature aggregation processing on each of the historical data to obtain at least one aggregated feature corresponding to each of the objects, and combine each of the aggregated features into an aggregated feature set; extract a candidate aggregated feature subset from the aggregated feature set; A training processing module, configured to perform training processing on the initialized ensemble tree model based on the candidate aggregated feature subset to obtain a trained ensemble tree model and training process data, where the training process data is data generated during the classification processing of the candidate aggregated feature subset by the ensemble tree model; perform training processing on the initialized prediction model based on the training process data to obtain a trained prediction model, where the trained prediction model is used to determine the importance index of each of the aggregated features; A prediction processing module, configured to perform prediction processing on the trained prediction model based on the aggregated feature set to obtain the importance index of each of the aggregated features; determine a target aggregated feature subset based on the importance index of each of the aggregated features, where the target aggregated feature subset is used to perform classification processing on the multiple objects by calling the trained ensemble tree model.
15. An electronic device for data processing, characterized in that, The electronic device for data processing includes: A memory, configured to store computer-executable instructions; A processor, configured to implement the data processing method according to any one of claims 1 to 13 when executing the computer-executable instructions or computer program stored in the memory.
16. A computer-readable storage medium stores computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer program, when executed by the processor, implement the data processing method according to any one of claims 1 to 13.
17. A computer program product, comprising computer-executable instructions or a computer program, characterized in that, The computer-executable instructions or computer program, when executed by the processor, implement the data processing method according to any one of claims 1 to 13.