Data processing method and device, equipment, storage medium and product
By generating a training sample set containing batch modal information, the impact of multiple batch sample data sets in account behavior prediction is solved, the accuracy and training speed of the model are improved, and more efficient training effects are achieved.
Patent Information
- Application Number
- CN202410003528.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
In the account behavior prediction scenario, the prior art failed to effectively utilize the information amount of multiple batches of sample data sets during the training process, resulting in slow convergence speed and low accuracy of the model, especially in the case of large distribution differences between batches, the training effect was poor.
By generating the training sample set, the account behavior information and batch mode information containing the sample data, the neural network is used to train the model to adapt to the modality of different batches, reduce the mutual influence between batches, and improve the training effect.
It improves the accuracy and training speed of account behavior prediction models, and can obtain models with good prediction capabilities faster.
Smart Images

Figure CN120256943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly relates to a data processing method, apparatus, device, storage medium, and product. Background Art
[0002] Account behavior prediction is an important topic in the development of artificial intelligence in related technologies. Learning technologies such as machine learning and deep learning can be used to train an account behavior prediction model. In the specific learning process, the process of model training is not achieved overnight, but a continuous iterative training process. Batch training can be implemented based on batch training data to achieve continuous iteration as a whole.
[0003] In the account behavior prediction scenario, in the related technology during the training process, multiple batches of sample data are continuously obtained, and the sample data of each batch are treated uniformly for training, which reduces the model convergence speed and also reduces the accuracy of account behavior prediction. Summary of the Invention
[0004] Embodiments of this application provide a data processing method, apparatus, device, storage medium, and product, which can reduce the mutual influence between different batches of sample data sets and improve the training effect. For the training process of the account behavior prediction model, it can improve the accuracy of the account behavior prediction model and the training speed.
[0005] According to one aspect of the embodiments of this application, a data processing method is provided. The method includes:
[0006] Obtain multiple batches of sample data sets, where the sample data in each sample data set includes account behavior information;
[0007] Generate a training sample set according to the multiple batches of sample data sets. Each training sample in the training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information, where the batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the training sample belongs;
[0008] Train a preset model based on the training sample set to obtain an account behavior prediction model.
[0009] According to one aspect of the embodiments of this application, a data processing apparatus is provided. The apparatus includes:
[0010] A data acquisition module, configured to obtain multiple batches of sample data sets, where the sample data in each sample data set includes account behavior information;
[0011] A data processing module, configured to generate a training sample set according to the multiple batches of sample data sets. Each training sample in the training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information, where the batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the training sample belongs; and, based on the training sample set, train a preset model to obtain an account behavior prediction model.
[0012] According to one aspect of the embodiments of the present application, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above data processing method.
[0013] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided. At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the above data processing method.
[0014] According to one aspect of the embodiments of the present application, a computer program product is provided. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to cause the computer device to execute to implement the above data processing method.
[0015] The technical solution provided by the embodiments of the present application can bring the following beneficial effects:
[0016] An embodiment of the present application provides a data processing method. This data processing method can generate a training sample set according to different batches of sample data sets and the number of batches of the overall sample data set. Each training sample in this training sample set not only includes the data information of the sample data corresponding to this training sample, but also includes the batch information corresponding to this sample data and the information corresponding to the number of batches of the sample population. In fact, an embodiment of the present application proposes a concept of batch modality. The batch information corresponding to the sample data corresponding to each of the foregoing training samples and the information corresponding to the number of batches of the sample population can be summarized as batch modality. This enables the training samples corresponding to sample data in different batches to have different batch modalities, guiding a preset model to adaptively learn the training samples with different batch modalities during the training process, thereby reducing the mutual influence between sample data sets in different batches and improving the training effect. For the training process of an account behavior prediction model, it can improve the accuracy of the account behavior prediction model and increase the training speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 is a schematic diagram of an application program running environment provided by an embodiment of the present application;
[0019] Figure 2 is a flowchart of a data processing method provided by an embodiment of the present application;
[0020] Figure 3 is a schematic flowchart of a method for generating a training sample set provided by an embodiment of the present application;
[0021] Figure 4 is a schematic flowchart of a method for generating a single training sample provided by an embodiment of the present application;
[0022] Figure 5 is a schematic diagram of a training sample generation process provided by an embodiment of the present application;
[0023] Figure 6 is a schematic diagram of a neural network structure provided by an embodiment of the present application;
[0024] Figure 7 is a block diagram of a data processing device provided by an embodiment of the present application;
[0025] Figure 8It is a structural block diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0026] Before introducing the method embodiments provided by the present application, relevant terms or nouns that may be involved in the method embodiments of the present application are briefly introduced first, so as to facilitate the understanding of those skilled in the art of the present application.
[0027] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science, which attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0028] Artificial intelligence technology is an interdisciplinary subject, involving a wide range of fields, including both hardware-level technologies and software-level technologies. Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0029] Machine Learning (ML) is an interdisciplinary subject in multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how a computer simulates or implements human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve its own performance. Machine learning is the core of artificial intelligence and the fundamental way to make a computer intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning.
[0030] Deep learning: The concept of deep learning stems from the research on artificial neural networks. A multi-layer perceptron with multiple hidden layers is a deep learning structure. Deep learning forms more abstract high-level representations of attribute categories or features by combining low-level features to discover the distributed feature representations of data.
[0031] Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or local area network to achieve data computing, storage, processing, and sharing. Cloud technology is the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used on demand, and is flexible and convenient. Cloud computing technology will become an important support. The back-end services of technical network systems require a large amount of computing and storage resources, such as video websites, picture-based websites, and more portal websites. With the highly developed and applied Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the back-end system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. This can only be achieved through cloud computing.
[0032] Media content: Media content refers to various forms of information and content presented on various media platforms. It includes: Text-based content: including short essays, long articles, columns, etc., which is the most traditional form of media content. Picture-based content: including pictures, posters, comics, etc., which can be published through platforms such as social media and communities, and is usually used to convey information or express emotions. Video-based content: including short videos, live broadcasts, etc., which can be published through short video platforms or live broadcast platforms. Audio-based content: including broadcasts, podcasts, etc., which are usually published through audio platforms and mainly target users' auditory needs. Interactive content: including Q&A, voting, surveys, etc., which can be published through Q&A platforms and can promote interaction and communication between the media and the audience. In addition, media content also includes various forms such as audio, video, animation, charts, and images, and can present content such as news reports, current affairs reviews, entertainment programs, and advertising. Different media platforms and forms are suitable for different audience groups, and media content is also constantly changing and innovating.
[0033] Exposure click behavior: In the case of exposing a piece of media content, the behavior of triggering a click on that media content. The click behavior represents a shallow interaction between a user and the application presenting the media content. The click behavior can trigger the application to display more details or interaction information of the media content, etc.
[0034] Click conversion behavior: In the case of clicking on a piece of media content, the behavior of triggering a conversion on that media content. Click conversion represents a deeper interaction between a user and the application presenting the media content. For example, it can be to purchase the item corresponding to the media content, participate in the activities related to the media content, or download the application software corresponding to the media content, register and log in to the application software corresponding to the media content, etc.
[0035] Exposure conversion behavior: The behavior of triggering conversion for a media content when the media content is exposed. If the exposure click behavior and the click conversion behavior occur successively for a media content, it can be considered that an exposure conversion behavior for the media content has occurred.
[0036] Before specifically elaborating on the embodiments of the present application, the relevant technical background related to the embodiments of the present application is introduced to facilitate the understanding of those skilled in the art of the present application.
[0037] Neural network training is an important topic in the development of artificial intelligence in related technologies. The models using learning technologies such as machine learning and deep learning are essentially neural network models, and such models need to learn various knowledge and capabilities through the way of neural network training. In the specific learning process, the process of neural network training is not achieved overnight, but a continuous iterative training process. In the process of neural network training, batch training can be realized based on batch training data, and continuous iteration can be achieved as a whole.
[0038] Taking the account behavior prediction scenario as an example, in order to obtain a model with the ability to predict account behavior, the related technology first determines a neural network and trains the neural network until it finally has the ability to predict account behavior. During the training process, multiple batches of sample data are continuously obtained, and then, without considering the macro distribution differences in batches presented by different batches of sample data, the sample data of each batch are treated uniformly for neural network training. Specifically, it is to directly mix the sample data sets of multiple batches and perform training. If the distribution differences between the sample data sets of different batches are small, the training effect of this training method is acceptable. However, in practical applications, there are often cases where the distribution differences between the sample data sets of different batches are large. In this case, the influence of the distribution differences between the sample data of different batches on the neural network training effect cannot be ignored, which will reduce the convergence speed of the neural network and also reduce the accuracy of account behavior prediction.
[0039] In fact, in the actual scenario of account behavior prediction, in many cases, it is impossible to obtain sufficient sample data including account behavior information at one time, and only the method of obtaining a batch of sample data sets at intervals can be used to obtain sample data in batches. However, due to the time span, there will be significant differences in the distribution of sample data sets in different batches. In some cases, the effect of training using a mixture of sample data sets in different batches is even worse than that of training with a single batch of sample data sets. This is caused by the large distribution differences in sample data sets in different batches, which lead to interference. However, the information volume of multi-batch sample data sets is significantly higher than that of single-batch sample data sets. If the multi-batch sample data sets can be reasonably utilized and the influence of the distribution differences between batches of sample data sets can be reduced, it is obviously more beneficial to improve the training effect. Specifically, it can accelerate the convergence of the neural network, quickly obtain an account behavior prediction model, and enable the account behavior prediction model to have good account behavior prediction ability.
[0040] In view of this, an embodiment of the present application provides a data processing method. This data processing method can generate a training sample set according to different batches of sample data sets and the total number of batches of the sample data sets. Each training sample in this training sample set not only includes the data information of the sample data corresponding to this training sample, but also includes the batch information corresponding to this sample data and the information corresponding to the total number of batches of the sample population. In fact, an embodiment of the present application proposes a concept of batch modality. The batch information corresponding to the sample data corresponding to each training sample and the information corresponding to the total number of batches of the sample population included in each training sample can be summarized as batch modality. In this way, the training samples corresponding to the sample data in different batches have different batch modalities, guiding the neural network to adaptively learn the training samples with different batch modalities during the training process, thereby reducing the mutual influence between sample data sets in different batches and improving the training effect. For the training process of the account behavior prediction model, it can improve the accuracy of the account behavior prediction model and the training speed.
[0041] To make the purpose, technical solution, and advantages of the present application clearer, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings.
[0042] Please refer to Figure 1 , which shows a schematic diagram of an application program running environment provided by an embodiment of the present application. This application program running environment may include: a terminal 10 and a server 20.
[0043] The terminal 10 includes, but is not limited to, electronic devices such as mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, game consoles, e-book readers, multimedia playback devices, and wearable devices. An application program client can be installed in the terminal 10.
[0044] In the embodiments of the present application, the above application can be any application that can provide data processing services. Typically, the application can be an account behavior service type or a media content recommendation type application. Of course, in addition to the account behavior service type or the media content recommendation type application, services that rely on data processing can also be provided in other types of applications. For example, news applications, social applications, interactive entertainment applications, browser applications, shopping applications, content sharing applications, virtual reality (VR) applications, augmented reality (AR) applications, etc. The embodiments of the present application do not make any limitations in this regard. The embodiments of the present application do not make any limitations in this regard. Optionally, a client of the above application runs in the terminal 10.
[0045] The server 20 is used to provide background services for the client of the application in the terminal 10. For example, the server 20 can be the background server of the above application. The server 20 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network, content distribution network), and big data and artificial intelligence platforms. Optionally, the server 20 provides background services for the applications in multiple terminals 10 at the same time.
[0046] Optionally, the terminal 10 and the server 20 can communicate with each other through the network 30. The terminal 10 and the server 20 can be directly or indirectly connected through wired or wireless communication methods, and the present application does not make any limitations in this regard.
[0047] Please refer to Figure 2 , which shows a flowchart of a data processing method provided by an embodiment of the present application. This method can be applied to a computer device. The above computer device refers to an electronic device with data calculation and processing capabilities. For example, the execution entity of each step can be Figure 1 the server 20 in the application running environment shown. This method can include the following steps:
[0048] S201. Obtain multiple batches of sample data sets, and the sample data in each of the above sample data sets includes account behavior information.
[0049] In the embodiments of the present application, the specific number of batches is not limited, and it can be a positive integer greater than or equal to 2. For each batch, the sample data of the sample data set of this batch are all the recorded account behavior information. The account behavior information in the embodiments of the present application generally refers to the general term of the information that can be used to train the account behavior prediction model. The embodiments of the present application do not limit the specific content of the account behavior information, but the account behavior information should include the information related to the account behavior that the account behavior prediction model can predict.
[0050] For example, if it is desired that the trained account behavior prediction model has the ability to predict exposure click behavior, then the account behavior information must include information on whether a certain account performs exposure click behavior. If it is desired that the trained account behavior prediction model has the ability to predict click conversion behavior, then the account behavior information must include information on whether a certain account performs click conversion behavior. If it is desired that the trained account behavior prediction model has the ability to predict exposure conversion behavior, then the account behavior information must include information on whether a certain account performs exposure conversion behavior.
[0051] In one implementation manner, the above-mentioned account behavior information includes at least one of the following: object historical behavior data, object attribute data. The object refers to the executor of the historical behavior or the description object of the attribute data, and can be simply understood that the object corresponds to the account in the account behavior information. The embodiments of the present application specifically state that all the account behavior information used in the embodiments of the present application is desensitized data and is legal usage information fully authorized by each relevant entity.
[0052] The object historical behavior data refers to the historical behavior of a certain object (account), for example, whether a certain item is purchased, whether certain activities are participated in, etc. The object attribute data refers to the inherent static attributes of the object (account), or the object portrait. For example, the age, gender, native place, educational background, personal interests, behavior habits, preference tendencies, etc. of the user corresponding to the account. The object historical behavior data and the object attribute data can be considered as the basic information required for performing account behavior prediction, and the embodiments of the present application do not limit their specific content.
[0053] S202. Generate a training sample set according to the above-mentioned multiple batches of sample data sets. Each training sample in the above-mentioned training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information. The above-mentioned batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the above-mentioned training sample belongs;
[0054] In the embodiments of the present application, each of the above sample data in each of the above sample data sets has the same number of dimensions. For example, if the sample data is a 5-dimensional data or has 5-dimensional features, it can be considered that the number of dimensions of the sample data is 5. Please refer to Figure 3 , which shows a schematic flowchart of a method for generating a training sample set provided by an embodiment of the present application. The method includes:
[0055] S301. Obtain a target dimension interval of the training samples in the training sample set based on the product of the above total number of batches and the above number of dimensions;
[0056] Suppose a total of 3 batches of sample data sets are obtained in step S201, then the total number of batches is 3, and each sample data has 5 dimensions, so there are 3 * 5 = 15 dimensions in the target dimension interval.
[0057] S302. For any sample data in any of the above sample data sets, generate a training sample corresponding to the sample data according to the batch corresponding to the sample data and the above target dimension interval; wherein, each dimension in the first dimension interval corresponding to the batch where the sample data is located in the above training sample is used to record the account behavior information in the sample data, and null information is recorded in each dimension in the second dimension interval, and the second dimension interval is an interval formed by the dimensions other than the first dimension interval in the above target dimension interval.
[0058] The null information in the embodiments of the present application is actually a kind of information occupying a position, which does not have an actual meaning and is not an effective information used when the preset model makes predictions. The null information can be set to 0, or any other number that does not represent an actual meaning. For any sample data in any of the above sample data sets, the embodiments of the present application adopt the same method to generate the corresponding training sample. Please refer to Figure 4 , which shows a schematic flowchart of a method for generating a single training sample provided by an embodiment of the present application. The generation method includes:
[0059] S401. Perform N dimension-based splicings on the sample data itself to obtain spliced data, where N represents the total number of batches;
[0060] Copy the sample data N times, and then perform splicing from the dimension perspective. Please refer to Figure 5 , which shows a schematic diagram of the training sample generation process provided by an embodiment of the present application. Suppose there are three batches, then N = 3, and the sample data has 5 dimensions. By performing horizontal splicing from the dimension perspective, spliced data with 15 dimensions can be obtained, where dimensions 1 - 5, 6 - 10, and 11 - 15 respectively correspond to the data content of the original 5 dimensions of the original sample data.
[0061] S402. Determine the first - dimension interval corresponding to the above - mentioned sample data within the above - mentioned target - dimension interval according to the batch k corresponding to the above - mentioned sample data. The dimension identifier of the above - mentioned first - dimension interval is greater than or equal to (k - 1)*d and less than (k + 1)*d, where d represents the number of dimensions of the above - mentioned sample data;
[0062] For example, for the sample data of the first batch, dimensions 1 - 5 are the first - dimension interval; for the sample data of the second batch, dimensions 6 - 10 are the first - dimension interval; for the sample data of the third batch, dimensions 11 - 15 are the first - dimension interval.
[0063] S403. Retain the data in the above - mentioned first - dimension interval in the above - mentioned spliced data, and set the data in other dimensions in the above - mentioned spliced data to empty data.
[0064] The embodiments of the present application do not limit the specific implementation manner of retaining the data in the above - mentioned first - dimension interval in the above - mentioned spliced data and setting the data in other dimensions in the above - mentioned spliced data to empty data. For example, the data in other dimensions (the second - dimension interval) in the above - mentioned spliced data can be directly set to empty. For the case where the total number of batches is N and the dimension of the sample data is d, first splice them into a piece of data, that is, the spliced data. Suppose a sample data x belongs to the k - th batch of data, x i represents the value of the i - th - dimensional feature of the sample data. Then, based on the following formula, the N*d - dimensional training sample x' corresponding to the sample data x is obtained, and the value of each dimension in x' is:
[0065] Or a mask can also be generated based on step S402. Multiplying the mask by the spliced data can also achieve the effect of step S403. For example, for the sample data of the first batch, the mask is "111110000000000", and multiplying this mask by the spliced data can also achieve the effect of step S403.
[0066] In a specific implementation manner, please refer to Figure 5 , because there are three batches of data, three channels are set. The number of dimensions of each channel is equal to the number of dimensions of the sample data, which is 5 dimensions. For each sample data, first determine which batch of data it belongs to, and then put the value into the corresponding channel, and set the other channels to empty. x represents the feature value. Each original data is 5 - dimensional, and after conversion, it becomes 15 - dimensional. After such processing, each channel can respectively represent the distribution of the data in the corresponding batch, thereby enabling the preset model to learn the distribution of different - batch data. This is more discriminative than the original method where the three batches of data are mixed together and only 5 - dimensional.
[0067] S203. Train a preset model based on the above training sample set to obtain an account behavior prediction model.
[0068] To enable the preset model to fully adapt to the data content of the training sample set modified in the embodiments of the present application, so as to fully learn batch modality information and account behavior information, the embodiments of the present application provide a preset model composed of a neural network, and also design the structure of the neural network. This structure is a network structure specially designed in the embodiments of the present application to fully learn the training samples including batch modality information. This structure can not only obtain a multi-objective model - an account behavior prediction model that can be compatible with multiple distribution data sets through training, but also enable multiple data sets with different distributions to be trained together without sacrificing the training effect. This neural network can learn the hidden meanings corresponding to different positions and the processing methods for different batches of data.
[0069] In one implementation, please refer to Figure 6 , which shows a schematic diagram of the neural network structure provided by the embodiments of the present application. Now, the technical terms in this neural network structure are explained as follows:
[0070] Input: The input of the neural network;
[0071] Channel: The channel of the training sample;
[0072] Embedding: In a neural network, the Embedding operation is a method of converting high-dimensional and sparse input data into low-dimensional and dense vectors. It can map discrete categorical data or continuous numerical data into a pre-defined embedding space, so that data points of the same category are closer in the space, and data points of different categories are farther apart in the space. In a recommendation system, the Embedding operation can map discrete categorical data such as users and items into a low-dimensional embedding space, so that the similarity or matching degree between users and items can be calculated. For example, in collaborative filtering, we can represent users and items as vectors respectively, and measure the degree of interest of users in items by calculating the cosine similarity between the user vector and the item vector.
[0073] MLP: The MLP model (Multi-Layer Perceptron) is a multi-layer feedforward neural network, which consists of an input layer, hidden layers (there can be multiple layers), and an output layer. The characteristic of the MLP model is that each layer of neurons is fully interconnected with the neurons in the next layer, and there are no connections within the same layer or cross-layer connections between neurons. Such a neural network structure is usually called a "multi-layer feedforward neural network". The MLP model can be applied to various fields, such as natural language processing, recommendation systems, image recognition, etc.
[0074] FM: The FM model is short for Factorization Machine, which is a machine learning model used in recommendation systems. The FM model can introduce cross-term features by performing second-order combinations between features, thereby improving the model's prediction ability. In data with high sparsity, the FM model can better estimate the relationships between features. Compared with traditional linear models, the FM model has an additional part for feature combination at the back. The FM model improves the model score by combining pairwise features and introducing cross-term features.
[0075] DNN model: The DNN neural network is a deep neural network model designed to solve non-linear regression problems. It has the advantages of depth and capacity and can adapt to various complex non-linear models.
[0076] Multi-task: Multi-task.
[0077] Pctr: Prediction result of the exposure click task.
[0078] Pcvr: Prediction result of the click conversion task.
[0079] Concat: The concat operation in a neural network is a common feature fusion method that can combine feature information from different levels or sources, thereby improving the performance and performance of the model. It can concatenate two or more feature tensors along a certain dimension to form a larger feature tensor. This operation can increase the dimension and diversity of features, promote the interaction and integration between features, and thus improve the performance of the model.
[0080] In one implementation, the above preset model is a neural network, including a shared embedding information extraction layer, N feature extraction layers connected in parallel to the shared embedding information extraction layer, where N represents the above total batch number, and a feature concatenation layer connected to the N feature extraction layers. Figure 6 It can be seen that this embedding information extraction layer can be Figure 6 the Embedding layer in Figure 6 and this feature extraction layer can be the MLP layer connected to the Embedding layer in Figure 6 and this feature concatenation layer can be the concat layer in
[0081] Correspondingly, training the preset model based on the above training sample set to obtain an account behavior prediction model includes: for the training samples in the above training sample set, inputting the above training samples into the above shared embedding information extraction layer to obtain initial sample features; cutting the above initial sample features to obtain N first feature segments, each feature segment having d dimensions, where d represents the number of dimensions of the above sample data; inputting each of the above first feature segments into the corresponding feature extraction layer to obtain the second feature segments output by each of the above feature extraction layers; inputting the above second feature segments into the above feature splicing layer to obtain spliced sample features; predicting sample behaviors based on the above spliced sample features to obtain an account behavior prediction result; and adjusting the parameters of the above neural network based on the difference between the above account behavior prediction result and the actual account behavior recorded in the above training samples to obtain the above account behavior prediction model.
[0082] The initial sample features can be understood as the features output by the Figure 6 Embedding layer in, and the second feature segment is understood as the feature output by the MLP layer, and the spliced sample features are the features output by the concat layer. Of course, there are many choices for the specific method of adjusting the neural network parameters and the criterion for stopping parameter adjustment in the embodiments of the present application, and the embodiments of the present application will not elaborate on this.
[0083] Still taking the example of generating training samples based on the sample data with 3 batch dimensions of 5 mentioned above, the generated training samples have 15 dimensions, 3 channels, and each channel has d dimensions, that is, each channel is a complete cutting result. Since the first feature of channel 1, the first feature of channel 2, and the first feature of channel 3 are the same feature, only the data batches are different, so the underlying Embedding is shared. The same applies to the features in other positions. After Embedding processing, an Embedding vector including three channels is obtained, and the first feature segments obtained by splitting it are respectively input into three MLP networks to extract the information of each channel, and then the second feature segments output by the three MLP networks are concatenated by concat to obtain spliced sample features. The splicing formula is as follows:
[0084] Y concat = concat(MLP(Embedding(X channel1 )) + MLP(Embedding(X channel2 )) + MLP(Embedding(X channel3 ))), where Embedding(X channel1) represents the first feature segment corresponding to the first channel, Embedding(X channel2 ) represents the first feature segment corresponding to the second channel, Embedding(X channel3 ) represents the first feature segment corresponding to the third channel. Y concat represents the concatenated sample features.
[0085] In another embodiment, the above neural network further includes a feature cross layer and an information extraction layer that are connected in parallel to the above feature concatenation layer, and a fusion layer that is connected to both the above feature cross layer and the above information extraction layer. Please refer to Figure 6 , where the feature cross layer and the information extraction layer can be an FM layer and a DNN layer respectively. Correspondingly, predicting the sample behavior based on the above concatenated sample features to obtain the account behavior prediction result includes: respectively inputting the above concatenated sample features into the above feature cross layer and the above information extraction layer to correspondingly obtain the first sample feature and the second sample feature; inputting the above first sample feature and the above second sample feature into the above fusion layer to obtain the fused sample feature; predicting the sample behavior based on the above fused sample feature to obtain the account behavior prediction result.
[0086] The concatenated sample features are respectively input into the FM and DNN layers for processing to obtain the first sample feature and the second sample feature that combine the three channels. The FM is good at feature crossing, while the DNN is good at conventional information extraction. The outputs of the two are input into the fusion layer to obtain the fused sample feature. In one embodiment, the fusion layer performs feature fusion through an activation function. In one embodiment, the DNN layer is actually composed of multiple MLP layers. The FM layer and the DNN layer respectively perform the following operations: Y FM =FM(Y concat ), Y DNN =MLP(MLP(MLP(Y concat ))).
[0087] It should be particularly noted that Figure 6 the structure of the neural network shown is only an example of more neural networks provided by the embodiments of the present application, and does not constitute a limitation on the structure of the neural network of the embodiments of the present application. Figure 6 Some structures or variant structures in Figure 6 can still follow the inventive concept of the embodiments of the present application to achieve the training effect of the embodiments of the present application. Moreover, the embodiments of the present application do not limit the specific structure of a certain functional layer. For example, the shared information extraction layer can also be other structures besides Figure 6 the Embedding shown, as long as the same or similar technical purposes can be achieved.
[0088] In one embodiment, the neural network includes an exposure click behavior prediction layer and a click conversion behavior prediction layer, and both the exposure click behavior prediction layer and the click conversion behavior prediction layer are connected to the fusion layer. Please refer to Figure 6 , the boxed part on the left is the exposure click behavior prediction layer, and the boxed part on the right is the click conversion behavior prediction layer. Both the click conversion behavior prediction layer and the exposure click behavior prediction layer can be implemented and constructed through a multi-level MLP.
[0089] Correspondingly, the above account behavior prediction results include exposure click behavior prediction results, click conversion behavior prediction results, and exposure conversion behavior prediction results; predicting the sample behavior based on the above fusion sample features to obtain the account behavior prediction results, including: inputting the above fusion sample features into the exposure click behavior prediction layer to obtain the above exposure click behavior prediction results; inputting the above fusion sample features into the click conversion behavior prediction layer to obtain the above click conversion behavior prediction results; and obtaining the above exposure conversion behavior prediction results based on the product of the above exposure click behavior prediction results and the above click conversion behavior prediction results.
[0090] The above account real behaviors include exposure click real behaviors, click conversion real behaviors, and exposure conversion real behaviors. Correspondingly, adjusting the parameters of the above neural network based on the difference between the above account behavior prediction results and the account real behaviors recorded in the training samples to obtain the above account behavior prediction model, including: obtaining a first loss according to the difference between the above exposure click real behavior and the above exposure click behavior prediction results; obtaining a second loss according to the difference between the above click conversion real behavior and the above click conversion behavior prediction results; obtaining a third loss according to the difference between the above exposure conversion real behavior and the above exposure conversion behavior prediction results; and adjusting the parameters of the above neural network according to the above first loss, the above second loss, and the above third loss to obtain the above account behavior prediction model.
[0091] This embodiment of the present application does not limit the specific function formulas of the first loss, the second loss, and the third loss, as long as the corresponding differences can be quantified. For example, the cross-entropy loss in related technologies can be used for quantification. This embodiment of the present application does not limit the method of adjusting the parameters of the above neural network according to the above first loss, the above second loss, and the above third loss. For example, a weighted sum operation can be performed on the above first loss, the above second loss, and the above third loss to obtain a total loss, and then the gradient descent method can be used for parameter adjustment based on the total loss. Of course, the weights are not limited and can be set according to the actual situation. In one implementation manner, when the weights are all 1, the total loss formula is as follows:
[0092] Loss = Loss ctr + Loss cvr+Loss ctrcvr
[0093] Obviously, Loss, Loss ctr , Loss cvr , Loss ctrcvr They respectively represent the total loss, the first loss, the second loss, and the third loss. The function formula of this total loss reflects that when the neural network is trained, three objectives are trained together to jointly iterate the entire neural network. Compared with a single ctrcvr objective, it can take into account the factors of pctr and pcvr and more accurately predict the pctcvr of whether a certain media content will be converted on the premise of being exposed.
[0094] After obtaining the account behavior prediction model, the first data required for account behavior prediction can be obtained; the above first data is concatenated N times based on dimensions by itself to obtain the second data, where N represents the total number of batches; the above second data is input into the above account behavior prediction model to obtain the account behavior prediction result.
[0095] This first data is the basic information required for predicting account behavior, and the relevant content of this basic information has been described above and will not be elaborated. The concatenation method is also the same as above, and the difference is that there will be no step of retaining some dimensional data and then setting some dimensions to null. Continuing with the previous example, a total of 3 batches of sample data sets are obtained during the training process, N is 3, the dimension of the first data is 5, and the content of the first data in 5 dimensions can be identified as abcde. Then the content of the second data in 15 dimensions can be expressed as abcdeabcdeabcde. Inputting this second data into the account behavior prediction model can obtain the account behavior prediction result. Specifically, it can predict the probability that the account pointed to by the second data triggers a conversion behavior when the media data pointed to by the second data is exposed.
[0096] The following is an embodiment of the apparatus of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the embodiment of the apparatus of the present application, please refer to the method embodiment of the present application.
[0097] Please refer to Figure 7 , which shows a block diagram of a data processing apparatus provided by an embodiment of the present application. This apparatus has the function of implementing the above data processing method, and the above function can be implemented by hardware or by hardware executing corresponding software. This apparatus can be a computer device or can be set in a computer device. This apparatus can include:
[0098] A data acquisition module 701, configured to acquire multiple batches of sample data sets, and each sample data in the above sample data sets includes account behavior information;
[0099] A data processing module 702 is configured to generate a training sample set according to the above-mentioned multiple batches of sample data sets. Each training sample in the training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information. The batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the training sample belongs. Further, the data processing module 702 is configured to train a preset model based on the training sample set to obtain an account behavior prediction model.
[0100] In one embodiment, each of the sample data in each of the above-mentioned sample data sets has the same number of dimensions. The data processing module 702 is configured to perform the following operations:
[0101] Based on the product of the total number of batches and the number of dimensions, obtain the target dimension range of the training samples in the training sample set;
[0102] For any sample data in any of the above-mentioned sample data sets, generate a training sample corresponding to the sample data according to the batch corresponding to the sample data and the target dimension range;
[0103] Wherein, each dimension in the first dimension range corresponding to the batch where the sample data is located in the training sample is used to record the account behavior information in the sample data, and empty information is recorded in each dimension in the second dimension range. The second dimension range is the range formed by the dimensions other than the first dimension range in the target dimension range.
[0104] In one embodiment, the data processing module 702 is configured to perform the following operations:
[0105] Perform N times of dimension-based splicing on the sample data itself to obtain spliced data, where N represents the total number of batches;
[0106] According to the batch k corresponding to the sample data, determine the first dimension range corresponding to the sample data within the target dimension range. The dimension identifier of the first dimension range is greater than or equal to (k - 1)*d and less than (k + 1)*d, where d represents the number of dimensions of the sample data;
[0107] Retain the data in the first dimension range in the spliced data, and set the data in other dimensions in the spliced data to empty data.
[0108] In one embodiment, the data processing module 702 is configured to perform the following operations:
[0109] Obtain first data required for account behavior prediction;
[0110] Perform N times of dimension-based splicing on the first data itself to obtain second data, where N represents the total number of batches;
[0111] Input the above second data into the above account behavior prediction model to obtain the account behavior prediction result.
[0112] In one embodiment, the above preset model is a neural network, including a shared embedding information extraction layer, N feature extraction layers connected in parallel to the shared embedding information extraction layer, where N represents the above total number of batches, a feature splicing layer connected to the N feature extraction layers, and the above data processing module 702 is used to perform the following operations:
[0113] For the training samples in the above training sample set, input the above training samples into the above shared embedding information extraction layer to obtain initial sample features;
[0114] Cut the above initial sample features to obtain N first feature segments, each feature segment having d dimensions, where d represents the number of dimensions of the above sample data;
[0115] Input each of the above first feature segments into the corresponding feature extraction layer to obtain the second feature segments output by each of the above feature extraction layers;
[0116] Input the above second feature segments into the above feature splicing layer to obtain spliced sample features;
[0117] Predict the sample behavior based on the above spliced sample features to obtain the account behavior prediction result;
[0118] Adjust the parameters of the above neural network based on the difference between the above account behavior prediction result and the actual account behavior recorded in the above training samples to obtain the above account behavior prediction model.
[0119] In one embodiment, the above neural network further includes a feature cross layer and an information extraction layer connected in parallel to the feature splicing layer, and a fusion layer connected to both the feature cross layer and the information extraction layer. The above data processing module 702 is used to perform the following operations:
[0120] Input the above spliced sample features into the above feature cross layer and the above information extraction layer respectively to obtain the first sample feature and the second sample feature correspondingly;
[0121] Input the above first sample feature and the above second sample feature into the above fusion layer to obtain the fused sample feature;
[0122] Predict the sample behavior based on the above fused sample feature to obtain the account behavior prediction result.
[0123] In one embodiment, the neural network includes an exposure click behavior prediction layer and a click conversion behavior prediction layer. Both the exposure click behavior prediction layer and the click conversion behavior prediction layer are connected to the fusion layer. The account behavior prediction results include an exposure click behavior prediction result, a click conversion behavior prediction result, and an exposure conversion behavior prediction result. The data processing module 702 is configured to perform the following operations:
[0124] Input the fused sample features into the exposure click behavior prediction layer to obtain the exposure click behavior prediction result;
[0125] Input the fused sample features into the click conversion behavior prediction layer to obtain the click conversion behavior prediction result;
[0126] Based on the product of the exposure click behavior prediction result and the click conversion behavior prediction result, obtain the exposure conversion behavior prediction result.
[0127] In one embodiment, the account real behavior includes an exposure click real behavior, a click conversion real behavior, and an exposure conversion real behavior. The data processing module 702 is configured to perform the following operations:
[0128] According to the difference between the exposure click real behavior and the exposure click behavior prediction result, obtain the first loss;
[0129] According to the difference between the click conversion real behavior and the click conversion behavior prediction result, obtain the second loss;
[0130] According to the difference between the exposure conversion real behavior and the exposure conversion behavior prediction result, obtain the third loss;
[0131] According to the first loss, the second loss, and the third loss, adjust the parameters of the neural network to obtain the account behavior prediction model.
[0132] It should be noted that for the device provided in the above embodiment, when implementing its functions, only the above division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0133] Please refer to Figure 8 , which shows the structural block diagram of a computer device provided in an embodiment of the present application. This computer device can be a server for executing the above data processing method. Specifically:
[0134] The computer device 1000 includes a Central Processing Unit (CPU) 1001, a system memory 1004 including a Random Access Memory (RAM) 1002 and a Read Only Memory (ROM) 1003, and a system bus 1005 connecting the system memory 1004 and the central processing unit 1001. The computer device 1000 also includes a Basic Input / Output System (I / O (Input / Output) system) 1006 that helps transfer information between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1015.
[0135] The Basic Input / Output System 1006 includes a display 1008 for displaying information and input devices 1009 such as a mouse and a keyboard for user input of information. The display 1008 and the input devices 1009 are both connected to the central processing unit 1001 through an input / output controller 1010 connected to the system bus 1005. The Basic Input / Output System 1006 may also include an input / output controller 1010 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 1010 also provides output to a display screen, a printer, or other types of output devices.
[0136] The mass storage device 1007 is connected to the central processing unit 1001 through a mass storage controller (not shown) connected to the system bus 1005. The mass storage device 1007 and its associated computer-readable medium provide non-volatile storage for the computer device 1000. That is, the mass storage device 1007 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0137] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will understand that computer storage media is not limited to the above several types. The above system memory 1004 and mass storage device 1007 can be collectively referred to as memory.
[0138] According to various embodiments of the present application, the computer device 1000 can also be run by a remote computer on the network through a network such as the Internet. That is, the computer device 1000 can be connected to the network 1012 through the network interface unit 1011 connected to the system bus 1005, or rather, the network interface unit 1011 can also be used to connect to other types of networks or remote computer systems (not shown).
[0139] The above memory further includes a computer program, which is stored in the memory and is configured to be executed by one or more processors to implement the above data processing method.
[0140] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction, at least one program, a code set, or an instruction set is stored in the above storage medium. When the at least one instruction, the at least one program, the code set, or the instruction set is executed by a processor, the above data processing method is implemented.
[0141] Specifically, the data processing method includes:
[0142] Obtain sample data sets of multiple batches, and the sample data in each of the above sample data sets includes account behavior information;
[0143] Generate a training sample set based on the above-mentioned sample data sets of multiple batches. Each training sample in the above training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information. The above batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the above training sample belongs;
[0144] Train a preset model based on the above training sample set to obtain an account behavior prediction model.
[0145] In one embodiment, each of the above sample data in each of the above sample data sets has the same number of dimensions. The above generating a training sample set based on the above-mentioned sample data sets of multiple batches includes:
[0146] Based on the product of the above total number of batches and the above number of dimensions, obtain the target dimension range of the training samples in the above training sample set;
[0147] For any sample data in any of the above sample data sets, generate a training sample corresponding to the above sample data according to the batch corresponding to the above sample data and the above target dimension range;
[0148] Among them, each dimension in the first dimension range corresponding to the batch where the above sample data is located in the above training sample is used to record the account behavior information in the above sample data, and empty information is recorded in each dimension in the second dimension range. The second dimension range is the range formed by the dimensions other than the first dimension range in the above target dimension range.
[0149] In one embodiment, the above generating a training sample corresponding to the above sample data according to the batch corresponding to the above sample data and the above target dimension range includes:
[0150] Perform N times of dimension-based splicing on the above sample data itself to obtain spliced data, where N represents the total number of batches;
[0151] According to the batch k corresponding to the above sample data, determine the first dimension range corresponding to the above sample data within the above target dimension range. The dimension identifiers of the first dimension range are greater than or equal to (k - 1)*d and less than (k + 1)*d, where d represents the number of dimensions of the above sample data;
[0152] Retain the data in the first dimension range in the above spliced data, and set the data in other dimensions in the above spliced data to empty data.
[0153] In one embodiment, after the above training a preset model based on the above training sample set to obtain an account behavior prediction model, the method further includes:
[0154] Obtain the first data required for account behavior prediction;
[0155] Perform N times of dimension-based concatenation on the above first data itself to obtain second data, where N represents the total number of batches;
[0156] Input the above second data into the above account behavior prediction model to obtain an account behavior prediction result.
[0157] In one embodiment, the above preset model is a neural network, including a shared embedding information extraction layer, N feature extraction layers connected in parallel to the above shared embedding information extraction layer, where N represents the above total number of batches, a feature concatenation layer connected to the above N feature extraction layers, and training the above preset model based on the above training sample set to obtain an account behavior prediction model, including:
[0158] For the training samples in the above training sample set, input the above training samples into the above shared embedding information extraction layer to obtain initial sample features;
[0159] Cut the above initial sample features to obtain N first feature segments, each feature segment having d dimensions, where d represents the number of dimensions of the above sample data;
[0160] Input each of the above first feature segments into the corresponding feature extraction layer to obtain second feature segments output by each of the above feature extraction layers;
[0161] Input each of the above second feature segments into the above feature concatenation layer to obtain concatenated sample features;
[0162] Predict sample behavior based on the above concatenated sample features to obtain an account behavior prediction result;
[0163] Adjust the parameters of the above neural network based on the difference between the above account behavior prediction result and the account real behavior recorded in the above training samples to obtain the above account behavior prediction model.
[0164] In one embodiment, the above neural network further includes a feature cross layer and an information extraction layer connected in parallel to the above feature concatenation layer, and a fusion layer connected to both the above feature cross layer and the above information extraction layer. The above predicting sample behavior based on the above concatenated sample features to obtain an account behavior prediction result includes:
[0165] Input the above concatenated sample features into the above feature cross layer and the above information extraction layer respectively to obtain a first sample feature and a second sample feature correspondingly;
[0166] Input the above first sample feature and the above second sample feature into the above fusion layer to obtain a fused sample feature;
[0167] Predict sample behavior based on the above fused sample feature to obtain an account behavior prediction result.
[0168] In one embodiment, the neural network includes an exposure click behavior prediction layer and a click conversion behavior prediction layer, and both the exposure click behavior prediction layer and the click conversion behavior prediction layer are connected to the fusion layer; the account behavior prediction results include exposure click behavior prediction results, click conversion behavior prediction results, and exposure conversion behavior prediction results;
[0169] Predicting the sample behavior based on the above-mentioned fused sample features to obtain the account behavior prediction results, including:
[0170] Inputting the above-mentioned fused sample features into the exposure click behavior prediction layer to obtain the above-mentioned exposure click behavior prediction results;
[0171] Inputting the above-mentioned fused sample features into the click conversion behavior prediction layer to obtain the above-mentioned click conversion behavior prediction results;
[0172] Based on the product of the above-mentioned exposure click behavior prediction results and the above-mentioned click conversion behavior prediction results, obtain the above-mentioned exposure conversion behavior prediction results.
[0173] In one embodiment, the above-mentioned account real behavior includes exposure click real behavior, click conversion real behavior, and exposure conversion real behavior. Adjusting the parameters of the above-mentioned neural network based on the difference between the above-mentioned account behavior prediction results and the account real behavior recorded in the training samples to obtain the above-mentioned account behavior prediction model, including:
[0174] Obtain the first loss according to the difference between the above-mentioned exposure click real behavior and the above-mentioned exposure click behavior prediction results;
[0175] Obtain the second loss according to the difference between the above-mentioned click conversion real behavior and the above-mentioned click conversion behavior prediction results;
[0176] Obtain the third loss according to the difference between the above-mentioned exposure conversion real behavior and the above-mentioned exposure conversion behavior prediction results;
[0177] Adjust the parameters of the above-mentioned neural network according to the above-mentioned first loss, the above-mentioned second loss, and the above-mentioned third loss to obtain the above-mentioned account behavior prediction model.
[0178] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0179] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above data processing method.
[0180] It should be understood that the "plurality" mentioned in this article refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described in this article only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the order of the numbers. For example, two steps with different numbers are executed simultaneously, or two steps with different numbers are executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0181] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0182] In addition, in the specific implementation manner of the present application, when it comes to data related to user information, etc., when the above embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0183] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included within the protection scope of the present application.
Claims
1. A data processing method, characterized in that, The method includes: Obtaining sample data sets of multiple batches, where the sample data in each sample data set includes account behavior information; Generating a training sample set according to the sample data sets of the multiple batches, where each training sample in the training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information, and the batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the training sample belongs; Training a preset model based on the training sample set to obtain an account behavior prediction model.
2. The method according to claim 1, characterized in that, The sample data in each of the sample data sets has the same number of dimensions. The generating a training sample set according to the sample data sets of the multiple batches includes: Obtaining a target dimension range of the training samples in the training sample set based on the product of the total number of batches and the number of dimensions; For any sample data in any of the sample data sets, generating a training sample corresponding to the sample data according to the batch corresponding to the sample data and the target dimension range; Wherein, the dimensions in the first dimension range corresponding to the batch where the sample data is located in the training sample are used to record the account behavior information in the sample data, and empty information is recorded in the dimensions in the second dimension range, and the second dimension range is the range formed by the dimensions except the first dimension range in the target dimension range.
3. The method according to claim 2, wherein The generating a training sample corresponding to the sample data according to the batch corresponding to the sample data and the target dimension range includes: Performing N times of dimension-based splicing on the sample data itself to obtain spliced data, where N represents the total number of batches; According to the batch k corresponding to the sample data, determining a first dimension range corresponding to the sample data in the target dimension range, where the dimension identifiers of the first dimension range are greater than or equal to (k - 1)*d and less than (k + 1)*d, and d represents the number of dimensions of the sample data; Retaining the data in the first dimension range in the spliced data and setting the data in other dimensions in the spliced data as empty data.
4. The method according to any one of claims 1 to 3, characterized in that, After training the preset model based on the training sample set to obtain an account behavior prediction model, the method further includes: Obtaining first data required for account behavior prediction; Performing N times of dimension-based splicing on the first data itself to obtain second data, where N represents the total number of batches; Inputting the second data into the account behavior prediction model to obtain an account behavior prediction result.
5. The method according to any one of claims 1 to 3, characterized in that, The preset model is a neural network, including a shared embedding information extraction layer, N feature extraction layers connected in parallel to the shared embedding information extraction layer, where N represents the total number of batches, and a feature splicing layer connected to the N feature extraction layers. The training the preset model based on the training sample set to obtain an account behavior prediction model includes: For the training samples in the training sample set, inputting the training samples into the shared embedding information extraction layer to obtain initial sample features; Cutting the initial sample features to obtain N first feature segments, and each feature segment has d dimensions, where d represents the number of dimensions of the sample data; Input each of the first feature segments into a corresponding feature extraction layer to obtain a second feature segment output by each of the feature extraction layers; Input each of the second feature segments into the feature splicing layer to obtain a spliced sample feature; Predict sample behavior based on the spliced sample feature to obtain an account behavior prediction result; Adjust the parameters of the neural network based on the difference between the account behavior prediction result and the actual account behavior recorded in the training sample to obtain the account behavior prediction model.
6. The method according to claim 5, characterized in that, The neural network further includes a feature cross layer and an information extraction layer that are connected to the feature splicing layer in parallel, and a fusion layer that is connected to both the feature cross layer and the information extraction layer. The predicting sample behavior based on the spliced sample feature to obtain an account behavior prediction result includes: Input the spliced sample feature into the feature cross layer and the information extraction layer respectively to obtain a first sample feature and a second sample feature correspondingly; Input the first sample feature and the second sample feature into the fusion layer to obtain a fused sample feature; Predict sample behavior based on the fused sample feature to obtain an account behavior prediction result.
7. The method according to claim 6, characterized in that, The neural network includes an exposure click behavior prediction layer and a click conversion behavior prediction layer, and both the exposure click behavior prediction layer and the click conversion behavior prediction layer are connected to the fusion layer; the account behavior prediction result includes an exposure click behavior prediction result, a click conversion behavior prediction result, and an exposure conversion behavior prediction result; The predicting sample behavior based on the fused sample feature to obtain an account behavior prediction result includes: Input the fused sample feature into the exposure click behavior prediction layer to obtain the exposure click behavior prediction result; Input the fused sample feature into the click conversion behavior prediction layer to obtain the click conversion behavior prediction result; Obtain the exposure conversion behavior prediction result based on the product of the exposure click behavior prediction result and the click conversion behavior prediction result.
8. The method according to claim 7, wherein The actual account behavior includes an exposure click actual behavior, a click conversion actual behavior, and an exposure conversion actual behavior. The adjusting the parameters of the neural network based on the difference between the account behavior prediction result and the actual account behavior recorded in the training sample to obtain the account behavior prediction model includes: Obtain a first loss according to the difference between the exposure click actual behavior and the exposure click behavior prediction result; Obtain a second loss according to the difference between the click conversion actual behavior and the click conversion behavior prediction result; Obtain a third loss according to the difference between the exposure conversion actual behavior and the exposure conversion behavior prediction result; Adjust the parameters of the neural network according to the first loss, the second loss, and the third loss to obtain the account behavior prediction model.
9. A data processing device, characterized in that, The device includes: A data acquisition module for acquiring multiple batches of sample data sets, and the sample data in each sample data set includes account behavior information; A data processing module, configured to generate a training sample set according to the multiple batches of sample data sets, where each training sample in the training sample set includes the account behavior information in the corresponding sample data and the corresponding batch modality information, and the batch modality information indicates the total number of batches and the batch of the sample data set to which the sample data corresponding to the training sample belongs; and, based on the training sample set, train a preset model to obtain an account behavior prediction model.
10. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the data processing method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, At least one instruction, at least one program, a code set or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the data processing method according to any one of claims 1 to 8.
12. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions so that the computer device executes to implement the data processing method according to any one of claims 1 to 8.