Data processing method, device and equipment
By interactively processing the features of target objects and target resources, high-order semantic features are generated, which solves the problem of insufficient accuracy in click-through rate estimation and achieves more efficient click-through rate estimation.
Patent Information
- Application Number
- CN202210039937.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-13
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2042-01-13
AI Technical Summary
The accuracy of click-through rate estimation in existing technologies is insufficient, which affects the accuracy of personalized push notifications and user experience.
By obtaining multiple features to be processed of the target object and target resource, performing outer product processing to generate a feature interaction vector, and performing click-through rate estimation processing, high-order semantic features are learned to improve the accuracy of click-through rate estimation.
The accuracy of target objects' click-through rate estimation for target resources is improved, and information of interest can be obtained more effectively.
Smart Images

Figure CN114385917B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence, and in particular to a data processing method, apparatus, device, storage medium, and computer program product. Background Art
[0002] Click-through rate (CTR) prediction is a technique for estimating the probability that a user will click on a resource. CTR prediction is a crucial component of personalized push notifications. It determines whether to push a resource to a user based on the estimated probability of a user clicking on it. The accuracy of CTR prediction directly impacts the accuracy of personalized push notifications, which in turn affects user experience and resource exposure. Therefore, improving the accuracy of CTR prediction is a current research priority. Summary of the Invention
[0003] Embodiments of the present application provide a data processing method, apparatus, device, storage medium, and computer program product, which can improve the accuracy of estimating the click rate of a target object for a target resource.
[0004] In one aspect, an embodiment of the present application provides a data processing method, comprising:
[0005] Acquire a plurality of features to be processed, the plurality of features to be processed comprising object features of a target object and resource features of a target resource to be pushed to the target object;
[0006] Performing outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtaining a feature interaction vector corresponding to each feature to be processed in the multiple features to be processed, wherein the feature interaction vector corresponding to each feature to be processed is a high-order semantic feature of each feature to be processed;
[0007] A click rate estimation process is performed on the feature interaction vectors corresponding to the respective features to be processed to obtain an estimated click rate of the target object for the target resource, where the estimated click rate is used to indicate a probability of the target object clicking on the target resource.
[0008] In one aspect, an embodiment of the present application provides a data processing device, comprising:
[0009] an acquiring unit, configured to acquire a plurality of features to be processed, wherein the plurality of features to be processed include object features of a target object and resource features of a target resource to be pushed to the target object;
[0010] a processing unit, configured to perform outer product processing on the plurality of features to be processed to perform feature interaction on the plurality of features to be processed, and obtain a feature interaction vector corresponding to each of the plurality of features to be processed, wherein the feature interaction vector corresponding to each of the features to be processed is a high-order semantic feature of each of the features to be processed;
[0011] The processing unit is further used to perform click rate estimation processing on the feature interaction vectors corresponding to the various features to be processed to obtain the estimated click rate of the target object for the target resource. The estimated click rate is used to indicate: the probability of the target object clicking on the target resource.
[0012] In one aspect, an embodiment of the present application provides a data processing device, characterized in that the data processing device includes an input interface and an output interface, and further includes:
[0013] a processor adapted to implement one or more instructions; and
[0014] A computer storage medium stores one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the above-mentioned data processing method.
[0015] On the one hand, an embodiment of the present application provides a computer storage medium, characterized in that computer program instructions are stored in the computer storage medium, and when the computer program instructions are executed by a processor, they are used to execute the above-mentioned data processing method.
[0016] On the one hand, an embodiment of the present application provides a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a data processing device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions. When the computer instructions are executed by the processor, they are used to execute the above-mentioned data processing method.
[0017] In an embodiment of the present application, after obtaining multiple features to be processed, including resource features of a target resource and object features of a target object, the data processing device can perform outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, thereby obtaining a feature interaction vector corresponding to each of the multiple features to be processed; and perform click-through rate estimation processing on the feature interaction vector corresponding to each feature to be processed to obtain an estimated click-through rate of the target object for the target resource; wherein the estimated click-through rate is used to indicate the probability of the target object clicking on the target resource. By performing outer product processing on multiple features to be processed, high-order semantic features of each feature to be processed can be learned, thereby improving the accuracy of the target object's click-through rate estimation for the target resource and efficiently obtaining information of interest. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0019] Figure 1 This is a structural diagram of a data processing system provided in an embodiment of the present application;
[0020] Figure 2 This is a flow chart of a data processing method provided in an embodiment of the present application;
[0021] Figure 3 This is a schematic diagram of obtaining a feature interaction vector corresponding to each feature to be processed provided by an embodiment of the present application;
[0022] Figure 4 Schematic diagram of a projection matrix and the corresponding relationship between scaling weights and features to be processed provided in an embodiment of the present application;
[0023] Figure 5 This is a schematic diagram of obtaining a feature interaction vector corresponding to the nth feature to be processed provided by an embodiment of the present application;
[0024] Figure 6 This is a schematic diagram of a click-through rate estimation based on a data processing model provided in an embodiment of the present application;
[0025] Figure 7 This is a flow chart of a method for training a data processing model provided in an embodiment of the present application;
[0026] Figure 8 This is a schematic diagram of the overall process of training an initial image processing model provided by an embodiment of the present application;
[0027] Figure 9 is a schematic diagram of a training initial image processing model provided in an embodiment of the present application;
[0028] Figure 10 is a structural diagram of a data processing device provided in an embodiment of the present application;
[0029] Figure 11 It is a structural diagram of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0031] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also involves studying the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making.
[0032] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0033] In order to improve the accuracy of click-through rate estimation, an embodiment of the present application provides a data processing solution based on machine learning in the field of artificial intelligence. After obtaining multiple features to be processed including resource features of the target resource and object features of the target object, the multiple features to be processed can be subjected to outer product processing to obtain feature interaction vectors corresponding to each of the multiple features to be processed; then, click-through rate estimation processing is performed on the feature interaction vectors corresponding to each feature to be processed to obtain the estimated click-through rate of the target object for the target resource.
[0034] The above-mentioned data processing scheme can be executed by a data processing device, wherein the data processing device can be a terminal device, which can include but is not limited to smart phones, tablet computers, laptops, desktop computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, smart wearable devices, etc.; it can also be a server, for example, it can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or it can be a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as basic cloud computing services such as big data and artificial intelligence platforms. The above-mentioned data processing scheme can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.
[0035] Based on the above data processing solution, the present application embodiment provides a data processing system, see Figure 1 , is a structural diagram of a data processing system provided in an embodiment of the present application. Figure 1 The data processing system shown may include a data processing device 101 and a terminal device 102. The data processing device 101 may be a server, for example, an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms; the terminal device 102 may include any one or more of a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart car-mounted device, and a smart wearable device. The data processing device 101 and the terminal device 102 may be directly or indirectly connected to each other through wired or wireless communication, which is not limited in this application.
[0036] In one embodiment, a target application may be running on terminal device 102. The target application may be any application that provides resource push services, such as a social application, a news application, a music application, or a video playback application. The target application may provide resource push services to a target object corresponding to terminal device 102. The resources pushed by the target application are resources related to the services provided by the target application. For example, if the target application is a news application, the target application may push news, current affairs commentary articles, etc. to the target object; if the target application is a video playback application, the target application may push movies, variety shows, documentaries, etc. to the target object. Data processing device 101 is a data processing device corresponding to the target application and may provide service support for the target application, specifically, resource push service support for the target application. For example, if the target application is a news application, data processing device 101 is the data processing device corresponding to the news application and may provide service support for the news application; for another example, if the target application is a video playback application, data processing device 101 is the data processing device corresponding to the video playback application and may provide service support for the video playback application. The target object may be any user who uses the services provided by the target application.
[0037] In one embodiment, the target object can use the services provided by the target application through its terminal device 102, such as clicking to view resources pushed by the target application. After obtaining multiple to-be-processed features, including resource features of the target resource and object features of the target object, the data processing device 101 can perform outer product processing on the multiple to-be-processed features to obtain feature interaction vectors corresponding to each of the multiple to-be-processed features. The data processing device 101 then performs click-through rate estimation processing on the feature interaction vectors corresponding to each to-be-processed feature to obtain an estimated click-through rate of the target object for the target resource. Furthermore, the data processing device 101 can determine whether to push the target resource to the target object based on the estimated click-through rate. If the data processing device 101 determines that the target resource should be pushed to the target object, the data processing device 101 can send the target resource to the terminal device 102 so that the target object can click to view the pushed target resource. For example, if the target resource is an advertisement resource to be pushed in the target application, the estimated click-through rate of each object in the target application for the advertisement resource can be used to determine which objects in the target application the advertisement resource should be pushed to, thereby increasing the exposure rate of the advertisement resource among the pushed objects. For example, in a smart transportation scenario, the terminal device 102 may be an in-vehicle terminal, and the target application may be an application for pushing traffic information. If the target resource is traffic information to be pushed, the data processing device 101 may determine whether to push the traffic information to the in-vehicle terminal based on the target subject's estimated click-through rate for the traffic information, thereby enabling the target subject to understand the traffic information and improve their driving experience. Furthermore, the data processing device 101 may determine which traffic information to push to the target subject based on the target subject's estimated click-through rate for various traffic information. For example, the traffic information with the highest estimated click-through rate may be pushed to the target subject to improve the driving experience.
[0038] In one embodiment, the object features of the target object included in the multiple features to be processed can be extracted by the data processing device 101 based on the original object feature data generated by the target object in the target application obtained from the terminal device 102, that is, the terminal device 102 can send the original object feature data generated by the target object in the target application to the data processing device 101, and the data processing device 101 can extract the object features from the original object feature data.
[0039] In this application, when it comes to data generated by the target object in the target application, such as the original object feature data generated by the target object in the target application, when the embodiments of this application are applied to specific products or technologies, they are all subject to user permission or consent, and the extraction, use and processing of relevant data comply with local laws and regulations. For example, before obtaining the data generated by the target object in the target application, such as the original object feature data generated by the target object in the target application, the data processing device 101 can send an authorization agreement for obtaining the data generated by the target object in the target application to the terminal device 102 of the target object; when the target object agrees to the authorization agreement, the data processing device 101 can obtain the data generated by the target object in the target application, otherwise, the data processing device 101 cannot obtain the data generated by the target object in the target application.
[0040] Based on the above data processing solution, the present application embodiment provides a data processing method. Figure 2 , which is a flow chart of a data processing method provided in an embodiment of the present application. Figure 2 The data processing method shown can be executed by a data processing device. Figure 2 The data processing method shown may include the following steps:
[0041] S201, obtaining multiple features to be processed.
[0042] The multiple features to be processed include object features of a target object, which can be any user using the services provided by the target application; the multiple features to be processed also include resource features of a target resource, which can be any resource provided by the target application; further, the target resource can be a resource to be pushed to the target object, that is, the multiple features to be processed can include object features of the target object and resource features of the target resource to be pushed to the target object. For example, if the target application is a video playback application, the target resource can be a movie to be pushed to the target object by the target application; the click-through rate of the target object clicking on the target resource is used to indicate the probability of the target object clicking on the target resource.
[0043] In one embodiment, the resource characteristics of the target resource are used to describe the target resource, for example, they can be used to describe the resource identifier of the target resource, the name of the target resource, the resource type of the target resource, etc. The object characteristics of the target object are used to describe the target object, for example, they can be used to describe the user identification (UID) of the target object, the nickname of the target object, the age of the target object, the city where the target object resides, etc. Furthermore, the object characteristics of the target object can also be used to describe the context in which the target object is located; for example, they can be used to describe the page information of the target page currently located in the target application, for example, they can be used to describe the page identifier of the target page currently located, the page type of the target page, etc.; for example, they can also be used to describe the access information generated by the target object in the target application before accessing the target page, for example, they can include the page information of a preset number of pages visited by the target object before accessing the target page, as well as the access order when accessing each page, etc. Among them, the preset number is set in advance according to specific needs. For example, it can be the page information of all pages visited by the target object before accessing the target page during a single use of the target application; it can also be the page information of multiple pages visited by the target object before accessing the target page during a single use of the target application, etc.; the single use process of the target object in the target application refers to the process from the target object opening the target application to closing the target application.
[0044] In one embodiment, in different business scenarios, when realizing the click-through rate estimation of the target object clicking on the target resource, the resource characteristics of the target resource and the object characteristics of the target object may be different. The embodiment of the present application does not limit the feature content specifically included in the resource characteristics of the target resource and the object characteristics of the target object. As long as the click-through rate estimation of the target object clicking on the target resource is realized based on the method provided in the embodiment of the present application, it is within the protection scope of the embodiment of the present application. For example, in a business scenario of movie recommendation, it is necessary to estimate the click-through rate of the target object clicking on the target movie (i.e., the target resource) based on the historical rating of the target object for the movie, that is, it is necessary to estimate the probability of the target object clicking on the target movie based on the historical rating of the target object for the movie; then, the resource characteristics of the target resource can be used to indicate the movie identification of the target movie, the movie name of the target movie, the movie type of the target movie, etc.; the object characteristics of the target object can be used to describe the UID of the target object, the nickname of the target object, the age of the target object, the resident city of the target object, etc.; further, it can be used to describe the movie information of the rated movie that the target object has generated a historical rating (for example, including movie identification, movie name, movie type, etc.), and can also be used to describe the historical rating of the target object for the rated movie.
[0045] In one embodiment, the resource characteristics of the target resource can be extracted from the original resource characteristic data of the target resource, which can be uploaded to the data processing device by the business personnel of the target application; the object characteristics of the target object can be extracted from the original object characteristic data of the target object, which can be obtained by the data processing device from the terminal device of the target object. The resource characteristics of the target resource are vectorized characteristics, that is, when the data processing device extracts the resource characteristics of the target resource from the original resource characteristic data of the target resource, it is necessary to vectorize the original resource characteristic data of the target resource. For example, the original resource characteristic data of the target resource indicates that the resource type of the target resource is resource type 1, that is, the original resource characteristic data of the target resource includes "resource type 1". If the resource type of the resource includes 3 categories, and the corresponding relationships between them and the vector elements are {resource type 1, resource type 2, resource type 3} respectively, then the resource characteristics corresponding to the original resource characteristic data can be expressed as {1, 0, 0}. The object features of the target object are features represented by vectorization, that is, when the data processing device extracts the object features of the target object from the original object feature data of the target object, it is necessary to vectorize the original object feature data of the target object.
[0046] S202 , performing outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtaining a feature interaction vector corresponding to each of the multiple features to be processed.
[0047] Among them, the feature interaction vector corresponding to each feature to be processed is the high-order semantic feature of each feature to be processed.
[0048] In one embodiment, the number of the multiple features to be processed is N, where N is an integer greater than 1; the data processing device performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed to obtain a feature interaction vector corresponding to each of the multiple features to be processed, which may include: traversing the N features to be processed, and in the target projection space corresponding to the nth feature to be processed among the N features to be processed, performing outer product processing on each feature to be processed and the nth feature to be processed to perform feature interaction on each feature to be processed and the nth feature to be processed in the target projection space to obtain a feature interaction vector corresponding to the nth feature to be processed, where n is a positive integer less than or equal to N. Figure 3As shown, it is a schematic diagram of obtaining a feature interaction vector corresponding to each feature to be processed provided in an embodiment of the present application. Assume that N=5, that is, the number of features to be processed is 5, and the features to be processed are feature 1 to be processed, feature 2 to be processed, feature 3 to be processed, feature 4 to be processed and feature 5 to be processed. As shown in the mark 301, it is the interaction process between each feature to be processed and each feature to be processed, and the feature interaction vector corresponding to feature 1 to be processed, the feature interaction vector corresponding to feature 2 to be processed, the feature interaction vector corresponding to feature 3 to be processed, the feature interaction vector corresponding to feature 4 to be processed and the feature interaction vector corresponding to feature 5 to be processed are obtained respectively.
[0049] Among them, each of the N features to be processed corresponds to a projection space, and the projection space corresponding to the feature to be processed is used to perform feature interaction between each feature to be processed and the feature to be processed. A set of projection matrices and a set of scaling weights are maintained based on the projection space corresponding to the feature to be processed. The set of projection matrices maintained based on the projection space corresponding to the feature to be processed includes: multiple projection matrices required for each feature to be processed in multiple features to be processed to perform feature interaction with the feature to be processed to generate a feature interaction vector corresponding to the feature to be processed, and the number of projection matrices is the same as the number of features to be processed; the set of scaling weights maintained based on the projection space corresponding to the feature to be processed includes: multiple scaling weights required for each feature to be processed in multiple features to be processed to perform feature interaction with the feature to be processed to generate a feature interaction vector corresponding to the feature to be processed, and the number of scaling weights is the same as the number of features to be processed.
[0050] Taking the nth feature to be processed among N features to be processed as an example, a set of projection matrices and a set of scaling weights are maintained based on the target projection space corresponding to the nth feature to be processed. The set of projection matrices maintained based on the target projection space corresponding to the nth feature to be processed includes: each feature to be processed among the N features to be processed performs feature interaction with the nth feature to be processed to generate the N projection matrices required for the feature interaction vector corresponding to the nth feature to be processed; the set of scaling weights maintained based on the projection space corresponding to the nth feature to be processed includes: each feature to be processed among the N features to be processed performs feature interaction with the nth feature to be processed to generate the N scaling weights required for the feature interaction vector corresponding to the nth feature to be processed. That is, a set of projection matrices maintained based on the target projection space corresponding to the nth feature to be processed includes N projection matrices, and a set of scaling weights maintained includes N scaling weights; the i-th projection matrix and the i-th scaling weight in the N projection matrices are used to adjust the result obtained by the feature interaction between the i-th feature to be processed and the n-th feature to be processed, and then, based on each projection matrix and each scaling weight, the result obtained by the feature interaction between each feature to be processed and the n-th feature to be processed is adjusted to obtain the feature interaction vector corresponding to the n-th feature to be processed. Furthermore, the total number of projection matrices maintained based on the N projection spaces corresponding to the N features to be processed is N*N, and the number of matrices of scaling weights maintained is N*N.
[0051] In one embodiment, F (*) Indicates a feature to be processed, that is, F (n) Indicates the nth feature to be processed among N features to be processed, and F (i) Represents the i-th feature to be processed among N features to be processed; it can be used It represents the projection matrix corresponding to the interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed, and is expressed as Indicates the scaling weight corresponding to the interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed. For example, Figure 4 As shown in the figure, a schematic diagram of the corresponding relationship between a projection matrix and scaling weights and features to be processed provided by an embodiment of the present application, assuming that N=5, taking n=2 as an example, in the feature space corresponding to the second feature to be processed, the projection matrices corresponding to the feature interaction between each feature to be processed in the N features to be processed and the second feature to be processed are respectively The scaling weights are
[0052] In one embodiment, the dimension of the projection matrix corresponding to the feature interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed is d i *d n , that is, the projection matrix is a d i OK, d n A matrix of columns; where d i is the dimension of the embedding vector corresponding to the i-th feature to be processed, d n is the dimension of the embedding vector corresponding to the nth feature to be processed; the embedding vector corresponding to the ith feature to be processed is obtained by performing feature embedding processing on the ith feature to be processed, and the embedding vector corresponding to the nth feature to be processed is obtained by performing feature embedding processing on the nth feature to be processed; further, the embedding vector corresponding to any feature to be processed among the N features to be processed can be obtained by performing feature embedding processing on any feature to be processed; performing feature embedding processing on the features to be processed can convert high-dimensional sparse vectors into dense vectors, realize dimensionality reduction processing of the features to be processed, save processing resources, transform the processing of the features to be processed into the processing of the embedding vectors corresponding to the features to be processed, and reduce the data in the processing process.
[0053] In a specific implementation, taking the example of obtaining the feature interaction vector corresponding to the nth feature to be processed among N features to be processed, the data processing device performs outer product processing on each feature to be processed and the nth feature to be processed in the target projection space corresponding to the nth feature to be processed among the N features to be processed, so as to perform feature interaction between each feature to be processed and the nth feature to be processed in the target projection space to obtain the feature interaction vector corresponding to the nth feature to be processed. This may include: obtaining the outer product of the embedding vector corresponding to the i-th feature to be processed among the N features to be processed and the embedding vector corresponding to the n-th feature to be processed in the target projection space to obtain a first processing result; adjusting the first processing result to obtain a feature interaction sub-vector of the feature interaction between the i-th feature to be processed and the n-th feature to be processed; and summing the feature interaction sub-vectors of the feature interaction between each feature to be processed and the n-th feature to be processed to obtain a feature interaction vector corresponding to the n-th feature to be processed. Among them, the embedding vector corresponding to any of the N features to be processed is obtained by performing feature embedding processing on any of the features to be processed, and any feature to be processed is the i-th feature to be processed or the n-th feature to be processed, and i is a positive integer less than or equal to N; the first processing result is adjusted, and the feature interaction subvector obtained by the i-th feature to be processed and the n-th feature to be processed for feature interaction is the target projection matrix (i.e., ) and the target scaling weight (i.e. ) obtained.
[0054] In a specific implementation, the data processing device adjusts the first processing result to obtain a feature interaction sub-vector for feature interaction between the i-th feature to be processed and the n-th feature to be processed, which may include: determining the target projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed based on the correspondence between the feature to be processed, the feature to be processed used for feature interaction with the feature to be processed, and the projection matrix; determining the target scaling weight corresponding to the i-th feature to be processed and the n-th feature to be processed based on the correspondence between the feature to be processed, the feature to be processed used for feature interaction with the feature to be processed, and the scaling weight; performing matrix adjustment processing on the first processing result based on the target projection matrix and the target scaling weight to obtain a second processing result; and performing vector conversion processing on the second processing result to obtain a feature interaction sub-vector for feature interaction between the i-th feature to be processed and the n-th feature to be processed.
[0055] Among them, the data processing device determines the target projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed based on the correspondence between the feature to be processed, the feature to be processed for feature interaction with the feature to be processed, and the projection matrix: the projection matrix corresponding to the feature interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space, that is, Based on the correspondence between the feature to be processed, the feature to be processed for feature interaction with the feature to be processed, and the scaling weight, the target scaling weight corresponding to the i-th feature to be processed and the n-th feature to be processed is determined as follows: the scaling weight corresponding to the feature interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space, that is,
[0056] In one embodiment, the data processing device obtains a feature interaction subvector for the feature interaction between the i-th feature to be processed and the n-th feature to be processed based on the target projection space corresponding to the n-th feature to be processed, which can be given by Formula 1:
[0057]
[0058] Among them, v i Represents the i-th feature to be processed F (i) The corresponding embedding vector, v n Indicates the nth feature to be processed F (n) The corresponding embedding vector; Represents the i-th feature to be processed F (i) The corresponding embedding vector v i and the nth feature to be processed F (n) The corresponding embedding vector vn The outer product between them is the first processing result; d i is the i-th feature to be processed F (i) The corresponding embedding vector v i Dimensions, Indicates dimension is 1*d i A vector whose elements are all 1. Since the i-th feature to be processed F (i) The corresponding embedding vector v i The dimension is d i , the nth feature to be processed F (n) The corresponding embedding vector v n The dimension is d n ; Then, the i-th feature to be processed F (i) The corresponding embedding vector v i and the nth feature to be processed F (n) The corresponding embedding vector v (n) The outer product between is a dimension d i *d n ⊙ represents the multiplication of the corresponding elements between the two matrices, that is, Represents the projection matrix With the matrix The corresponding elements are multiplied, because the projection matrix The dimension is d i *d n , the scaling weight is a scalar parameter, so the second processing result is For a dimension d i *d n Matrix. Using a dimension of 1*d i The all-1 vector (i.e. ) is to adjust the projection matrix and the scaling weight to obtain a dimension d i *d n The matrix (i.e., the second processing result) is vector-converted to obtain the dimension of the embedding vector corresponding to the nth feature to be processed (i.e., d n ) are equal to each other (i.e., the feature interaction subvector of the feature interaction between the i-th feature to be processed and the n-th feature to be processed).
[0059] In one embodiment, the data processing device sums the feature interaction subvectors of the feature interaction between each feature to be processed and the nth feature to be processed, obtained in the target projection space corresponding to the nth feature to be processed, to obtain a feature interaction vector corresponding to the nth feature to be processed. The feature interaction vector corresponding to the nth feature to be processed can be given by Formula 2:
[0060]
[0061] Among them, φ n (v) represents the feature interaction vector corresponding to the nth feature to be processed.
[0062] In one embodiment, the data processing device obtains the feature interaction vector corresponding to each feature to be processed based on outer product processing of multiple features to be processed, and can obtain the high-order semantic space corresponding to each feature to be processed. The high-order semantic characteristics of each feature to be processed can be fully learned, so that the feature interaction vector corresponding to each feature to be processed obtained based on the outer product processing has the high-order semantic characteristics of each feature to be processed, thereby improving the accuracy of click-through rate prediction.
[0063] In another embodiment, taking the example of obtaining a feature interaction vector corresponding to the nth feature to be processed among N features to be processed, the data processing device performs outer product processing on each feature to be processed and the nth feature to be processed in the target projection space corresponding to the nth feature to be processed among the N features to be processed, so as to perform feature interaction between each feature to be processed and the nth feature to be processed in the target projection space to obtain a feature interaction vector corresponding to the nth feature to be processed, which may include: performing vector conversion processing on each feature to be processed among the N features to be processed and the scaling weight corresponding to the nth feature to be processed to obtain N scaling weight vectors, and splicing the N scaling weight vectors to obtain a spliced scaling weight vector; using the spliced scaling weight vector to re-adjust the spliced embedding vector obtained by splicing the embedding vectors corresponding to each feature to be processed to obtain a third processing result; performing matrix adjustment processing on the third processing result based on the splicing projection matrix to obtain an implicit vector; performing vector element interaction processing on the embedding vector corresponding to the nth feature to be processed and the implicit vector to obtain a feature interaction vector corresponding to the nth feature to be processed. Among them, the scaling weights corresponding to each feature to be processed and the nth feature to be processed are: the scaling weights corresponding to the feature interaction between each feature to be processed and the nth feature to be processed in the target projection space; taking the i-th feature to be processed among N features to be processed as an example, the scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed are: the scaling weights corresponding to the feature interaction between the i-th feature to be processed and the n-th feature to be processed in the target projection space, that is, The embedding vector corresponding to each feature to be processed is obtained by performing feature embedding processing on each feature to be processed; the splicing projection matrix is: the matrix obtained by splicing each feature to be processed and the projection matrix corresponding to the nth feature to be processed, that is, the matrix obtained by splicing the projection matrices corresponding to the feature interaction between each feature to be processed in the target projection space and the nth feature to be processed.
[0064] In a specific implementation, taking the i-th feature to be processed among N features to be processed as an example, when the data processing device performs vector conversion processing on the i-th feature to be processed among N features to be processed and the scaling weight corresponding to the n-th feature to be processed, it can be based on a dimension of 1*d i The all-1 vector (i.e. ) Perform vector conversion processing on the scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed to obtain a scaling weight vector, the vector elements of which are all the scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed (i.e. ), the dimension of the scaling weight vector is 1*d i ; That is, the scaling weight vector is a 1-row, d i Column, vector elements are Furthermore, the data processing device performs splicing processing on the N scaling weight vectors, and the obtained splicing scaling weight vector can be given by Formula 3:
[0065]
[0066] in, represents the splicing scaling weight vector; Indicates that the scaling weights corresponding to the first feature to be processed and the nth feature to be processed among the N features to be processed are converted into vectors, and the resulting scaling weight vector has a dimension of d1. Indicates that the scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed among the N features to be processed are vectorized, and the resulting scaling weight vector has a dimension of d i , Indicates that the scaling weight corresponding to the Nth feature to be processed and the nth feature to be processed are vectorized and converted to obtain a scaling weight vector with a dimension of d. N ; Since the dimension of the embedding vector corresponding to the i-th feature to be processed is d i , so the dimension of the spliced scaling weight vector is the same as the dimension of the spliced embedding vector obtained by splicing the embedding vectors corresponding to the features to be processed, which is d. ...represents a concatenation operation.
[0067] In one embodiment, the concatenated embedding vector obtained by concatenating the embedding vectors corresponding to the features to be processed can be given by Formula 4:
[0068] v=[v1,…,v i ,…,v N ] (4)
[0069] Among them, v represents the concatenated embedding vector, v1 represents the embedding vector corresponding to the first feature to be processed among N features to be processed, and v i represents the embedding vector corresponding to the i-th feature to be processed, v N represents the embedding vector corresponding to the Nth feature to be processed, ... represents the concatenation operation.
[0070] In one embodiment, the spliced projection matrix refers to a matrix obtained by splicing the projection matrices corresponding to each feature to be processed and the nth feature to be processed, that is, a matrix obtained by splicing the projection matrices corresponding to the feature interactions between each feature to be processed and the nth feature to be processed in the target projection space. The spliced projection matrix can be given by Formula 5:
[0071]
[0072] in, Represents the projection matrix corresponding to the first feature to be processed and the nth feature to be processed, that is, the projection matrix corresponding to the feature interaction between the first feature to be processed and the nth feature to be processed in the target projection space, with a dimension of d1*d n ; Represents the projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed, with a dimension of d i *d n ; Represents the Nth feature to be processed and the projection matrix corresponding to the nth feature to be processed, with a dimension of d N *d n , Represents the splicing projection matrix, dimension is d*d n ,…represents the splicing operation.
[0073] In one embodiment, matrix adjustment processing is performed on the third processing result based on the splicing projection matrix, and the obtained implicit vector can be given by Formula 6:
[0074]
[0075] Among them, ⊙ represents the corresponding element-wise multiplication between two matrices or vectors, v represents the concatenated embedding vector, represents the splicing scaling weight vector, Indicates the third processing result; represents the splicing projection matrix; g n Represents the latent vector, the dimension of the latent vector is the same as the dimension of the embedding vector corresponding to the nth feature to be processed.
[0076] Furthermore, the feature interaction vector corresponding to the nth feature to be processed can be given by Formula 7:
[0077]
[0078] Among them, ⊙ represents the corresponding element multiplication between two matrices or vectors, φ n (v) represents the feature interaction vector corresponding to the nth feature to be processed, v n Represents the embedding vector corresponding to the nth feature to be processed.
[0079] See also Figure 5 , which is a schematic diagram of obtaining a feature interaction vector corresponding to the nth feature to be processed provided by an embodiment of the present application; wherein the concatenated embedding vector can be as shown by 501, and as shown by 502, it is the first vector element (expressed as) in the embedding vector corresponding to the i-th feature to be processed. ); re-adjust the splicing embedding vector using the splicing scaling weight vector as shown in 503 to obtain the third processing result. The first vector element in the i-th scaling weight vector can be as shown in 504 (represented as ); Based on the splicing projection matrix as shown in the mark 505, the third processing result is subjected to matrix adjustment processing to obtain the implicit vector as shown in the mark 506. The dimension of the implicit vector is the same as the dimension of the embedding vector corresponding to the nth feature to be processed. The first vector element in the implicit vector can be expressed as As shown in mark 507, the embedding vector corresponding to the n-th feature to be processed as shown in mark 508 is subjected to vector element interaction processing with the latent vector as shown in mark 506 to obtain a feature interaction vector corresponding to the n-th feature to be processed, as shown in mark 509.
[0080] In one embodiment, the number of multiple features to be processed is N, and there are H channels in the target projection space corresponding to the nth feature to be processed among the N features to be processed, where N is an integer greater than 1, H is a positive integer, and n is a positive integer less than or equal to N; the data processing device performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed to obtain a feature interaction vector corresponding to each feature to be processed among the multiple features to be processed, which may include: traversing the N features to be processed, and performing outer product processing on each feature to be processed and the nth feature to be processed in each channel in the target projection space to perform feature interaction on each feature to be processed and the nth feature to be processed in each channel in the target projection space to obtain a feature interaction vector corresponding to the nth feature to be processed in each channel; and combining the feature interaction vectors corresponding to the nth feature to be processed in each channel to obtain a feature interaction vector corresponding to the nth feature to be processed. Since the feature interaction vector corresponding to the nth feature to be processed is obtained by combining the feature interaction vectors corresponding to the nth feature to be processed in each channel, the number of channels can be set according to different needs. Generally speaking, when H is set to a value in the range of [2,6], the feature interaction vectors corresponding to each feature to be processed are relatively better, that is, the high-order semantic features of each feature to be processed can be more fully learned. By setting H channels in the projection space corresponding to each feature to be processed, and for each projection space corresponding to each feature to be processed, the feature interaction vector corresponding to each projection space (i.e., the feature interaction vector corresponding to each feature to be processed) is calculated based on H groups of projection matrices and H groups of scaling weights maintained by the H channels in each projection space; the ambiguity of each feature to be processed can be fully learned, thereby increasing the accuracy of click-through rate estimation.
[0081] Among them, H groups of projection matrices and H groups of scaling weights are maintained based on the target projection space corresponding to the n-th feature to be processed, that is, one group of projection matrices and one group of scaling weights are maintained for each of the H channels in the target projection space corresponding to the n-th feature to be processed; the one group of projection matrices and one group of scaling weights maintained for each of the H channels in the target projection space corresponding to the n-th feature to be processed have similar characteristics to the one group of projection matrices and one group of scaling weights maintained based on the target projection space corresponding to the n-th feature to be processed. In a specific implementation, the data processing device performs outer product processing on each feature to be processed and the nth feature to be processed in each channel in the target projection space, so as to perform feature interaction between each feature to be processed and the nth feature to be processed in each channel in the target projection space, and obtains a feature interaction vector corresponding to the nth feature to be processed in each channel. This is similar to the above-mentioned process of performing outer product processing on each feature to be processed and the nth feature to be processed in the target projection space corresponding to the nth feature to be processed among N features to be processed, so as to perform feature interaction between each feature to be processed and the nth feature to be processed in the target projection space, and obtain a feature interaction vector corresponding to the nth feature to be processed, and will not be repeated here.
[0082] Furthermore, the data processing device combines and processes the feature interaction vectors corresponding to the nth feature to be processed in each channel to obtain the feature interaction vector corresponding to the nth feature to be processed. The combination function can be any function that can realize the combination of multiple vectors. For example, it can be: a function that indicates that the feature interaction vectors corresponding to the nth feature to be processed in each channel are summed element by element; that is, the first vector element of the feature interaction vector corresponding to the nth feature to be processed in each channel can be summed to obtain the first vector element of the feature interaction vector corresponding to the nth feature to be processed; the second vector element of the feature interaction vector corresponding to the nth feature to be processed in each channel is summed to obtain the second vector element of the feature interaction vector corresponding to the nth feature to be processed; each vector element of the feature interaction vector corresponding to the nth feature to be processed is obtained based on this method, and then the feature interaction vector corresponding to the nth feature to be processed is obtained. For example, it can be: a function that indicates that the feature interaction vector corresponding to the nth feature to be processed in each channel is averaged element by element; that is, the first vector element of the feature interaction vector corresponding to the nth feature to be processed in each channel can be summed and then averaged to obtain the first vector element of the feature interaction vector corresponding to the nth feature to be processed; the second vector element of the feature interaction vector corresponding to the nth feature to be processed in each channel is summed and then averaged to obtain the second vector element of the feature interaction vector corresponding to the nth feature to be processed; each vector element of the feature interaction vector corresponding to the nth feature to be processed is obtained based on this method, and then the feature interaction vector corresponding to the nth feature to be processed is obtained.
[0083] In one embodiment, if in the hth channel in the target projection space, each feature to be processed is processed with the nth feature to be processed by outer product processing, the feature interaction vector corresponding to the nth feature to be processed in the hth channel is expressed as Then the feature interaction vector corresponding to the nth feature to be processed can be given by Formula 8:
[0084]
[0085] Among them, φ n (v) represents the feature interaction vector corresponding to the nth feature to be processed, f represents the combination function, Represents the feature interaction vector corresponding to the nth feature to be processed in the first channel, The feature interaction vector corresponding to the nth feature to be processed in the Hth channel.
[0086] In one embodiment, the data processing method provided by the embodiment of the present application can be implemented based on a data processing model, which can include an embedding layer, a feature interaction layer, a forward fully connected layer, and a normalization layer; the parameters used to implement the data processing method provided by the embodiment of the present application are model parameters of the data processing model, for example, the projection matrix and the scaling weight are both model parameters of the data processing model, and the data processing model can be obtained based on training the initial data processing model. Figure 6 As shown, a schematic diagram of click-through rate estimation based on a data processing model provided in an embodiment of the present application is provided. The data processing device can perform outer product processing on multiple features to be processed through the embedding layer and the feature interaction layer in the data processing model to perform feature interaction on the multiple features to be processed, and obtain a feature interaction vector corresponding to each of the multiple features to be processed; the feature interaction vector corresponding to each feature to be processed is processed through the forward fully connected layer and the normalization layer in the data processing model to obtain the estimated click rate of the target object for the target resource.
[0087] In a specific implementation, when the data processing device performs outer product processing on a plurality of features to be processed through the embedding layer and the feature interaction layer in the data processing model, and then obtains the feature interaction vector corresponding to each of the plurality of features to be processed, the embedding layer in the data processing model can be used to perform feature embedding processing on each feature to be processed to obtain the embedding vector corresponding to each feature to be processed; then, the feature interaction layer in the data processing model can be used to obtain the feature interaction vector corresponding to each feature to be processed based on the embedding vector corresponding to each feature to be processed. The specific process of obtaining the feature interaction vector corresponding to each feature to be processed has been described in detail above and will not be repeated here. Furthermore, the data processing device performs click-through rate estimation processing on the feature interaction vector corresponding to each feature to be processed through the forward fully connected layer and the normalization layer in the data processing model to obtain the estimated click-through rate of the target object for the target resource, that is, to implement the relevant process of step S203.
[0088] S203 , performing click-through rate estimation processing on the feature interaction vector corresponding to each feature to be processed to obtain an estimated click-through rate of the target object for the target resource.
[0089] The estimated click-through rate is used to indicate the probability that the target object clicks the target resource.
[0090] In one embodiment, the data processing device performs click-through rate estimation processing on the feature interaction vectors corresponding to each feature to be processed to obtain the estimated click-through rate of the target object for the target resource, which may include: performing full connection processing on the feature interaction vectors corresponding to each feature to be processed to obtain prediction vectors corresponding to multiple features to be processed; and normalizing the prediction vectors corresponding to multiple features to be processed to obtain the estimated click-through rate of the target object for the target resource. Wherein, the data processing device performs full connection processing on the feature interaction vectors corresponding to each feature to be processed to obtain prediction vectors corresponding to multiple features to be processed, which may be achieved through a forward fully connected layer in the data processing model, and the data processing device performs normalization processing on the prediction vectors corresponding to multiple features to be processed to obtain the estimated click-through rate of the target object for the target resource, which may be achieved through a normalization layer in the data processing model. Wherein, the activation function of the forward fully connected layer is a sigmoid function, which can map the prediction vectors corresponding to multiple features to be processed to a probability value in the range of [0,1].
[0091] In one embodiment, click-through rate estimation processing is performed on the feature interaction vector corresponding to each feature to be processed, and the estimated click-through rate of the target object to the target resource can be given by Formula 9:
[0092] y=σ(w([φ 1 (v),…,φ n (v),…,φ N (v)]+b) (9)
[0093] Among them, φ 1 (v) represents the feature interaction vector corresponding to the first feature to be processed, φ n (v) represents the feature interaction vector corresponding to the nth feature to be processed, φ N (v) represents the feature interaction vector corresponding to the Nth feature to be processed; w represents the weight parameter in the forward fully connected layer, b represents the bias term in the forward fully connected layer, represents the activation function of the normalization layer of the data processing model, and y represents the estimated click rate of the target object to the target resource.
[0094] In one embodiment, the data processing device can determine whether to push the target resource to the target object based on the estimated click-through rate; if the data processing device determines that the target resource needs to be pushed to the target object, the target resource can be sent to the terminal device of the target object so that the target object can click to view the pushed target resource.
[0095] In an embodiment of the present application, after obtaining multiple features to be processed, including resource features of a target resource and object features of a target object, the data processing device can perform outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, thereby obtaining a feature interaction vector corresponding to each of the multiple features to be processed; and perform click-through rate estimation processing on the feature interaction vector corresponding to each feature to be processed to obtain an estimated click-through rate of the target object for the target resource; wherein the estimated click-through rate is used to indicate the probability of the target object clicking on the target resource. By performing outer product processing on multiple features to be processed, high-order semantic features of each feature to be processed can be learned, thereby improving the accuracy of the target object's click-through rate estimation for the target resource and efficiently obtaining information of interest.
[0096] The above data processing method can be implemented through a data processing model. The data processing model can be obtained by training an initial data processing model. The initial data processing model includes an embedding layer, a feature interaction layer, a forward fully connected layer, and a normalization layer. The model parameters of the initial data processing model and the data processing model are different. Based on this, the embodiment of the present application provides a method for training a data processing model. See Figure 7 , which is a flow chart of a training method for a data processing model provided in an embodiment of the present application. Figure 7 The training method of the data processing model shown can be executed by a data processing device, or by any other electronic device that can implement data processing model training. The embodiment of the present application is explained using a data processing device as an example. Figure 7 The training method of the data processing model shown may include the following steps:
[0097] S701: Obtain training samples.
[0098] The training sample includes multiple training features and sample labels; the multiple training features include resource features of the training resource and object features of the training object. The sample label indicates whether the training object clicked on the training resource. The resource features of the training resource and object features of the training object included in the multiple training features are identical to the resource features of the target resource and object features of the target object included in the multiple features to be processed, and are not further described here.
[0099] In one embodiment, different data processing models can be trained for training samples in different business scenarios to achieve the prediction of the click-through rate of target objects clicking on target resources in different business scenarios. For example, in a movie recommendation business scenario, it is necessary to estimate the click-through rate of the target object clicking on the target movie (i.e., the target resource) based on the target object's historical rating of the movie, that is, it is necessary to estimate the probability of the target object clicking on the target movie based on the target object's historical rating of the movie; then, it is necessary to train the initial data processing model based on the training object's historical rating of the movie to obtain the data processing model. Based on this, when training the initial data processing model in the movie recommendation business scenario, the resource features of the training resource can be used to indicate the movie ID of the training movie, the movie name of the training movie, the movie type of the training movie, etc.; the object features of the training object can be used to describe the UID of the training object, the nickname of the training object, the age of the training object, the resident city of the training object, etc.; further, it can be used to describe the movie information of the rated movie for which the training object has generated historical ratings (for example, including movie ID, movie name, movie type, etc.), and can also be used to describe the historical ratings of the rated movie by the training object.
[0100] In one embodiment, the training samples may be extracted from historical data generated by users of the target application in the target application; the training object may be any user who is different from the target object and uses the services provided by the target application, and the training resources may be any resources that the target application can provide, such as resources that the target application has pushed to users of the target application. In another embodiment, the training samples may be obtained from commonly used data sets for model training in the field of click-through rate prediction; further, when training data processing models for implementing click-through rate prediction in different business scenarios, the training samples may be obtained from commonly used data sets for model training in the corresponding business scenarios in the field of click-through rate prediction; for example, if in a movie recommendation business scenario, it is necessary to train the initial data processing model based on the historical ratings of the training object for the movie, and obtain a model that can predict the target object's click on the target movie (i.e., the target object's historical ratings of the movie) based on the target object's historical ratings of the movie. A data processing model is used to estimate the click-through rate of a target resource. Training samples can be obtained from the MovieLens dataset, a common dataset used for model training in the movie recommendation scenario in the field of click-through rate estimation. A data item in the MovieLens dataset records the historical ratings of the rated movies by the object, which are rated according to a 5-star system and incremented by half a star. That is, the historical ratings of a rated movie can only be 0.5 stars, 1 star, 1.5 stars, 2 stars, 2.5 stars, 3 stars, 3.5 stars, 4 stars, 4.5 stars, and 5 stars.
[0101] In one embodiment, if a data processing model is trained for achieving click-through rate prediction in the context of a movie recommendation business scenario, when a data processing device obtains training samples from the MovieLens dataset, it can first determine the training data from the dataset; then, it can preprocess the training data to obtain the training samples. In a specific implementation, the data processing device can determine the training data from the dataset based on random sampling; further, the data processing device can randomly extract a certain amount of data from the dataset as training data, for example, randomly extracting 80% of the data from the dataset as training data, randomly extracting 10% of the data as verification data, and randomly extracting 10% of the data as test data, wherein the training data, verification data, and test data are all different.
[0102] Furthermore, when the data processing device preprocesses the training data to obtain training samples, it can determine the sample label of the training sample determined by the training data based on the historical rating of the training resource by the training subject corresponding to the training data. Specifically, if the historical rating of the training resource by the training subject is greater than or equal to a preset rating threshold, the sample label of the training sample is determined to indicate that the training subject clicked on the training sample, that is, the training sample is a positive sample; if the historical rating of the training resource by the training subject is less than the preset rating threshold, the sample label of the training sample is determined to indicate that the training subject did not click on the training sample, that is, the training sample is a negative sample. For example, the preset rating threshold can be 3 stars.
[0103] In one embodiment, when the data processing device pre-processes the training data and obtains the object characteristics of the training objects in the training samples and the resource characteristics of the training resources, it is necessary to vectorize the training data. Furthermore, when the data processing device vectorizes the training data, it can first standardize the numerical data included in the training data, and then vectorize the data after the standardization. Optionally, the data processing device can use a series of standardization methods to standardize the numerical data. For example, the mean-variance standardization method can be used to standardize the numerical data, or the logarithm-based standardization method can be used to standardize the numerical data, and so on. For example, when the logarithm-based standardization method is used, the numerical data greater than 2 can be standardized so that the data after the standardization is the result of taking the logarithm of the numerical data greater than 2 with a base of 2 (i.e., a2=log2(a1), where a1 represents the numerical data greater than 2 and a2 is the data after the standardization).
[0104] See also Figure 8, is a schematic diagram of the overall process of training an initial image processing model provided in an embodiment of the present application, wherein the data processing device determines training data from a data set used for model training, pre-processes the training data, and obtains training samples; initializes the model parameters of the initial data processing model through a hypercomplex product parameterization strategy; performs outer product processing on multiple training features included in the training sample through the initial data processing model to perform feature interaction on the multiple training features, and obtains a feature interaction vector corresponding to each of the multiple training features; performs click-through rate estimation processing on the feature interaction vector corresponding to each training feature, and obtains the training estimated click-through rate of the training object for the training resource. Among them, when the data processing device performs outer product processing on multiple training features included in the training sample through the initial data processing model to perform feature interaction on the multiple training features, and obtains a feature interaction vector corresponding to each of the multiple training features, the model parameters used are the model parameters obtained by initializing the model parameters of the initial data processing model through the hypercomplex product parameterization strategy.
[0105] S702 : Initializing model parameters of the initial data processing model through a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the plurality of training features.
[0106] Among them, the embedding vector corresponding to each training feature is obtained by performing feature embedding processing on each training feature.
[0107] In one embodiment, the number of multiple training features is N, and the model parameters of the initial data processing model include a training projection matrix, where N is an integer greater than 1; the data processing device can train the model parameters of the initial data processing model based on the training samples, and construct the data processing model based on the trained model parameters; that is, the model parameters in the data processing model are obtained based on the training of the model parameters in the initial data processing model, and the projection matrix in the data processing model is obtained based on the training of the training projection matrix in the initial data processing model. Therefore, a projection matrix included in the model parameters of the data processing model has a similar data structure and characteristics to a projection matrix included in the initial data processing model. Since the projection matrix in the model parameters of the data processing model is introduced in detail above, the data structure and characteristics of the training projection matrix will not be repeated here. In a specific implementation, the data processing device initializes the model parameters of the initial data processing model through a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in a plurality of training features, which may include: traversing N training features, and initializing the splicing training projection matrix in the model parameters through a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature in the N training features and the dimension of the training splicing embedding vector; splitting the splicing training projection matrix to obtain N training projection matrices; wherein the training splicing embedding vector is obtained by splicing the embedding vectors corresponding to each training feature, and the N training projection matrices are used to obtain the feature interaction vector corresponding to the kth training feature.
[0108] In one embodiment, F′ (*) Represents a training feature, that is, F′ (k) Indicates the kth training feature among N training features, and F′ (j) Represents the jth training feature among N training features; it can be used It represents the training projection matrix corresponding to the feature interaction between the jth training feature and the kth training feature in the training projection space corresponding to the kth training feature, that is, Represents the training projection matrix corresponding to the j-th training feature and the k-th training feature; v′ j Denotes the embedding vector corresponding to the jth training feature, with v′ k Represents the embedding vector corresponding to the kth training feature.
[0109] Furthermore, the training concatenated embedding vector obtained by concatenating the embedding vectors corresponding to each training feature can be expressed as: v′=[v′1,…,v′ j ,…,v′ N]; where v′ represents the training concatenated embedding vector, v′1 represents the embedding vector corresponding to the first training feature among N training features, and v′ j Represents the embedding vector corresponding to the jth training feature, v′ N represents the embedding vector corresponding to the Nth training feature, ... represents the concatenation operation; the dimension of the embedding vector corresponding to the jth training feature is d j , the dimension of the embedding vector corresponding to the kth training feature is d k , the training spliced embedding vector obtained by splicing the embedding vectors corresponding to each training feature is d,
[0110] In one embodiment, the projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed in the model parameters of the data processing model is: the projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed; the projection matrix corresponding to the j-th feature to be processed and the k-th feature to be processed in the model parameters of the data processing model is: the projection matrix corresponding to the j-th feature to be processed and the k-th feature to be processed in the target projection space corresponding to the k-th feature to be processed; the projection matrix corresponding to the j-th feature to be processed and the k-th feature to be processed in the model parameters of the data processing model is obtained by training based on the training projection matrix corresponding to the j-th training feature and the k-th training feature in the initial data processing model; according to the training requirements, it is required to initialize the training projection matrix corresponding to the j-th training feature and the k-th training feature The dimension is d j *d k ; The splicing projection matrix corresponding to the target projection space corresponding to the kth feature to be processed in the model parameters of the data processing model corresponds to the splicing training projection matrix corresponding to the training projection space corresponding to the kth training feature in the model parameters of the initial data processing model, that is, the splicing training projection matrix is the matrix obtained after splicing each training feature and the training projection matrix corresponding to the kth training feature, that is, the matrix obtained after splicing each training feature in the training projection space corresponding to the kth training feature and the training projection matrix corresponding to the kth training feature when the training features interact with each other; therefore, according to the training requirements, it can be seen that the dimension of the initialized splicing training projection matrix is required to be d*d kBased on this, the data processing device can initialize the splicing training projection matrix in the model parameters through a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector; and then split the splicing training projection matrix to obtain N training projection matrices, which are used to obtain the feature interaction vector corresponding to the kth training feature.
[0111] In a specific implementation, the data processing device initializes the splicing training projection matrix in the model parameters based on the dimension of the embedding vector corresponding to the kth training feature in the N training features and the dimension of the training splicing embedding vector through the hypercomplex product parameterization strategy, which may include: based on the dimension d of the embedding vector corresponding to the kth training feature k And the dimension d of the training spliced embedding vector, initialize B projection parameter matrices; the b-th projection parameter matrix in the B projection parameter matrices is the first parameter matrix with the initialization dimension of b*b and the initialization dimension of The Kronecker product of the second parameter matrix of is calculated; the B projection parameter matrices are summed to obtain the spliced training projection matrix. B can be set based on specific requirements. Optionally, the Xavier mode can be used to initialize the model parameters.
[0112] If the first parameter matrix in the b-th projection parameter matrix is Indicates that the second parameter matrix in the b-th projection parameter matrix is Represented; then the splicing training projection matrix can be shown as formula 10:
[0113]
[0114] in, represents the splicing training projection matrix, represents the Kronecker product of the first parameter matrix in the b-th projection parameter matrix and the second parameter matrix in the b-th projection parameter matrix; since the Kronecker product of any two matrices is a block matrix, that is, the dimension of the Kronecker product of the matrix X1 with dimension x1*x2 and the matrix X2 with dimension x3*x4 is x1x3*x2x4; if x1*x2 is 3*4 and x3*x4 is 5*6, then the dimension of the Kronecker product of the matrix X1 and the matrix X2 is 15*24 (that is, (3·5)*(4·6)). Therefore, the dimension of the Kronecker product of the first parameter matrix in the b-th projection parameter matrix and the second parameter matrix in the b-th projection parameter matrix is d*d k ; The dimension of the splicing training projection matrix is d*d kThe requirements of ; Since the splicing training projection matrix can be the matrix obtained by splicing the training projection matrices corresponding to each training feature and the k-th training feature, that is Therefore, the spliced training projection matrix can be split according to the dimensions required by each training projection matrix to obtain N training projection matrices, and the training projection matrices corresponding to the j-th training feature and the k-th training feature in the model parameters of the initial data processing model can be obtained.
[0115] In one embodiment, the model parameters of the initial data processing model are initialized by a hypercomplex product parameterization strategy, so that when the initial data processing model is subsequently trained, the training projection matrix in the initial data processing model can be learned from the hypercomplex space, the interaction between the real component and the imaginary component can be learned, and the vector outer product can be generalized to a higher-dimensional real space; and, because the Kronecker product of the first parameter matrix in the b-th projection parameter matrix and the second parameter matrix in the b-th projection parameter matrix reuses the matrix parameters in the first parameter matrix in the b-th projection parameter matrix and the matrix parameters in the second parameter matrix in the b-th projection parameter matrix, the model parameters of the initial data processing model can be reduced to The number of model parameters that need to be trained can be reduced to that of conventional training methods. Since the number of model parameters used in the training process of the initial data processing model is reduced, the model parameters of the data processing model obtained based on the training of the initial data processing model are reduced, which can reduce the deployment cost of the data processing model and make the deployment of the data processing model easier; and since the model parameters of the data processing model are reduced, computing resources are saved.
[0116] In one embodiment, when the training projection space corresponding to each training feature is set to include H channels, that is, when estimating the click-through rate of the target object for the target resource, the target projection space corresponding to each feature to be processed includes H channels; the initialization of a set of training projection matrices maintained for each channel is the same as the initialization process of maintaining a set of training projection matrices for a training projection space mentioned above, and will not be repeated here; wherein, H can be set according to specific needs. If the training sample is obtained from a MovieLens dataset in a movie recommendation business scenario, then when H is set to 3, that is, the training projection space corresponding to each training feature includes 3 channels, and maintaining 3 sets of training projection matrices has a better training effect.
[0117] S703 , performing outer product processing on the multiple training features through the initial data processing model to perform feature interaction on the multiple training features, and obtaining a feature interaction vector corresponding to each training feature in the multiple training features.
[0118] Among them, the feature interaction vector corresponding to each training feature is the high-order semantic feature of each training feature.
[0119] S704 , performing click-through rate estimation processing on the feature interaction vector corresponding to each training feature to obtain the training estimated click-through rate of the training object for the training resource.
[0120] The training estimated click rate is used to indicate: the probability that the training subject clicks on the training resource;
[0121] Steps S703 to S704 are similar to the above-mentioned steps S202 to S203 and are not described again here.
[0122] S705: Based on the training estimated click rate and sample labels, the initial data processing model is trained to obtain a data processing model.
[0123] See also Figure 9 , which is a schematic diagram of a training initial image processing model provided in an embodiment of the present application. The data processing device can perform outer product processing on multiple training features through the embedding layer and the feature interaction layer in the initial data processing model to perform feature interaction on the multiple training features, and obtain a feature interaction vector corresponding to each training feature in the multiple training features; the feature interaction vector corresponding to each training feature is processed by the forward fully connected layer and the normalization layer in the initial data processing model to obtain a training estimated click rate of the training object for the training resource; the data processing device trains the initial data processing model based on the training estimated click rate and the sample label to obtain a data processing model.
[0124] In one embodiment, the data processing device trains the initial data processing model based on the training estimated click-through rate and sample labels to obtain the data processing model, which may include: determining the loss value of the loss function based on the training estimated click-through rate and sample labels; training the initial data processing model in the direction of reducing the loss value to obtain the data processing model. Optionally, when training the initial data processing model, the Adam optimization algorithm in the optimization algorithm can be used to solve the model parameters in the initial image processing model, and the error back propagation algorithm can be used to perform training optimization. When the initial training model is trained based on multiple training samples, the loss function can be given by Formula 11:
[0125]
[0126] Among them, Z is the number of training samples, z is the independent variable of the training sample; y′ z is the sample label of the zth training sample in the Z training samples, is the training estimated click rate corresponding to the zth training sample obtained based on processing the zth training sample.
[0127] In one embodiment, the model parameters of the initial data processing model can be initialized using a hypercomplex product parameterization strategy, and then the initial data processing model can be trained based on the training samples to obtain a data processing model. Optionally, the feature interaction layer in the obtained data processing model can be applied in a modular manner to other existing models for achieving click-through rate estimation, which can reduce the deployment cost of existing models that apply the feature interaction layer of the data processing model. For example, it can be applied to a field-aware factorization machine (FFM) model or a weighted field-aware factorization machine (FwFM) model.
[0128] In an embodiment of the present application, after obtaining a training sample including multiple training features and sample labels, the data processing device can initialize the model parameters of the initial data processing model through a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the multiple training features; and then train the initial data processing model based on the initialized model parameters. The adoption of the hypercomplex product parameterization strategy can learn the training projection matrix in the initial data processing model from the hypercomplex space, learn the interaction between the real component and the imaginary component, and generalize the vector outer product to a higher-dimensional real space; it can reduce the number of model parameters used in the training process of the initial data processing model, thereby reducing the model parameters of the data processing model obtained based on the training of the initial data processing model, reducing the deployment cost of the data processing model, and making the deployment of the data processing model easier; and due to the reduction of the model parameters of the data processing model, computing resources are saved.
[0129] Based on the above data processing method embodiment, the present application embodiment provides a data processing device. Figure 10 , is a structural diagram of a data processing device provided in an embodiment of the present application, and the data processing device may include an acquisition unit 1001 and a processing unit 1002. Figure 10 The data processing device shown can run the following units:
[0130] An acquiring unit 1001 is configured to acquire a plurality of features to be processed, wherein the plurality of features to be processed include object features of a target object and resource features of a target resource to be pushed to the target object;
[0131] The processing unit 1002 is configured to perform outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtain a feature interaction vector corresponding to each of the multiple features to be processed, where the feature interaction vector corresponding to each of the features to be processed is a high-order semantic feature of each of the features to be processed;
[0132] The processing unit 1002 is further used to perform click rate estimation processing on the feature interaction vectors corresponding to the various features to be processed to obtain an estimated click rate of the target object for the target resource. The estimated click rate is used to indicate: the probability of the target object clicking on the target resource.
[0133] In one embodiment, the number of the plurality of features to be processed is N, where N is an integer greater than 1;
[0134] When the processing unit 1002 performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed and obtains a feature interaction vector corresponding to each of the multiple features to be processed, the processing unit 1002 specifically performs the following operations:
[0135] The N features to be processed are traversed, and in a target projection space corresponding to an n-th feature to be processed among the N features to be processed, outer product processing is performed on each feature to be processed and the n-th feature to be processed, so as to perform feature interaction on each feature to be processed and the n-th feature to be processed in the target projection space, and a feature interaction vector corresponding to the n-th feature to be processed is obtained, where n is a positive integer less than or equal to N.
[0136] In one embodiment, the processing unit 1002 performs outer product processing on each of the N features to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed, so as to perform feature interaction between each of the features to be processed and the n-th feature to be processed in the target projection space, and specifically performs the following operations when obtaining a feature interaction vector corresponding to the n-th feature to be processed:
[0137] In the target projection space, the outer product of the embedding vector corresponding to the i-th feature to be processed among the N features to be processed and the embedding vector corresponding to the n-th feature to be processed is obtained to obtain a first processing result; the embedding vector corresponding to any one of the N features to be processed is obtained by performing feature embedding processing on the any one of the N features to be processed, where the any one of the features to be processed is the i-th feature to be processed or the n-th feature to be processed, and i is a positive integer less than or equal to N;
[0138] Performing adjustment processing on the first processing result to obtain a feature interaction subvector for feature interaction between the i-th feature to be processed and the n-th feature to be processed;
[0139] The feature interaction subvectors of the feature interactions between each of the features to be processed and the n-th feature to be processed are summed to obtain a feature interaction vector corresponding to the n-th feature to be processed.
[0140] In one embodiment, when the processing unit 1002 adjusts the first processing result to obtain a feature interaction subvector for the feature interaction between the i-th feature to be processed and the n-th feature to be processed, the processing unit 1002 specifically performs the following operations:
[0141] Determine a target projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between a feature to be processed, a feature to be processed for feature interaction with the feature to be processed, and a projection matrix;
[0142] Determining target scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between a feature to be processed, a feature to be processed for feature interaction with the feature to be processed, and scaling weights;
[0143] performing matrix adjustment processing on the first processing result based on the target projection matrix and the target scaling weight to obtain a second processing result;
[0144] The second processing result is subjected to vector conversion processing to obtain a feature interaction sub-vector for the feature interaction between the i-th feature to be processed and the n-th feature to be processed.
[0145] In one embodiment, the number of the plurality of features to be processed is N, and there are H channels in the target projection space corresponding to the nth feature to be processed among the N features to be processed, where N is an integer greater than 1, H is a positive integer, and n is a positive integer less than or equal to N;
[0146] When the processing unit 1002 performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed and obtains a feature interaction vector corresponding to each of the multiple features to be processed, the processing unit 1002 specifically performs the following operations:
[0147] Traversing the N features to be processed, performing outer product processing on each feature to be processed and the nth feature to be processed in each channel in the target projection space, so as to perform feature interaction on each feature to be processed and the nth feature to be processed in each channel in the target projection space, and obtaining a feature interaction vector corresponding to the nth feature to be processed in each channel;
[0148] The feature interaction vectors corresponding to the n-th feature to be processed in each channel are combined and processed to obtain the feature interaction vector corresponding to the n-th feature to be processed.
[0149] In one embodiment, performing outer product processing on the plurality of features to be processed to obtain the estimated click rate of the target object for the target resource is achieved by a data processing model, wherein the data processing model is obtained by training an initial data processing model;
[0150] The acquisition unit 1001 is further configured to acquire a training sample, wherein the training sample includes a plurality of training features and a sample label; the plurality of training features include resource features of a training resource and object features of a training object, and the sample label indicates whether the training object clicks on the training resource;
[0151] The processing unit 1002 is further configured to initialize the model parameters of the initial data processing model using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature among the multiple training features; the embedding vector corresponding to each training feature is obtained by performing feature embedding processing on each training feature;
[0152] The processing unit 1002 is further configured to perform outer product processing on the multiple training features using the initial data processing model to perform feature interaction on the multiple training features, thereby obtaining a feature interaction vector corresponding to each of the multiple training features, where the feature interaction vector corresponding to each training feature is a high-order semantic feature of the each training feature;
[0153] The processing unit 1002 is further configured to perform click-through rate estimation processing on the feature interaction vectors corresponding to the respective training features to obtain a training estimated click-through rate of the training subject for the training resource, wherein the training estimated click-through rate indicates a probability of the training subject clicking on the training resource.
[0154] The processing unit 1002 is further configured to train the initial data processing model based on the training estimated click rate and the sample labels to obtain the data processing model.
[0155] In one embodiment, the number of the plurality of training features is N, the model parameters of the initial data processing model include a training projection matrix, and N is an integer greater than 1;
[0156] When the processing unit 1002 initializes the model parameters of the initial data processing model using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the multiple training features, the processing unit 1002 specifically performs the following operations:
[0157] Traversing the N training features, initializing the splicing training projection matrix in the model parameters by a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector; the training splicing embedding vector is obtained by splicing the embedding vectors corresponding to the respective training features;
[0158] The spliced training projection matrix is split to obtain N training projection matrices, and the N training projection matrices are used to obtain a feature interaction vector corresponding to the k-th training feature.
[0159] In one embodiment, the dimension of the embedding vector corresponding to the kth training feature is d k , the dimension of the training concatenated embedding vector is d;
[0160] The processing unit 1002 specifically performs the following operations when initializing the splicing training projection matrix in the model parameters using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector:
[0161] Based on the dimension d of the embedding vector corresponding to the k-th training feature k And the dimension d of the training splicing embedding vector, initialize B projection parameter matrices; the b-th projection parameter matrix in the B projection parameter matrices is the first parameter matrix with an initialized dimension of b*b and the initialized dimension of Kronecker product of the second parameter matrix;
[0162] The B projection parameter matrices are summed to obtain the splicing training projection matrix.
[0163] According to one embodiment of the present application, Figure 2 as well as Figure 7 The steps involved in the data processing method shown can be Figure 10 The data processing apparatus shown in FIG. 1 is executed by each unit. For example, Figure 2 Step S201 shown can be performed by Figure 10 The acquisition unit 1001 in the data processing device shown is executed, Figure 2 Steps S202 to S203 shown in FIG. Figure 10 The processing unit 1002 in the data processing device shown in FIG. Figure 7 Step S701 shown can be performed by Figure 10 The acquisition unit 1001 in the data processing device shown is executed, Figure 7 Steps S702 to S705 shown in FIG. Figure 10 The processing unit 1002 in the data processing device shown is executed.
[0164] According to another embodiment of the present application, Figure 10 The various units in the data processing apparatus shown can be separately or all merged into one or several other units to constitute, or a certain unit (or units) therein can also be split into multiple smaller units in function to constitute, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above-mentioned units are divided based on logical functions. In actual applications, the functions of a unit can also be implemented by multiple units, or the functions of multiple units can be implemented by one unit. In other embodiments of the present application, the data processing apparatus divided based on logical functions can also include other units. In actual applications, these functions can also be implemented with the assistance of other units, and can be implemented by the collaboration of multiple units.
[0165] According to another embodiment of the present application, the program can be executed by running on a general computing device such as a computer including a central processing unit (CPU), a random access memory (RAM), a read-only memory (ROM) and other processing elements and storage elements. Figure 2 as well as Figure 7 A computer program (including program code) for each step of the corresponding method shown in FIG. Figure 10 The data processing device shown in and the data processing method of the embodiment of the present application are implemented. The computer program can be recorded on, for example, a computer-readable storage medium, and loaded into the above-mentioned computing device through the computer-readable storage medium and run therein.
[0166] In an embodiment of the present application, after the acquisition unit 1001 acquires multiple features to be processed, including resource features of the target resource and object features of the target object, the processing unit 1002 can perform outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtain a feature interaction vector corresponding to each of the multiple features to be processed; perform click-through rate estimation processing on the feature interaction vector corresponding to each feature to be processed to obtain an estimated click-through rate of the target object for the target resource; wherein the estimated click-through rate is used to indicate: the probability of the target object clicking on the target resource. By performing outer product processing on multiple features to be processed, the high-order semantic features of each feature to be processed can be learned, thereby improving the accuracy of the target object's click-through rate estimation for the target resource and efficiently acquiring information of interest.
[0167] Based on the above data processing method embodiment and data processing device embodiment, the present application also provides a data processing device. Figure 11 , is a structural diagram of a data processing device provided in an embodiment of the present application. Figure 11 The data processing device shown may include at least a processor 1101, an input interface 1102, an output interface 1103, and a computer storage medium 1104. The processor 1101, the input interface 1102, the output interface 1103, and the computer storage medium 1104 may be connected via a bus or other means.
[0168] Computer storage medium 1104 may be stored in a memory of a data processing device. Computer storage medium 1104 is used to store a computer program, which includes program instructions. Processor 1101 is used to execute the program instructions stored in computer storage medium 1104. Processor 1101 (or CPU (Central Processing Unit)) is the computing core and control core of the data processing device. It is suitable for implementing one or more instructions, and is specifically suitable for loading and executing one or more instructions to implement the above-mentioned data processing method flow or corresponding functions.
[0169] The embodiment of the present application also provides a computer storage medium (Memory), which is a memory device in a data processing device for storing programs and data. It is understandable that the computer storage medium here can include both the built-in storage medium in the terminal and, of course, the extended storage medium supported by the terminal. The computer storage medium provides a storage space that stores the operating system of the terminal. In addition, one or more instructions suitable for being loaded and executed by the processor 1101 are also stored in the storage space. These instructions can be one or more computer programs (including program codes). It should be noted that the computer storage medium here can be a high-speed random access memory (RAM) memory, or a non-volatile memory (non-volatile memory), such as at least one disk storage; optionally, it can also be at least one computer storage medium located away from the aforementioned processor.
[0170] In one embodiment, the processor 1101 and the input interface 1102 can load and execute one or more instructions stored in the computer storage medium to implement the above-mentioned Figure 2 as well as Figure 7 In the corresponding steps of the method in the data processing method embodiment, in a specific implementation, one or more instructions in the computer storage medium are loaded by the processor 1101 and the input interface 1102 and execute the following steps:
[0171] An input interface 1102 is used to obtain a plurality of features to be processed, wherein the plurality of features to be processed include object features of a target object and resource features of a target resource to be pushed to the target object;
[0172] Processor 1101 is configured to perform outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtain a feature interaction vector corresponding to each of the multiple features to be processed, where the feature interaction vector corresponding to each of the features to be processed is a high-order semantic feature of each of the features to be processed;
[0173] The processor 1101 is further configured to perform click rate estimation processing on the feature interaction vectors corresponding to the features to be processed to obtain an estimated click rate of the target object for the target resource, wherein the estimated click rate indicates the probability of the target object clicking on the target resource.
[0174] In one embodiment, the number of the plurality of features to be processed is N, where N is an integer greater than 1;
[0175] When the processor 1101 performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed and obtains a feature interaction vector corresponding to each of the multiple features to be processed, the processor 1101 specifically performs the following operations:
[0176] The N features to be processed are traversed, and in a target projection space corresponding to an n-th feature to be processed among the N features to be processed, outer product processing is performed on each feature to be processed and the n-th feature to be processed, so as to perform feature interaction on each feature to be processed and the n-th feature to be processed in the target projection space, and a feature interaction vector corresponding to the n-th feature to be processed is obtained, where n is a positive integer less than or equal to N.
[0177] In one embodiment, the processor 1101 performs outer product processing on each of the N features to be processed and the n-th feature to be processed in the target projection space corresponding to the n-th feature to be processed, so as to perform feature interaction on each of the features to be processed and the n-th feature to be processed in the target projection space, and obtains a feature interaction vector corresponding to the n-th feature to be processed, specifically performing the following operations:
[0178] In the target projection space, the outer product of the embedding vector corresponding to the i-th feature to be processed among the N features to be processed and the embedding vector corresponding to the n-th feature to be processed is obtained to obtain a first processing result; the embedding vector corresponding to any one of the N features to be processed is obtained by performing feature embedding processing on the any one of the N features to be processed, where the any one of the features to be processed is the i-th feature to be processed or the n-th feature to be processed, and i is a positive integer less than or equal to N;
[0179] Performing adjustment processing on the first processing result to obtain a feature interaction subvector for feature interaction between the i-th feature to be processed and the n-th feature to be processed;
[0180] The feature interaction subvectors of the feature interactions between each of the features to be processed and the n-th feature to be processed are summed to obtain a feature interaction vector corresponding to the n-th feature to be processed.
[0181] In one embodiment, when the processor 1101 adjusts the first processing result to obtain a feature interaction subvector for the feature interaction between the i-th feature to be processed and the n-th feature to be processed, the processor 1101 specifically performs the following operations:
[0182] Determine a target projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between a feature to be processed, a feature to be processed for feature interaction with the feature to be processed, and a projection matrix;
[0183] Determining target scaling weights corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between a feature to be processed, a feature to be processed for feature interaction with the feature to be processed, and scaling weights;
[0184] performing matrix adjustment processing on the first processing result based on the target projection matrix and the target scaling weight to obtain a second processing result;
[0185] The second processing result is subjected to vector conversion processing to obtain a feature interaction sub-vector for the feature interaction between the i-th feature to be processed and the n-th feature to be processed.
[0186] In one embodiment, the number of the plurality of features to be processed is N, and there are H channels in the target projection space corresponding to the nth feature to be processed among the N features to be processed, where N is an integer greater than 1, H is a positive integer, and n is a positive integer less than or equal to N;
[0187] When the processor 1101 performs outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed and obtains a feature interaction vector corresponding to each of the multiple features to be processed, the processor 1101 specifically performs the following operations:
[0188] Traversing the N features to be processed, performing outer product processing on each feature to be processed and the nth feature to be processed in each channel in the target projection space, so as to perform feature interaction on each feature to be processed and the nth feature to be processed in each channel in the target projection space, and obtaining a feature interaction vector corresponding to the nth feature to be processed in each channel;
[0189] The feature interaction vectors corresponding to the n-th feature to be processed in each channel are combined and processed to obtain the feature interaction vector corresponding to the n-th feature to be processed.
[0190] In one embodiment, performing outer product processing on the plurality of features to be processed to obtain the estimated click rate of the target object for the target resource is achieved by a data processing model, wherein the data processing model is obtained by training an initial data processing model;
[0191] The input interface 1102 is further used to obtain a training sample, wherein the training sample includes a plurality of training features and a sample label; the plurality of training features include resource features of a training resource and object features of a training object, and the sample label is used to indicate whether the training object clicks on the training resource;
[0192] The processor 1101 is further configured to initialize model parameters of the initial data processing model using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the multiple training features; the embedding vector corresponding to each training feature is obtained by performing feature embedding processing on each training feature;
[0193] The processor 1101 is further configured to perform outer product processing on the multiple training features using the initial data processing model to perform feature interaction on the multiple training features, thereby obtaining a feature interaction vector corresponding to each of the multiple training features, where the feature interaction vector corresponding to each training feature is a high-order semantic feature of the each training feature;
[0194] The processor 1101 is further configured to perform click-through rate estimation processing on the feature interaction vectors corresponding to the respective training features to obtain a training estimated click-through rate of the training subject for the training resource, wherein the training estimated click-through rate indicates a probability of the training subject clicking on the training resource.
[0195] The processor 1101 is further configured to train the initial data processing model based on the training estimated click-through rate and the sample labels to obtain the data processing model.
[0196] In one embodiment, the number of the plurality of training features is N, the model parameters of the initial data processing model include a training projection matrix, and N is an integer greater than 1;
[0197] When the processor 1101 initializes the model parameters of the initial data processing model using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the multiple training features, the processor 1101 specifically performs the following operations:
[0198] Traversing the N training features, initializing the splicing training projection matrix in the model parameters by a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector; the training splicing embedding vector is obtained by splicing the embedding vectors corresponding to the respective training features;
[0199] The spliced training projection matrix is split to obtain N training projection matrices, and the N training projection matrices are used to obtain a feature interaction vector corresponding to the k-th training feature.
[0200] In one embodiment, the dimension of the embedding vector corresponding to the kth training feature is d k , the dimension of the training concatenated embedding vector is d;
[0201] The processor 1101 specifically performs the following operations when initializing the splicing training projection matrix in the model parameters using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector:
[0202] Based on the dimension d of the embedding vector corresponding to the k-th training feature k And the dimension d of the training splicing embedding vector, initialize B projection parameter matrices; the b-th projection parameter matrix in the B projection parameter matrices is the first parameter matrix with an initialized dimension of b*b and the initialized dimension of Kronecker product of the second parameter matrix;
[0203] The B projection parameter matrices are summed to obtain the splicing training projection matrix
[0204] The present invention provides a computer program product or a computer program. The computer program product includes a computer program stored in a computer storage medium. A processor of a data processing device reads the computer program from the computer storage medium and executes the computer program, so that the data processing device performs the above-mentioned Figure 2 as well as Figure 7 The computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0205] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: include: Acquire a plurality of features to be processed, the plurality of features to be processed comprising object features of a target object and resource features of a target resource to be pushed to the target object; Performing outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed, and obtaining a feature interaction vector corresponding to each feature to be processed in the multiple features to be processed, wherein the feature interaction vector corresponding to each feature to be processed is a high-order semantic feature of each feature to be processed; Performing click-through rate estimation processing on the feature interaction vectors corresponding to the features to be processed to obtain an estimated click-through rate of the target object for the target resource, where the estimated click-through rate indicates the probability of the target object clicking on the target resource; The method further comprises: Based on the dimensions of the embedding vectors corresponding to each training feature in a plurality of training features, the model parameters of the initial data processing model are initialized using a hypercomplex product parameterization strategy; the embedding vectors corresponding to each training feature are obtained by performing feature embedding processing on each training feature; the plurality of training features include resource features of the training resources and object features of the training objects; the hypercomplex product parameterization strategy is to learn the training projection matrix in the initial data processing model from a hypercomplex space, learn the interaction between real components and imaginary components, and generalize the vector outer product to a higher-dimensional real space; Through the initial data processing model, outer product processing is performed on the multiple training features to perform feature interaction on the multiple training features, and a feature interaction vector corresponding to each training feature in the multiple training features is obtained, and the feature interaction vector corresponding to each training feature is a high-order semantic feature of each training feature.
2. The method according to claim 1, wherein The number of the plurality of features to be processed is N, where N is an integer greater than 1; The performing outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed to obtain a feature interaction vector corresponding to each feature to be processed in the multiple features to be processed includes: The N features to be processed are traversed, and in a target projection space corresponding to an n-th feature to be processed among the N features to be processed, outer product processing is performed on each feature to be processed and the n-th feature to be processed, so as to perform feature interaction on each feature to be processed and the n-th feature to be processed in the target projection space, and a feature interaction vector corresponding to the n-th feature to be processed is obtained, where n is a positive integer less than or equal to N.
3. The method according to claim 2, wherein The process of performing outer product processing on each of the features to be processed and the nth feature to be processed in the target projection space corresponding to the nth feature to be processed among the N features to be processed, so as to perform feature interaction on each of the features to be processed and the nth feature to be processed in the target projection space, and obtaining a feature interaction vector corresponding to the nth feature to be processed, includes: In the target projection space, the outer product of the embedding vector corresponding to the i-th feature to be processed among the N features to be processed and the embedding vector corresponding to the n-th feature to be processed is obtained to obtain a first processing result; the embedding vector corresponding to any one of the N features to be processed is obtained by performing feature embedding processing on the any one of the N features to be processed, where the any one of the features to be processed is the i-th feature to be processed or the n-th feature to be processed, and i is a positive integer less than or equal to N; The first processing result is adjusted to obtain a feature interaction subvector for feature interaction between the i-th feature to be processed and the n-th feature to be processed; including: determining a target projection matrix corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between the feature to be processed, the feature to be processed for feature interaction with the feature to be processed, and a projection matrix; determining a target scaling weight corresponding to the i-th feature to be processed and the n-th feature to be processed based on a correspondence between the feature to be processed, the feature to be processed for feature interaction with the feature to be processed, and a scaling weight; performing matrix adjustment on the first processing result based on the target projection matrix and the target scaling weight to obtain a second processing result; performing vector conversion on the second processing result to obtain a feature interaction subvector for feature interaction between the i-th feature to be processed and the n-th feature to be processed; The feature interaction subvectors of the feature interactions between each of the features to be processed and the n-th feature to be processed are summed to obtain a feature interaction vector corresponding to the n-th feature to be processed.
4. The method according to any one of claims 1 to 3, wherein The number of the plurality of features to be processed is N, and there are H channels in the target projection space corresponding to the nth feature to be processed among the N features to be processed, where N is an integer greater than 1, H is a positive integer, and n is a positive integer less than or equal to N; The performing outer product processing on the multiple features to be processed to perform feature interaction on the multiple features to be processed to obtain a feature interaction vector corresponding to each feature to be processed in the multiple features to be processed includes: Traversing the N features to be processed, performing outer product processing on each feature to be processed and the nth feature to be processed in each channel in the target projection space, so as to perform feature interaction on each feature to be processed and the nth feature to be processed in each channel in the target projection space, and obtaining a feature interaction vector corresponding to the nth feature to be processed in each channel; The feature interaction vectors corresponding to the n-th feature to be processed in each channel are combined and processed to obtain the feature interaction vector corresponding to the n-th feature to be processed.
5. The method according to claim 1, wherein The estimated click rate of the target object for the target resource is obtained by performing outer product processing on the plurality of features to be processed through a data processing model, wherein the data processing model is obtained based on training of an initial data processing model; The method further comprises: Acquire a training sample, wherein the training sample includes a plurality of training features and a sample label; the sample label is used to indicate whether the training subject clicks on the training resource; Performing click-through rate estimation processing on the feature interaction vectors corresponding to the respective training features to obtain a training estimated click-through rate of the training subject for the training resource, wherein the training estimated click-through rate is used to indicate a probability of the training subject clicking on the training resource; Based on the training estimated click-through rate and the sample labels, the initial data processing model is trained to obtain the data processing model.
6. The method according to claim 5, wherein The number of the plurality of training features is N, the model parameters of the initial data processing model include a training projection matrix, and N is an integer greater than 1; The step of initializing the model parameters of the initial data processing model by using a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to each training feature in the multiple training features includes: Traversing the N training features, initializing the splicing training projection matrix in the model parameters by a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector; the training splicing embedding vector is obtained by splicing the embedding vectors corresponding to the respective training features; The spliced training projection matrix is split to obtain N training projection matrices, and the N training projection matrices are used to obtain a feature interaction vector corresponding to the k-th training feature.
7. The method according to claim 6, wherein The dimension of the embedding vector corresponding to the kth training feature is , the dimension of the training concatenated embedding vector is ; The step of initializing the splicing training projection matrix in the model parameters by a hypercomplex product parameterization strategy based on the dimension of the embedding vector corresponding to the kth training feature among the N training features and the dimension of the training splicing embedding vector includes: Based on the dimension of the embedding vector corresponding to the k-th training feature And the dimensions of the training concatenated embedding vector ,initialization projection parameter matrix; The projection parameter matrix The projection parameter matrix is initialized to have a dimension of The first parameter matrix and the dimensions of the initialization are Kronecker product of the second parameter matrix; The The projection parameter matrices are summed to obtain the splicing training projection matrix.
8. A data processing device, characterized in that: include: an acquiring unit, configured to acquire a plurality of features to be processed, wherein the plurality of features to be processed include object features of a target object and resource features of a target resource to be pushed to the target object; a processing unit, configured to perform outer product processing on the plurality of features to be processed to perform feature interaction on the plurality of features to be processed, and obtain a feature interaction vector corresponding to each of the plurality of features to be processed, wherein the feature interaction vector corresponding to each of the features to be processed is a high-order semantic feature of each of the features to be processed; The processing unit is further configured to perform click-through rate estimation processing on the feature interaction vectors corresponding to the respective features to be processed, to obtain an estimated click-through rate of the target object for the target resource, wherein the estimated click-through rate indicates a probability of the target object clicking on the target resource; The processing unit is further configured to initialize model parameters of the initial data processing model using a hypercomplex product parameterization strategy based on the dimensions of the embedding vectors corresponding to each of the multiple training features; the embedding vectors corresponding to each of the training features are obtained by performing feature embedding processing on the respective training features; the multiple training features include resource features of the training resources and object features of the training objects; the hypercomplex product parameterization strategy is to learn the training projection matrix in the initial data processing model from the hypercomplex space, learn the interaction between the real components and the imaginary components, and generalize the vector outer product to a higher-dimensional real space; The processing unit is further used to perform outer product processing on the multiple training features through the initial data processing model to perform feature interaction on the multiple training features to obtain feature interaction vectors corresponding to each training feature in the multiple training features, where the feature interaction vectors corresponding to each training feature are high-order semantic features of each training feature.
9. A data processing device, characterized in that: The data processing device includes an input interface and an output interface, and further includes: a processor adapted to implement one or more instructions; and A computer storage medium storing one or more instructions, wherein the one or more instructions are suitable for being loaded by the processor and executing the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Behavior relationship determination method and device, computer equipment and readable storage medium
CN111459781A