Information Processing Method, Apparatus, Computer Device, and Storage Medium
By introducing an automatic feature cross-section layer into the double tower model, the feature cross-section between the user side and the information side is realized, the information interaction and expression ability is improved, the problem of limited interaction between user features and information features in the double tower model is solved, and the information processing effect is improved.
Patent Information
- Application Number
- CN202211196845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-28
AI Technical Summary
In the existing dual-tower model, the user tower structure and the information tower structure are naturally separated, resulting in limited interaction between user characteristics and information characteristics in the neural network, affecting the sorting effect of the to-process information.
An automatic feature cross-section layer is introduced. By processing the feature cross-section on the user side and the information side, the cross-expression vectors on the user side and the information side are generated, and they are spliced with the user tower vector and information tower vector of the double tower structure to form a target vector and improve information interaction and expression capabilities.
Through the introduction of the automatic feature cross-section layer, the information interaction between the user side and the information side is improved, the processing effect of target information is improved, and the expression ability of the double tower structure is enhanced.
Smart Images

Figure CN115438802B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of Internet technologies, and particularly to an information processing method, apparatus, computer device, and storage medium. Background Art
[0002] In the information flow scenario, the recommendation algorithm is a core technology for optimizing the content ecosystem and improving the user experience. In large-scale information flow recommendation scenarios, the recommendation system often consists of a funnel structure of recall, rough ranking, fine ranking, and re-ranking. At each layer, candidate content is filtered layer by layer through the design of algorithms and strategies, and finally high-quality content is pushed and presented to users.
[0003] In related technologies, the Deep Structured Semantic Model (DSSM), also known as the twin tower model, is commonly used as the service model in the recall and rough ranking stages to obtain the sorting scores of the information to be processed.
[0004] However, due to the natural separation of the user tower structure and the information tower structure in the twin tower model, the interaction between user features and information features in the neural network only occurs at the top of the tower, which limits the expression ability of the model, thus affecting the sorting effect of the information to be processed, and further resulting in poor processing effect of the information to be processed. Summary of the Invention
[0005] Embodiments of the present application provide an information processing method, apparatus, device, and storage medium, which can enhance the information interaction between the user side and the information side, improve the expression ability of the original twin tower structure, and thus improve the processing effect of target information. The technical solution is as follows:
[0006] On the one hand, an information processing method is provided, and the method includes:
[0007] Obtain a first pooling vector of a target user; the first pooling vector is a vector generated in the process of processing the user features of the target user through the twin tower structure of an information processing model to obtain the user tower vector of the target user;
[0008] Input the first pooling vector into the feature automatic cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature automatic cross layer;
[0009] Concatenate the first cross-equivalent vector and the user tower vector to obtain a first target vector of the target user;
[0010] Process the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result; the target manner is determined based on a target processing manner for the target information; the second target vector is obtained after the information processing model processes the information features of the target information.
[0011] Based on the vector processing result, process the target information in the target processing manner.
[0012] On the other hand, an information processing device is provided, and the device includes:
[0013] A first acquisition module, configured to acquire a first pooling vector of a target user; the first pooling vector is a vector generated in the process of obtaining a user tower vector of the target user by processing the user features of the target user through a two-tower structure of an information processing model.
[0014] A second acquisition module, configured to input the first pooling vector into a feature auto-cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer.
[0015] A first splicing module, configured to splice the first cross-equivalent vector and the user tower vector to obtain a first target vector of the target user.
[0016] A vector processing module, configured to process the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result; the target manner is determined based on a processing manner for the target information; the second target vector is obtained after the information processing model processes the information features of the target information.
[0017] An information processing module, configured to process the target information based on the vector processing result.
[0018] In a possible implementation manner, the first pooling vector includes K-dimensional features of the target user in n feature domains; n and K are positive integers.
[0019] The second acquisition module includes:
[0020] A first feature acquisition sub-module, configured to perform bitwise summation on the K-dimensional features in n feature domains of the first pooling vector to obtain a first summation feature.
[0021] A first construction sub-module, configured to construct the first cross-equivalent vector based on the first summation feature and the first pooling vector.
[0022] In a possible implementation, the apparatus further includes:
[0023] A third acquisition module, configured to acquire a second pooling vector of the target information; the second pooling vector is a vector generated in the process of processing the information features of the target information through the twin tower structure to obtain the information tower vector of the target information;
[0024] A fourth acquisition module, configured to input the second pooling vector into the feature auto-cross layer to obtain a second cross-equivalent vector corresponding to the target information output by the feature auto-cross layer;
[0025] A second splicing module, configured to splice the second cross-equivalent vector and the information tower vector to obtain the second target vector of the target information.
[0026] In a possible implementation, the second pooling vector includes K-dimensional features of the target information in m feature domains; m and K are positive integers;
[0027] The fourth acquisition module includes:
[0028] A second feature acquisition sub-module, configured to perform bitwise summation on the K-dimensional features in m feature domains of the second pooling vector to obtain a second summation feature;
[0029] A second construction sub-module, configured to construct the second cross-equivalent vector based on the second summation feature and the second pooling vector.
[0030] In a possible implementation, the first cross-equivalent vector is [U, 1, 0.5 * (U ⊙ U - sum(X_user 2 ))]; the second cross-equivalent vector is [I, 0.5 * (I ⊙ I - sum(X_item 2 ))];
[0031] where U represents the first summation feature, X_user represents the first pooling vector, I represents the second summation feature, X_item represents the second pooling vector, 1 represents a vector with a dimension of 1 and a vector value of 1, and ⊙ represents the inner product between vectors.
[0032] In a possible implementation, the vector processing module includes:
[0033] A similarity calculation sub-module, configured to calculate the cosine similarity between the first target vector and the second target vector when the target processing method for the target information is recall;
[0034] A position determination sub-module, configured to determine the position of the target information in the recall sequence based on the cosine similarity; at least two pieces of information to be recalled are included in the recall sequence;
[0035] The information processing module is configured to recall the target information when the position of the target information in the recall sequence meets the recall condition.
[0036] In a possible implementation manner, the vector processing module includes:
[0037] A dot product calculation sub-module, configured to perform a dot product calculation on the first target vector and the second target vector to obtain a dot product result;
[0038] A score acquisition sub-module, configured to process the dot product result through an activation function to obtain the sorting score of the target information;
[0039] The information processing module is configured to determine the position of the target information in the rough ranking sequence based on the sorting score; at least two pieces of rough ranking information are included in the rough ranking sequence.
[0040] On the other hand, a computer device is provided, which includes a processor and a memory. The memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the above information processing method.
[0041] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored, and the computer program is loaded and executed by a processor to implement the above information processing method.
[0042] On the other hand, a computer program product is provided, which includes at least one computer program, and the computer program is loaded and executed by a processor to implement the information processing methods provided in the above various optional implementation manners.
[0043] The technical solution provided by this application may include the following beneficial effects:
[0044] The information processing method provided by the embodiments of the present application introduces a feature automatic cross layer on the basis of the original dual-tower model structure. The feature automatic cross layer can obtain a first cross-equivalent vector on the user side, which is used to represent the user-side part of the feature interaction result between the user side and the information side, and is concatenated with the user tower vector obtained after the dual-tower structure processes the user features to form a first target vector on the user side. And a second target vector on the information side is obtained synchronously or asynchronously. Through the processing of the first target vector and the second target vector, a vector processing result is obtained to realize the processing of the target information. By introducing the feature automatic cross layer, the information interaction between the user side and the information side can be improved, and the expression ability of the original dual-tower structure can be enhanced, thereby improving the processing effect of the target information.
[0045] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Brief Description of the Drawings
[0046] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0047] Figure 1 It shows a schematic structural diagram of an information processing model involved in the information processing method provided by an exemplary embodiment of the present application;
[0048] Figure 2 It shows a flowchart of the information processing method provided by an exemplary embodiment of the present application;
[0049] Figure 3 It is a framework diagram of an information processing model generation and information processing shown according to an exemplary embodiment;
[0050] Figure 4 It shows a flowchart of the training method of the information processing model provided by an exemplary embodiment of the present application;
[0051] Figure 5 It shows a flowchart of the information processing method provided by an exemplary embodiment of the present application;
[0052] Figure 6 It shows a schematic diagram of the first target vector and the second target vector in an exemplary embodiment implemented by the present application;
[0053] Figure 7 It shows a block diagram of the information processing device provided by an exemplary embodiment of the present application;
[0054] Figure 8 It is a structural block diagram of a computer device shown according to an exemplary embodiment;
[0055] Figure 9 is a block diagram of a computer device shown according to an exemplary embodiment. Detailed implementation manners
[0056] Here, the exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present application as detailed in the appended claims.
[0057] Figure 1 shows a schematic structural diagram of an information processing model involved in an information processing method provided by an exemplary embodiment of the present application, as Figure 1 shown, the information processing model includes a two-tower structure 101 and a feature auto-crossing layer 102.
[0058] The two-tower structure 101 includes two decoupled parts, namely, a user tower structure and an information tower structure, and the neural networks of the two parts are independent of each other; wherein, the user tower structure is used to process the received user features, and the information tower structure is used to process the received information features.
[0059] The feature auto-crossing layer 102 is used to enhance the feature interaction of the two-tower structure; the feature auto-crossing layer 102 can have various implementation forms. For example, it can be combined after manually screening features, or implemented by machine learning algorithms or deep learning networks. In the embodiments of the present application, the feature crossing layer can be based on the principle of factorization machine (FM), and the feature crossing result after feature crossing is equivalently disassembled into two parts on the user side and the information side, so as to be spliced corresponding to the results of the user tower vector and the information tower vector output by the two-tower structure 101, to obtain a first target vector on the user side and a second target vector on the information side, so that the computer device processes the first target vector and the second target vector to obtain the sorting score of the target information. The feature crossing result is a processing result obtained by performing feature crossing processing on the intermediate features obtained when the user tower processes the user features and the intermediate features obtained when the information tower processes the information features.
[0060] Among them, the derivation process of equivalently disassembling the feature crossing result obtained after feature crossing into two parts on the user side and the information side is as follows:
[0061] The original formula of FM is:
[0062]
[0063] where x iDenote the one-hot encoded features, which take values of 0 or 1, n represents the number of features, and w l Denote the feature x i The coefficient of the first-order term of, and v i Denote the feature x i The corresponding latent vector, <, > represents the dot product, b represents the global bias, and the second-order cross part of the above formula can be equivalently transformed into:
[0064]
[0065] Among them, k represents the dimension of the latent vector v. From the perspective of feature crossing, after dividing the features into the user side and the information side, Contains the feature crossings between the user side and the user side, the user side and the information side, and the information side and the information side. Denote the matrix obtained by sum-pooling the user-side features after passing through the embedding layer as X_user, and the matrix obtained by sum-pooling the information-side features after passing through the embedding layer as X_item. Denote the cross-equivalent vector of the user side as U, and the cross-equivalent vector of the information side as I, then:
[0066]
[0067]
[0068] Among them, ⊙ represents the inner product between vectors, sum(X_user 2 ), sum(X_item 2 ) represent the sum of the squares of the matrix elements of the user side and the information side respectively. Therefore, the feature crossing result between the user side and the information side can be equivalently decoupled into the form of the inner product of the user-side vector and the information-side vector. Based on this, in the embodiments of the present application, the above formula derivation results can be directly used to calculate the cross-equivalent vector of the user side and the cross-equivalent vector of the information side separately through the feature automatic crossing layer, and then vector processing is performed based on the cross-equivalent vector of the user side and the cross-equivalent vector of the information side, and the purpose of enhancing the feature interaction of the dual tower structure can still be achieved. Since the user side and the information side are still in a decoupled relationship during the processing, the online service time consumption will not increase compared with only using the dual tower structure for processing, and the real-time calculation requirements can be met.
[0069] Since in the embodiments of the present application, the process of obtaining the cross-equivalent vector of the user side and the cross-equivalent vector of the information side can be decoupled, therefore, the process of obtaining the first target vector of the user side and the second target vector of the information side can be executed synchronously or asynchronously, Figure 2The flowchart of the information processing method provided by an exemplary embodiment of the present application is shown. This information processing method can be executed by a computer device, which can be equipped with an information processing model as shown in Figure 1 . The computer device can be implemented as a server or a terminal. As shown in Figure 2 , the information processing method may include the following steps:
[0070] Step 210: Obtain the first pooling vector of the target user; the first pooling vector is a vector generated during the process of processing the user features of the target user through the two-tower structure of the information processing model to obtain the user tower vector of the target user.
[0071] Optionally, when the computer device receives an online request sent by the terminal device corresponding to the target user, it can extract features from the user information of the target user to obtain the user features of the target user. Schematically, the user features can be features constructed by the recommendation system based on the user information of the target user.
[0072] After receiving the user features of the target user, the computer device inputs the user features into the two-tower structure of the information processing model for processing; further, the computer device inputs the user features into the user tower of the two-tower structure for processing to obtain the user tower vector of the target user; and during the process of the user tower processing the user features, the first pooling vector of the target user can be generated.
[0073] Step 220: Input the first pooling vector into the feature auto-cross layer of the information processing model to obtain the first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer.
[0074] The first cross-equivalent vector is used to represent the feature vector on the user side decoupled from the feature cross result; the feature cross result refers to the processing result obtained by performing feature cross processing on the first pooling vector on the user side and the second pooling vector on the information side.
[0075] Step 230: Concatenate the first cross-equivalent vector and the user tower vector to obtain the first target vector of the target user.
[0076] Among them, the dimension of the first target vector is the sum of the dimension of the first cross-equivalent vector and the dimension of the user tower vector.
[0077] Step 240: Process the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result; the target manner is determined based on the target processing manner of the target information; the second target vector is obtained after the information processing model processes the information features of the target information.
[0078] Among them, the process of obtaining the second target vector of the target information can be executed synchronously with the process of obtaining the first target vector of the target user, or can be executed asynchronously.
[0079] When the target processing method for the target information is recall, this target method is used to determine the matching degree between the target information and the target user. That is to say, the vector processing result is used to indicate the matching degree between the target information and the target user; when the target processing method for the target information is rough ranking, this target method is used to determine the ranking score of the target information. That is to say, the vector processing result is the ranking score of the target information.
[0080] Step 250, based on the vector processing result, process the target information according to the target processing method.
[0081] In summary, the information processing method provided in the embodiments of the present application introduces a feature automatic cross layer on the basis of the original dual tower model structure. The feature automatic cross layer can obtain the first cross-equivalent vector on the user side. The first cross-equivalent vector is used to represent the user side part of the feature interaction result between the user side and the information side, and is concatenated with the user tower vector obtained after processing the user features by the dual tower structure to form the first target vector on the user side, and synchronously or asynchronously obtains the second target vector on the information side. Through the processing of the first target vector and the second target vector, the obtained vector processing result is used to process the target information. By introducing the feature automatic cross layer, the information interaction between the user side and the information side can be improved, and the expression ability of the original dual tower structure can be enhanced, thereby improving the processing effect of the target information.
[0082] The solution involved in the present application includes an information processing model generation stage and an information processing stage. Figure 3 It is a framework diagram of information processing model generation and information processing shown according to an exemplary embodiment. As Figure 3 shown, in the information processing model generation stage, the information processing model generation device 310 obtains an information processing model through a pre-set training sample data set (including user features of sample users, information features of sample information, and score labels corresponding to the sample information). In the information processing stage, the information processing device 320 processes the user features of the input target user and the information features of the target information based on the information processing model, obtains the first target vector of the target user and the second target vector of the target information, and processes the first target vector and the second target vector based on the target processing method for the target information to obtain a vector processing result, so as to process the target information according to the vector processing result according to the target processing method.
[0083] Among them, the above-mentioned information processing model generation device 310 and information processing device 320 may be computer devices. For example, the computer device may be a fixed computer device such as a personal computer or a server, or the computer device may also be a mobile computer device such as a tablet computer or an e-book reader.
[0084] Optionally, the above-mentioned information processing model generation device 310 and information processing device 320 may be the same device, or the information processing model generation device 310 and information processing device 320 may also be different devices. Moreover, when the information processing model generation device 310 and information processing device 320 are different devices, the information processing model generation device 310 and information processing device 320 may be devices of the same type. For example, both the information processing model generation device 310 and information processing device 320 may be servers; or the information processing model generation device 310 and information processing device 320 may also be devices of different types. For example, the information processing device 320 may be a personal computer or a terminal, while the information processing model generation device 310 may be a server, etc. The embodiments of the present application do not limit the specific types of the information processing model generation device 310 and information processing device 320.
[0085] Figure 4 The flowchart of the training method of the information processing model provided by an exemplary embodiment of the present application is shown. This method can be executed by a computing device, and the computer device can implement an information processing model generation device as Figure 3 shown, as Figure 4 shown, the training method of this information processing model may include the following steps:
[0086] Step 410, obtain a training sample data set, and this training sample set includes user features of sample users, information features of sample information, and score labels corresponding to the sample information.
[0087] Step 420, respectively process the user features of the sample users and the information features of the sample information through the two-tower structure of the information processing model to obtain the user tower vector of the sample users and the information tower vector of the sample information.
[0088] Step 430, respectively process the first pooling vector of the sample users and the second pooling vector of the sample information through the feature automatic cross layer of the information processing model to obtain the first cross-equivalent vector of the sample users and the second cross-equivalent vector of the sample information; among them, the first pooling vector of the sample users and the second pooling vector of the sample information are intermediate vectors generated during the process of respectively processing the user features of the sample users and the information features of the sample information by the two-tower structure.
[0089] Step 440: Concatenate the user tower vector of the sample user with the first cross-equivalent vector of the sample user, and the information tower vector of the sample information with the second pooling vector of the sample information respectively, to obtain the first target vector of the sample user and the second target vector of the sample information.
[0090] Step 450: Calculate the dot product of the first target vector of the sample user and the second target vector of the sample information to obtain the dot product result.
[0091] Step 460: Process the dot product result through an activation function to obtain the ranking score of the sample information.
[0092] Step 470: Update the parameters of the information processing model based on the ranking score of the sample information and the score label corresponding to the sample information, so as to train the information processing model.
[0093] In the embodiment of the present application, the computer device can calculate the loss function based on the ranking score of the sample information and the score label corresponding to the sample information, and perform backpropagation according to the calculation result of the loss function to update the parameters in each structure of the information processing model.
[0094] Repeat the above process of updating the parameters in each structure of the information processing model based on the loss function through the combination of the user features of different sample users and the information features of the sample information until the information processing model converges, so as to obtain the trained information processing model.
[0095] Figure 5 The flowchart of the information processing method provided by an exemplary embodiment of the present application is shown. This information processing method can be executed by a computer device, which can be equipped with an information processing model as shown in Figure 1 The computer device can be implemented as a server or a terminal, as shown in Figure 5 The information processing method may include the following steps:
[0096] Step 510: Obtain the first pooling vector of the target user; the first pooling vector is a vector generated in the process of processing the user features of the target user through the two-tower structure of the information processing model to obtain the user tower vector of the target user.
[0097] As shown in Figure 1 The two-tower structure includes a user tower structure and an information tower structure. Among them, the user tower structure is composed of a first embedding layer, a first sum pooling layer, and a first deep network, and the information tower structure is composed of a second embedding layer, a second sum pooling layer, and a second deep network. The user tower structure and the information tower structure are decoupled from each other.
[0098] After receiving the user characteristics of the target user, the computer device inputs the user characteristics into the user tower structure in the two-tower structure of the information processing model to process the user characteristics through the user tower structure.
[0099] Among them, if the user characteristics have n feature domains, and each feature domain contains several features, after processing the user characteristics through the first embedding layer, an embedding vector of dimension K corresponding to each feature can be obtained, where n and K are both positive integers.
[0100] By performing bitwise summation on the K-dimensional embedding vectors corresponding to each feature in the same feature domain through the first sum pooling layer, the first pooling vector X_user of the target user is obtained, and the dimension is [n, K].
[0101] Among them, given an n-dimensional embedding vector x and an n-dimensional embedding vector y, the formula for bitwise summation is as follows:
[0102] x = [x1, x2,..., x n
[0103] y = [y1, y2,..., y n
[0104] x + y = [x1 + y1, x2 + y2,..., x n + y n
[0105] Input the first pooling vector of the target user into the first deep network to obtain the user tower vector output by the first deep network. The first deep network can be composed of several fully connected neural (MLP, Multilayer Perceptron) networks with the same structure.
[0106] The dimension of the user tower feature is T dimensions, and T is a positive integer.
[0107] Step 520: Input the first pooling vector into the feature auto-cross layer of the information processing model to obtain the first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer.
[0108] The first cross-equivalent vector is used to represent the feature vector on the user side decoupled from the feature cross result.
[0109] As can be seen from the process of obtaining the user tower vector in step 510, the computer device can obtain the first pooling vector from the output of the first sum pooling layer, and the first pooling vector contains the K-dimensional features of the target user in n feature domains.
[0110] Based on this, the process of the feature automatic cross layer obtaining the first cross equivalent vector corresponding to the target user can be implemented as follows:
[0111] Sum the K-dimensional features on n feature domains in the first pooling vector bit by bit to obtain the first summation feature;
[0112] Based on the first summation feature and the first pooling vector, construct the first cross equivalent vector.
[0113] Among them, the dimension of the first summation feature vector is K-dimensional. In the embodiments of the present application, the first summation feature is denoted as U.
[0114] Based on the derivation process of decoupling the feature cross result, after obtaining the first summation feature, combining the first pooling vector can obtain 0.5*(U⊙U - sum(X_user 2 ))
[0115] The first cross equivalent vector constructed based on the first summation feature and the first pooling vector can be: [U, 1, 0.5*(U⊙U - sum(X_user 2 ))], where U represents the first summation feature, X_user represents the first pooling vector, 1 represents a vector with a dimension of 1 and a vector value of 1, and ⊙ represents the inner product between vectors.
[0116] The dimension of the first cross equivalent vector is K + 2 dimensions.
[0117] Step 530, concatenate the first cross equivalent vector and the user tower vector to obtain the first target vector of the target user.
[0118] The dimension of the first target vector is the sum of the dimension of the first cross equivalent vector and the dimension of the user tower vector, that is, T + K + 2 dimensions.
[0119] Step 540, obtain the second pooling vector of the target information; this second pooling vector is a vector generated during the process of obtaining the information tower vector of the target information by processing the information features of the target information through a two-tower structure.
[0120] The information features of the target information can be the features obtained after feature extraction of the relevant content of the target information; after receiving the information features of the target information, the computer device inputs the information features into the information tower structure in the two-tower structure of the information processing model to process the information features through the information tower structure.
[0121] Among them, if there are m feature domains in the information feature, and each feature domain contains several features, after processing the information feature through the second embedding layer, an embedding vector of dimension K corresponding to each feature can be obtained, where both m and K are positive integers, and the values of m and n can be the same or different.
[0122] By performing bitwise summation on the K-dimensional embedding vectors corresponding to each feature in the same feature domain through the second summation pooling layer, a second pooling vector X_item of the target information is obtained, and the dimension is [m, K].
[0123] Input the second pooling vector of the target information into the second deep network to obtain an information tower vector output by the second deep network. The first deep network can be composed of several fully connected neural networks with the same structure.
[0124] The dimension of this information tower feature is T dimensions, and T is a positive integer.
[0125] Step 550: Input the second pooling vector into the feature auto-cross layer to obtain a second cross-equivalent vector corresponding to the target information output by the feature auto-cross layer.
[0126] This second cross-equivalent vector is used to represent the feature vector on the information side decoupled from the feature cross result.
[0127] As can be seen from the process of obtaining the information tower vector in step 540, the computer device can obtain the second pooling vector from the output of the second summation pooling layer, and this second pooling vector contains the K-dimensional features of the target information on m feature domains.
[0128] Based on this, the process of the feature auto-cross layer obtaining the second cross-equivalent vector corresponding to the target information can be implemented as follows:
[0129] Perform bitwise summation on the K-dimensional features on the m feature domains in the second pooling vector to obtain a second summation feature;
[0130] Construct a second cross-equivalent vector based on the second summation feature and the second pooling vector.
[0131] Among them, the dimension of the second summation feature vector is K dimensions. In the embodiment of the present application, the second summation feature is denoted as I.
[0132] Based on the derivation process of decoupling the feature cross result, after obtaining the second summation feature, combining the second pooling vector can obtain 0.5*(I⊙I - sum(X_item 2 ))
[0133] The second cross-equivalent vector constructed based on the second summation feature and the second pooling vector can be: [I, 0.5*(I⊙I - sum(X_item 2 ))), 1], where I represents the second summation feature, X_item represents the second pooling vector, 1 represents a vector with a dimension of 1 and a vector value of 1, and ⊙ represents the inner product between vectors.
[0134] The dimension of the second cross-equivalent vector is K + 2 dimensions.
[0135] Step 560, concatenate the second cross-equivalent vector and the information tower vector to obtain the second target vector of the target information.
[0136] The dimension of the second target vector is the sum of the dimension of the second cross-equivalent vector and the dimension of the information tower vector, that is, T + K + 2 dimensions.
[0137] Figure 6 Shows a schematic diagram of the first target vector and the second target vector in an exemplary embodiment of the present application. As Figure 6 shown, the first target vector 610 is composed of the user tower vector and the first cross-equivalent vector, with a dimension of T + K + 2 dimensions; the second target vector 620 is composed of the information tower vector and the second cross-equivalent vector, and the dimension is also T + K + 2 dimensions.
[0138] Since the processing processes of the user side and the information side by the information processing model can be decoupled, therefore, during online services, considering the large number of information, the process of obtaining the first target vector on the user side and the process of obtaining the second target vector on the information side can be executed asynchronously. Among them, the process of obtaining the first target vector on the user side can be performed when a recommendation request sent by the user side is received; the process of obtaining the second target vector on the information side can be performed offline, and the second target vectors of each target information obtained are stored.
[0139] Optionally, when the target processing method for the target information is different, the storage location of the second target vector is different; schematically, when the target processing method for the target information is recall, the second target vector of the target information is stored in the similarity search calculation library.
[0140] Among them, the similarity search calculation library can be the Faiss library to complete the screening of the information to be recalled according to the similarity between the user side and the information side.
[0141] When the target processing method for the target information is rough ranking, the second target vector of the target information is stored offline in the database based on the information primary key of the target information.
[0142] Among them, the information primary key may be the information ID of the target information, and the key is the second target vector of the target information, so that the second target vector of the target information can be indexed from the database based on the information primary key.
[0143] Step 570, process the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result; the target manner is determined based on the target processing manner of the target information.
[0144] When the target processing manner of the target information is recall, processing the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result can be implemented as:
[0145] Calculate the cosine similarity between the first target vector and the second target vector;
[0146] Determine the position of the target information in the recall sequence based on the cosine similarity; the recall sequence contains at least two information to be recalled.
[0147] That is to say, in the recall stage, the target information is any one of the information to be recalled, and the computer device can sort each information to be recalled according to the matching degree between each information to be recalled and the target user, that is, determine the position of each information to be recalled in the recall sequence.
[0148] Alternatively, the computer device can also determine the probability that the target information is recalled based on the cosine similarity; schematically, a similarity threshold is set. If the pre-similarity of the target information is greater than or equal to the similarity threshold, it is recalled, that is, the probability of being recalled is 1; if the pre-similarity of the target information is less than the similarity threshold, it is not recalled, that is, the probability of being recalled is 0.
[0149] When the target processing manner of the target information is rough ranking, processing the first target vector and the second target vector of the target information in a target manner to obtain a vector processing result can be implemented as:
[0150] Perform a dot product calculation on the first target vector and the second target vector to obtain a dot product result;
[0151] Process the dot product result through an activation function to obtain the ranking score of the target information.
[0152] At this time, the target information is any one of the information entering the rough ranking stage.
[0153] Among them, assuming an n-dimensional embedding vector x and an n-dimensional embedding vector y, the formula for the dot product calculation can be expressed as:
[0154] x = [x1, x2,..., x n
[0155] y = [y1, y2, ..., y n
[0156] <x, y> = x1y1, x2y2, ..., x n y n
[0157] Step 580. Process the target information according to the target processing method based on the vector processing result.
[0158] When the target processing method for the target information is recall, processing the target information according to the target processing method based on the vector processing result can be implemented as:
[0159] Recall the target information when the position of the target information in the recall sequence meets the recall condition.
[0160] Alternatively, recall the target information when the recall probability of the target information is 1.
[0161] When the target processing method for the target information is rough ranking, the process of processing the first target vector and the second target vector according to the target method to obtain the vector processing result can be implemented as:
[0162] Determine the position of the target information in the rough ranking sequence based on the ranking score; the rough ranking sequence contains at least two rough ranking information.
[0163] In summary, the information processing method provided by the embodiments of the present application adds an embedding second-order cross structure, that is, a feature automatic cross layer, on the basis of the dual tower model. Through this feature automatic cross layer, the feature cross between the user side and the information side is realized, the information interaction between the user side and the information side is improved, and the expression ability of the dual tower structure is improved, thereby improving the processing effect of the target information.
[0164] At the same time, through derivation, the forward propagation algorithm of the feature automatic cross layer is transformed, so that the calculation processes on the user side and the information side are decoupled and designed in the form of calculating the inner product of the vectors on the user side and the information side. Thus, while improving the model expression ability, the high-efficiency service performance of the model online is ensured.
[0165] It should be noted that this application involves relevant data such as user information. When applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data must comply with the relevant laws and regulations and standards of relevant countries and regions.
[0166] Figure 7 Shows a block diagram of an information processing device provided by an exemplary embodiment of the present application. This information processing device is used to implement asFigure 2 or Figure 5 all or part of the steps of any embodiment, such as Figure 7 as shown, the information processing device includes:
[0167] A first acquisition module 710, configured to acquire a first pooling vector of a target user; the first pooling vector is a vector generated in the process of processing the user features of the target user through a two-tower structure of an information processing model to obtain a user tower vector of the target user;
[0168] A second acquisition module 720, configured to input the first pooling vector into a feature auto-cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer;
[0169] A first splicing module 730, configured to splice the first cross-equivalent vector and the user tower vector to obtain a first target vector of the target user;
[0170] A vector processing module 740, configured to process the first target vector and a second target vector of target information in a target manner to obtain a vector processing result; the target manner is determined based on a processing manner of the target information; the second target vector is obtained after the information processing model processes information features of the target information;
[0171] An information processing module 750, configured to process the target information based on the vector processing result.
[0172] In a possible implementation manner, the first pooling vector includes K-dimensional features of the target user in n feature domains; n and K are positive integers;
[0173] The second acquisition module 720 includes:
[0174] A first feature acquisition sub-module, configured to perform bitwise summation on the K-dimensional features in n feature domains of the first pooling vector to obtain a first summation feature;
[0175] A first construction sub-module, configured to construct the first cross-equivalent vector based on the first summation feature and the first pooling vector. [[ID=!34]]
[0176] In a possible implementation manner, the device further includes:
[0177] A third acquisition module, configured to acquire a second pooling vector of target information; the second pooling vector is a vector generated in the process of processing information features of the target information through the two-tower structure to obtain an information tower vector of the target information;
[0178] A fourth acquisition module, configured to input the second pooling vector into the feature auto-cross layer to obtain a second cross-equivalent vector corresponding to the target information output by the feature auto-cross layer;
[0179] A second splicing module, configured to splice the second cross-equivalent vector and the information tower vector to obtain a second target vector of the target information.
[0180] In a possible implementation, the second pooling vector includes K-dimensional features of the target information in m feature domains; m and K are positive integers;
[0181] The fourth acquisition module includes:
[0182] A second feature acquisition sub-module, configured to perform bitwise summation on the K-dimensional features in m feature domains of the second pooling vector to obtain a second summation feature;
[0183] A second construction sub-module, configured to construct the second cross-equivalent vector based on the second summation feature and the second pooling vector.
[0184] In a possible implementation, the first cross-equivalent vector is [U, 1, 0.5 * (U ⊙ U - sum(X_user 2 ))]; the second cross-equivalent vector is [I, 0.5 * (I ⊙ I - sum(X_item 2 ))], 1];
[0185] Wherein, U represents the first summation feature, X_user represents the first pooling vector, I represents the second summation feature, X_item represents the second pooling vector, 1 represents a vector with a dimension of 1 and a vector value of 1, and ⊙ represents the inner product between vectors.
[0186] In a possible implementation, the vector processing module 740 includes:
[0187] A similarity calculation sub-module, configured to calculate the cosine similarity between the first target vector and the second target vector when the target processing method for the target information is recall;
[0188] A position determination sub-module, configured to determine the position of the target information in the recall sequence based on the cosine similarity; at least two pieces of information to be recalled are included in the recall sequence;
[0189] The information processing module 750 is configured to recall the target information when the position of the target information in the recall sequence meets the recall condition.
[0190] In a possible implementation, the vector processing module 740 includes:
[0191] A dot product calculation sub-module, configured to perform a dot product calculation on the first target vector and the second target vector to obtain a dot product result;
[0192] A fraction acquisition sub-module, configured to process the dot product result through an activation function to obtain a sorting fraction of the target information;
[0193] The information processing module 750 is configured to determine the position of the target information in the rough sorting sequence based on the sorting fraction; at least two rough sorting information are included in the rough sorting sequence.
[0194] In summary, the information processing device provided in the embodiments of the present application introduces a feature automatic cross layer on the basis of the original dual tower model structure. The feature automatic cross layer can obtain a first cross-equivalent vector on the user side, and the first cross-equivalent vector is used to represent the user side part in the feature interaction result between the user side and the information side, and is concatenated with the user tower vector obtained after processing the user features by the dual tower structure to form a first target vector on the user side, and synchronously or asynchronously obtains a second target vector on the information side. Through the processing of the first target vector and the second target vector, the obtained vector processing result realizes the processing of the target information. By introducing the feature automatic cross layer, the information interaction between the user side and the information side can be improved, and the expression ability of the original dual tower structure can be improved, thereby improving the processing effect of the target information.
[0195] Figure 8 FIG. shows a block diagram of the structure of a computer device 800 shown in an exemplary embodiment of the present application. The computer device can be implemented as the information processing model generation device or the information processing device in the above solution of the present application. The computer device 800 includes a central processing unit (CPU) 801, a system memory 804 including a random access memory (RAM) 802 and a read-only memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 also includes a mass storage device 806 for storing an operating system 809, a client 810, and other program modules 811.
[0196] Without loss of generality, the computer-readable medium may include a computer storage medium and a communication medium. The computer storage medium includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. The computer storage medium includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM), flash memory or other solid-state memory technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will understand that the computer storage medium is not limited to the above several types. The above-mentioned system memory 804 and mass storage device 806 can be collectively referred to as memory.
[0197] According to various embodiments of the present application, the computer device 800 may also run on a remote computer on the network through a network such as the Internet. That is, the computer device 800 may be connected to the network 808 through the network interface unit 807 connected to the system bus 805, or in other words, the network interface unit 807 may also be used to connect to other types of networks or remote computer systems (not shown).
[0198] The memory also includes at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, at least one program, code set or instruction set is stored in the memory, and the central processing unit 801 implements all or part of the steps in the information processing method shown in the above various embodiments by executing the at least one instruction, at least one program, code set or instruction set.
[0199] Figure 9 The block diagram of the computer device 900 shown in an exemplary embodiment of the present application is shown. The computer device 900 may be implemented as the above-mentioned information processing model generation device or information processing device, such as: smart phones, tablet computers, notebook computers, desktop computers, smart watches, and televisions. The computer device 900 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0200] Generally, the computer device 900 includes: a processor 901 and a memory 902.
[0201] In some embodiments, the computer device 900 may further optionally include: a peripheral device interface 903 and at least one peripheral device. The processor 901, the memory 902, and the peripheral device interface 903 may be connected via a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 903 via a bus, signal lines, or a circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 904, a display screen 905, a camera assembly 906, an audio circuit 907, and a power supply 908.
[0202] In some embodiments, the computer device 900 further includes one or more sensors 909. The one or more sensors 909 include, but are not limited to: an acceleration sensor 910, a gyroscope sensor 911, a pressure sensor 912, an optical sensor 913, and a proximity sensor 914.
[0203] Those skilled in the art can understand that Figure 9 the structure shown in does not constitute a limitation on the computer device 900, and it may include more or fewer components than shown in the figure, combine certain components, or adopt a different component arrangement.
[0204] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by the processor to implement all or part of the steps in the above information processing method. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage device, etc.
[0205] In an exemplary embodiment, a computer program product is further provided. The computer program product includes at least one computer program, and the computer program is loaded and executed by the processor to implement all or part of the steps in the above Figure 2 、 Figure 4 or Figure 5 any of the information processing methods shown in any of the embodiments.
[0206] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, and these variations, uses, or adaptations follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the claims.
[0207] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An information processing method, characterized in that, The method includes: Obtaining a first pooling vector of a target user; the first pooling vector is a vector generated in the process of processing the user features of the target user through the two-tower structure of an information processing model to obtain a user tower vector of the target user; Inputting the first pooling vector into a feature auto-cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer; Concatenating the first cross-equivalent vector and the user tower vector to obtain a first target vector of the target user; Processing the first target vector and a second target vector of target information in a target manner to obtain a vector processing result; the target manner is determined based on a target processing manner for the target information; the second target vector is obtained after the information processing model processes the information features of the target information; Processing the target information in the target processing manner based on the vector processing result.
2. The method according to claim 1, wherein The first pooling vector includes K-dimensional features of the target user in n feature domains; n and K are positive integers; The step of inputting the first pooling vector into a feature auto-cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature auto-cross layer includes: Performing a bitwise summation on the K-dimensional features in n feature domains of the first pooling vector to obtain a first summation feature; Constructing the first cross-equivalent vector based on the first summation feature and the first pooling vector.
3. The method according to claim 2, wherein The method further includes: Obtaining a second pooling vector of target information; the second pooling vector is a vector generated in the process of processing the information features of the target information through the two-tower structure to obtain an information tower vector of the target information; Inputting the second pooling vector into the feature auto-cross layer to obtain a second cross-equivalent vector corresponding to the target information output by the feature auto-cross layer; Concatenating the second cross-equivalent vector and the information tower vector to obtain the second target vector of the target information.
4. The method according to claim 3, wherein The second pooling vector includes K-dimensional features of the target information in m feature domains; m and K are positive integers; The step of inputting the second pooling vector into the feature auto-cross layer to obtain a second cross-equivalent vector corresponding to the target information output by the feature auto-cross layer includes: Performing a bitwise summation on the K-dimensional features in m feature domains of the second pooling vector to obtain a second summation feature; Constructing the second cross-equivalent vector based on the second summation feature and the second pooling vector.
5. The method according to claim 4, characterized in that, The first cross-equivalent vector is ; The second cross-equivalent vector is ; Wherein, U represents the first summation feature, X_user represents the first pooling vector, I represents the second summation feature, X_item represents the second pooling vector, and 1 represents a vector with a dimension of 1 and a vector value of 1. represents the inner product between vectors.
6. The method according to claim 1, characterized in that When the target processing manner for the target information is recall, the step of processing the first target vector and the second target vector in a target manner to obtain a vector processing result includes: Calculating the cosine similarity between the first target vector and the second target vector; Determining the position of the target information in a recall sequence based on the cosine similarity; the recall sequence includes at least two pieces of information to be recalled; Based on the vector processing result, process the target information according to the target processing method: When the position of the target information in the recall sequence meets the recall condition, recall the target information.
7. The method according to claim 1, characterized in that When the target processing method for the target information is rough ranking, processing the first target vector and the second target vector according to the target method to obtain a vector processing result, including: Perform a dot product calculation on the first target vector and the second target vector to obtain a dot product result; Process the dot product result through an activation function to obtain the ranking score of the target information; Based on the vector processing result, process the target information according to the target processing method: Based on the ranking score, determine the position of the target information in the rough ranking sequence; the rough ranking sequence contains at least two rough ranking information.
8. An information processing apparatus, characterized in that, The device includes: A first acquisition module, configured to acquire a first pooled vector of a target user; the first pooled vector is a vector generated in the process of processing the user feature of the target user through the two-tower structure of an information processing model to obtain the user tower vector of the target user; A second acquisition module, configured to input the first pooled vector into the feature automatic cross layer of the information processing model to obtain a first cross-equivalent vector corresponding to the target user output by the feature automatic cross layer; A first splicing module, configured to splice the first cross-equivalent vector and the user tower vector to obtain a first target vector of the target user; A vector processing module, configured to process the first target vector and a second target vector of the target information according to a target method to obtain a vector processing result; the target method is determined based on the processing method for the target information; the second target vector is obtained after the information processing model processes the information feature of the target information; An information processing module, configured to process the target information based on the vector processing result.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the information processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement the information processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Feature preprocessing method and device, electronic equipment and storage medium
CN114677170A
Method and apparatus for generating user tag, storage medium and computer device
WO2020207196A1