Cross feature processing method and device, model training method and device, equipment and medium
By extracting features and performing vector cross processing on interactive data in e-commerce live streaming scenarios, calculating gating weights and filtering important cross features, the problem of feature interaction modeling in e-commerce live streaming is solved, and the accuracy and efficiency of the model are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-19
- Publication Date
- 2026-04-07
AI Technical Summary
In the high dynamic and high sparsity scenarios of e-commerce live streaming, how can we effectively model the interaction relationships between features to improve model performance?
By extracting features from the interaction data, forming vector cross pairs, calculating the gating weights of the vector cross pairs, filtering out important vector cross pairs, obtaining feature selection results, and finally determining the interaction detection results between objects.
It improves the effectiveness of cross features, enhances the generalization ability of the model, reduces noise and computational redundancy, and improves the accuracy of cross detection results.
Smart Images

Figure CN121808674A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing, specifically to the fields of deep learning, artificial intelligence and large model technology, and in particular to a cross-feature processing, model training method, apparatus, device and medium. Background Technology
[0002] In recent years, with the deep integration of e-commerce and live streaming technologies, the popularity of e-commerce live streaming has continued to rise.
[0003] In highly dynamic and sparse scenarios such as e-commerce live streaming, effectively modeling the interaction relationships between features is the core issue for improving model performance. Summary of the Invention
[0004] This disclosure provides a method, apparatus, device, and medium for cross-feature processing and model training.
[0005] According to one aspect of this disclosure, a cross-feature processing method is provided, comprising:
[0006] Feature extraction is performed on each feature in the interactive data to obtain the feature vector of each feature; the interactive data includes: at least one feature of a first object and at least one feature of a second object;
[0007] The aforementioned feature vectors are combined to form at least one vector cross pair;
[0008] Based on each of the aforementioned feature vectors, calculate the gating weights for each of the aforementioned vector cross pairs;
[0009] The feature selection results are obtained by filtering each vector cross pair according to the gating weight of each vector cross pair;
[0010] Based on the feature vectors and feature selection results, the interaction detection results between the first object and the second object are determined.
[0011] According to one aspect of this disclosure, a method for training a cross-feature processing model is provided, comprising:
[0012] The feature extraction module is used to extract features from each feature in the interactive data to obtain feature vectors for each feature; the interactive data includes: at least one feature of a first object and at least one feature of a second object;
[0013] The vector crossing module is used to combine the feature vectors to form at least one vector crossing pair;
[0014] The gating weight calculation module is used to calculate the gating weight of each vector cross pair based on each of the feature vectors.
[0015] The feature selection module is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result;
[0016] The interaction prediction module is used to determine the interaction detection result between the first object and the second object based on the feature vectors and feature selection results.
[0017] According to one aspect of this disclosure, a cross-feature processing apparatus is provided, comprising:
[0018] The feature extraction module is used to extract features from each feature in the interactive data to obtain feature vectors for each feature; the interactive data includes: at least one feature of a first object and at least one feature of a second object;
[0019] The vector crossing module is used to combine the feature vectors to form at least one vector crossing pair;
[0020] The gating weight calculation module is used to calculate the gating weight of each vector cross pair based on each of the feature vectors.
[0021] The feature selection module is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result;
[0022] The interaction prediction module is used to determine the interaction detection result between the first object and the second object based on the feature vectors and feature selection results.
[0023] According to one aspect of this disclosure, a training apparatus for a cross-feature processing model is provided, comprising:
[0024] A cross-feature processing model is used to extract features from each feature in the training samples to obtain feature vectors for each feature; the training samples include: at least one feature of a first object, at least one feature of a second object, and interaction events between the first object and the second object;
[0025] The cross-feature processing model is used to combine the feature vectors to form at least one vector cross pair.
[0026] The cross-feature processing model is used to calculate the gating weights of each of the vector cross pairs based on each of the feature vectors.
[0027] The cross-feature processing model is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result;
[0028] The cross-feature processing model is used to determine the interaction detection result between the first object and the second object based on each feature vector and the feature selection result;
[0029] The model parameter tuning module is used to adjust the parameters of the cross-feature processing model based on the difference between the interaction detection results and the interaction events between the first object and the second object in the training samples.
[0030] According to another aspect of this disclosure, an electronic device is provided, comprising:
[0031] At least one processor; and
[0032] A memory communicatively connected to the at least one processor; wherein,
[0033] The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the cross-feature processing method or the training method of the cross-feature processing model as described in any embodiment of this disclosure.
[0034] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the cross-feature processing method or the training method of the cross-feature processing model described in any embodiment of this disclosure.
[0035] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the cross-feature processing method or the training method of the cross-feature processing model described in any embodiment of this disclosure.
[0036] The embodiments disclosed herein can improve the effectiveness of cross features and enhance the model generalization ability of sparse features.
[0037] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0038] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:
[0039] Figure 1 This is a flowchart of a cross-feature processing method disclosed in an embodiment of this disclosure;
[0040] Figure 2 This is a flowchart of a training method for a cross-feature processing model disclosed in an embodiment of this disclosure;
[0041] Figure 3 This is a schematic diagram of the structure of a cross-feature processing model disclosed in an embodiment of this disclosure;
[0042] Figure 4 This is a schematic diagram of the cross-feature processing device disclosed in the embodiments of this disclosure;
[0043] Figure 5 This is a schematic diagram of the structure of a training device for a cross-feature processing model disclosed in an embodiment of this disclosure;
[0044] Figure 6 This is a block diagram of an electronic device according to the cross-feature processing method or cross-feature processing model training method disclosed in the embodiments of this disclosure. Detailed Implementation
[0045] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0046] Figure 1 This is a flowchart of a cross-feature processing method disclosed in an embodiment of this disclosure. This embodiment can be applied to situations where cross-feature information between the features of interactive objects needs to be processed. The method of this embodiment can be executed by a cross-feature processing device, which can be implemented in software and / or hardware and specifically configured in an electronic device with a certain data processing capability. The electronic device can be a terminal device or a server. The terminal device can include mobile phones, tablets, laptops, personal computers, vehicle terminals, or wearable devices, etc., wherein wearable devices can include watches or smart glasses, etc.
[0047] S101. Extract features from each feature in the interactive data to obtain feature vectors for each feature; the interactive data includes at least one feature of the first object and at least one feature of the second object.
[0048] Interaction data can refer to the information of two objects that can interact. Interaction data includes multiple features, with at least features of the first object and at least features of the second object. Features can be the object's raw data, or they can be obtained by encrypting or extracting features from the raw data. Features describe the information of the object itself. Feature extraction is independent of each feature; for each feature, feature extraction is performed to obtain its feature vector.
[0049] In some embodiments, the features are embedded, mapping them to dense vectors. This is based on the following formula for f... i Feature extraction is performed on the features to obtain the feature vector e. i :
[0050]
[0051] Among them, f i The feature is the i-th feature, e i Let be the i-th feature vector, N be the total number of features, and d be the embedding dimension, for example, d is 16 or 32.
[0052] The first object can interact with the second object. The first object may include users and devices. The second object may include users, devices, products, web pages, and information, etc. In some embodiments, the first object is a user in the live stream, and the second object is an item in the live stream. The characteristics of the first object include at least the identifying characteristics of the first object. The characteristics of the second object include at least the identifying characteristics of the second object. The identifying characteristics are used to distinguish different objects.
[0053] S102. Combine each of the feature vectors to form at least one vector cross pair.
[0054] In this process, feature vectors are combined pairwise to form vector cross pairs. The interaction between two objects is often influenced by the combined effects of multiple features. To extract the relationships between these features, feature vectors can be crossed. Obtaining the weights of feature crosses reflects the importance of specific feature combinations, enhancing interpretability. For scenarios involving user identifier feature coefficients, crossing user identifier features with item identifier features yields user-item combinations that can be collected from historical records. This allows for the extraction of more realistic feature information, enabling model prediction based on richer feature content and improving the accuracy of the output results.
[0055] S103. Calculate the gating weights of each vector cross pair based on each of the aforementioned feature vectors.
[0056] In this process, a gating weight is calculated for each vector cross pair based on all the feature vectors. Each feature vector represents the global context. The gating weight can be understood as a gating signal used to determine whether a vector cross pair is important and whether to retain it for subsequent cross calculations. The gating weights of the vector cross pairs are determined based on the feature vectors representing the global context, allowing the gating weights to dynamically adapt to different scenarios (the scenario corresponding to the global context) or tasks (the task corresponding to the global context).
[0057] In some embodiments, the feature vectors can be fused, and the gating weights can be calculated based on the fusion result and the gating parameters. The gating parameters can be pre-trained.
[0058] S104. Based on the gating weights of each vector cross pair, filter each vector cross pair to obtain the feature selection result.
[0059] Specifically, based on the gating weights of the vector cross pairs, important and unimportant vector cross pairs can be identified. Only important vector cross pairs can be retained, while unimportant ones can be discarded. The feature selection result is then calculated based on the important vector cross pairs. The feature selection result can refer to retaining feature cross combinations that are effective for the current task.
[0060] S105. Based on the feature vectors and feature selection results, determine the interaction detection results between the first object and the second object.
[0061] Each feature vector provides the features of the two objects themselves. The feature selection result is used to adaptively activate effective feature cross combinations according to the task. The interaction detection result can refer to the detection result of interaction between the first object and the second object. In some embodiments, the interaction detection result can include the probability that the first object interacts with the second object, the type of interaction between the first object and the second object, the probability that the first object and the second object have an interaction relationship, and the direction of interaction between the first object and the second object, etc. In some embodiments, such as in an e-commerce live streaming scenario, the first object is the audience user in the live stream, and the second object is the item in the live stream; the interaction detection result can be the click rate of the audience user on the item.
[0062] In some embodiments, the feature vectors and feature selection results can be fused, and the fused result can be nonlinearly mapped to obtain the interactive detection result.
[0063] In some embodiments, after obtaining the interaction detection results, the interaction detection results of the first object with multiple second objects can be obtained. The corresponding second objects are then sorted according to each interaction detection result, and information from at least one second object with the highest interaction probability is pushed to the first object.
[0064] The cross-feature processing method of this disclosure can be executed by a pre-trained cross-feature processing model. Specifically, the cross-feature processing method includes: extracting features from each feature in the interaction data using the cross-feature processing model to obtain feature vectors for each feature; the interaction data includes: at least one feature of a first object and at least one feature of a second object; combining the feature vectors using the cross-feature processing model to form at least one vector cross pair; calculating the gating weights of each vector cross pair using the cross-feature processing model based on each feature vector; filtering each vector cross pair using the cross-feature processing model based on the gating weights to obtain a feature selection result; and determining the interaction detection result between the first object and the second object using the cross-feature processing model based on the feature vectors and the feature selection result.
[0065] In some embodiments, feature extraction is performed on each feature in the interaction data through the embedding layer in the cross-feature processing model to obtain feature vectors for each feature; the interaction data includes at least one feature of a first object and at least one feature of a second object; the feature vectors are combined by the gating layer in the cross-feature processing model to form at least one vector cross pair; the gating layer in the cross-feature processing model calculates the gating weights of each vector cross pair based on each feature vector; the gating layer in the cross-feature processing model filters each vector cross pair based on the gating weights to obtain feature selection results; the prediction layer in the cross-feature processing model determines the interaction detection result between the first object and the second object based on each feature vector and the feature selection result.
[0066] According to the technical solution of this disclosure, multiple features in the interactive data are combined to obtain vector cross pairs, and the gating weights of each vector cross pair are calculated based on each feature vector. The feature pairs are then selected based on the gating weights to obtain feature selection results. This achieves the filtering of important cross feature combinations based on the global context, reduces noise and computational redundancy caused by invalid crosses, reduces latency, and retains valid cross terms, thereby improving the accuracy of cross detection results.
[0067] In an optional embodiment, calculating the gating weights of each vector cross pair based on each feature vector includes: fusing each feature vector to obtain a context vector; and performing a nonlinear mapping on the context vector using the gating parameters of each vector cross pair to obtain the gating weights of the vector cross pair.
[0068] The feature vectors are fused to obtain the feature representation of the global context, i.e., the context vector. The context vector is used for global perception and information fusion. Furthermore, the features in the interactive data are independent of each other, can belong to different domains, and can be features determined by data from different modalities. By fusing the feature vectors, multimodal fusion and cross-domain fusion can be achieved. The gating weights of the vector cross pairs are used to describe whether the vector cross pair is activated.
[0069] The gating parameters can be pre-trained. The gating parameters correspond to vector cross pairs. The gating parameters are used to perform a linear transformation on the context vector, resulting in a linear transformation result. An activation function is then used to perform a non-linear mapping on the linear transformation result, yielding the gating weights.
[0070] The gating weight g can be calculated using the following formula. ij :
[0071]
[0072] Among them, g ij For the i-th eigenvector e i and the j-th eigenvector e j The gating weights of the resulting vector cross pairs, g ij ∈[0,1],w ij and b ij It's a gating parameter, w ij b is the gating weight parameter. ij Here, σ is the gate bias parameter, h is the sigmoid activation function, and h is the gating bias parameter. ctx This is the context vector.
[0073] As can be seen, by fusing the feature vectors to obtain the context vector, and using vector cross-mapping to perform nonlinear mapping on the context vector with the corresponding gating parameters to obtain the gating weight, the gating signal can be generated by using the context vector with fused global features. This can generate the gating signal based on dynamically changing global features and can be adapted to cross-feature processing tasks in different scenarios.
[0074] In an optional embodiment, fusing the feature vectors to obtain a context vector includes: performing at least one nonlinear transformation on each feature vector to obtain a context vector; the nonlinear transformation includes: linear combination and nonlinear mapping.
[0075] The feature vectors can be input into a Multi-Layer Perceptron (MLP), which performs at least one nonlinear transformation on each feature vector. The MLP can be pre-trained. It consists of an input layer, hidden layers, and an output layer. The number of nonlinear transformations corresponds to the number of hidden layers. The hidden layers perform linear transformations on each feature vector and fuse the results to obtain a fused result. This fused result is then subjected to a nonlinear transformation to obtain an intermediate result. If the MLP has only one hidden layer, the intermediate result output from the first and last hidden layer is used as the context vector. If the MLP has at least two hidden layers, nonlinear transformations can be performed on the intermediate results, and the intermediate result output from the last hidden layer is used as the context vector.
[0076] Nonlinear transformations can be performed based on the following formula:
[0077]
[0078] Where, d ctx This represents the dimension of the context vector.
[0079] It is evident that by performing at least one nonlinear transformation on each feature vector, noise features can be suppressed, effective features can be enhanced, global features can be compressed and enhanced, the representativeness of context vectors can be improved, and a gating signal can be generated. Based on a deep understanding of the global state, effective vector cross pairs can be activated to achieve the effect of adaptive activation.
[0080] In an optional embodiment, filtering each vector cross pair based on its gating weights to obtain a feature selection result includes: for each vector cross pair, when the gating weights of the vector cross pair meet the filtering conditions, performing element-wise multiplication of the two feature vectors in the vector cross pair to obtain the vector product of the vector cross pair; using the product of the vector product of the vector cross pair and the gating weights as the conditional cross result of the vector cross pair; when the gating weights of the vector cross pair meet the filtering conditions, determining that the conditional cross result of the vector cross pair is zero; and summing the conditional cross results of each vector cross pair to obtain the feature selection result.
[0081] The selection criteria are used to retain important vector cross pairs. The filtering criteria are used to remove unimportant vector cross pairs. In some embodiments, the selection criteria may be a gate weight greater than a preset weight threshold; the filtering criteria may be a gate weight less than or equal to a weight threshold. Element-wise multiplication can preserve the information of vector cross pairs in a fine-grained manner. The conditional cross result can refer to the feature interaction information dynamically adjusted and calibrated according to the context. The conditional cross result of a vector cross pair that meets the selection criteria is the product of the vector product of the vector cross pair and the gate weight, indicating the feature interaction information of important vector cross pairs. The conditional cross result of a vector cross pair that meets the filtering criteria is zero, indicating that the vector cross pair has been removed. The feature selection result obtained by summing the conditional cross results of each vector cross pair represents the retention of only the feature interaction information of important vector cross pairs, thus achieving the selection of important vector cross pairs and the exclusion of unimportant vector cross pairs.
[0082] The conditional crossover result c can be calculated based on the following formula. ij :
[0083]
[0084] τ is the weight threshold.
[0085] In some embodiments, the conditional cross results of each vector cross pair can be directly summed to obtain the feature selection result. In some embodiments, the conditional cross results of each vector cross pair can be weighted and summed to obtain the feature selection result, and the weights can be determined based on their own content or context.
[0086] The feature selection result x can be calculated based on the following formula. cross :
[0087]
[0088] As can be seen, by performing element-wise multiplication on the feature vectors in the vector cross pairs retained by the gating weights, complex interaction information can be captured, improving processing accuracy. Furthermore, redundant vector cross pairs are eliminated, allowing only a few cross terms to be activated, reducing noise interference and improving computational efficiency.
[0089] In an optional embodiment, determining the interaction detection result between the first object and the second object based on the feature vectors and feature selection results includes: obtaining an input vector formed by concatenating the feature vectors; concatenating the input vector and the feature selection results to obtain a fused feature; and performing a nonlinear mapping on the fused feature to obtain the interaction detection result between the first object and the second object.
[0090] The feature vectors can be concatenated using the following formula to generate the input vector x:
[0091]
[0092] Where D is the dimension of the input vector.
[0093] The input vector x and the feature selection result x can be concatenated using the following formula. cross Generate fused feature z:
[0094] z=Concat(x,x cross )
[0095] In some embodiments, the fused feature z can be nonlinearly mapped using the following formula to obtain the interactive detection result y:
[0096]
[0097] Among them, w out To output the weight parameters, b out This is the output bias parameter. out and b out It can be obtained through training.
[0098] In some embodiments, an MLP can be used to perform at least one nonlinear transformation on the fused features to obtain the interaction detection results.
[0099] It is evident that by performing nonlinear mapping on the fused features obtained by concatenating the feature vectors and feature selection results, interactive detection results can be obtained. Prediction can be made based on all features and the selected vector cross pairs. All features can be used as the basis, and the effectiveness information provided by the selected feature cross pairs can be combined to improve the comprehensiveness and effectiveness of the features. In turn, nonlinear mapping can obtain interactive detection results, making the interactive detection results more accurate.
[0100] In an optional embodiment, the features of the first object include: user identification features and user attribute features; the features of the second object include: interaction object identification features and interaction object attribute features.
[0101] In this framework, the first object is the user, and the second object is the user's interaction object. For example, the interaction object could be an item. Items can be real or virtual. User identification features can refer to characteristics that distinguish different users. User attribute features can refer to stable user characteristics, such as attributes describing a user's inherent nature, state, preferences, and group affiliation. Interaction object identification features can refer to characteristics that distinguish different interaction objects. Interaction object attribute features can refer to stable interaction object characteristics, such as attributes describing the interaction object's inherent properties, external appearance, and functionality.
[0102] It is evident that by defining the characteristics of the first object and the second object, we can focus on effective information related to the interaction, avoid noise interference, adapt to the interaction scenarios between users and interactive objects, and improve the accuracy of interaction detection results.
[0103] In an optional embodiment, the characteristics of the second object further include: scene source characteristics.
[0104] The scene source feature is used to limit the scene from which the feature originates. When the cross-feature processing task is triggered under a specific scene, scene-related information is obtained and the scene source feature is determined.
[0105] For example, the scenario source feature is the information feed two-hop feature. An information feed two-hop can refer to the behavioral data generated immediately after a user clicks on the first piece of content in the feed (one hop), at a related or next location (such as continuing to browse, clicking on a second recommended piece of content, dwelling, or swiping). In fact, compared to one-hop scenarios, two-hop scenarios exhibit higher sparsity in features across almost all dimensions. Information feed two-hop features can include features definite from the main feed page, the discovery page, the tag navigation bar, and push notifications. In one example, the first object is the viewers in the live stream, the second object is the items in the live stream, and the scenario source feature is the live stream that the user jumps to via the information feed two-hop.
[0106] It is evident that by limiting the interactive data to include scene source features, we can better adapt the dynamic construction of gating signals to the scene, filter out features that are suitable for the scene, increase the flexibility of the gating signals, accurately adapt to diverse scenes, and significantly improve detection performance.
[0107] Figure 2 This is a flowchart of a training method for a cross-feature processing model disclosed in an embodiment of this disclosure. This embodiment can be applied to the training of a cross-feature processing model, which is used to process the cross-information between the features of interactive objects. The method of this embodiment can be executed by a cross-feature processing device, which can be implemented in software and / or hardware and specifically configured in an electronic device with a certain data processing capability. The electronic device can be a terminal device or a server. The terminal device can include mobile phones, tablets, laptops, personal computers, vehicle terminals, or wearable devices, etc., wherein wearable devices can include watches or smart glasses, etc.
[0108] S201. Through a cross-feature processing model, feature extraction is performed on each feature in the training sample to obtain the feature vector of each feature; the training sample includes: at least one feature of the first object, at least one feature of the second object, and the interaction event between the first object and the second object.
[0109] Feature vectors for each feature are obtained by extracting features from the training samples through the embedding layer in the cross-feature processing model. An interaction event can refer to an observable behavioral instance between a first object and a second object, reflecting the relationship and interaction between the two objects. An interaction event can be empty, and the number of interaction events can be at least one. For example, if the first object is a user and the second object is an item, the interaction event could be a user clicking on an item.
[0110] The acquisition of interaction events is authorized by the user. Alternatively, the interaction events may be processed by encrypting or feature-processing the data before analysis.
[0111] S202. Using the cross-feature processing model, the feature vectors are combined to form at least one vector cross pair.
[0112] The feature vectors are combined by the gating layer in the cross-feature processing model to form at least one vector cross pair.
[0113] S203. Using the cross-feature processing model, calculate the gating weights of each vector cross pair based on each feature vector.
[0114] In the cross-feature processing model, the gating layer calculates the gating weights of each vector cross pair based on each feature vector.
[0115] S204. Using the cross-feature processing model, each vector cross pair is screened according to the gating weights of each vector cross pair to obtain the feature selection result.
[0116] The feature selection result is obtained by filtering each vector cross pair by the gating layer in the cross feature processing model according to the gating weight of each vector cross pair.
[0117] S205. Using the cross-feature processing model, determine the interaction detection result between the first object and the second object based on the feature vectors and feature selection results.
[0118] The interaction detection result between the first object and the second object is determined by the prediction layer in the cross-feature processing model based on the feature vectors and feature selection results.
[0119] S206. Adjust the parameters of the cross-feature processing model based on the difference between the interaction detection results and the interaction events between the first object and the second object in the training samples.
[0120] The interaction detection result can be the probability of an interaction event occurring between the first object and the second object. If the interaction event between the first object and the second object exists in the training sample, the corresponding ground truth probability is determined to be 100%. If the interaction event between the first object and the second object does not exist in the training sample, the corresponding ground truth probability is determined to be 0. The difference is the difference between the probability value determined in the training sample and the probability value in the interaction detection result.
[0121] In some embodiments, the first object is a user, the second object is an item, and the interaction event can be a user clicking on an item. Correspondingly, the interaction detection result is the click-through rate (CTR). Based on the user-item interaction events in the training samples, the user's click-through rate is calculated. Specifically, the click-through rate is the ratio between the number of times the user clicks on the item and the number of times the item is displayed. The difference is the difference between the click-through rate determined in the training samples and the click-through rate in the interaction detection result.
[0122] The parameters of the cross-feature processing model are optimized to reduce discrepancies. The cross-feature processing model is considered to have completed training when its accuracy on the validation set exceeds a preset accuracy threshold, or when the discrepancies converge or reach a minimum.
[0123] In some optional embodiments, calculating the gating weights of each vector cross pair based on each feature vector includes: fusing each feature vector to obtain a context vector; and performing a nonlinear mapping on the context vector using the gating parameters of the vector cross pair to obtain the gating weights of the vector cross pair.
[0124] In some optional embodiments, fusing the feature vectors to obtain a context vector includes: performing at least one nonlinear transformation on each feature vector to obtain the context vector; the nonlinear transformation includes: linear combination and nonlinear mapping.
[0125] In some optional embodiments, filtering each vector cross pair based on its gating weights to obtain a feature selection result includes: for each vector cross pair, when the gating weights of the vector cross pair meet the filtering conditions, performing element-wise multiplication of the two feature vectors in the vector cross pair to obtain the vector product of the vector cross pair; using the product of the vector product of the vector cross pair and the gating weights as the conditional cross result of the vector cross pair; when the gating weights of the vector cross pair meet the filtering conditions, determining that the conditional cross result of the vector cross pair is zero; and summing the conditional cross results of each vector cross pair to obtain the feature selection result.
[0126] In some optional embodiments, determining the interaction detection result between the first object and the second object based on the feature vectors and feature selection results includes: obtaining an input vector formed by concatenating the feature vectors; concatenating the input vector and the feature selection results to obtain a fused feature; and performing a nonlinear mapping on the fused feature to obtain the interaction detection result between the first object and the second object.
[0127] In some optional embodiments, the features of the first object include: user identification features and user attribute features; the features of the second object include: interaction object identification features and interaction object attribute features.
[0128] In some optional embodiments, the characteristics of the second object further include: scene source characteristics.
[0129] According to the technical solution disclosed herein, by introducing a sample-adaptive neural gating mechanism, the high-order feature cross paths that are effective for the current sample can be adaptively activated without sacrificing expressive power. This can be generalized to new scenarios or large-scale sparse feature scenarios, avoiding noise and computational redundancy caused by invalid crosses, improving the prediction accuracy of the cross feature processing model, reducing computation and storage overhead, reducing overfitting, and enhancing business controllability and interpretability.
[0130] In a scenario, such as Figure 3 As shown:
[0131] Cross-feature processing models can include embedding layers, gating layers, and prediction layers. The embedding layer extracts features from multiple features within the same training sample. The encoding results of the same feature are summed element-wise to obtain the feature vector e for that feature. i The feature vectors of each feature are concatenated to form the input vector x, which is then standardized. The standardized input vector is fed into the NeuronGate Layer. The NeuronGate Layer consists of fully connected layers that fuse the feature vectors to obtain the context vector h. ctx For each vector cross pair e formed by combining the feature vectors pairwise. i and e j The context vector is input into the fully connected layer in the gated layer to be gated according to the gating parameters w of the vector cross pairs. ij and b ij A nonlinear mapping is performed on the context vectors to obtain the gating weights g of the vector cross pairs. ij For each vector cross pair, the conditional cross layer in the gating layer multiplies the two feature vectors of the cross pair element-wise to obtain the vector product of that cross pair. Gating weights are then used to select the vector product of the cross pair to obtain the conditional cross result c.ij In the gated layer, a cross-aggregation layer sums (or weights pools) all activated vectors across pairs to obtain the feature selection result x. cross The input vector x and the feature selection result x cross The input is fed into the prediction layer, which can be an MLP layer. The prediction layer takes the input vector x and the feature selection result x as input. cross Perform at least one nonlinear transformation to determine the interaction detection results between the first and second objects, such as the user's click rate on the item.
[0132] The embodiments disclosed herein introduce a lightweight gating unit to perform conditional activation and sparsification selection of feature cross paths, which significantly improves training stability and inference efficiency while ensuring the model's expressive power.
[0133] According to embodiments of this disclosure, Figure 4 This is a structural diagram of the cross-feature processing device in an embodiment of this disclosure. This embodiment is applicable to situations where cross-feature information between the features of interactive objects is processed. The device is implemented in software and / or hardware and is specifically configured in an electronic device with a certain data processing capability.
[0134] like Figure 4 The illustrated cross-feature processing device 400 includes: a feature extraction module 401, a vector cross-processing module 402, a gating weight calculation module 403, a feature selection module 404, and an interactive prediction module 405. Among these,
[0135] The feature extraction module 401 is used to extract features from each feature in the interactive data to obtain feature vectors for each feature; the interactive data includes: at least one feature of a first object and at least one feature of a second object;
[0136] The vector crossing module 402 is used to combine the feature vectors to form at least one vector crossing pair;
[0137] The gating weight calculation module 403 is used to calculate the gating weight of each vector cross pair based on each of the feature vectors.
[0138] The feature selection module 404 is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result;
[0139] The interaction prediction module 405 is used to determine the interaction detection result between the first object and the second object based on the feature vectors and feature selection results.
[0140] The cross-feature processing method of this disclosure can be executed by a pre-trained cross-feature processing model. Specifically, the cross-feature processing method includes: extracting features from each feature in the interaction data using the cross-feature processing model to obtain feature vectors for each feature; the interaction data includes: at least one feature of a first object and at least one feature of a second object; combining the feature vectors using the cross-feature processing model to form at least one vector cross pair; calculating the gating weights of each vector cross pair using the cross-feature processing model based on each feature vector; filtering each vector cross pair using the cross-feature processing model based on the gating weights to obtain a feature selection result; and determining the interaction detection result between the first object and the second object using the cross-feature processing model based on the feature vectors and the feature selection result.
[0141] Optionally, the gating weight calculation module 403 includes:
[0142] The feature vector fusion unit is used to fuse the feature vectors to obtain a context vector;
[0143] A nonlinear processing unit is used to perform a nonlinear mapping on the context vector for each vector cross pair using the gating parameters of the vector cross pair, so as to obtain the gating weight of the vector cross pair.
[0144] Optional, the feature vector fusion unit includes:
[0145] The feature vector aggregation subunit is used to perform at least one nonlinear transformation on each of the feature vectors to obtain a context vector; the nonlinear transformation includes linear combination and nonlinear mapping.
[0146] Optionally, the feature selection module 404 includes:
[0147] An important combination processing unit is used to perform element-wise multiplication of the two feature vectors in each vector cross pair when the gating weight of the vector cross pair meets the screening conditions, so as to obtain the vector product of the vector cross pair.
[0148] The conditional crossover calculation unit is used to take the product of the vector product of the vector crossover pair and the gating weight as the conditional crossover result of the vector crossover pair;
[0149] The unimportant combination processing unit is used to determine that the conditional cross result of the vector cross pair is zero when the gating weights of the vector cross pair meet the filtering conditions.
[0150] The combined aggregation unit is used to sum the conditional cross results of each vector cross pair to obtain the feature selection result.
[0151] Optionally, the interaction prediction module 405 includes:
[0152] The feature vector concatenation unit is used to obtain the input vector formed by concatenating the feature vectors;
[0153] The selection result fusion unit is used to concatenate the input vector and the feature selection result to obtain the fused feature;
[0154] The interaction prediction unit is used to perform nonlinear mapping on the fused features to obtain the interaction detection results between the first object and the second object.
[0155] Optionally, the features of the first object include: user identification features and user attribute features; the features of the second object include: interaction object identification features and interaction object attribute features.
[0156] Optionally, the features of the second object may also include: scene source features.
[0157] The aforementioned cross-feature processing device can execute the cross-feature processing method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the cross-feature processing method.
[0158] According to embodiments of this disclosure, Figure 5 This is a structural diagram of the training device for the cross-feature processing model in this embodiment of the disclosure. This embodiment is applicable to the training of the cross-feature processing model, which is used to process the cross-information between the features of interactive objects. The device is implemented in software and / or hardware and is specifically configured in an electronic device with certain data processing capabilities.
[0159] like Figure 5 The training device 500 for a cross-feature processing model shown includes: a cross-feature processing model 501 and a model parameter tuning module 502. Wherein,
[0160] A cross-feature processing model 501 is used to extract features from each feature in the training samples to obtain feature vectors for each feature; the training samples include: at least one feature of a first object, at least one feature of a second object, and interaction events between the first object and the second object;
[0161] The cross-feature processing model 501 is used to combine the feature vectors to form at least one vector cross pair.
[0162] The cross-feature processing model 501 is used to calculate the gating weights of each of the vector cross pairs based on each of the feature vectors.
[0163] The cross-feature processing model 501 is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result;
[0164] The cross-feature processing model 501 is used to determine the interaction detection result between the first object and the second object based on each feature vector and the feature selection result;
[0165] The model parameter tuning module 502 is used to adjust the parameters of the cross-feature processing model based on the difference between the interaction detection results and the interaction events between the first object and the second object in the training samples.
[0166] According to the technical solution disclosed herein, by introducing a sample-adaptive neural gating mechanism, the high-order feature cross paths that are effective for the current sample can be adaptively activated without sacrificing expressive power. This can be generalized to new scenarios or large-scale sparse feature scenarios, avoiding noise and computational redundancy caused by invalid crosses, improving the prediction accuracy of the cross feature processing model, reducing computation and storage overhead, reducing overfitting, and enhancing business controllability and interpretability.
[0167] The training apparatus for the aforementioned cross-feature processing model can execute the training method for the cross-feature processing model provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the training method for the cross-feature processing model.
[0168] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0169] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0170] Figure 6 A schematic area diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein. The electronic device may be an autonomous vehicle.
[0171] like Figure 6As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required by the instructions of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0172] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of displays, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0173] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as cross-feature processing methods. For example, in some embodiments, the cross-feature processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the cross-feature processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform cross-feature processing methods by any other suitable means (e.g., by means of firmware).
[0174] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard objects (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0175] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / instructions specified in the flowcharts and / or area diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0176] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0177] To provide interaction with the first object, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the first object; and a keyboard and pointing device (e.g., a mouse or trackball) through which the first object provides input to the computer. Other types of devices can also be used to provide interaction with the first object; for example, feedback provided to the first object can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the first object can be received in any form (including sound input, voice input, or tactile input).
[0178] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a first-object computer having a graphical first-object interface or a web browser, through which the first object can interact with the implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication (e.g., a communication network) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0179] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem that addresses the management difficulties and weak business scalability inherent in traditional physical hosting and VPS services. Servers can also be servers for distributed systems or servers integrated with blockchain technology.
[0180] Artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.
[0181] Cloud computing refers to a technology system that enables access to a shared pool of physical or virtual resources via a network. These resources can include servers, instruction sets, networks, software, applications, and storage devices, and can be deployed and managed on demand and in a self-service manner. Cloud computing technology can provide efficient and powerful data processing capabilities for applications such as artificial intelligence and blockchain, as well as for model training.
[0182] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.
[0183] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for processing cross features, comprising: Feature extraction is performed on each feature in the interactive data to obtain the feature vector of each feature; The interactive data includes: at least one feature of the first object and at least one feature of the second object; The aforementioned feature vectors are combined to form at least one vector cross pair; Based on each of the aforementioned feature vectors, calculate the gating weights for each of the aforementioned vector cross pairs; The feature selection results are obtained by filtering each vector cross pair according to the gating weight of each vector cross pair; Based on the feature vectors and feature selection results, the interaction detection results between the first object and the second object are determined.
2. The method according to claim 1, wherein, The step of calculating the gating weights of each vector cross pair based on each of the feature vectors includes: The aforementioned feature vectors are fused to obtain a context vector; For each vector cross pair, the context vector is nonlinearly mapped using the gating parameters of the vector cross pair to obtain the gating weights of the vector cross pair.
3. The method according to claim 2, wherein, The process of fusing the feature vectors to obtain the context vector includes: Each of the aforementioned feature vectors undergoes at least one nonlinear transformation to obtain a context vector; the nonlinear transformation includes linear combination and nonlinear mapping.
4. The method according to claim 1, wherein, The step of filtering each vector cross pair according to the gating weights of each vector cross pair to obtain the feature selection result includes: For each of the vector cross pairs, when the gating weights of the vector cross pairs meet the screening conditions, the two feature vectors in the vector cross pairs are multiplied element by element to obtain the vector product of the vector cross pairs. The product of the vectors in the vector crossover pair and the product of the gating weights is taken as the conditional crossover result of the vector crossover pair. When the gating weights of the vector cross pair meet the filtering conditions, the conditional cross result of the vector cross pair is determined to be zero. The conditional cross results of each vector cross pair are summed to obtain the feature selection result.
5. The method according to claim 1, wherein, The step of determining the interaction detection result between the first object and the second object based on the feature vectors and feature selection results includes: Obtain the input vector formed by concatenating the aforementioned feature vectors; The input vector and the feature selection result are concatenated to obtain the fused feature; The fusion features are nonlinearly mapped to obtain the interaction detection results between the first object and the second object.
6. The method according to claim 1, wherein, The features of the first object include: user identification features and user attribute features; the features of the second object include: interaction object identification features and interaction object attribute features.
7. The method according to claim 6, wherein, The characteristics of the second object also include: scene source characteristics.
8. A training method for a cross-feature processing model, comprising: By using a cross-feature processing model, feature vectors of each feature are obtained by extracting features from the training samples. The training samples include: at least one feature of a first object, at least one feature of a second object, and interaction events between the first object and the second object; The cross-feature processing model is used to combine the feature vectors to form at least one vector cross pair. Using the cross-feature processing model, the gating weights of each vector cross pair are calculated based on each feature vector. The cross-feature processing model filters each vector cross pair according to the gating weights of each vector cross pair to obtain the feature selection result. Based on the cross-feature processing model, the interaction detection result between the first object and the second object is determined according to the feature vectors and feature selection results. The parameters of the cross-feature processing model are adjusted based on the difference between the interaction detection results and the interaction events between the first object and the second object in the training samples.
9. A cross-feature processing device, comprising: The feature extraction module is used to extract features from each feature in the interactive data to obtain the feature vector of each feature; The interactive data includes: at least one feature of the first object and at least one feature of the second object; The vector crossing module is used to combine the feature vectors to form at least one vector crossing pair; The gating weight calculation module is used to calculate the gating weight of each vector cross pair based on each of the feature vectors. The feature selection module is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result; The interaction prediction module is used to determine the interaction detection result between the first object and the second object based on the feature vectors and feature selection results.
10. A training device for a cross-feature processing model, comprising: A cross-feature processing model is used to extract features from each feature in the training samples to obtain the feature vector of each feature; The training samples include: at least one feature of a first object, at least one feature of a second object, and interaction events between the first object and the second object; The cross-feature processing model is used to combine the feature vectors to form at least one vector cross pair. The cross-feature processing model is used to calculate the gating weights of each of the vector cross pairs based on each of the feature vectors. The cross-feature processing model is used to filter each vector cross pair according to the gating weight of each vector cross pair to obtain the feature selection result; The cross-feature processing model is used to determine the interaction detection result between the first object and the second object based on each feature vector and the feature selection result; The model parameter tuning module is used to adjust the parameters of the cross-feature processing model based on the difference between the interaction detection results and the interaction events between the first object and the second object in the training samples.
11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the cross-feature processing method according to any one of claims 1-7, or the training method for the cross-feature processing model according to claim 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the cross-feature processing method according to any one of claims 1-7, or the training method of the cross-feature processing model according to claim 8.
13. A computer program product comprising a computer program that, when executed by a processor, implements the cross-feature processing method according to any one of claims 1-7, or the training method for the cross-feature processing model according to claim 8.