Fine rate prediction method and system fusing factorization feature interaction

By integrating the deep win rate prediction model (DeepFM-Rank) based on factorization feature interactions, the problem of insufficient low-order feature interactions and high-order expressive capabilities in online advertising real-time bidding systems is solved, achieving efficient and accurate win rate prediction, and is suitable for low-latency, high-throughput advertising systems.

CN120996274APending Publication Date: 2025-11-21GUANGZHOU TAIDONG TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511139829.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-14
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing win rate prediction models suffer from problems in online advertising real-time bidding systems, such as missing low-order feature interactions, limited high-order expressive power, heavy reliance on feature engineering, imbalance between training and inference performance, and lack of robustness to data sparsity, resulting in poor model performance.

Method used

The DeepFM-Rank model, which integrates factorization feature interaction, is used to output the win probability by mapping high-dimensional sparse features to dense embedding vectors and using multi-branch feature modeling, including explicit second-order feature interaction branches, nonlinear modeling branches and residual connection branches, combined with feature fusion mechanism.

Benefits of technology

It improves prediction accuracy and model stability, making it suitable for deployment in low-latency, high-throughput advertising systems. It also enhances the ability to fit complex behavioral patterns and improves the robustness of the model in scenarios with missing features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120996274A_ABST
    Figure CN120996274A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a winning rate prediction method and system fusing factorization feature interaction, and the method comprises the steps: obtaining high-dimensional sparse features of bidding data, and mapping the high-dimensional sparse features into dense embedded vectors; performing multi-branch feature modeling on the dense embedded vector to obtain multi-branch output information; performing feature fusion on the multi-branch output information to obtain a feature fusion result; wherein the feature fusion comprises the steps of summing the output of the second branch and the output of the third branch, and splicing the summed feature with the output of the first branch; and carrying out Sigmoid transformation and clip operation constraint on a feature fusion result to output a winning rate probability. According to the scheme, multi-dimensional balance of prediction precision, model stability and engineering efficiency is effectively guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing. More particularly, the present application relates to a win rate prediction method and system fusing factorization feature interaction. BACKGROUND

[0002] With the rapid development of the Internet advertising industry, real-time bidding (RTB) has become the mainstream mechanism of programmatic advertising. In RTB, the advertising platform needs to evaluate and bid for multiple advertisements within hundreds of milliseconds, and win rate prediction (WRP) as a key link directly affects the advertising cost of advertisers and the platform revenue. Therefore, building an efficient, accurate and deployable win rate prediction model has become one of the core problems of advertising system optimization.

[0003] Traditional win rate prediction methods usually use logistic regression or embedding + deep neural network (DNN) structure to model high-dimensional sparse advertising features.

[0004] Currently, in the online advertising real-time bidding system, in order to improve the accuracy and real-time performance of win rate prediction, the following modeling techniques are widely used in the industry:

[0005] 1) Logistic Regression model. Logistic Regression has the advantages of simple model structure and fast reasoning speed, and is widely used in early RTB systems. However, this model can only model linear feature relationships and is difficult to capture complex nonlinear interactions, so the prediction effect is limited.

[0006] 2) Embedding + multi-layer perceptron (DNN) structure. After embedding mapping of high-dimensional sparse features, input into deep neural network for high-order feature interaction learning, which is one of the current mainstream methods (such as Wide&Deep, DeepCTR). Although this method improves the model expression ability, it often ignores the explicit interaction between low-order features, resulting in insufficient information utilization.

[0007] 3) Models based on hand-crafted feature cross. Some systems design feature combination cross terms manually, and then input them into shallow or deep models for learning. However, this method relies on the experience of domain experts, has high feature engineering cost, and lacks generalization and automation capabilities.

[0008] 4) Factorization Machine (FM) model. FM can explicitly model second-order feature cross and is suitable for sparse data scenarios, but has limited expression ability and is difficult to capture high-order feature relationships, so the model performance has bottlenecks when used alone.

[0009] Although current mainstream win rate prediction methods have improved model performance to some extent, they still have the following significant drawbacks:

[0010] (1) Lack of low-order feature interactions. Traditional DNN models mainly rely on multi-layer nonlinear structures to learn feature interactions, making it difficult to explicitly model low-order combination relationships between features. Information is easily weakened in deep structures, causing the model to fail to fully utilize the cross information of basic features.

[0011] (2) Limited high-order expressive power. Although methods such as FM can effectively model second-order feature interactions, they lack nonlinear transformation capabilities and are difficult to capture complex high-order feature patterns. Their performance is limited by the dimension and structure of feature interactions.

[0012] (3) Heavy reliance on feature engineering. Some models rely on manually constructed cross features to enhance their expressive power. This approach is not only costly and inefficient, but also lacks automatic generalization ability and is difficult to adapt to the rapidly changing online advertising environment.

[0013] (4) Imbalance between training and inference performance. Although deep model structures have strong expressive power, they have a large number of parameters and high computational complexity, which can easily lead to problems such as unstable training, slow convergence and high online inference latency, affecting the response speed of the real-time bidding system.

[0014] (5) Sensitive to data sparsity and missing features. In practical applications, due to incomplete user behavior data and frequent changes in advertising context, the model faces problems of missing features and data sparsity. Existing models are prone to performance degradation and lack robust design in such situations.

[0015] Therefore, while these methods possess strong nonlinear expressive power, they lack the ability to model feature interactions, especially when low-order feature interactions are not explicitly considered, which can limit model performance. Furthermore, deep structures are prone to overfitting or training instability, particularly performing poorly in advertising scenarios with missing features or sparse data distribution.

[0016] Therefore, how to construct a win rate prediction model that can effectively model low-order interactions between features, has deep nonlinear expression capabilities, and is efficient and deployable in engineering implementation has become an important technical challenge that urgently needs to be solved in the current online advertising industry. Summary of the Invention

[0017] To address the aforementioned technical problems of low efficiency and poor performance in win rate prediction, this invention provides solutions in the following aspects.

[0018] In a first aspect, the present application provides a win rate prediction method of fusion factorization feature interaction, comprising: obtaining high-dimensional sparse features of bidding data, and mapping the high-dimensional sparse features into dense embedding vectors; performing multi-branch feature modeling on the dense embedding vectors to obtain multi-branch output information, wherein a first branch includes an explicit second-order feature interaction branch, a second branch includes a high-order nonlinear modeling branch, and a third branch includes a residual connection branch; performing feature fusion on the multi-branch output information to obtain a feature fusion result; wherein the feature fusion includes summing the output of the second branch and the output of the third branch, and splicing the summed features with the output of the first branch; performing Sigmoid transformation and clip operation constraint on the feature fusion result to output a win rate probability.

[0019] In one embodiment, obtaining high-dimensional sparse features of bidding data and mapping the high-dimensional sparse features into dense embedding vectors comprises: inputting various discrete category features of the bidding data into an embedding layer to obtain dense embedding vectors corresponding to the various discrete category features; and merging the dense embedding vectors to respectively output a three-dimensional tensor and a flattened vector.

[0020] In one embodiment, the modeling process of the first branch comprises: extracting embedding vectors of all fields from the three-dimensional tensor, and calculating a sum and a sum of squares of all embedding vectors; subtracting the square of the sum of all embedding vectors from all sum of squares to obtain a second-order interaction feature output; and mapping the second-order interaction feature to a first low-dimensional vector of a set length by using a full connection layer to obtain the output of the first branch.

[0021] In one embodiment, the calculation formula of the second-order interaction feature output is:

[0022]

[0023] The calculation formula of the output of the first branch is:

[0024] F fm =ReLU(W fm ·FM+b fm )

[0025] In the formula, FM represents the second-order interaction feature output, v i represents the embedding vector of the i-th field, represents the sum of all field embeddings, represents the sum of squares of the embedding vectors of all fields, F fm represents the first low-dimensional vector, W fm represents the weight matrix of the first low-dimensional vector, b fm represents the bias vector, and ReLU() represents the rectified linear unit activation function, ReLU(x) = max(0, x).

[0026] In one embodiment, the modeling process of the second branch includes: performing a multi-layer nonlinear transformation on the flattened vector, and outputting a high-order feature.

[0027] In one embodiment, the modeling process of the third branch includes: performing a linear transformation on the flattened vector to obtain a second low-dimensional vector compressed to a set length; and performing a nonlinear activation using a ReLU function to obtain a residual branch output vector.

[0028] In one embodiment, the calculation formula of the output of the third branch is:

[0029] F res =ReLU(W res ·E flat +b res )

[0030] In the formula, F res represents the residual branch output vector, W res represents the weight matrix of the residual connection, E flat represents the flattened vector, b res represents the bias vector of the residual layer, and ReLU() represents the activation function.

[0031] In one embodiment, the calculation formula of the process of adding and concatenating is:

[0032] F deep =F dnn +F res

[0033] F concat =concat(F deep ,F fm )

[0034] In the formula, F dnn represents the output of the second branch, F res represents the output of the third branch, F deep represents the deep feature representation after weighted fusion of the two, F concat represents the final feature fusion result, concat() represents the concatenation operation, and F fm represents the output of the first branch.

[0035] In one embodiment, the calculation formula of the Sigmoid transformation and clip operation constraint on the feature fusion result is:

[0036]

[0037] In the formula, P represents the win rate probability, with a value range of (0, 1), W out represents the weight matrix of the output layer, and bout represents an output layer bias term, and σ() represents a Sigmoid function defined as: represents an operation for limiting a value within a specified interval, and ε represents a very small positive number for preventing gradient instability caused by an output of 0 or 1.

[0038] In a second aspect, the present application also provides a win rate prediction system fusing factorization feature interaction, comprising: a processor; a memory storing computer program instructions, when the computer program instructions are executed by the processor, realizing a win rate prediction method fusing factorization feature interaction according to one or more embodiments of the first aspect.

[0039] The present application has the beneficial effects that: according to the scheme of the present application, by uniformly embedding and mapping the input high-dimensional sparse features to obtain a dense representation, and introducing three parallel sub-modules for multi-level feature modeling, a design framework is formed with the idea of "embedding representation + feature interaction modeling + multi-branch fusion" as the core, which combines the shallow feature cross and deep feature extraction capability in the structure design, effectively improving the multi-dimensional balance of prediction accuracy, model stability and engineering efficiency. At the same time, in order to effectively integrate the multi-branch output information, the present application designs a flexible feature fusion mechanism to fuse the results of multi-branch modeling, and then outputs the final win rate prediction value through the logistic regression layer. This fusion method not only realizes the complementarity of shallow and deep information, but also maintains the simplicity of the model structure, which is suitable for deployment in low-latency and high-throughput advertising systems.

[0040] Further, in the present application, the factorization machine branch is used to explicitly model the second-order interaction relationship between features, effectively capturing the shallow dependency information between features. The deep neural network (DNN) branch is used to extract high-order nonlinear feature combinations to improve the fitting ability of the model to complex behavior patterns. At the same time, the residual connection branch is used to maintain the linear transformation information of the original features, thereby improving the robustness of the model in the feature missing scenario.

[0041] Further, the present application also realizes the complementarity of shallow and deep information through the fusion process, and maintains the simplicity of the model structure, which is suitable for deployment in low-latency and high-throughput advertising systems. BRIEF DESCRIPTION OF DRAWINGS

[0042] The above and other objects, features and advantages of the exemplary embodiments of the present application will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0043] Figure 1A flowchart illustrating a win rate prediction method based on a fusion factorization feature interaction according to an embodiment of the present invention is shown.

[0044] Figure 2 A structural diagram of a win rate prediction system based on fusion factorization feature interaction according to an embodiment of the present invention is shown. Detailed Implementation

[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0046] This invention addresses the shortcomings of win rate prediction models in online real-time bidding (RTB) systems, including insufficient low-order feature interaction modeling, limited high-order expressive power, lack of robustness to data sparsity, and high online inference overhead. It proposes a DeepFM-Rank model that integrates factorization feature interactions. This model's structural design combines shallow feature cross-referencing with deep feature extraction capabilities, aiming to achieve a multi-dimensional balance between prediction accuracy, model stability, and engineering efficiency. Furthermore, this invention utilizes a feature fusion mechanism based on multi-branch modeling results and outputs the final win rate prediction value through logistic regression, achieving complementarity between shallow and deep information and effectively improving the applicability of this win rate prediction method in advertising bidding systems.

[0047] Based on the context of this invention, the technical solution of this invention can be applied to tasks such as click-through rate prediction and conversion rate estimation in personalized recommendations, improving recommendation accuracy by leveraging feature interaction capabilities. In financial data with high-dimensional sparse features, it can be used for risk modeling scenarios such as default rate prediction and fraud detection.

[0048] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0049] Figure 1 A flowchart is shown for a win rate prediction method 100 based on fused factorization feature interactions according to an embodiment of the present invention.

[0050] like Figure 1 As shown, in step S101, high-dimensional sparse features of the bidding data are obtained, and the high-dimensional sparse features are mapped to dense embedding vectors. In some embodiments, various discrete category features of the bidding data are input into the embedding layer to obtain dense embedding vectors corresponding to various discrete category features; the dense embedding vectors are merged to output three-dimensional tensors and flattened vectors respectively.

[0051] At step S102, the dense embedding vectors are subjected to multi-branch feature modeling to obtain multi-branch output information, wherein the first branch includes an explicit second-order feature interaction branch, the second branch includes a high-order nonlinear modeling branch, and the third branch includes a residual connection branch.

[0052] The modeling process of the first branch includes: extracting embedding vectors of all fields from the three-dimensional tensor, and calculating the sum and square sum of all embedding vectors; subtracting the square of the sum of all embedding vectors from all square sums to obtain second-order interaction feature output; mapping the second-order interaction feature to a first low-dimensional vector of a set length by using a fully connected layer to obtain the output of the first branch.

[0053] The calculation formula of the second-order interaction feature output is:

[0054]

[0055] The calculation formula of the output of the first branch is:

[0056] F fm =ReLU(W fm ·FM+b fm )∈R batch_size×32

[0057] In the formula, FM represents the second-order interaction feature output, v i represents the embedding vector of the i-th field, v i ∈R embed_dim , represents the sum of all field embeddings, represents the square sum of the embedding vectors of all fields, F fm represents the first low-dimensional vector, with a dimension of [batch_size, 32], W fm represents the weight matrix of the first low-dimensional vector, b fm represents the bias vector, with a dimension of

[32] , ReLU() represents the rectified linear unit activation function, ReLU(x) = max(0, x). batch_size is the number of samples. embed_dim is the embedding dimension of each feature.

[0058] The modeling process of the second branch includes: performing multi-layer nonlinear transformation on the flattened vector to output high-order features.

[0059] The modeling process of the third branch includes: performing linear transformation on the flattened vector to obtain a second low-dimensional vector compressed to a set length; performing nonlinear activation by using a ReLU function to obtain a residual branch output vector. The calculation formula of the output of the third branch is:

[0060] F res =ReLU(W res ·Eflat +b res )∈R batch_size×32

[0061] where F res represents the residual branch output vector with dimension [batch_size, 32], W res represents the weight matrix of the residual connection with dimension [n embed_dim, 32], E flat represents the flattened vector, and b res represents the bias vector of the residual layer with dimension

[32] , and ReLU() represents the activation function.

[0062] At step S103, the multi-branch output information is fused to obtain a feature fusion result. The feature fusion includes adding the output of the second branch and the output of the third branch, and splicing the added features with the output of the first branch.

[0063] The calculation formula of the adding and splicing process is:

[0064] F deep = F dnn +F res ∈R batch_size×32

[0065] F concat = concat(F deep ,F fm )∈R batch_size×64

[0066] where F dnn represents the output of the second branch with dimension [batch_size, 32], F res represents the output of the third branch, F deep represents the deep feature representation after weighted fusion of the two, F concat represents the final feature fusion result with dimension [batch_size, 64], concat() represents the splicing operation, and F fm represents the output of the first branch.

[0067] At step S104, the feature fusion result is subjected to Sigmoid transformation and clip operation constraint to output the win rate probability. The calculation formula of the Sigmoid transformation and clip operation constraint of the feature fusion result is:

[0068]

[0069] where represents the win rate probability, with a value range of (0, 1), W out represents the weight matrix of the output layer.out represents an output layer bias term, and represents a Sigmoid function defined as: represents an operation for limiting a numerical value in a specified interval, and represents a very small positive number, which is used to prevent gradient instability caused by output being 0 or 1.

[0070] Further, the above model structure can be combined with an attention module such as a Transformer, a weighted feature interaction mechanism is introduced, and the high-order feature modeling capability is further enhanced.

[0071] Next, the scheme of the application will be described in detail in combination with a specific implementation process.

[0072] The win rate prediction model of the application adopts a fusion structure composed of an embedding layer, an explicit second-order feature interaction branch (hereinafter referred to as an FM branch), a high-order nonlinear modeling branch (hereinafter referred to as a DNN branch), a residual connection branch (hereinafter referred to as a residual branch), and a feature fusion and output module. The overall implementation process is as follows:

[0073] 1. Feature input and embedding processing.

[0074] 1.1, Input data format: dictionary type features: Dict[str, Tensor], where the key is the slot name, and the value is the tensor [batch_size, 1]. Each slot represents a discrete category feature, usually ID encoded.

[0075] 1.2, Embedding processing: encode each slot into a dense vector through ChannelEmbeddingLayers (channel embedding layer), shape: emb i ∈R batch_size×embed_dim .

[0076] After merging all embedding vectors, (1) a three-dimensional tensor form is obtained: used for FM branch calculation. E FM ∈R batch _size×field_num×embed_dim . field_num is the number of feature fields (such as user ID,

[0077] advertisement ID, ad position ID, etc. discrete features), and embed_dim is the embedding dimension of each feature. (2) flat embedding form: used for DNN branch and residual branch. E flat =concat(emb1,emb2,…,emb n )∈R batch_size×(field_num·embed_dim) .

[0078] 2. FM branch (second-order feature interaction). FM branch has low computation overhead, which is suitable for edge device deployment, and can achieve real-time prediction and fast response. FM branch is used to explicitly model the interaction information between the embedding vectors of any two fields. The calculation process is as follows:

[0079] 2.1, let v i ∈R embed_dim represent the embedding vector of the i-th field, and the FM output is:

[0080]

[0081] wherein: represents the sum of all field embeddings; FM∈R batch_size×embed_dim represents the output of the FM branch. The process is to sum along the feature domain dimension first, and then square element by element and sum.

[0082] 2.2, project to a unified dimension through a fully connected layer: F fm = ReLU(W fm · FM + b fm )∈R batch_size×32 . Wherein, F fm represents the output feature of the FM branch, with a dimension of [batch_size, 32]; W fm represents the weight matrix of the FM output mapping, with a dimension of [embed_dim, 32]; b fm represents the bias vector, with a dimension of

[32] ; ReLU() represents the rectified linear unit activation function, ReLU(x) = max(0, x). In the output result, the 32-dimensional feature vector of each sample contains the following information: explicit second-order interaction relationship between all feature domains; compressed representation through ReLU activation; lightweight low-order feature cross signature.

[0083] 3. DNN branch (high-order feature modeling).

[0084] This branch is used to learn the nonlinear high-order combination relationship between features, and the structure is a five-layer fully connected network with ReLU activation function, combined with batch normalization (BatchNorm) and regularization (Dropout).

[0085] 3.1, the input of the DNN branch is: E flat = concat(v1, v2, …, v n ). Wherein, E flat represents the flattened vector, with a dimension of [batch_size, n·embed_dim]; v i represents the embedding vector of the i-th field; concat() represents the vector concatenation operation along the dimension.

[0086] 3.2, The layer structure is in turn: Dense (512) → BN → Dropout (0.3) → Dense (256) → Dense (128) → Dense (64) → Dense (32). Wherein Dense layer is a fully connected layer, Dense (512) means that the vector is mapped to a 512-dimensional high-order feature space, and the rest is the same. BN (Batch Normalization) is a technique for deep neural networks, which normalizes the input data distribution to speed up training, alleviate gradient vanishing / explosion problems, and improve model stability. Dropout means random inactivation, which makes neurons "randomly inactivated". Specifically, during training, a certain probability of randomly "masking" the output of a portion of neurons, making the value 0. Dropout (0.3) means that the dropout rate of this layer is 0.3. Since this training process is prior art, it will not be described here.

[0087] 3.3, Final output:

[0088] F dnn ∈R batch_size×32

[0089] In the formula, F dnn represents the output of the high-order nonlinear modeling branch.

[0090] 4. Residual branch (linear transformation enhancement). In order to improve the stability and robustness of the model, a residual branch is introduced to perform linear transformation on the embedded features.

[0091] 4.1, The input of the residual branch is the flattened tensor, that is, E flat .

[0092] 4.2, The specific operation is:

[0093] F res = ReLU (W res ·E flat +b res ) ∈R batch_size×32

[0094] Where: F res represents the output vector of the residual branch, with a dimension of [batch_size, 32]; W res represents the weight matrix of the residual connection, with a dimension of [n·embed_dim, 32]; b res represents the bias vector of the residual layer, with a dimension of

[32] . This branch can help the model retain basic expression ability when features are missing or data is abnormal.

[0095] 5. Feature fusion and prediction output. The three-branch fusion constitutes the final feature representation, and the win rate prediction value is output through the sigmoid function.

[0096] 5.1, First add the DNN branch output and the residual branch output:

[0097] F deep +F dnn +F res ∈R batch_size×32

[0098] Where: F dnn represents the output of the DNN branch, with a dimension of [batch_size, 32]; F res represents the residual branch output; F deep represents the deep feature representation after weighted fusion of the two.

[0099] 5.2, The result obtained after addition is spliced with the FM branch output:

[0100] F concat = concat(F deep , F fm ) ∈R batch_size×64

[0101] Where: F concat represents the final feature fusion result, with a dimension of [batch_size, 64]; concat() represents the splicing operation.

[0102] 5.3, Through the sigmoid output layer:

[0103]

[0104] Where: represents the win rate probability of the prediction output, with a value range of (0, 1); W out represents the weight matrix of the output layer, with a dimension of [64, 1]; b out represents the output layer bias term, with a dimension of [1]; σ() represents the Sigmoid function, defined as:

[0105] 5.4, To avoid the influence of extreme numerical values on stability, use the clip operation to constrain:

[0106]

[0107] Where, clip() represents the operation of limiting the numerical value within the specified interval; ∈ represents a very small positive number, used to prevent the output from being 0 or 1, which causes gradient instability, commonly set to 1 × 10 -6 ; For the final output of the win rate prediction value, the dimension is [batch_size].

[0108] The above scheme of the present application, by explicitly modeling low-order feature interaction and deeply mining high-order nonlinear relationship, the model more accurately captures the matching degree between users and advertisements in win rate prediction, effectively enhances the prediction accuracy. By introducing the residual structure, effectively alleviate the performance decline problem caused by feature missing or data sparsity, improve the stability of the model. FM branch calculation cost is low, the overall structure is light, suitable for real-time bidding scene with high response speed requirement. At the same time, the multi-branch combination is flexible, realizes the end-to-end training, improves the development efficiency and system maintainability.

[0109] Figure 2 A component structure diagram of a win rate prediction system fusing factorization feature interaction according to an embodiment of the present application is shown.

[0110] The present application also provides a win rate prediction system fusing factorization feature interaction, as shown in Figure 2 The system includes a processor and a memory, and the memory stores computer program instructions, which realize the win rate prediction method fusing factorization feature interaction when executed by the processor.

[0111] The system also includes a communication bus and a communication interface and other components familiar to those skilled in the art, the setting and function of which are known in the art, so here is not repeated.

[0112] In this document, the terms "computer-readable medium" and "storage medium" are used to generally refer to any tangible media or storage means capable of storing the program for use by or in connection with an instruction execution system, apparatus, or device. By way of example, and not limitation, such computer-readable or storage media can include any suitable media, including, but not limited to, any suitable magnetic, optical, or electrical storage media, such as, but not limited to, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), and the like, or any other medium which can be used to store the desired information and which can be accessed by an application, module, or both. Any such computer-readable medium can be part of a device or accessible or connectable thereto. Any application or module described herein can be implemented using computer-readable / executable instructions which can be stored or otherwise held by such computer-readable media.

[0113] While the present application has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the application is not limited to the disclosed embodiments. Numerous alternative changes, modifications, and substitutions are intended to be constructively embraced by the application and fall within its spirit and scope. It is to be understood that in the process of rendering the present application into a practice, various alternative means and / or materials can be employed, and certain changes can be made, without departing from the spirit and scope of the present application.

Claims

1. A win rate prediction method that integrates factorization feature interactions, characterized in that, include: Obtain high-dimensional sparse features from bidding data and map these features into dense embedding vectors. The dense embedding vector is modeled with multi-branch features to obtain multi-branch output information, wherein the first branch includes an explicit second-order feature interaction branch, the second branch includes a higher-order nonlinear modeling branch, and the third branch includes a residual connection branch. The multi-branch output information is fused to obtain the feature fusion result; the feature fusion includes summing the outputs of the second branch and the third branch, and concatenating the summed features with the output of the first branch; The feature fusion results are subjected to Sigmoid transformation and clip operation constraints to output the win probability.

2. The win rate prediction method based on the fusion of factorization feature interactions according to claim 1, characterized in that, Obtain high-dimensional sparse features from the bidding data and map these features into dense embedding vectors, including: Input the various discrete category features of the bidding data into the embedding layer to obtain the dense embedding vectors corresponding to the various discrete category features; Densely embedded vectors are merged to output a 3D tensor and a flattened vector, respectively.

3. The win rate prediction method based on the fusion of factorization feature interactions according to claim 2, characterized in that, The modeling process for the first branch includes: Extract the embedding vectors of all fields from the 3D tensor and calculate the sum and sum of squares of all embedding vectors; Subtract the square of the sum of all embedded vectors from all sums of squares to obtain the second-order interactive feature output; The second-order interactive features are mapped to a first low-dimensional vector of a set length using a fully connected layer to obtain the output of the first branch.

4. The win rate prediction method based on the fusion of factorization feature interactions according to claim 3, characterized in that, The formula for calculating the second-order interactive feature output is: The formula for calculating the output of the first branch is: F fm =ReLU(W fm ·FM+b fm ) In the formula, FM represents the second-order interactive feature output, v i This represents the embedding vector of the i-th field. This indicates that all fields are summed. F represents the sum of squares of the embedding vectors of all fields. fm Let W represent the first low-dimensional vector. fm Let b be the weight matrix of the first low-dimensional vector. fm Represents the bias vector, and ReLU() represents the modified linear unit activation function, ReLU(x) = max(0,x).

5. The win rate prediction method based on the fusion of factorization feature interactions according to claim 1, characterized in that, The modeling process for the second branch includes: Perform multi-level nonlinear transformations on the flattened vector to output higher-order features.

6. The win rate prediction method based on the fusion of factorization feature interactions according to claim 1, characterized in that, The modeling process for the third branch includes: A linear transformation is performed on the flattened vector to obtain a second low-dimensional vector compressed to a set length; The ReLU function is used for nonlinear activation to obtain the residual branch output vector.

7. The win rate prediction method based on the fusion of factorization feature interactions according to claim 6, characterized in that, The formula for calculating the output of the third branch is: F res =ReLU(W res ·HAVE BEEN flat +b res ) In the formula, F res W represents the residual branch output vector. res E represents the weight matrix of the residual connection. flat Let b represent the flattened vector. res This represents the bias vector of the residual layer, and ReLU() represents the activation function.

8. The win rate prediction method based on the fusion of factorization feature interactions according to claim 1, characterized in that, The calculation formulas for the addition and splicing process are as follows: F deep =F dnn +F res F concat =concat(F deep ,F fm ) In the formula, F dnn F represents the output of the second branch. res F represents the output of the third branch. deep F represents the deep feature representation after weighted fusion of the two. concat This represents the final feature fusion result; `concat()` represents the concatenation operation; F fm This indicates the output of the first branch.

9. The win rate prediction method based on the fusion of factorization feature interactions according to claim 1, characterized in that, The formula for calculating the Sigmoid transform and clip operation constraints of the feature fusion results is as follows: In the formula, W represents the winning probability, with a value ranging from (0,1). out b represents the weight matrix of the output layer. out This represents the output layer bias term, and σ() represents the Sigmoid function, defined as: The clip(·) operation restricts the value to a specified range. ε represents a very small positive number and is used to prevent gradient instability caused by outputting 0 or 1.

10. A win rate prediction system that integrates factorization feature interaction, characterized in that, include: processor; A memory storing computer program instructions that, when executed by the processor, implement a win rate prediction method according to any one of claims 1-9, which integrates factorization feature interactions.