Platform user behavior characteristic transfer learning method oriented to advertisement putting

By constructing a behavior sequence encoding subnetwork and a contextual attention subnetwork, and combining self-supervised learning and domain adaptive adversarial training, the problems of insufficient feature representation of user behavior sequences and low accuracy of cross-platform feature migration are solved, achieving precise and high-yield advertising.

CN120634644AInactive Publication Date: 2025-09-12NANJING HEYI COMM EQUIP CO LTD

Patent Information

Application Number
CN202510822345.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing advertising delivery technologies have problems such as insufficient representation of user behavior sequence features, low accuracy of cross-platform feature migration, and lack of privacy protection measures, resulting in weak model generalization ability and poor delivery results.

Method used

By constructing a behavior sequence encoding subnetwork and a contextual attention subnetwork, and combining self-supervised learning and domain adaptive adversarial training, we can achieve deep migration and precise alignment of user behavior features, including data preprocessing, behavior sequence encoding, contextual feature fusion and migration training, and generate migration feature vectors for advertising delivery decisions.

Benefits of technology

It achieves efficient use of cross-platform user behavior data, improves the accuracy of user behavior feature migration, reduces advertising costs, and achieves precise and high-yield advertising.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634644A_ABST
    Figure CN120634644A_ABST
Patent Text Reader

Abstract

The invention discloses an advertisement putting-oriented platform user behavior feature transfer learning method. The method comprises the following steps: acquiring a source domain behavior sequence set and a target domain behavior sequence set; constructing a user behavior representation model based on the source domain behavior sequence set, and training by using a self-supervised prediction task to obtain a source domain behavior feature encoder; carrying out migration training on the target domain behavior sequence set by taking the network parameter of the source domain behavior feature encoder as an initialization weight to obtain a target domain behavior feature encoder; and generating a migration feature vector for the newest behavior sequence of the active user of the target platform, inputting the migration feature vector into the advertisement putting decision model, and outputting a sorting and bidding result of the to-be-put advertisement. According to the method, the bottleneck problems of data sparsity, weak feature generalization ability, insufficient privacy protection and poor delivery effect in traditional cross-platform feature migration are solved, the user behavior feature migration precision is greatly improved, and the advertisement delivery cost is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence data mining and precise advertising delivery, and in particular to a method for transferring learning of platform user behavior characteristics for advertising delivery. Background Art

[0002] With the rapid development of Internet technology and the continuous expansion of e-commerce platforms, precise and personalized advertising has gradually become a key means for platforms to improve conversion rates and user experience. In recent years, advertising strategies based on machine learning and deep learning technologies have made significant progress. Through in-depth mining and modeling of users' historical behavioral characteristics, a significant improvement in advertising accuracy has been achieved. Among them, transfer learning methods have gradually attracted attention because they can effectively solve the problem of scarce user behavior data on new platforms and new fields.

[0003] CN102508859A discloses an advertising classification method based on web page features, which uses transfer learning to align web page features and advertising features in a common feature space to improve advertising classification accuracy. However, this technical solution only focuses on the common space mapping of static web page and advertising features, ignores the dynamic behavior sequence information of users in the advertising delivery platform, and fails to use the user's historical behavior sequence to effectively infer advertising interests and preferences, resulting in insufficient accuracy of feature representation.

[0004] CN110400169A discloses an information push method that achieves accurate prediction of target users in a small number of sample scenarios by migrating a click prediction model to a sharing prediction model. Although it utilizes a deep learning model, its migration training process is limited to a simple migration from a single click prediction to a sharing prediction task. It does not fully consider the deep alignment of data in different behavioral domains and has difficulty capturing the subtle differences in user behavior sequences on different platforms. In addition, the above method lacks sufficient consideration of data privacy protection and does not effectively desensitize the privacy issues of multi-platform user data, thereby limiting the efficient and secure use of cross-platform behavioral data.

[0005] The feature transfer methods used in existing technologies mostly focus on establishing simple mapping relationships, ignoring the deep contextual associations between different behavior sequences and the effective representation of serialized information, which limits the generalization ability and prediction accuracy of the model; therefore, existing advertising delivery technologies have problems such as insufficient representation of user behavior sequence features, low accuracy of cross-platform feature transfer, and lack of privacy protection measures; and the present invention proposes a platform user behavior feature transfer learning method for advertising delivery to address the above problems. By constructing a behavior sequence encoding subnetwork and a contextual attention subnetwork, the method realizes deep migration and precise alignment of behavior sequence features through self-supervised learning and domain adaptive adversarial training. Summary of the Invention

[0006] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract of the specification and the title of the invention of this application to avoid blurring the purpose of this section, the abstract of the specification and the title of the invention, and such simplifications or omissions cannot be used to limit the scope of the invention.

[0007] In view of the above existing problems, the present invention is proposed.

[0008] To solve the above technical problems, the present invention provides the following technical solutions: historical user behavior logs from the source platform and real-time user behavior logs from the target platform are obtained, and a source domain behavior sequence set and a target domain behavior sequence set are obtained by performing deduplication, serialization, timestamp standardization, and privacy desensitization preprocessing on the two obtained behavior logs;

[0009] Based on the source domain behavior sequence set, a user behavior representation model is constructed, which includes a behavior sequence encoding subnetwork and a context attention subnetwork. A source domain behavior feature encoder is obtained by training using a self-supervised prediction task.

[0010] Using the network parameters of the source domain behavior feature encoder as initialization weights, combined with the minimum mean difference regularization term and the adversarial domain discriminator loss, transfer training is performed on the target domain behavior sequence set to obtain the target domain behavior feature encoder aligned in the common feature space;

[0011] The target domain behavior feature encoder is called to generate a migration feature vector for the latest behavior sequence of active users on the target platform, the migration feature vector is input into the advertising delivery decision model, and the ranking and bidding results of the advertisements to be delivered are output.

[0012] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, the source domain behavior sequence set and the target domain behavior sequence set are obtained by performing deduplication, serialization, timestamp standardization and privacy desensitization preprocessing on the two obtained behavior logs, respectively, including:

[0013] For the acquired behavior logs, based on the user ID, behavior type, and preset time threshold, only the earliest behavior log will be retained for the same type of behavior that occurs continuously within a short period of time by the same user.

[0014] The retained behavior logs are grouped according to user IDs and arranged in chronological order according to the occurrence of the behaviors to form an orderly user behavior sequence;

[0015] Convert the different log time formats between the source and target platforms to Coordinated Universal Time and standardize them in milliseconds.

[0016] The user ID, device ID and fields containing sensitive information are hashed to achieve desensitization of private data, thereby obtaining the source domain behavior sequence set and the target domain behavior sequence set respectively.

[0017] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, the source platform historical user behavior log includes at least the user's behavior information on the source platform, including ad display, ad click, ad page browsing, ad page dwell time, product collection, adding to shopping cart, purchase order, payment completion, order cancellation, and post-sale return, as well as the ad page context and user attribute information of the historical behavior;

[0018] The real-time user behavior log of the target platform includes at least the user's real-time advertising display, advertising clicks, page browsing, page dwelling, product intention behavior, shopping cart addition, payment attempts, and interruption behavior information of uncompleted payment on the target platform, and includes the user's real-time page context information and device type information.

[0019] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, a user behavior representation model containing a behavior sequence encoding subnetwork and a contextual attention subnetwork is constructed based on the source domain behavior sequence set, including:

[0020] Construct a behavior sequence encoding subnetwork, input the user behavior sequence into the bidirectional self-attention mechanism in sequence, and generate a contextual encoding representation of each behavior by learning the contextual dependencies of the behavior;

[0021] Construct a contextual attention sub-network. By fusing user historical behavior encoding, page context features, and user static profile information, the gated attention mechanism is used to dynamically adjust the weights of information from different sources to obtain the final user behavior representation vector that integrates multi-source information.

[0022] The generated user behavior representation vector is used as the input for the next self-supervised prediction task to complete the model building process.

[0023] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, a source domain behavior feature encoder is obtained by training using a self-supervised prediction task, including:

[0024] The historical behavior data of the user behavior sequence is used to construct the next behavior category prediction task, and the user behavior representation model is trained to predict the type of the user's next behavior to capture the evolution law within the behavior sequence;

[0025] Randomly selecting some behaviors in the user behavior sequence for masking, and training the user behavior representation model to reconstruct the masked behavior content, so as to enhance the ability of the user behavior representation model to capture long-range behavior dependencies;

[0026] The losses of the above prediction tasks are superimposed according to the set ratio, and the user behavior representation model parameters are iteratively optimized through the back propagation algorithm until the loss value converges;

[0027] The network parameters of the user behavior representation model after training are fixed to the parameters of the source domain behavior feature encoder.

[0028] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, the step of obtaining the target domain behavior feature encoder includes:

[0029] Copying all network parameters of the source domain behavior feature encoder to the target domain behavior feature encoder as initial parameters;

[0030] While keeping the parameters of the source domain behavioral feature encoder frozen, we introduce a domain discriminator to perform adversarial training with the goal of minimizing the feature difference between the target domain and the source domain.

[0031] Unfreeze the target domain behavioral feature encoder parameters and further optimize the target domain behavioral feature encoder by combining the minimum mean difference regularization term, so that the feature distributions of the source domain and the target domain gradually converge;

[0032] When the inter-domain feature difference index reaches the preset convergence standard, the training of the target domain behavior feature encoder is completed.

[0033] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, the transfer training specifically includes:

[0034] The training process is organized in batches, alternating between source domain self-supervised tasks and target domain adversarial tasks.

[0035] In each target domain training phase, when target domain samples lack accurate labels, a pseudo-label generation strategy based on feature vector similarity judgment is adopted to automatically generate labels for samples whose confidence exceeds a set threshold and add them to the training set;

[0036] By gradually reducing the proportion of source domain samples in training and correspondingly increasing the proportion of target domain samples, the dominant role of target domain data is enhanced;

[0037] The inter-domain feature difference index is evaluated in real time, and the training convergence condition is that the inter-domain difference index meets the set threshold.

[0038] As a preferred solution of the platform user behavior feature transfer learning method for advertising delivery described in the present invention, calling the target domain behavior feature encoder to generate a transfer feature vector for the latest behavior sequence of active users on the target platform includes:

[0039] Acquire the most recent continuous behavior sequence of the target platform user in real time, and dynamically truncate the length of the behavior sequence to a standard length according to the input requirements of the target domain behavior feature encoder;

[0040] The adjusted behavior sequence is input into the trained target domain behavior feature encoder for transfer feature analysis to generate the user's transfer feature vector.

[0041] As a preferred solution of the platform user behavior feature transfer learning method for advertising placement described in the present invention, the transfer feature vector is input into the advertising placement decision model, and the ranking and bidding results of the advertisements to be placed are output, including:

[0042] After combining the user migration feature vector with the ad feature information, it is input into the ad placement decision model for multi-task learning to predict the ad click-through rate and conversion rate respectively.

[0043] Calculate the comprehensive revenue score of each candidate ad based on the click-through rate and conversion rate prediction results and the ad's weight coefficient;

[0044] Sort the candidate advertisements in descending order according to the comprehensive revenue scores to form an advertisement recommendation list;

[0045] Based on the advertising budget and historical delivery data, determine the advertising bid, and finally output the ranking and bidding results.

[0046] Beneficial effects of the present invention: The present invention realizes the efficient utilization of cross-platform user behavior data through data preprocessing technology, deep self-supervised behavior representation model, adaptive transfer learning mechanism and refined advertising delivery decision strategy, breaking through the bottleneck problems of data sparsity, weak feature generalization ability, insufficient privacy protection and poor delivery effect in traditional cross-platform feature migration, greatly improving the accuracy of user behavior feature migration, reducing advertising delivery costs, and realizing precise and high-profit advertising delivery. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort. Among them:

[0048] Figure 1 This is a flow chart of the platform user behavior feature transfer learning method for advertising delivery shown in the present invention. DETAILED DESCRIPTION

[0049] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, but not all of the embodiments.

[0050] Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without making any creative work should fall within the scope of protection of the present invention.

[0051] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0052] According to an embodiment of the present invention, Figure 1 The flowchart shown is a method for transferring learning of platform user behavior features for advertising delivery, which specifically includes the following steps:

[0053] S1. Obtain historical user behavior logs from the source platform and real-time user behavior logs from the target platform. Perform deduplication, serialization, timestamp standardization, and privacy masking preprocessing on the two behavior logs to obtain the source domain behavior sequence set and the target domain behavior sequence set. Note that:

[0054] For the acquired behavior logs, based on the user ID, behavior type, and preset time threshold, for the same type of behavior that occurs consecutively by the same user within a short period of time (e.g., 30 seconds), only the earliest behavior log is retained;

[0055] The retained behavior logs are grouped according to user IDs and arranged in chronological order according to the occurrence of the behaviors to form an orderly user behavior sequence;

[0056] Convert the different log time formats between the source and target platforms to Coordinated Universal Time and standardize them in milliseconds.

[0057] The user ID, device ID and fields containing sensitive information are hashed to achieve desensitization of private data, thereby obtaining the source domain behavior sequence set and the target domain behavior sequence set respectively.

[0058] In an optional implementation, the source platform's historical user behavior log includes at least information on the user's behavior on the source platform, including ad display, ad clicks, ad page browsing, ad page dwell time, product collection, adding to shopping cart, purchase order, payment completion, order cancellation, and post-sales return, as well as the ad page context and user attribute information of the historical behavior.

[0059] In an optional embodiment, the real-time user behavior log of the target platform includes at least the user's real-time advertising display, advertising clicks, page browsing, page dwelling, product intention behavior, adding to shopping cart, payment attempts, and interruption behavior information of uncompleted payment on the target platform, and includes the user's real-time page context information and device type information.

[0060] As an example, the source platform log service interface is synchronously called to pull historical behavior logs covering the past 30 days, and the real-time log stream of the target platform is subscribed to. The two types of original logs are uniformly written into the distributed object storage, the source identifier and the collection batch number are recorded, and the logs are grouped according to the user identifier UID. The records in each group are arranged in ascending chronological order to obtain a complete user behavior time series; and empty behavior placeholders are filled for users whose sequence length is less than L, and those whose length exceeds L are truncated to the nearest L items to ensure the consistency of the subsequent model input dimensions.

[0061] Exemplarily, the user identifier UID is generated by any stable key of the platform account number, CookieID, or IMEI, for example: UID = "user_12345", which is hashed to obtain UID = "8e9f...a6".

[0062] As an example, the privacy desensitization processing in this embodiment specifically includes the following steps:

[0063] Use an irreversible one-way hash algorithm (such as SHA-256) to generate a 256-bit digest for the user ID (UID) and device ID (DID). For example, perform SHA-256 on the string "device_abc" to generate "9b74…e6";

[0064] Perform value generalization on structured sensitive fields such as names and mobile phone numbers (for example, retain the last four digits of a mobile phone number, "****6789");

[0065] Regular NER joint detection is applied to entity objects such as addresses and email addresses in free text, and they are replaced with placeholders to ensure that personal identities cannot be restored after log desensitization; for example, the entity in the text "Zhongguancun, Haidian District, Beijing" is replaced with "[ADDR]".

[0066] As an example, the mathematical expressions of the source domain behavior sequence set and the target domain behavior sequence set finally obtained in this step are as follows:

[0067]

[0068] Among them, UID is the user ID after hashing. represents the behavior-time tuple sequence of the i-th user in the source domain sorted by time, is the behavior-time tuple sequence of the jth user in the target domain sorted by time, N s ,N t are the number of valid users in the source domain and the target domain after deduplication and filtering, L s ,L t is the final sequence length of each user, satisfying L min ≤L {s,t} ≤L max .

[0069] It should be noted that this step achieves unified organization and structured representation of cross-platform heterogeneous data by obtaining historical user behavior logs from the source platform and real-time user behavior logs from the target platform, and performing deduplication, serialization, timestamp standardization, and privacy desensitization preprocessing on the obtained logs. This can effectively eliminate data redundancy and confusion, and reduce the negative impact of noise data on subsequent model training.

[0070] Furthermore, privacy desensitization is used to protect user privacy, eliminate data security risks, and improve the flowability and actual availability of user data between platforms, ultimately achieving efficient, secure, and standardized data processing results.

[0071] S2. Based on the source domain behavior sequence set, a user behavior representation model is constructed, which includes a behavior sequence encoding subnetwork and a contextual attention subnetwork. The source domain behavior feature encoder is trained using the self-supervised prediction task. Note that:

[0072] Construct a behavior sequence encoding subnetwork, input the user behavior sequence into the bidirectional self-attention mechanism in sequence, and generate a contextual encoding representation of each behavior by learning the contextual dependencies of the behavior;

[0073] Construct a contextual attention sub-network. By fusing user historical behavior encoding, page context features, and user static profile information, the gated attention mechanism is used to dynamically adjust the weights of information from different sources to obtain the final user behavior representation vector that integrates multi-source information.

[0074] The generated user behavior representation vector is used as the input for the next self-supervised prediction task to complete the model building process.

[0075] Leveraging historical behavioral data from user behavior sequences, we construct a next-step behavior category prediction task. This task trains a user behavior representation model to predict the type of the user's next behavior, capturing the evolutionary patterns within the behavior sequence. For example, if a user's historical sequence follows an increasing conversion chain of "display → click → add to cart → payment," the next-step prediction task forces the model to learn this high-conversion path and predict the "payment" trend with high confidence when observing "display → click → add to cart" in a new sequence.

[0076] Randomly select some behaviors in the user behavior sequence for masking, and train the user behavior representation model to reconstruct the masked behavior content to enhance the user behavior representation model's ability to capture long-range behavior dependencies;

[0077] The losses of the above prediction tasks are superimposed according to the set ratio, and the user behavior representation model parameters are iteratively optimized through the back propagation algorithm until the loss value converges;

[0078] The network parameters of the user behavior representation model after training are fixed to the parameters of the source domain behavior feature encoder.

[0079] It is not difficult to understand that the user behavior representation model after training is used as the final source domain behavior feature encoder. For example, in this embodiment, the mathematical expression of the source domain behavior feature encoder is as follows:

[0080]

[0081] Among them, Z s is the final output vector of the source domain behavior feature encoder, whose dimension is d, G(·) is the layer normalization function, Θ is the upper limit of the observation window in seconds, P is the number of self-attention heads, is the p-th linear transformation matrix, is the Swish activation function, is the information related to the p-th head attention, X(t) is the behavior input vector at time t, Q is the number of samples in the batch, β>0 is the temperature scaling coefficient, ω(Y q ) is the noise contrast function, for the qth negative sample feature Y q Calculate the Euclidean norm to reduce the influence of outliers.

[0082] As an example, the behavior sequence encoding subnetwork is configured as a 4-layer Transformer-Encoder with 8 attention heads in each layer; the context attention subnetwork is configured as a two-layer gated multi-source fusion. The first layer calculates the attention weights of the user static portrait vector and the contextual encoding representation of the behavior. The second layer calculates the mutual attention between the page context vector and the output of the previous layer and uses a Sigmoid threshold to suppress low-weight features.

[0083] It should be noted that this step achieves the effective capture and learning of the deep contextual relationships and behavioral evolution laws of user behavior sequences by obtaining the source domain behavioral feature encoder, so as to more accurately represent the user's true behavioral intentions and interest preferences; at the same time, the self-supervised learning method does not require a large amount of manual labeling, which significantly reduces the data labeling cost and training threshold, and ultimately achieves the effect of accurate modeling of behavioral characteristics and cost-effective control.

[0084] S3. Using the network parameters of the source domain behavior feature encoder as the initial weights, combined with the minimum mean difference regularization term and the adversarial domain discriminator loss, transfer training is performed on the target domain behavior sequence set to obtain the target domain behavior feature encoder aligned in the common feature space. Among them, the following points need to be explained in this step:

[0085] Copy all network parameters of the source domain behavior feature encoder to the target domain behavior feature encoder as initial parameters;

[0086] While keeping the parameters of the source domain behavioral feature encoder frozen, we introduce a domain discriminator to perform adversarial training with the goal of minimizing the feature difference between the target domain and the source domain.

[0087] Unfreeze the target domain behavioral feature encoder parameters and further optimize the target domain behavioral feature encoder by combining the minimum mean difference regularization term, so that the feature distributions of the source domain and the target domain gradually converge;

[0088] When the inter-domain feature difference index reaches the preset convergence standard, the training of the target domain behavior feature encoder is completed.

[0089] Further, transfer training includes:

[0090] The training process is organized in batches, alternating between source domain self-supervised tasks and target domain adversarial tasks.

[0091] In each target domain training phase, when target domain samples lack accurate labels, a pseudo-label generation strategy based on feature vector similarity judgment is adopted to automatically generate labels for samples whose confidence exceeds a set threshold (e.g., 0.88) and add them to the training set.

[0092] By gradually reducing the proportion of source domain samples in training and correspondingly increasing the proportion of target domain samples, the dominant role of target domain data is enhanced;

[0093] The inter-domain feature difference index is evaluated in real time, and the training convergence condition is that the inter-domain difference index meets the set threshold.

[0094] It should be further explained that the preset convergence criteria include: when the discriminator accuracy is ≤0.6 and the MMD (minimum homogeneous difference regularization is performed in odd batches in the alternating batch strategy) distance is ≤0.03, the adversarial update is stopped and only one round of self-supervised fine-tuning in the target domain is retained to stabilize the aligned feature space. If there is no significant fluctuation in the discriminator accuracy and MMD distance within 5 consecutive cycles (such as the change does not exceed 0.005), the transfer training is determined to have converged, and all current trainable parameters are saved and recorded as the target domain behavior feature encoder.

[0095] As an example, the mathematical expression formula of the target domain behavior feature encoder is:

[0096]

[0097] Among them, t is the final output vector of the target domain behavior feature encoder, Λ is the learnable scaling factor, λ is the lower limit of the target domain time integration in seconds, μ is the upper limit of the integration in seconds, A is the number of attention heads, Q a is the linear projection matrix of the a-th head, tanh(·) is the hyperbolic tangent activation function, φ a is the gating coefficient of the a-th head, K(τ) is the representation of the target domain behavior input tensor at time τ after being mapped by the frozen source weights, J0(·) is the zero-order Bessel function, χ is the domain difference adjustment coefficient, ρ is the integral variable, B is the number of batch samples, η is the singular value root index, σ is the temperature factor to suppress the outlier effect, γ b is the bth source domain reference vector used to calculate the normalized denominator energy term.

[0098] It should be noted that this step achieves deep feature alignment of source and target domain data in the common feature space by obtaining the target domain behavior feature encoder, effectively alleviating the sparsity and cold start problems of user behavior data on the target platform.

[0099] Preferably, domain adaptive training enables the target domain encoder to quickly adapt to the distribution differences of user characteristics of the target platform, effectively improving the generalization performance and migration accuracy of the target domain model, and improving the precise migration of user behavior characteristics across platforms.

[0100] S4. Call the target domain behavior feature encoder to generate a migration feature vector based on the latest behavior sequence of active users on the target platform, input the migration feature vector into the advertising placement decision model, and output the ranking and bidding results of the ads to be placed.

[0101] Obtain the most recent continuous behavior sequence of the target platform user in real time, and dynamically truncate the behavior sequence length to a standard length based on the input requirements of the target domain behavior feature encoder;

[0102] The adjusted behavior sequence is input into the trained target domain behavior feature encoder for transfer feature analysis to generate the user's transfer feature vector;

[0103] After combining the user migration feature vector with the ad feature information, it is input into the ad placement decision model for multi-task learning to predict the ad click-through rate and conversion rate respectively.

[0104] Calculate the comprehensive revenue score of each candidate ad based on the click-through rate and conversion rate prediction results and the ad's weight coefficient;

[0105] Sort candidate ads in descending order according to their comprehensive revenue scores to form an ad recommendation list;

[0106] Based on the advertising budget and historical delivery data, the advertising bid is determined, and the ranking and bidding results of the ads to be delivered are finally output.

[0107] As an example, the mathematical expression formula of the advertisement placement decision model in this embodiment is as follows:

[0108]

[0109] in, is the output vector of the advertising decision model. Its first dimension corresponds to the predicted click-through rate, and the second dimension corresponds to the predicted conversion rate. The value range is (0,1). Ψ is the migration feature vector of the current target user, Φ i is the static feature vector of candidate advertisement i, is the Sigmoid-Softmax hybrid normalization function, Ω is the upper limit of the high-order interaction integral to control the accumulation depth, K is the number of interaction kernels, W k is the kth learnable projection matrix, σ(·) is the GELU activation function, A k is the kth Hadamard filter kernel, α and β are the lower and upper limits of the L1 norm regularized integral interval for users, γ is the user activity amplification coefficient, λ is the normalized root index, H is the number of candidate ads in the same batch, τ is the temperature attenuation constant, which is used to adjust the influence of feature distance, Φ h is the feature vector of the hth advertisement in the batch;

[0110] The closer the value is to 1, the higher the predicted probability of the ad being clicked or converted; conversely, if Values ​​below 0.05 are considered cold interest.

[0111] As an example, according to the click revenue r set by the advertiser c and conversion revenue v , calculate the comprehensive revenue score of candidate ads:

[0112] Scorei =r c CTR i +r v CVR i -θσ i

[0113] Among them, θ is the risk penalty coefficient, CTR i is the predicted click-through rate of the i-th candidate advertisement, CVR i is the predicted conversion rate of the i-th candidate advertisement, and σ is the uncertainty;

[0114] Score i Arrange in descending order to form a list rank , check the consumed budget B in turn i Daily budget The difference between the two, remove excess ads and update the list;

[0115] A second-order sawtooth price adjustment strategy is used for reserved ads, and the calculation formula is as follows:

[0116]

[0117] Among them, ξ∈(0,1) is the discount degree, It is a flexible adjustment item. i When the price is lower than the platform's reserve price Bid0, it will be automatically set to Bid0.

[0118] For example, define r c =1.2 yuan / click, r v = 8 yuan / conversion, θ = 0.1, and the comprehensive revenue score and ranking-bid example are shown in Table 1 below:

[0119] Table 1. Comprehensive benefit score and ranking - bid table

[0120] <![CDATA[Ad i ]]> <![CDATA[CTR i ]]> <![CDATA[CVR i ]]> <![CDATA[σ i ]]> <![CDATA[Score i ]]> <![CDATA[List rank ]]> <![CDATA[Bid i (Yuan)]]> A1 0.12 0.04 0.05 0.646 1 1.80 A2 0.09 0.03 0.03 0.471 2 1.55 A3 0.15 0.01 0.07 0.370 3 1.44

[0121] The corresponding output result is: 〈A1,List rank 1,Bid 1.80〉,〈A2,List rank 2,Bid 1.55〉,〈A3,List rank 3, Bid 1.44 ;

[0122] Among them, Ad i is the i-th candidate advertisement.

[0123] It should be noted that this step generates a migration feature vector for the latest behavior sequence of active users on the target platform by calling the target domain behavior feature encoder, and inputs the migration feature vector into the advertising delivery decision model to output the ranking and bidding results of the ads to be delivered, thereby achieving accurate prediction of the target users' real-time advertising interests and potential behavioral trends; by combining the correlation evaluation of user migration characteristics with advertising content and the click-through conversion prediction model, a more accurate and targeted advertising ranking and bidding strategy is formed, thereby improving the conversion rate and input-output ratio of advertising delivery, and ultimately achieving the effect of improving the efficiency of the advertising platform and user experience.

[0124] The aforementioned preprocessing method of the user behavior log can be performed using methods and means in the prior art, and will not be described in detail in this example.

[0125] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A platform user behavior feature transfer learning method for advertising delivery, characterized by: include: Obtain historical user behavior logs from the source platform and real-time user behavior logs from the target platform. Perform deduplication, serialization, timestamp standardization, and privacy masking preprocessing on the two behavior logs to obtain the source domain behavior sequence set and the target domain behavior sequence set respectively. Based on the source domain behavior sequence set, a user behavior representation model is constructed, which includes a behavior sequence encoding subnetwork and a context attention subnetwork. A source domain behavior feature encoder is obtained by training using a self-supervised prediction task. Using the network parameters of the source domain behavior feature encoder as initialization weights, combined with the minimum mean difference regularization term and the adversarial domain discriminator loss, transfer training is performed on the target domain behavior sequence set to obtain the target domain behavior feature encoder aligned in the common feature space; The target domain behavior feature encoder is called to generate a migration feature vector for the latest behavior sequence of active users on the target platform, the migration feature vector is input into the advertising delivery decision model, and the ranking and bidding results of the advertisements to be delivered are output.

2. The platform user behavior feature transfer learning method for advertising delivery according to claim 1 is characterized in that: The two types of behavior logs are pre-processed by deduplication, serialization, timestamp standardization, and privacy desensitization to obtain a source domain behavior sequence set and a target domain behavior sequence set, respectively, including: For the acquired behavior logs, based on the user ID, behavior type, and preset time threshold, only the earliest behavior log will be retained for the same type of behavior that occurs continuously within a short period of time by the same user. The retained behavior logs are grouped according to user IDs and arranged in chronological order according to the occurrence of the behaviors to form an orderly user behavior sequence; Convert the different log time formats between the source and target platforms to Coordinated Universal Time and standardize them in milliseconds. The user ID, device ID and fields containing sensitive information are hashed to achieve desensitization of private data, thereby obtaining the source domain behavior sequence set and the target domain behavior sequence set respectively.

3. The platform user behavior feature transfer learning method for advertising delivery according to claim 1 or 2, characterized in that: The source platform's historical user behavior logs include at least information about users' behavior on the source platform, including ad display, ad clicks, ad page views, ad page dwell time, product collection, adding to shopping carts, purchases, payment completion, order cancellations, and post-sale returns, as well as the ad page context and user attribute information of the historical behavior. The real-time user behavior log of the target platform includes at least the user's real-time advertising display, advertising clicks, page browsing, page dwelling, product intention behavior, shopping cart addition, payment attempts, and interruption behavior information of uncompleted payment on the target platform, and includes the user's real-time page context information and device type information.

4. The platform user behavior feature transfer learning method for advertising delivery according to claim 1 or 2, characterized in that: Based on the source domain behavior sequence set, a user behavior representation model including a behavior sequence encoding subnetwork and a context attention subnetwork is constructed, including: Construct a behavior sequence encoding subnetwork, input the user behavior sequence into the bidirectional self-attention mechanism in sequence, and generate a contextual encoding representation of each behavior by learning the contextual dependencies of the behavior; Construct a contextual attention sub-network. By fusing user historical behavior encoding, page context features, and user static profile information, the gated attention mechanism is used to dynamically adjust the weights of information from different sources to obtain the final user behavior representation vector that integrates multi-source information. The generated user behavior representation vector is used as the input for the next self-supervised prediction task to complete the model building process.

5. The platform user behavior feature transfer learning method for advertising delivery according to claim 4 is characterized in that: The source domain behavior feature encoder is trained using the self-supervised prediction task, including: The historical behavior data of the user behavior sequence is used to construct the next behavior category prediction task, and the user behavior representation model is trained to predict the type of the user's next behavior to capture the evolution law within the behavior sequence; Randomly selecting some behaviors in the user behavior sequence for masking, and training the user behavior representation model to reconstruct the masked behavior content, so as to enhance the ability of the user behavior representation model to capture long-range behavior dependencies; The losses of the above prediction tasks are superimposed according to the set ratio, and the user behavior representation model parameters are iteratively optimized through the back propagation algorithm until the loss value converges; The network parameters of the user behavior representation model after training are fixed to the parameters of the source domain behavior feature encoder.

6. The platform user behavior feature transfer learning method for advertising delivery according to claim 5 is characterized in that: The step of obtaining a target domain behavior feature encoder includes: Copying all network parameters of the source domain behavior feature encoder to the target domain behavior feature encoder as initial parameters; While keeping the parameters of the source domain behavioral feature encoder frozen, we introduce a domain discriminator to perform adversarial training with the goal of minimizing the feature difference between the target domain and the source domain. Unfreeze the target domain behavioral feature encoder parameters and further optimize the target domain behavioral feature encoder by combining the minimum mean difference regularization term, so that the feature distributions of the source domain and the target domain gradually converge; When the inter-domain feature difference index reaches the preset convergence standard, the training of the target domain behavior feature encoder is completed.

7. The platform user behavior feature transfer learning method for advertising delivery according to claim 6 is characterized in that: The migration training specifically includes: The training process is organized in batches, alternating between source domain self-supervised tasks and target domain adversarial tasks. In each target domain training phase, when target domain samples lack accurate labels, a pseudo-label generation strategy based on feature vector similarity judgment is adopted to automatically generate labels for samples whose confidence exceeds a set threshold and add them to the training set; By gradually reducing the proportion of source domain samples in training and correspondingly increasing the proportion of target domain samples, the dominant role of target domain data is enhanced; The inter-domain feature difference index is evaluated in real time, and the training convergence condition is that the inter-domain difference index meets the set threshold.

8. The platform user behavior feature transfer learning method for advertising delivery according to claim 6 is characterized in that: Calling the target domain behavior feature encoder to generate a migration feature vector for the latest behavior sequence of active users on the target platform includes: Acquire the most recent continuous behavior sequence of the target platform user in real time, and dynamically truncate the length of the behavior sequence to a standard length according to the input requirements of the target domain behavior feature encoder; The adjusted behavior sequence is input into the trained target domain behavior feature encoder for transfer feature analysis to generate the user's transfer feature vector.

9. The platform user behavior feature transfer learning method for advertising delivery according to claim 8 is characterized in that: The migration feature vector is input into the advertisement placement decision model, and the ranking and bidding results of the advertisements to be placed are output, including: After combining the user migration feature vector with the ad feature information, it is input into the ad placement decision model for multi-task learning to predict the ad click-through rate and conversion rate respectively. Calculate the comprehensive revenue score of each candidate ad based on the click-through rate and conversion rate prediction results and the ad's weight coefficient; Sort the candidate advertisements in descending order according to the comprehensive revenue scores to form an advertisement recommendation list; Based on the advertising budget and historical delivery data, determine the advertising bid, and finally output the ranking and bidding results.

Citation Information

Patent Citations

  • Advertisement classification method and device based on webpage characteristic

    CN102508859A

  • Information pushing method, device and equipment

    CN110400169A

Cited By

  • Cross-border intelligent advertisement putting optimization method and system based on reinforcement learning

    CN121329506A

  • Cross-border intelligent advertisement putting optimization method and system based on reinforcement learning

    CN121329506B