Advertisement pushing method and device, equipment and storage medium

By generating user and consumption behavior vectors through mapping auxiliary networks and memory networks, and combining attention mechanisms and sequence encoders, the accuracy and diversity of advertising push on e-commerce platforms are achieved, thereby improving the advertising marketing conversion rate.

CN120823002AInactive Publication Date: 2025-10-21深圳市信诚数字科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510884254.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The advertising push effect of e-commerce platforms has not met expectations, the algorithm accuracy is not well matched with user needs, and traditional optimization models are unable to cope with the challenges of the dynamic market environment, resulting in low traffic conversion efficiency.

Method used

User data is processed through a preset mapping-assisted network to generate user portrait vectors, user personalization vectors, and consumption behavior vectors. Combined with the user memory network and the consumption memory network, the attention mechanism and sequence encoder are used to enhance feature representation, perform multi-dimensional feature fusion matching, and determine the target push advertisements.

Benefits of technology

It improves the accuracy and diversity of advertising push, solves the problems of insufficient expression of user interests and single recommendations, and improves the conversion rate of advertising marketing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823002A_ABST
    Figure CN120823002A_ABST
Patent Text Reader

Abstract

The invention provides an advertisement pushing method and device, equipment and a storage medium, and the method comprises the steps: processing user data through a preset mapping auxiliary network, and obtaining a user portrait vector, a user personalized vector, a commodity vector and a consumption behavior vector; determining a user portrait enhancement vector according to a preset user memory network, the user portrait vector and the user personalized vector; generating a personalized enhancement vector according to the commodity vector, a preset user memory network and a user portrait vector; determining a consumption clustering center in a preset consumption memory network according to the commodity vector, and inputting an attention mechanism and a sequence encoder to obtain a consumption behavior enhancement vector; and matching the user portrait enhancement vector, the personalized enhancement vector and the consumption behavior enhancement vector with a preset candidate advertisement, and determining a target push advertisement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of big data technology, and in particular to an advertisement push method, apparatus, device, and storage medium. Background Art

[0002] Currently, the lack of advertising effectiveness on e-commerce platforms has become a common pain point across the industry. The core conflict lies in the mismatch between algorithmic accuracy and user needs. While the collaborative filtering and deep learning recommendation models (such as the Twin Towers model) that platforms rely on can generate recommendations based on historical behavior, they are susceptible to interference from "overly precise traffic." This includes misjudged recommendations caused by casual user browsing or label confusion caused by fraudulent order manipulation, resulting in low traffic conversion efficiency. Furthermore, the complexity of user behavior exacerbates the difficulty of advertising. In non-active search scenarios, casual browsing and multiple product comparisons significantly reduce purchasing decision efficiency. Existing solutions often focus on optimizing algorithms and operational strategies. For example, A / B testing can be used to optimize ad creative design (e.g., price tag visibility and celebrity endorsement strategies), or tools (such as Lingxing ERP and Lalimao Fingerprint Browser) can be used to manage ad bidding and prevent account association. However, the dynamic market environment presents new challenges: sudden changes in competing product strategies, platform algorithm adjustments, and a decreasing user fatigue threshold make traditional static optimization models unsustainable. Summary of the Invention

[0003] The present application provides an advertising push method, apparatus, device and storage medium for improving the push accuracy of advertisements and increasing the conversion rate of advertising marketing.

[0004] In a first aspect, an embodiment of the present application provides an advertisement push method, the method comprising: Process user data through a preset mapping auxiliary network to obtain user portrait vectors, user personalization vectors, product vectors, and consumer behavior vectors; Determining a user portrait enhancement vector based on a preset user memory network, the user portrait vector, and the user personalization vector; Generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector; Determine the consumption cluster center in the preset consumption memory network according to the product vector, and input the attention mechanism and sequence encoder to obtain the consumption behavior enhancement vector; The user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector are matched with preset candidate advertisements to determine a target push advertisement.

[0005] In a second aspect, an embodiment of the present application provides an advertisement push device, the device comprising: The data mapping module is used to process user data through a preset mapping auxiliary network to obtain user portrait vectors, user personalization vectors, product vectors, and consumption behavior vectors; A first enhancement module, configured to determine a user portrait enhancement vector based on a preset user memory network, the user portrait vector, and the user personalization vector; A second enhancement module is configured to generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector; A third enhancement module is configured to determine the consumption cluster center in a preset consumption memory network based on the product vector, and input the center into an attention mechanism and a sequence encoder to obtain a consumption behavior enhancement vector; The advertisement matching module is used to match the user portrait enhancement vector, the personalization enhancement vector and the consumption behavior enhancement vector with preset candidate advertisements to determine the target push advertisement.

[0006] In a third aspect, an embodiment of the present application provides an electronic device, the electronic device including a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the advertisement pushing method as described in any one of the embodiments of the present application when executing the computer program.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor enables the processor to implement the advertising push method as described in any one of the embodiments of the present application.

[0008] The embodiment of the present application provides an advertisement push method, apparatus, device and storage medium, the method comprising: processing user data through a preset mapping auxiliary network to obtain a user portrait vector, a user personalization vector, a product vector and a consumption behavior vector; determining a user portrait enhancement vector based on a preset user memory network, a user portrait vector and a user personalization vector; generating a personalization enhancement vector based on the product vector, the preset user memory network and the user portrait vector; determining a consumption cluster center in a preset consumption memory network based on the product vector, and inputting the attention mechanism and the sequence encoder to obtain a consumption behavior enhancement vector; matching the user portrait enhancement vector, the personalization enhancement vector and the consumption behavior enhancement vector with the preset candidate advertisements to determine the target push advertisement. In the above method, a unified representation of the features of the data piece is achieved through a mapping auxiliary network, and the three memory networks are combined to enhance the user's static interests, personalization preferences and dynamic behavior features respectively, and the behavior evolution law is captured through the attention mechanism and the sequence encoder. The fusion matching mechanism of multi-dimensional features not only ensures the accuracy of advertisement push, but also maintains the diversity of recommendations, effectively solving the technical problems of insufficient expression of user interests and single recommendations. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic flow chart of an advertisement push method provided in an embodiment of the present application; Figure 2 A schematic block diagram of an advertisement push device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0011] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0012] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0013] It should also be understood that the terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0014] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0015] See also Figure 1 , Figure 1 This is a schematic flow chart of an advertisement push method provided by an embodiment of the present application. Figure 1 As shown, the specific steps of the advertisement pushing method include: S101-S105.

[0016] S101. Process user data through a preset mapping auxiliary network to obtain a user portrait vector, a user personalization vector, a product vector, and a consumption behavior vector.

[0017] Exemplarily, the pre-configured mapping auxiliary networks include a user-assisted network and an item-assisted network. The user-assisted network utilizes a three-layer fully connected neural network architecture with 256 input nodes, 128 hidden nodes, and 64 output nodes. The network receives user data, including demographic characteristics (age, gender, and region), historical behavior characteristics (clicks, favorites, and purchases), and contextual characteristics (time and device). During the feature engineering phase, categorical features are one-hot encoded and numerical features are normalized to generate user profile vectors and user-personalized vectors. The item-assisted network utilizes the same network architecture, taking in product feature information and outputting a 64-dimensional product vector. For each user's historical consumption behavior vector, the item-assisted network converts each behavior into a consumption behavior vector, sampling the most recent 100 behaviors to form a sequence. During training, a learning rate of 0.001 and a training subset size of 256 are set. The Adam optimizer is used for parameter updates. To enhance the model's generalization, a dropout mechanism is introduced during training, with a dropout rate of 0.3. L2 regularization is also used to prevent overfitting, with a regularization coefficient of 0.0001. For sparse features, an embedding layer is used for dimensionality reduction, with the embedding dimension set based on the logarithm of the feature cardinality. To address the temporal characteristics of user behavior sequences, a time decay factor is added during the feature engineering phase, with a decay coefficient of 0.01. Negative sampling is used during model training, with a negative to positive sample ratio of 4:1. The loss function is calculated using cosine similarity.

[0018] S102: Determine a user portrait enhancement vector based on a preset user memory network, a user portrait vector, and a user personalization vector.

[0019] Exemplarily, 128 64-dimensional cluster center vectors are pre-initialized in the user memory network, and the vector elements obey a normal distribution with a mean of 0 and a variance of 0.01. The cosine similarity between the user portrait vector and all cluster centers is calculated to obtain a 128-dimensional similarity score. The top 16 cluster centers with the highest similarity are selected, and the cluster centers with negative similarity are replaced with zero vectors. The 16 similarity scores are normalized using the softmax function, and the temperature parameter is set to 0.1. The normalized weights are weighted summed with the corresponding cluster center vectors and concatenated with the user personalized vector to generate a 128-dimensional user portrait enhancement vector. The cluster centers are updated using an exponential moving average method with a momentum coefficient of 0.9, and are updated every 1000 iterations. In order to improve the representativeness of the cluster centers, a center point update strategy is introduced. When a cluster center is not selected by any user vector for 100 consecutive times, it is reset to the mean of the most recent 100 user vectors. When calculating similarity, we use a spherical distance metric, which better handles vector similarity calculations in high-dimensional spaces. We also set a time weighting factor for user behavior over different time periods, giving recent behavior a greater influence on cluster center updates. Furthermore, to prevent drastic changes in cluster centers, we set a maximum update step size of 0.1, and L2 normalized the cluster centers after updates.

[0020] S103: Generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector.

[0021] For example, the product vector is input into the user memory network, and the 24 cluster centers with the highest similarity are retrieved. The inner product of these cluster centers with the current user profile vector is calculated, and vectors corresponding to negative results are set to zero. A 64-dimensional personalized enhancement vector is maintained for each user in the personality memory network, initialized to a zero vector. The personalized enhancement vector is updated based on the retrieved cluster centers and the normalized similarity score. The update process uses a gated update mechanism, with the update intensity positively correlated with the user's activity level, which is calculated based on the frequency of their recent behavior. The update step size of the personalized enhancement vector is set to 0.1, and L2 regularization is applied after each update. To capture the dynamic changes in user interests, a forgetting mechanism is introduced to decay historical interest features, with the decay coefficient exponentially proportional to the time interval. When updating the personalized enhancement vector, the diversity of user behavior is taken into account. The degree of dispersion of user behavior is calculated using Shannon entropy. The higher the entropy value, the larger the update step size. Furthermore, to prevent the personalized enhancement vector from being dominated by extreme behaviors, maximum and minimum constraints are set to ensure that the vector elements are within the range [-1, 1]. During the training process, the interest drift speed of each user is recorded and the update step size is dynamically adjusted.

[0022] S104: Determine the consumption cluster center in the preset consumption memory network based on the product vector, and input the attention mechanism and sequence encoder to obtain the consumption behavior enhancement vector.

[0023] For example, 256 64-dimensional cluster centers are randomly initialized in the consumer memory network using Xavier initialization. For each of the user's most recent 100 consumer behavior vectors, the Euclidean distance between each vector and all cluster centers is calculated, and the eight centers with the smallest distances are selected. Attention weights are calculated using the softmax function, with a temperature parameter set to 0.2. Weighted summation is performed on the selected cluster centers to obtain an enhanced vector for each consumer behavior. These enhanced vectors are sequenced and input into a 6-layer Transformer encoder with a hidden layer dimension of 64, 4 attention heads, and a feedforward network dimension of 256, resulting in a 64-dimensional enhanced consumer behavior vector. Positional encoding information is incorporated into the sequence processing, using a sinusoidal positional encoding scheme. To handle variable-length sequences, an attention masking mechanism is used, with a masking threshold set to -1e9. Layer Normalization and residual connections are added after each Transformer layer, with a dropout rate set to 0.1. To enhance the model's ability to model long-term dependencies, relative position encoding is introduced into the self-attention mechanism, with a maximum distance set to 50. At the same time, the gated linear unit (GLU) activation function is used to improve the nonlinear expression ability of the model.

[0024] S105: Match the user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector with the preset candidate advertisements to determine the target push advertisement.

[0025] For example, the user profile enhancement vector, personalization enhancement vector, and consumer behavior enhancement vector are fed into a multi-head attention layer with eight heads. The attention weights of the three vectors are calculated, with the sum of the weights being 1. Pre-set candidate ads are vectorized, with the same dimensions as the user vector. The cosine similarity between the fused user vector and the ad vector is calculated, with a temperature parameter of 0.2. The ad with the highest similarity is selected as the target ad for push notification, with a similarity threshold of 0.5. User feedback on push ads, such as clicks and conversions, is recorded. This feedback data is then used to update the user's personalization enhancement vector after feature extraction, with an update cycle of 12 hours. To balance recommendation accuracy and diversity, an exploration factor is introduced, the size of which is positively correlated with the diversity of a user's historical clicks. When matching ads, a stratified sampling strategy is employed, classifying candidate ads into multiple tiers based on their historical click-through rates. The sampling probability of each tier is proportional to its average click-through rate. Furthermore, considering ad timeliness and budget constraints, the matching probability of ads that exceed budget or are about to expire is reduced. To prevent recommendation results from being dominated by popular ads, a reranking mechanism is implemented to adjust the matching score based on recent user feedback on similar ads. During the actual delivery process, the parameter configuration of the matching strategy is continuously optimized through A / B testing.

[0026] In order to more clearly introduce the technical solution of the present application, the technical solution of the present application will be introduced through specific embodiments below. It should be noted that the specific embodiments are used to expand the technical solution of the present application, but are not intended to limit the present application.

[0027] In some embodiments, user data is processed through a preset mapping auxiliary network to obtain a user portrait vector, a user personalization vector, a product vector, and a consumption behavior vector, including: S101-S104.

[0028] S101. Classify demographic features, historical behavior features, and context features in user data, and standardize them to obtain standardized features.

[0029] Exemplarily, user data is classified and processed. Demographic features include discrete features such as age, gender, region, occupation, and income. Historical behavior features include sequential features such as clicks, browsing, favorites, add-to-cart, and purchases. Contextual features include environmental features such as time, device type, and network environment. For discrete features, one-hot encoding is used to convert them into sparse vectors. The encoding dimension is set according to the number of feature values. Hash encoding is used for dimensionality reduction for features with more than 100 values. For continuous features, the min-max normalization method is used to map the values ​​to the [0, 1] interval. When processing outliers, upper and lower thresholds are set, and values ​​exceeding the threshold are truncated to the threshold position. For time features, periodic features such as hours, weeks, and months are extracted, and the sin-cos transformation is used for periodic encoding. For sequential features, a fixed-length time window is set to 30 days. Historical behaviors exceeding the time window are truncated, and sequences that are less than the window length are padded with zero vectors. During the standardization process, the mean and standard deviation of each feature are calculated. The z-score standardization method is used to transform the features into a standard normal distribution with a mean of 0 and a variance of 1, resulting in the standardized features. To ensure data quality, a missing value filling strategy is set: numerical features are filled with the median, and categorical features are filled with the mode.

[0030] S102. Input the standardized features into the first fully connected layer. The first fully connected layer is used to map input features of different dimensions to the hidden space of the same dimension, and introduce nonlinear transformation through the ReLU activation function.

[0031] For example, the first fully connected layer adopts a three-layer structure. The input layer dimension is set according to the total feature dimension, the middle hidden layer dimension is 512, and the output layer dimension is 256. Relu activation functions are used between each layer to introduce nonlinear transformations, with the slope of the activation function on the negative semi-axis set to 0.01 to prevent neuron death. To accelerate network convergence, a batch normalization layer is added after each fully connected layer, with the momentum parameter set to 0.9 and the epsilon parameter set to 1e-5. Weight initialization uses the He initialization method to ensure consistent input variance within each layer. The Adam optimizer is used during training with a learning rate of 0.001, beta1 parameters of 0.9, and beta2 parameters of 0.999. To prevent vanishing and exploding gradients, residual connections are added after each fully connected layer. When the output of the residual block does not match the input dimension, 1x1 convolution is used to align the dimensions. Furthermore, to enhance the model's feature extraction capabilities, an attention mechanism is introduced in the fully connected layers to calculate the correlation weights between different features. The number of attention heads is set to 8. For sparse features, an embedding layer is used in the input layer for dimensionality reduction. The embedding dimension is set to 16, and the number of parameters is reduced by sharing the embedding weights.

[0032] S103. Input the output result of the first fully connected layer into the second fully connected layer. The second fully connected layer is used to capture the high-order interaction relationship between features and randomly discard some neurons through the dropout mechanism to prevent overfitting.

[0033] For example, the second fully connected layer uses a four-layer structure, with 256, 128, 128, and 64 nodes per layer, respectively. Leaky ReLU activation functions are used between adjacent layers, with a negative slope of 0.2. A dropout layer with a dropout rate of 0.3 is placed after each fully connected layer. During training, 30% of neurons are randomly dropped, and during testing, weights are scaled by 0.7. To capture high-order interactions between features, a crossover layer is added between the second and third layers, with a crossover layer count of 3. Each crossover layer calculates second- and third-order interactions between features. The weight matrix uses Xavier initialization to ensure consistent variance between inputs and outputs at each layer. To enhance the model's expressiveness, a multi-head self-attention layer with 4 heads and a key-value pair dimension of 32 is added before the last fully connected layer. The scaled dot-product attention mechanism calculates feature correlations. During training, an annealing learning rate strategy is used, with an initial learning rate of 0.001 and a decay of 0.9 every 1000 iterations. At the same time, gradient clipping is used to limit the L2 norm of the gradient to less than 5 to prevent gradient explosion. To improve the generalization ability of the model, the mixup data augmentation technique is used during training, and the mixing parameter alpha is set to 0.2.

[0034] S104. The output of the second fully connected layer is input into independent fully connected layers in the user profile branch, user personalization branch, product feature branch, and consumer behavior branch. The fully connected layer in the user profile branch extracts static user interest features and outputs a fixed-dimensional user profile vector. The fully connected layer in the user personalization branch learns the user's personalized preferences and outputs a personalized user vector. The fully connected layer in the product feature branch extracts key product attributes and outputs a product vector. The fully connected layer in the consumer behavior branch models the user's dynamic behavior patterns and outputs a consumer behavior vector.

[0035] For example, the independent fully connected layers of the four branches use the same network architecture, consisting of two hidden layers with 64 and 32 nodes, respectively, and an output layer dimension of 16. The user profile branch adds an attention pooling layer after the fully connected layer to perform weighted aggregation of user static features. The attention weights are calculated using a trainable query vector. The user personalization branch uses a gating mechanism to control the contribution of different features to the output vector. The activation function of the gating unit is a sigmoid function. The product feature branch introduces a feature selection layer before the fully connected layer and uses L1 regularization to learn feature importance weights with a regularization coefficient of 0.01. The consumer behavior branch uses a temporal attention mechanism to assign different weights to behavioral features in different time windows, with a temporal decay factor of 0.1. The loss functions of the four branches adopt a multi-task learning framework, including classification loss, regression loss, and contrastive loss. Task weights are determined through grid search. To balance the gradient scales of different branches, gradient normalization is used to scale the gradients of each branch to the same magnitude. During training, an alternating training strategy is adopted, with one branch randomly selected for parameter update at each epoch. At the same time, an early stopping strategy is used to prevent overfitting. Training is stopped when the validation set loss does not decrease for five consecutive epochs. To improve the interpretability of the model, an attention visualization module is added to the output layer of each branch to display the importance scores of different features.

[0036] In some embodiments, determining a user portrait enhancement vector according to a preset user memory network, a user portrait vector, and a user personalization vector includes: S201-S206.

[0037] S201. Randomly initialize N cluster center vectors in the user memory network, and the dimension of each cluster center vector is the same as the dimension of the user portrait vector.

[0038] Exemplarily, N=128 cluster center vectors are initialized in the user memory network, and the dimension of each vector is 64 dimensions, which is consistent with the dimension of the user portrait vector. The initial value of the vector element adopts the Xavier initialization method, and obeys the uniform distribution with a mean of 0 and a variance of 2 / (input dimension + output dimension). To ensure the effectiveness of the cluster center, the initialized vector is L2 normalized. Each cluster center vector is assigned a unique identifier to facilitate subsequent update and retrieval operations. During the initialization process, the position of the initial cluster center is selected by the K-means++ algorithm to improve the dispersion of the cluster center. At the same time, a counter is maintained for each cluster center to record the number of times the center is selected for subsequent update strategy adjustments.

[0039] S202. Calculate the cosine similarity between the user portrait vector and the N cluster center vectors to obtain a similarity score matrix.

[0040] Exemplarily, the cosine similarity between the user portrait vector and the 128 cluster center vectors is calculated. The cosine similarity is calculated by dividing the vector dot product by the vector modulus, and the computational efficiency is improved through batch matrix operations. The dimension of the calculated similarity score matrix is ​​1x128, and each element in the matrix represents the degree of similarity between the user portrait vector and the corresponding cluster center vector, with a value range of [-1, 1]. To improve the calculation accuracy, double-precision floating-point numbers are used in the calculation process. For similarity scores that are too small (less than 1e-6), they are set to zero to avoid numerical instability. Vectorized operations are used in the similarity calculation process to reduce loop structures and improve calculation speed.

[0041] S203. Based on the similarity score matrix, the cluster center vector that is most similar to the user portrait vector is updated using an exponential moving weighted average method.

[0042] For example, based on the similarity score matrix, the cluster center vector with the highest similarity is updated. The update adopts the exponential moving weighted average method. During the update process, the learning rate is set to 0.01, and the vector is L2 normalized after each update. To prevent drastic changes in the cluster center, the maximum step size of a single update is set to 0.1. The update frequency is related to the frequency of user interaction. The cluster center of active users is updated more frequently. At the same time, the usage frequency statistics of each cluster center are maintained. When a center is not selected for 1000 consecutive times, it is reset to the mean of the last 1000 user vectors.

[0043] S204: Retrieve K1 cluster center vectors that are most similar to the user portrait vector in the user memory network, and set the cluster center vectors with negative similarity to the user portrait vector to zero vectors.

[0044] For example, the K1=16 cluster center vectors that are most similar to the user portrait vector are retrieved in the user memory network. The retrieval process uses cosine similarity sorting to select the 16 vectors with the highest similarity. For cluster center vectors with negative similarity, they are replaced with 64-dimensional zero vectors. The retrieval process uses a heap data structure for TopK selection, and the time complexity is O(NlogK1). In order to improve the retrieval efficiency, a nearest neighbor index structure is maintained, and the local sensitive hashing (LSH) technology is used to accelerate the similarity calculation. The retrieval result contains the ID of the cluster center vector and the corresponding similarity score. The replacement operation of the zero vector is implemented through a vector mask to ensure the efficiency of the calculation process.

[0045] S205. Perform softmax normalization on the similarity scores of the K1 cluster center vectors to obtain attention weights, and perform weighted summation on the K1 cluster center vectors based on the attention weights to obtain a preliminary enhanced vector.

[0046] For example, the similarity scores of the 16 retrieved cluster center vectors are softmax normalized, and the temperature parameter is set to 0.1. The normalized scores are used as attention weights, and the weight sum is 1. Based on the attention weights, the 16 cluster center vectors are weighted and summed to obtain a 64-dimensional preliminary enhancement vector. In the weighted summation process, matrix multiplication is used to implement batch calculations. To enhance the robustness of the model, Gaussian noise is added to the attention weights, and the standard deviation of the noise is set to 0.01. At the same time, the preliminary enhancement vector is L2 normalized to maintain the numerical stability of the vector. Double-precision floating-point numbers are used in actual operations to ensure the accuracy of numerical calculations.

[0047] S206: Concatenate the preliminary enhancement vector and the user personalized vector to obtain a user portrait enhancement vector.

[0048] For example, the 64-dimensional preliminary enhancement vector and the 64-dimensional user personalized vector are spliced ​​to obtain a 128-dimensional user portrait enhancement vector. The splicing operation uses vector concatenation to maintain the original information of the two vectors. In order to balance the importance of the two parts of information, the spliced ​​vector is normalized. During the splicing process, the alignment of the vectors is considered to ensure dimension matching. To improve the efficiency of subsequent processing, the splicing results are stored as continuous memory blocks. At the same time, the position information of the preliminary enhancement vector and the user personalized vector in the spliced ​​vector is recorded to facilitate subsequent feature extraction and update operations. After the splicing is completed, the vector is verified for validity to ensure that there are no invalid values ​​or abnormal values.

[0049] In some embodiments, generating a personalized enhancement vector based on the product vector, a preset user memory network, and a user portrait vector includes: S1031-S1034.

[0050] S1031. Retrieve the K2 cluster center vectors that are most similar to the product vector in the user memory network, and calculate the corresponding similarity scores.

[0051] For example, the cosine similarity metric is used to calculate the similarity between the product vector and the cluster center vectors in the user memory network. The similarity calculation formula is the dot product of the two vectors divided by the product of the vector moduli. A threshold of K2 is set to 5, and the five cluster center vectors with the highest similarity are selected and their corresponding similarity scores are recorded. The similarity score reflects the degree of match between the recommended product and the user's historical interests. A higher score indicates that the product more closely matches the user's preferences.

[0052] S1032. Calculate the inner product of K2 cluster center vectors and the user portrait vector, and set the cluster center vector with a negative inner product to a zero vector.

[0053] For example, the inner product value is calculated for each of the five selected cluster center vectors with the user profile vector. The inner product operation can capture the correlation between vectors. A positive value indicates that the two vectors point in similar directions, while a negative value indicates opposite directions. Cluster center vectors with negative inner products are replaced with zero vectors of the same dimension as the vector. Cluster center vectors with positive correlation are retained to avoid introducing negative interference.

[0054] S1033: In the preset personalized memory network, initialize a corresponding personalized enhancement vector for each user as a zero vector.

[0055] For example, in the personalized memory network, each user is assigned a zero vector with the same dimension as the product vector as the initial state of the personalized enhancement vector. The personalized memory network uses a key-value storage structure, with the user ID as the key and the corresponding zero vector as the value. Each component of the zero vector is 0, providing a basis for subsequent similarity score-based enhancement vector updates.

[0056] S1034: Update the personalized enhancement vector based on the similarity score and the non-zero vector in the preset personalized memory network.

[0057] For example, the remaining non-zero cluster center vectors are weighted and summed according to their similarity scores to update the corresponding user's personalized enhancement vector in the personalized memory network. The weight coefficient is set to the normalized value of the similarity score, ensuring that the contribution of each cluster center vector to the enhancement vector is proportional to its similarity. The updated personalized enhancement vector incorporates the user's historical interest characteristics and the current product characteristics.

[0058] In some embodiments, a consumption cluster center is determined in a preset consumption memory network based on the product vector, and an attention mechanism and a sequence encoder are input to obtain a consumption behavior enhancement vector, including: S1041-S1044.

[0059] S1041. Randomly initialize M cluster center vectors in the consumer memory network, where the dimension of each cluster center vector is the same as the dimension of the product vector.

[0060] For example, in the consumer memory network, a Gaussian distribution is used to randomly initialize M = 128 cluster center vectors. Each cluster center vector has a dimension of 256, consistent with the product vector. During initialization, the mean of the Gaussian distribution is set to 0 and the standard deviation is set to 0.01 to ensure that the initial values ​​of the cluster center vectors are distributed within a reasonable range. The cluster center vectors are normalized using the Xavier initialization method to keep the values ​​of each dimension within the range [-1, 1], providing a good numerical foundation for subsequent similarity calculations.

[0061] S1042: Calculate the similarity between each product vector in the consumption behavior vector and the M cluster center vectors, and select K3 cluster center vectors that are most similar to each product vector.

[0062] For example, for each product vector in the consumer behavior vector, we calculate its cosine similarity with the 128 cluster center vectors. Cosine similarity is calculated by dividing the vector inner product by the vector modulus, with a similarity range of [-1, 1]. Using a heap sort algorithm, we select the cluster center vectors corresponding to the largest K3=8 similarity values ​​from the 128 similarity values ​​to construct a set of neighboring cluster centers for the product vector, preparing for the calculation of attention weights.

[0063] S1043. Calculate the attention weight based on the attention mechanism and similarity, perform weighted summation on the K3 cluster center vectors, and obtain the behavior enhancement vector corresponding to each consumption behavior.

[0064] For example, the product vector is used as the query vector and the eight cluster center vectors as the key-value vectors. Attention scores are calculated using the scaled dot-product attention mechanism. The attention scores are normalized using the Softmax function and converted into attention weights, with the sum of the weights being 1. Each of the eight attention weights is multiplied by the corresponding cluster center vector and summed to generate a 256-dimensional behavior enhancement vector that incorporates information from multiple relevant cluster centers.

[0065] S1044: Input all behavior enhancement vectors in the consumption behavior vector into a sequence encoder to generate a consumption behavior enhancement vector.

[0066] For example, a bidirectional LSTM is used as a sequence encoder with a hidden layer dimension of 256. The behavior enhancement vectors in the consumer behavior vector are input into the encoder in chronological order. The bidirectional LSTM processes the sequence information in both the forward and reverse directions, capturing the temporal dependencies between behaviors. The forward and reverse hidden states are concatenated and mapped through a fully connected layer to generate a 256-dimensional consumer behavior enhancement vector, which contains the complete consumer behavior vector information of the user.

[0067] In some embodiments, the user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector are matched with preset candidate advertisements to determine the target push advertisement, including: S1051-S1055.

[0068] S1051. Calculate the importance weights of the user portrait enhancement vector, personalization enhancement vector, and consumer behavior enhancement vector through the attention mechanism.

[0069] For example, for the user portrait enhancement vector, personalization enhancement vector, and consumer behavior enhancement vector, linear transformations are used to map them to the query vector, key vector, and value vector space, respectively. The raw attention score is obtained by calculating the dot product of the query vector and the key vector and dividing it by the square root of the vector dimension for normalization. These raw scores are converted into probability distributions through the softmax function to obtain the importance weights of each vector. This weight calculation method based on the attention mechanism can adaptively adjust the importance of different features, allowing the model to dynamically adjust the contribution of each vector according to different scenarios.

[0070] S1052. Based on the importance weight, perform weighted fusion on the user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector to obtain a fused user vector.

[0071] Exemplarily, this step uses the importance weights calculated in S1051 to perform weighted fusion on the three enhanced vectors. Multiply the user portrait enhancement vector, the personalized enhancement vector, and the consumer behavior enhancement vector by their corresponding importance weights to obtain weighted vectors. Perform element-wise summation on these weighted vectors to generate a fused user vector. In order to maintain the expressive power of the vector, a residual connection mechanism is introduced in the fusion process to combine the original three enhanced vectors with the weighted fused vector. This fusion method not only retains the original information of each vector, but also highlights more important features through attention weights.

[0072] S1053: Convert the preset candidate advertisement into an advertisement vector through a mapping auxiliary network.

[0073] For example, the mapping-assisted network consists of two sub-networks: a user-assisted network and an item-assisted network. The item-assisted network is used when processing candidate ads. For each candidate ad, features such as its title, description, category, and price are extracted and processed through feature engineering before being fed into the item-assisted network. The item-assisted network, through a multi-layer neural network structure, converts these heterogeneous features into an ad vector of uniform dimension. This vector representation allows ad features to be compared with user vectors in the same vector space.

[0074] S1054: Calculate the matching score between the fused user vector and the advertisement vector, and introduce a temperature parameter to adjust the matching score.

[0075] For example, cosine similarity is used to calculate the similarity between two vectors, yielding a raw match score. A temperature parameter, τ, is introduced to adjust the match score by dividing the raw score by the temperature parameter. A larger temperature parameter yields a smoother distribution of match scores and more diverse recommendation results. A smaller temperature parameter yields more pronounced differences in match scores, resulting in a greater concentration of recommendations on highly matching ads. By adjusting the temperature parameter, a balance can be achieved between recommendation accuracy and diversity.

[0076] S1055: Based on the adjusted matching scores, select K candidate advertisements with the highest scores as target push advertisements.

[0077] Exemplarily, this step selects target ads based on the adjusted match scores. All candidate ads are sorted in descending order according to the adjusted match scores, and the K ads with the highest scores are selected as the target ads. During the selection process, historical performance metrics such as impressions and click-through rates are considered, and ads that have received a high number of impressions but performed poorly are appropriately prioritized. Furthermore, to prevent overly similar content from being pushed, a diversity constraint is incorporated into the selection process to ensure that the K selected ads are somewhat different. This selection mechanism ensures both the relevance of the pushed ads and the diversity of the recommendation results.

[0078] In some embodiments, calculating a matching score between the fused user vector and the advertisement vector, and introducing a temperature parameter to adjust the matching score, includes: S541 - S545 .

[0079] S541 . Perform L2 regularization on the fused user vector to obtain a standardized user vector.

[0080] Exemplarily, an L2 regularization algorithm is applied to the 512-dimensional fused user vector, the Euclidean norm of the vector is calculated, and each dimension value of the vector is divided by the norm to normalize the modulus of the vector to 1, thereby obtaining a standardized user vector.

[0081] S542: Perform L2 regularization on the advertisement vector to obtain a standardized advertisement vector.

[0082] Exemplarily, the same L2 regularization process is applied to the 512-dimensional advertisement vector, the Euclidean norm of the vector is calculated and normalized so that the modulus of the vector is standardized to 1, thereby obtaining a standardized advertisement vector.

[0083] S543: Calculate cosine similarity based on the normalized user vector and the normalized advertisement vector to obtain an initial matching score.

[0084] Exemplarily, the dot product of two normalized vectors is calculated. Since the vectors have been L2 regularized, the dot product result is directly equal to the cosine similarity value, which ranges from [-1, 1] to obtain the initial matching score.

[0085] S544: Scale the initial matching score using a preset temperature parameter to obtain a scaled matching score.

[0086] For example, the initial value of the temperature parameter T is set to 0.07, and the initial matching score is scaled through the following four sub-steps: (1) Numerical mapping is performed on the initial matching score, and the similarity value in the interval [-1, 1] is mapped to the interval [0, 1]. Specifically, the initial matching score is added by 1 and then divided by 2, so that the score value is limited to the range [0, 1], which is convenient for subsequent temperature adjustment. The mapped score maintains the relative size relationship of the original similarity. (2) A nonlinear transformation is performed on the mapped score based on the temperature parameter T. The mapped score is divided by the temperature parameter T. A smaller T value will amplify the difference between the scores, and a larger T value will weaken the difference between the scores. The choice of T value directly affects the model's sensitivity to similarity differences. When T approaches 0, the model will be more inclined to select the advertisement with the highest similarity. (3) An adaptive adjustment mechanism is introduced to dynamically adjust the temperature parameter. According to the average similarity level of the current batch of advertisement candidate sets, the T value is adjusted within the range [0.05, 0.1]. When the average similarity of candidate ads is high, the T value is appropriately increased to increase the diversity of ad selection; when the average similarity is low, the T value is reduced to highlight the similarity difference. (4) Apply an exponential function to transform the scaled score. Substitute the temperature-adjusted score into exp(score / T) to calculate the scaled matching score. The exponential transformation can further amplify the score difference, making ads with high similarity have a higher probability of being selected, while keeping the probability value always positive.

[0087] S545. According to the scaled matching score, a softmax function is used to perform normalization processing to obtain a matching score.

[0088] Exemplarily, the scaled matching scores are normalized by applying a softmax function, and the scores of all candidate ads are converted into a probability distribution, where the sum of the probability values ​​of each ad is 1.

[0089] See also Figure 2 , Figure 2 2 is a schematic block diagram of an advertisement pushing device provided in an embodiment of the present application, wherein the advertisement pushing device 200 is used to execute the aforementioned advertisement pushing method.

[0090] Among them, the server can be an independent server, a server cluster, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0091] like Figure 2 As shown, the advertisement pushing device 200 includes: a data mapping module 201 , a first enhancement module 202 , a second enhancement module 203 , a third enhancement module 204 and an advertisement matching module 205 .

[0092] The data mapping module 201 is used to process user data through a preset mapping auxiliary network to obtain a user portrait vector, a user personalization vector, a product vector and a consumption behavior vector.

[0093] The first enhancement module 202 is configured to determine a user portrait enhancement vector based on a preset user memory network, a user portrait vector, and a user personalization vector.

[0094] The second enhancement module 203 is configured to generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector.

[0095] The third enhancement module 204 is used to determine the consumption cluster center in the preset consumption memory network according to the product vector, and input the attention mechanism and sequence encoder to obtain the consumption behavior enhancement vector.

[0096] The advertisement matching module 205 is used to match the user portrait enhancement vector, the personalization enhancement vector and the consumption behavior enhancement vector with the preset candidate advertisements to determine the target push advertisement.

[0097] An embodiment of the present application provides an electronic device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement any one of the advertising push methods in the embodiments of the present application when executing the computer program.

[0098] An embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the processor implements any one of the advertising push methods in the embodiments of the present application.

[0099] The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. An advertisement push method, characterized in that: The method comprises: Process user data through a preset mapping auxiliary network to obtain user portrait vectors, user personalization vectors, product vectors, and consumer behavior vectors; Determining a user portrait enhancement vector based on a preset user memory network, the user portrait vector, and the user personalization vector; Generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector; Determine the consumption cluster center in the preset consumption memory network according to the product vector, and input the attention mechanism and sequence encoder to obtain the consumption behavior enhancement vector; The user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector are matched with preset candidate advertisements to determine a target push advertisement.

2. The advertisement pushing method according to claim 1, wherein: The determining of the user portrait enhancement vector according to the preset user memory network, the user portrait vector, and the user personalization vector includes: Randomly initializing N cluster center vectors in the user memory network, where the dimension of each cluster center vector is the same as the dimension of the user portrait vector; Calculate the cosine similarity between the user portrait vector and the N cluster center vectors to obtain a similarity score matrix; Based on the similarity score matrix, the cluster center vector that is most similar to the user portrait vector is updated by using an exponential moving weighted average method; Retrieving K1 cluster center vectors that are most similar to the user portrait vector in the user memory network, and setting the cluster center vectors with negative similarity to the user portrait vector to zero vectors; Performing softmax normalization on the similarity scores of the K1 cluster center vectors to obtain attention weights, and performing weighted summation on the K1 cluster center vectors based on the attention weights to obtain a preliminary enhanced vector; The preliminary enhancement vector is concatenated with the user personalized vector to obtain the user portrait enhancement vector.

3. The advertisement pushing method according to claim 1, wherein: Generating a personalized enhancement vector according to the product vector, the preset user memory network, and the user portrait vector includes: Retrieving the K2 cluster center vectors that are most similar to the product vector in the user memory network and calculating the corresponding similarity scores; Calculate the inner product of the K2 cluster center vectors and the user portrait vector, and set the cluster center vector with a negative inner product to a zero vector; In the preset personality memory network, a corresponding personalized enhancement vector is initialized as a zero vector for each user; The personalized enhancement vector is updated based on the similarity score and the non-zero vector in the preset personality memory network.

4. The advertisement pushing method according to claim 1, wherein: The process of determining the consumption cluster center in a preset consumption memory network based on the product vector and inputting the attention mechanism and sequence encoder to obtain a consumption behavior enhancement vector includes: Randomly initializing M cluster center vectors in the consumption memory network, where the dimension of each cluster center vector is the same as the dimension of the product vector; Calculate the similarity between each product vector in the consumption behavior vector and the M cluster center vectors, and select K3 cluster center vectors that are most similar to each product vector; Calculating attention weights based on the attention mechanism and the similarity, performing weighted summation on the K3 cluster center vectors, and obtaining a behavior enhancement vector corresponding to each consumption behavior; All behavior enhancement vectors in the consumption behavior vector are input into a sequence encoder to generate a consumption behavior enhancement vector.

5. The advertisement pushing method according to claim 1, wherein: The user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector are matched with preset candidate advertisements to determine a target push advertisement, including: Calculate the importance weights of the user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector through an attention mechanism; Based on the importance weights, weighted fusion is performed on the user portrait enhancement vector, the personalization enhancement vector, and the consumption behavior enhancement vector to obtain a fused user vector; Converting the preset candidate advertisement into an advertisement vector through the mapping auxiliary network; Calculating a matching score between the fused user vector and the advertisement vector, and introducing a temperature parameter to adjust the matching score; Based on the adjusted matching scores, the K candidate ads with the highest scores are selected as target push ads.

6. The advertisement pushing method according to claim 5, wherein: The calculating the matching score between the fused user vector and the advertisement vector, and introducing a temperature parameter to adjust the matching score, includes: Performing L2 regularization on the fused user vector to obtain a standardized user vector; Performing L2 regularization on the advertisement vector to obtain a standardized advertisement vector; Calculating cosine similarity based on the standardized user vector and the standardized advertisement vector to obtain an initial matching score; Scaling the initial matching score by a preset temperature parameter to obtain a scaled matching score; The scaled matching scores are normalized using a softmax function to obtain matching scores.

7. The advertisement pushing method according to claim 1, wherein: The user data is processed by the preset mapping auxiliary network to obtain the user portrait vector, user personalization vector, product vector and consumption behavior vector, including: Classifying demographic features, historical behavior features, and context features in the user data and normalizing them to obtain standardized features; The standardized features are input into the first fully connected layer, which is used to map the input features of different dimensions to the hidden space of the same dimension and introduce nonlinear transformation through the ReLU activation function; The output of the first fully connected layer is input into the second fully connected layer, which is used to capture high-order interactions between features and randomly discard some neurons through the dropout mechanism to prevent overfitting; The output results of the second fully connected layer are respectively input into the independent fully connected layers of the user portrait branch, user personalization branch, product feature branch and consumer behavior branch, wherein the independent fully connected layer of the user portrait branch is used to extract the user's static interest features and output a fixed-dimensional user portrait vector; the independent fully connected layer of the user personalization branch is used to learn the user's personalized preferences and output a user personalized vector; the independent fully connected layer of the product feature branch is used to extract the key attribute features of the product and output a product vector; the independent fully connected layer of the consumer behavior branch is used to model the user's dynamic behavior pattern and output a consumer behavior vector.

8. An advertisement pushing device, characterized in that: The advertisement pushing device comprises: The data mapping module is used to process user data through a preset mapping auxiliary network to obtain user portrait vectors, user personalization vectors, product vectors, and consumption behavior vectors; A first enhancement module, configured to determine a user portrait enhancement vector based on a preset user memory network, the user portrait vector, and the user personalization vector; A second enhancement module is configured to generate a personalized enhancement vector based on the product vector, the preset user memory network, and the user portrait vector; A third enhancement module is configured to determine the consumption cluster center in a preset consumption memory network based on the product vector, and input the center into an attention mechanism and a sequence encoder to obtain a consumption behavior enhancement vector; The advertisement matching module is used to match the user portrait enhancement vector, the personalization enhancement vector and the consumption behavior enhancement vector with preset candidate advertisements to determine the target push advertisement.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor; The memory is used to store computer programs; The processor is configured to execute the computer program and implement the advertisement pushing method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor implements the advertisement pushing method according to any one of claims 1 to 7.