Intelligent optimization of advertising placement method and system
Patent Information
- Application Number
- CN202611049578.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-15
- Publication Date
- 2026-09-15
AI Technical Summary
[0004]但是其在实际使用时,仍旧存在一些缺点,如依赖的用户行为数据仅限于历史点击和停留时长,忽略了用户滑动浏览的微观轨迹路径以及视线注视的热点分布,无法捕捉用户在屏幕上的注意力焦点与探索行为,导致对用户瞬时意图的刻画不够精准
Smart Images

Figure CN122760174A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of advertising delivery technology, and more specifically, to an intelligent optimization method and system for advertising delivery. Background Technology
[0002] With the rapid development of mobile internet and digital advertising, advertising delivery systems need to accurately identify potential audiences from a massive user base and maximize the effectiveness of advertising within a limited budget. However, users' actual intentions are often dynamic and highly contextualized. Relying solely on offline statistical data of historical behavior is insufficient to reflect the true attention and consumption tendencies at the current moment, thus limiting the accuracy and timeliness of advertising delivery.
[0003] Currently, common intelligent advertising methods typically combine user behavior sequence modeling with reinforcement learning. The specific implementation process is as follows: Collect users' historical click sequences and page dwell times within the information feed; extract user behavior feature vectors using recurrent neural networks or Transformer models; calculate the similarity between the ad's creative feature vector and the user behavior feature vector, selecting the ads with the highest similarity as a candidate set; then, based on a preset budget allocation strategy or simple bidding ranking rules, push the ads to the user's terminal; after the push, collect user click feedback to periodically update the user behavior model.
[0004] However, in practical use, it still has some shortcomings. For example, the user behavior data it relies on is limited to historical clicks and dwell time, ignoring the micro-trajectory path of user scrolling and the distribution of eye-focused hotspots. It cannot capture the user's attention focus and exploration behavior on the screen, resulting in an inaccurate portrayal of the user's instantaneous intent. Secondly, the ad push after similarity filtering uses a fixed ratio or simple sorting rules, failing to incorporate multi-dimensional constraints such as the remaining budget of the ad, real-time bid density, and the actual click-through rate of the previous time slot into a unified mathematical optimization framework. This can easily lead to problems such as low-quality ads occupying high-quality exposure resources and uneven budget consumption. In addition, the feedback signal is only used for offline batch updating of the model, failing to combine asynchronous incremental updates of the neural network with real-time correction of ad click-through rate statistics, thus limiting the timeliness of budget allocation and the response speed of closed-loop optimization. Summary of the Invention
[0005] In order to overcome the above-mentioned defects of the prior art, embodiments of the present invention provide an intelligent optimization advertising delivery method and system, which solves the problems mentioned in the background art through the following solutions.
[0006] To achieve the above objectives, the present invention provides the following technical solution: an intelligent optimization advertising delivery method, comprising S1: acquiring interactive behavior data and touch swipe trajectory sequences generated by user terminals, and collecting eye-tracking gaze heatmaps using a front-facing camera;
[0007] S2: Reconstruct the interactive behavior data into a behavior feature matrix, then input the behavior feature matrix into a convolutional neural network to extract the behavior feature vector; input the touch swipe trajectory sequence into the first convolutional neural network to extract the path feature vector; input the eye-tracking gaze heatmap into the second convolutional neural network to extract the gaze focus feature matrix; flatten the path feature vector and the gaze focus feature matrix and concatenate them, input them into a fully connected layer, and output the user's instantaneous intent vector;
[0008] S3: Read the available ads in the ad database, calculate the similarity between user features and ad features, and sort and filter them; preset a linear programming solver, input the current remaining budget value of the ad, the real-time bid density of the ad slot, and the actual click-through rate of the ad in the previous time slot, solve the total expected revenue and allocate the proportion, with the allocation proportion ranging from 0 to 1, and push the corresponding ad creative materials to the user terminal according to the allocation proportion.
[0009] S4: Within the first preset time threshold after the push is completed, collect user feedback signals on creative materials, concatenate the feedback signals with the user's instantaneous intent vector and creative feature vector to form a training sample, store it in the sample pool, and asynchronously trigger parameter updates of the first convolutional neural network, the second convolutional neural network and the fully connected layer. At the same time, adjust the actual click-through rate of the advertisement based on the feedback signals.
[0010] Preferably, each two-dimensional coordinate point of the touch sliding trajectory sequence is accompanied by a timestamp and a touch pressure value, and is divided into independent sliding segments with the user's hand-raising event as the dividing boundary. The coordinates of terminal screens of different resolutions are uniformly mapped to a standard grid coordinate system, and a corresponding trajectory density matrix is generated.
[0011] Preferably, the eye-tracking gaze heatmap is generated by mapping effective gaze points with a dwell time greater than a preset threshold from the gaze data collected by the front-facing camera to a standard grid, and the heat value of each grid is calculated by multiplying the gaze duration and gaze stability.
[0012] Preferably, the first convolutional neural network adopts a lightweight MobileNetV2 backbone structure, extracts intermediate feature maps from depth-separable convolutional layers, and outputs path feature vectors through global average pooling layers; the second convolutional neural network contains multiple sets of downsampled convolutional units, each set consisting of a convolutional layer, a batch normalization layer, and a ReLU activation layer, and outputs a gaze focus feature matrix.
[0013] Preferably, the interactive behavior data consists of historical click behavior sequences and page dwell time records, which are preprocessed and reconstructed into a behavior feature matrix. The reconstruction method is as follows: the behavior frequency sequence aggregated by hour is arranged according to "time dimension × behavior type dimension", the number of rows corresponds to continuous time windows, the number of columns corresponds to different types of user click behaviors, and the value of the matrix cell is the occurrence frequency of each type of behavior within the corresponding time window. The size of the behavior feature matrix is preset according to the number of time windows and the number of click behavior categories.
[0014] Preferably, the user features are obtained by concatenating the behavioral feature vector output by S2 with the user's instantaneous intent vector, and then mapped by a fully connected layer to an aligned feature vector with the same dimension as the ad feature vector; the cosine similarity between the aligned feature vector and the ad feature vectors of all ad placements is calculated, and the ad vectors are sorted in descending order of similarity to form a candidate ad set.
[0015] Preferably, the linear programming solver employs the simplex method, aiming to maximize the sum of the expected revenues of all candidate ads. The expected revenue for each ad is determined by the product of the ad's actual click-through rate in the previous time slice, the expected conversion value per click, and the allocation ratio. Constraints include:
[0016] Proportion constraint: The sum of the allocation proportions of all candidate ads is 1, and the allocation proportion of each ad is between 0 and 1.
[0017] Budget constraint: The estimated cost of a single ad within the current time frame will not exceed the smoothed rate of the current remaining budget;
[0018] Bid density constraint: Set bid percentile thresholds based on real-time bid density.
[0019] Preferably, the feedback signal includes whether a click event occurred, whether a preset conversion event was completed, and the duration of the ad page stay; the training sample is assembled by concatenating the user's instantaneous intent vector and the creative feature vector dimensions as the feature part, and using the encoded value of the feedback signal as the label part, wherein click or conversion is marked as a positive label, no interaction is marked as a negative label, and the duration of the page stay is used as the sample weight coefficient.
[0020] Preferably, the actual click-through rate is calculated using the exponential moving average method, with the smoothing coefficient set to a fixed value. It is obtained by comparing the observed click-through rate of the advertisement in the current time slice with the old actual click-through rate, where the weight of the old actual click-through rate is a fixed value of the smoothing coefficient. The corrected actual click-through rate is synchronously written back to the advertisement database for linear programming solution in the next time slice.
[0021] An intelligent optimized advertising delivery system includes a data acquisition module: acquiring a sequence of touch swipe trajectories, segmenting them into independent swipe segments based on hand-raising events, mapping them to a standard grid and generating a trajectory density matrix, calling the front-facing camera to acquire eye-tracking gaze heatmaps, mapping effective gaze points with a dwell time greater than a preset threshold to the same grid, and calculating the heat value of each grid by multiplying the gaze duration by the inverse of the pupil jitter amplitude, and normalizing it to a grayscale range;
[0022] The multimodal feature extraction module is configured to input the trajectory density matrix into the first 8 layers of a lightweight MobileNetV2 and output a path feature vector. It also inputs the eye-tracking gaze heatmap into the second convolutional neural network and outputs a gaze focus feature matrix. After flattening, the matrix is concatenated with the path feature vector and input into two fully connected layers to output a user instantaneous intent vector.
[0023] The ad recall and planning module is configured to concatenate the user's instantaneous intent vector and behavioral feature vector, map them through a fully connected layer, calculate the cosine similarity with all ad feature vectors, and then filter them. It uses a simplex linear programming solver to maximize the sum of the products of the actual click-through rate, the expected conversion value per click, and the allocation ratio of each ad. It is subject to ratio constraints, budget constraints, and bid density constraints, and outputs the allocation ratio.
[0024] Push and Feedback Update Module: Configured to push creative materials according to the allocated ratio. After the push, click events, conversion events and dwell time are collected as feedback signals. The user's instantaneous intent vector and creative feature vector are concatenated into features, feedback encoding is used as labels, and dwell time is used as weights. The concatenation is used as training samples and stored in the hot and cold partition sample pool. When the number of new samples in the hot zone reaches the preset threshold, the neural network parameters are asynchronously updated. Small batch stochastic gradient descent is used, and the actual click rate is corrected by the exponential moving average method. The data is written back to the database for the next time slice solution.
[0025] The technical effects and advantages of this invention are as follows:
[0026] 1. This invention constructs a multimodal user perception system by collecting touch swipe trajectory sequences from user terminals and eye-tracking gaze heatmaps from front-facing cameras, combined with interactive behavior data. It simultaneously captures the user's finger swipe path, screen gaze area, and long-term behavioral preferences. A first convolutional neural network and a second convolutional neural network are used to extract path feature vectors and gaze focus feature matrices, respectively. A fixed-dimensional instantaneous user intent vector is output through cascaded fully connected layers, significantly improving the accuracy of characterizing the user's current attention focus and true intent, providing a more reliable decision-making basis for subsequent ad recall and allocation.
[0027] 2. This invention introduces a linear programming solver after ad recall, using the remaining budget of the ad, real-time bid density, and actual click-through rate in the previous time slice as input constraints. With the goal of maximizing the total expected revenue of all candidate ads, it solves for the optimal allocation ratio of each ad, achieving mathematically optimal allocation under budget smoothing consumption and bid density constraints. This effectively avoids low-quality ads occupying high-quality exposure resources, improving the overall placement efficiency of ad slots and the return on investment for advertisers.
[0028] 3. This invention uses real-time feedback signals to concatenate the user's instantaneous intent vector, the ad creative feature vector, and feedback tags into training samples and stores them in a hot and cold partition sample pool. When the number of samples in the hot zone reaches a preset threshold, the parameters of the neural network are asynchronously updated. At the same time, the exponential moving average method is used to correct the actual click-through rate of the ad in real time and write it back to the database for solving the next time slice. This invention breaks through the limitations of offline batch updates and static CTR statistics in the prior art, realizes incremental optimization of model parameters and real-time response of allocation strategies, and greatly improves the adaptive ability of the ad delivery system to changes in user intent and fluctuations in market bidding. Attached Figure Description
[0029] Figure 1 This is a schematic diagram of the overall structure of the present invention;
[0030] Figure 2 This is a schematic diagram of the touch sliding trajectory sequence processing of the present invention;
[0031] Figure 3 This is a schematic diagram illustrating the generation of eye-tracking fixation heatmaps according to the present invention;
[0032] Figure 4 This is a schematic diagram illustrating the generation of the user's instantaneous intent vector according to the present invention;
[0033] Figure 5 This is a schematic diagram illustrating the linear programming solution for the advertising allocation ratio in this invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] like Figure 1 The method for intelligently optimizing ad delivery includes S1: acquiring interactive behavior data and touch swipe trajectory sequences generated by the user terminal, and collecting eye-tracking gaze heatmaps using the front-facing camera.
[0036] It should be noted that each two-dimensional coordinate point of the touch sliding trajectory sequence is accompanied by a timestamp and touch pressure value. It is divided into independent sliding segments with the user's hand lifting event as the dividing boundary, and the coordinates of terminal screens of different resolutions are uniformly mapped to the standard grid coordinate system, while generating the corresponding trajectory density matrix.
[0037] The eye-tracking gaze heatmap is generated by mapping effective gaze points with a dwell time greater than a preset threshold from the gaze data collected by the front-facing camera to a standard grid. The heatmap value of each grid is calculated by multiplying the gaze duration and gaze stability.
[0038] It should be further explained that S1 collects interaction behavior data through the built-in event tracking SDK in the terminal application. The data includes two types: historical click behavior sequences and page dwell time records.
[0039] Historical click behavior sequence: Collect all click events of users in the past 7 days. Each event is arranged in chronological order and includes three attributes: event timestamp, click object category, and page ID. In the preprocessing stage, invalid data with duplicate reports and abnormal timestamps are removed, and the data is aggregated into a structured behavior frequency sequence with an hourly granularity.
[0040] Page dwell time recording: Dwell time is split and counted by page functional blocks (content area, recommendation area, and advertising area) with a time precision of 100ms; in the preprocessing stage, abnormally long dwell times exceeding 300 seconds are truncated, and invalid durations generated by application background suspension are filtered out, and the final output is a dataset of dwell time of each block.
[0041] like Figure 2 As shown, raw touch data is obtained from the terminal touch driver layer at a sampling frequency of 60Hz. The raw data consists of two-dimensional coordinate points arranged in chronological order, with each coordinate point accompanied by a timestamp and touch pressure value.
[0042] The preprocessing first removes abrupt anomalies whose coordinates exceed the screen boundary and whose sampling interval is greater than 200ms. Then, using the user's hand-raising event as the segmentation boundary, continuous touch is split into independent sliding segments. Finally, the screen coordinates of terminals with different resolutions are uniformly mapped to a standard 224×224 grid coordinate system, and the standardized two-dimensional coordinate point sequence arranged in chronological order is retained. The corresponding trajectory density matrix is generated synchronously as a two-dimensional representation of the sequence to adapt to the input requirements of convolutional neural networks.
[0043] To further explain, the core of synchronously generating the corresponding trajectory density matrix as a two-dimensional representation of the sequence is to transform the one-dimensional coordinate point sequence into a two-dimensional matrix form that can be directly processed by the convolutional neural network. The specific logic is as follows:
[0044] First, divide the standard screen into a uniform grid of 224×224, with each grid corresponding to a cell in the matrix; count the number of times the touch trajectory passes through each grid / the cumulative dwell time, and fill the values into the corresponding positions in the matrix. Grids with no trajectory passing through are recorded as 0; the final two-dimensional numerical matrix is the trajectory density matrix.
[0045] like Figure 3 As shown, the system utilizes the front-facing camera of the terminal to collect user gaze data using facial landmark detection and pupil pose calculation algorithms, with a sampling frequency of 30Hz. The processing flow is as follows: First, invalid gaze points during blinking and rapid scanning are filtered out, retaining only valid gaze points with a dwell time ≥100ms. Then, all valid gaze points are mapped to a standard 224×224 grid. The heatmap value of each grid is calculated by weighting "gaze duration × gaze stability," where gaze stability is the reciprocal of pupil jitter amplitude; the smoother the gaze (smaller jitter amplitude), the higher the stability. Finally, the heatmap values are linearly normalized to the 0-255 grayscale range, generating a grayscale heatmap matrix, i.e., the eye-tracking gaze heatmap.
[0046] The data uniformly adopts the terminal system clock as the time reference, aligns the timestamps of interactive behavior data, touch swipe trajectory sequences, and eye-tracking gaze heatmaps within the same time window, and controls the time error within 10ms to ensure the spatiotemporal correspondence of multimodal data.
[0047] S2: Reconstruct the interactive behavior data into a behavior feature matrix, then input the behavior feature matrix into a convolutional neural network to extract the behavior feature vector; input the touch swipe trajectory sequence into the first convolutional neural network to extract the path feature vector; input the eye-tracking gaze heatmap into the second convolutional neural network to extract the gaze focus feature matrix; flatten the path feature vector and the gaze focus feature matrix and concatenate them, input them into a fully connected layer, and output the user's instantaneous intent vector.
[0048] like Figure 4 As shown, it should be specifically noted that the first convolutional neural network adopts a lightweight MobileNetV2 backbone structure, extracts intermediate feature maps from depthwise separable convolutional layers, and outputs path feature vectors through global average pooling layers; the second convolutional neural network contains multiple sets of downsampled convolutional units, each set consisting of a convolutional layer, a batch normalization layer, and a ReLU activation layer, and outputs a gaze focus feature matrix.
[0049] The interactive behavior data consists of historical click behavior sequences and page dwell time records. After preprocessing, it is reconstructed into a behavior feature matrix. The reconstruction method is as follows: the behavior frequency sequence aggregated by hour is arranged according to "time dimension × behavior type dimension". The number of rows corresponds to continuous time windows, the number of columns corresponds to different types of user click behaviors, and the value of the matrix cell is the occurrence frequency of each type of behavior within the corresponding time window. The size of the behavior feature matrix is preset according to the number of time windows and the number of click behavior categories.
[0050] It should be further explained that S2 reconstructs the interaction behavior data preprocessed by S1 into a 32×32 behavior feature matrix, which is then input into a convolutional neural network. The convolutional neural network contains two sets of convolutional units (each set consists of a 3×3 convolutional layer, a batch normalization layer, a ReLU activation layer, and a 2×2 max pooling layer). Finally, it outputs a 64-dimensional behavior feature vector through a global average pooling layer, which is used to characterize the user's long-term content preferences and category tendencies.
[0051] To further explain, reconstructing the 32×32 behavioral feature matrix involves converting one-dimensional structured behavioral data into a two-dimensional matrix format that can be processed by convolutional neural networks. Specifically, the hourly aggregated behavioral frequency sequence is arranged in two dimensions: "time dimension × behavior type dimension". The 32 rows correspond to 32 consecutive one-hour time windows, and the 32 columns correspond to 32 subcategories of user click behaviors. The value of each cell in the matrix represents the frequency of occurrence of each type of behavior within the corresponding time window. This transforms discrete behavioral statistics into a 32×32 class image matrix, and the convolutional network is used to extract the correlation features between time and behavior category.
[0052] To further explain, after two sets of convolutional units, the network outputs 64 two-dimensional feature maps (64 channels). Each feature map corresponds to a class of low-level behavioral features. Global average pooling averages all values in each feature map, and each feature map finally yields a value representing the overall strength of the feature. The 64 feature maps correspond to 64 output values, which, when concatenated in order, form a 64-dimensional behavioral feature vector.
[0053] The two-dimensional trajectory density matrix corresponding to the touch swipe trajectory sequence processed by S1 is input into the first convolutional neural network. The first convolutional neural network adopts a lightweight MobileNetV2 backbone structure, extracts the first 8 depth-separable convolutional layers for feature encoding, outputs a 7×7×128 intermediate feature map, and then outputs a 128-dimensional path feature vector through a global average pooling layer, which is used to characterize the rhythm, exploration range and attention movement trend of the user's swipe browsing.
[0054] The eye-tracking gaze heatmap generated by S1 is input into the second convolutional neural network, which contains multiple groups (e.g., 3 groups) of downsampled convolutional units (each group consists of a 3×3 convolutional layer, a batch normalization layer, and a ReLU activation layer), and outputs a 14×14×64 gaze focus feature matrix, which retains the spatial distribution information of the user's gaze points, corresponding to the screen area location that the user is focusing on.
[0055] The gaze focus feature matrix is flattened into a one-dimensional vector in row-major order, with a dimension of 14×14×64=12544. This vector is then concatenated with the 128-dimensional path feature vector in dimensional order to obtain a 12672-dimensional fused feature. The fused feature is then input into a fully connected layer, which consists of two layers: the first layer contains 2048 neurons, and the second layer contains 128 neurons. Both layers are configured with Dropout=0.3 to suppress overfitting. The final output is a 128-dimensional user instantaneous intent vector, with each dimension corresponding to a fine-grained real-time intent.
[0056] S3: Read the available ads in the ad database, calculate the similarity between user features and ad features, and sort and filter them; preset a linear programming solver, input the current remaining budget value of the ad, the real-time bid density of the ad slot, and the actual click-through rate of the ad in the previous time slot, solve for the total expected revenue and allocate the proportion, with the allocation proportion ranging from 0 to 1, and push the corresponding ad creative materials to the user terminal according to the allocation proportion.
[0057] like Figure 5 As shown, it should be specifically explained that the user features are obtained by concatenating the behavioral feature vector output by S2 with the user's instantaneous intent vector, and then mapped by a fully connected layer to an aligned feature vector with the same dimension as the ad feature vector; the cosine similarity between the aligned feature vector and the ad feature vectors of all ad placements is calculated, and the ad vectors are sorted in descending order of similarity to form a candidate ad set.
[0058] The linear programming solver employs the simplex method, aiming to maximize the sum of expected revenues from all candidate ads. The expected revenue for each ad is determined by the product of its actual click-through rate (CTR) in the previous time slice, its expected conversion value per click, and its allocation ratio. Constraints include:
[0059] Proportion constraint: The sum of the allocation proportions of all candidate ads is 1, and the allocation proportion of each ad is between 0 and 1.
[0060] Budget constraint: The estimated cost of a single ad within the current time frame will not exceed the smoothed rate of the current remaining budget;
[0061] Bid density constraint: Set bid percentile thresholds based on real-time bid density.
[0062] It should be further explained that S3 reads all ads in the delivery state from the advertising database. Each ad pre-generates a 128-dimensional ad feature vector (including four types of information: category attributes, creative tags, target audience, and conversion value) and a corresponding 128-dimensional creative feature vector of creative materials, which are stored in the advertising database.
[0063] User features are obtained by concatenating the 64-dimensional behavioral feature vector output by S2 and the 128-dimensional instantaneous user intent vector according to their dimensions, for a total of 192 dimensions, which simultaneously covers both long-term user preferences and real-time intent.
[0064] Inputting the 192-dimensional user features into a fully connected layer (without activation function, output dimension 128) yields an aligned feature vector, making the dimension consistent with the advertising feature vector (128 dimensions).
[0065] The cosine similarity between the aligned feature vector and the feature vectors of all available ads is calculated. The similarity value represents the degree of matching between the user and the corresponding ad. The higher the value, the stronger the matching. The ads are sorted in descending order of similarity value, and the top 50 ads are selected to enter the candidate pool. This greatly reduces the computational load of the subsequent linear programming. The weight parameters of the fully connected layer are optimized together with the asynchronous neural network parameters in S4.
[0066] The server comes pre-configured with a linear programming solver based on the simplex method. The core input parameters and definitions of the solver are as follows:
[0067] Budget parameter: The current remaining budget value for each candidate ad (the total budget set for the ad minus the cumulative amount spent up to the current moment);
[0068] Bidding parameters: The real-time bid density of the current ad slot, that is, the bid distribution data of all bidding ads in the current ad slot;
[0069] Performance parameters: The actual click-through rate of each ad within the previous 5-minute time slot.
[0070] Construct a linear programming model with the objective function of maximizing the total expected return:
[0071]
[0072] in, For the first The actual click-through rate of the ad in a given time frame. For the first The expected conversion value per click of an ad (historical clicks of the ad - conversion rate × preset value per conversion). For the first The allocation ratio of each advertisement, This represents the total number of candidate ads.
[0073] Set three types of constraints:
[0074] Proportional constraint: The sum of the allocation proportions of all candidate ads is 1, i.e. And the range of values for the allocation ratio of a single advertisement is: ;
[0075] Budget constraint: The estimated cost of a single ad within the current time frame shall not exceed the smoothed rate of the current remaining budget, to avoid sudden increases or decreases in the budget;
[0076] Bid density constraint: Based on real-time bid density, set bid percentile thresholds. Ads with bids below the 30th percentile have their allocation ratio capped at 0.05 to prevent low-quality ads from occupying high-quality exposure resources.
[0077] The solver outputs the optimal allocation ratio for each candidate ad (within the range of 0 to 1), completing the allocation calculation to maximize the total expected revenue.
[0078] Based on the calculated allocation ratio, the corresponding creative materials for the advertisements are pushed to the user's terminal: the first advertisement exposure position on the current page is allocated to the advertisement with the highest allocation ratio, and the exposure positions generated by the user's subsequent swipe refresh are randomly selected from the corresponding advertisements according to the allocation ratio for delivery.
[0079] S4: Within the first preset time threshold after the push is completed, collect user feedback signals on creative materials, concatenate the feedback signals with the user's instantaneous intent vector and creative feature vector to form a training sample, store it in the sample pool, and asynchronously trigger parameter updates of the first convolutional neural network, the second convolutional neural network and the fully connected layer. At the same time, adjust the actual click-through rate of the advertisement based on the feedback signals.
[0080] It should be specifically noted that the feedback signal includes whether a click event occurred, whether a preset conversion event was completed, and the duration of the ad page stay; the training sample is assembled as follows: the user's instantaneous intent vector and the creative feature vector are concatenated as the feature part, and the encoded value of the feedback signal is used as the label part, wherein click or conversion is marked as a positive label, no interaction is marked as a negative label, and the duration of the page stay is used as the sample weight coefficient.
[0081] The actual click-through rate (CTR) is calculated using the exponential moving average method, with the smoothing coefficient set to a fixed value. It is obtained by comparing the observed CTR of the ads in the current time slice with the old actual CTR, where the weight of the old actual CTR is a fixed value of the smoothing coefficient. The corrected actual CTR is synchronously written back to the ad database for use in the linear programming solution of the next time slice.
[0082] It should be further explained that the S4 sets the first preset duration threshold to 30 seconds. Testing has shown that this 30-second threshold represents the optimal balance between three factors: the real-time interactive behavior patterns of mobile users' information feeds, the accuracy of the correspondence between feedback signals and users' instantaneous intentions, and the completeness of feedback collection and the efficiency of closed-loop optimization. Timing begins when the ad creative is pushed to the user's device, and user feedback signals are continuously collected within the threshold duration. The core feedback signals include three categories: whether a click event occurred, whether a preset conversion event was completed, and the duration of time spent on the ad page. Additional data such as swipe amplitude and eye-tracking fixation time are only used as a basis for sample weight calculation and are not included in the core feedback signal category.
[0083] The user's instantaneous intent vector, ad creative feature vector, and feedback signal corresponding to this campaign are encoded and concatenated into a training sample: the feature part is the dimensional concatenation of the user's instantaneous intent vector and the ad creative feature vector, the label part is the encoded value of the feedback signal, click / conversion is marked as a positive label, no interaction is marked as a negative label, and page dwell time is used as the sample weight coefficient. The sample is stored in a distributed sample pool, which is divided into hot and cold zones: samples added within 24 hours are stored in the hot zone, and those older than 24 hours are moved to the cold zone. Model training prioritizes using samples from the hot zone to ensure a fast response to changes in user intent.
[0084] When the cumulative number of newly added samples in the hot zone of the sample pool reaches 128, an asynchronous parameter update task is triggered. The convolutional neural network used for behavioral feature extraction, which represents long-term user preferences, adopts an independent rhythm of daily full batch updates and is not included in the S4 asynchronous update scope. The parameter update adopts the mini-batch stochastic gradient descent algorithm, and the loss function is the weighted sum of cross-entropy loss (click / conversion classification task) and mean squared error loss (dwell time regression task). The update process is implemented through a parameter server architecture. The online inference service loads the old parameters to provide delivery services throughout the process, and the new parameters are switched atomically after training is completed without interrupting the online delivery chain.
[0085] To further explain, the mini-batch stochastic gradient descent algorithm takes "user multimodal fusion features + corresponding ad features" as input for each training sample, along with two sets of true labels: a classification label (whether a click / conversion occurred, 1 for occurrence and 0 for non-occurrence) and a regression label (the user's actual dwell time on the ad). Each update randomly selects a fixed number of samples (e.g., 32 / 64) from the feedback sample pool to form a mini-batch, which serves as the input for this training iteration. The features of the samples within the batch are input into the neural network. After layer-by-layer computation, the output layer outputs two sets of predicted values: the classification branch outputs a probability value of 0-1, representing the predicted probability of a user clicking / converting; the regression branch outputs a continuous value, representing the estimated user dwell time.
[0086] To further explain, the loss function calculates the deviation between the two tasks separately, and then combines them into a total loss according to preset weights:
[0087] The cross-entropy formula for binary classification is used to calculate the deviation between the actual click / conversion tag and the predicted probability, and the cross-entropy loss value is obtained.
[0088] The mean squared error loss value is obtained by averaging the squares of the differences between the actual stay duration and the predicted stay duration.
[0089] Total loss = classification weight × cross-entropy loss + regression weight × mean squared error loss. The weights are preset according to the business focus (e.g., classification weight 0.7, regression weight 0.3).
[0090] Using the total loss as a benchmark, the gradient value corresponding to each weight and bias parameter in the network is obtained by calculating backwards from the output layer to the input layer using the chain rule, thus clarifying the adjustment direction and magnitude of each parameter.
[0091] Following the mini-batch stochastic gradient descent rule, all trainable parameters of the network are updated synchronously with the formula: new parameters = original parameters - learning rate × batch average gradient. The learning rate is a preset step size coefficient used to control the magnitude of a single update. After completing this batch update, the above process is repeated after the next batch of samples is accumulated, thereby achieving asynchronous incremental optimization of the model.
[0092] Based on the feedback signals collected, the actual click-through rate of the advertisements is adjusted in real time using the exponential moving average method. The specific formula is as follows:
[0093]
[0094] smoothness coefficient Take 0.9, This represents the actual click-through rate for the current campaign period. This represents the historical actual click-through rate from the previous campaign period. This represents the observed click-through rate (clicks / impressions) for ads within the current time slot. The corrected actual click-through rate is synchronously written back to the ad database and directly used for linear programming solutions in the next time slot, replacing the traditional T+1 statistical method and improving the timeliness and accuracy of budget allocation.
[0095] An intelligent optimized advertising delivery system includes a data acquisition module: it collects touch swipe trajectory sequences, segments them into independent swipe segments based on hand-raising events, maps them to a standard grid and generates a trajectory density matrix, calls the front-facing camera to collect eye-tracking gaze heatmaps, maps effective gaze points with a dwell time greater than a preset threshold to the same grid, and calculates the heat value of each grid by multiplying the gaze duration by the inverse of the pupil jitter amplitude, and normalizes it to the grayscale range.
[0096] Multimodal feature extraction module: It is configured to input the trajectory density matrix into the first 8 layers of the lightweight MobileNetV2 and output the path feature vector. It inputs the eye-tracking gaze heatmap into the second convolutional neural network and outputs the gaze focus feature matrix. After flattening, it is concatenated with the path feature vector and input into two fully connected layers to output the user's instantaneous intent vector.
[0097] The ad recall and planning module is configured to concatenate the user's instantaneous intent vector and behavioral feature vector, map them through a fully connected layer, calculate the cosine similarity with all ad feature vectors, and then filter them. It uses a simplex linear programming solver to maximize the sum of the products of the actual click-through rate, the expected conversion value per click, and the allocation ratio for each ad. It is subject to ratio constraints, budget constraints, and bid density constraints, and outputs the allocation ratio.
[0098] Push and Feedback Update Module: Configured to push creative materials according to the allocated ratio. After the push, click events, conversion events and dwell time are collected as feedback signals. The user's instantaneous intent vector and creative feature vector are concatenated into features, feedback encoding is used as labels, and dwell time is used as weights. The concatenation is used as training samples and stored in the hot and cold partition sample pool. When the number of new samples in the hot zone reaches the preset threshold, the neural network parameters are asynchronously updated. Small batch stochastic gradient descent is used, and the actual click rate is corrected by the exponential moving average method. The data is written back to the database for the next time slice solution.
[0099] Secondly: The accompanying drawings of the embodiments disclosed in this invention only involve the structures involved in the embodiments disclosed in this invention. Other structures can refer to the general design. In the absence of conflict, the same embodiment and different embodiments of this invention can be combined with each other.
[0100] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for intelligent optimization of advertisement placement, characterized in that, include: S1: Acquire interactive behavior data and touch swipe trajectory sequences generated by the user terminal, and collect eye-tracking gaze heatmaps using the front-facing camera; S2: Reconstruct the interactive behavior data into a behavior feature matrix, then input the behavior feature matrix into a convolutional neural network to extract the behavior feature vector, input the touch swipe trajectory sequence into the first convolutional neural network to extract the path feature vector, input the eye-tracking gaze heatmap into the second convolutional neural network to extract the gaze focus feature matrix, flatten the path feature vector and gaze focus feature matrix and then concatenate them, input them into a fully connected layer, and output the user's instantaneous intent vector. S3: Read the available ads from the ad database, calculate the similarity between user characteristics and ad characteristics, and sort and filter them; The system uses a pre-defined linear programming solver. It takes as input the current remaining budget value of the ad, the real-time bid density of the ad slot, and the actual click-through rate of the ad in the previous time slot. It calculates the total expected revenue and allocates the proportion. The allocation proportion is between 0 and 1. The corresponding creative materials of the ad are pushed to the user's terminal according to the allocation proportion. S4: Within the first preset time threshold after the push is completed, collect user feedback signals on creative materials, concatenate the feedback signals with the user's instantaneous intent vector and creative feature vector to form a training sample, store it in the sample pool, and asynchronously trigger parameter updates of the first convolutional neural network, the second convolutional neural network and the fully connected layer. At the same time, adjust the actual click-through rate of the advertisement based on the feedback signals.
2. The intelligent optimized advertising delivery method according to claim 1, characterized in that: Each two-dimensional coordinate point of the touch sliding trajectory sequence is accompanied by a timestamp and touch pressure value. It is divided into independent sliding segments with the user's hand-raising event as the dividing boundary, and the coordinates of terminal screens of different resolutions are uniformly mapped to the standard grid coordinate system, while generating the corresponding trajectory density matrix.
3. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The eye-tracking gaze heatmap is generated by mapping effective gaze points with a dwell time greater than a preset threshold from the gaze data collected by the front-facing camera to a standard grid. The heatmap value of each grid is calculated by multiplying the gaze duration and gaze stability.
4. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The first convolutional neural network adopts a lightweight MobileNetV2 backbone structure, extracts intermediate feature maps from depthwise separable convolutional layers, and outputs path feature vectors through global average pooling layers; the second convolutional neural network contains multiple sets of downsampled convolutional units, each set consisting of a convolutional layer, a batch normalization layer, and a ReLU activation layer, and outputs a gaze focus feature matrix.
5. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The interactive behavior data consists of historical click behavior sequences and page dwell time records. After preprocessing, it is reconstructed into a behavior feature matrix. The reconstruction method is as follows: the behavior frequency sequence aggregated by hour is arranged according to "time dimension × behavior type dimension". The number of rows corresponds to continuous time windows, the number of columns corresponds to different types of user click behaviors, and the value of the matrix cell is the occurrence frequency of each type of behavior within the corresponding time window. The size of the behavior feature matrix is preset according to the number of time windows and the number of click behavior categories.
6. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The user features are obtained by concatenating the behavioral feature vector output by S2 with the user's instantaneous intent vector, and then mapping them through a fully connected layer to an aligned feature vector with the same dimension as the ad feature vector. The cosine similarity between the aligned feature vector and the ad feature vectors of all available ads is calculated, and the ad vectors are sorted in descending order of similarity to form a candidate ad set.
7. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The linear programming solver employs the simplex method, aiming to maximize the sum of expected revenues from all candidate ads. The expected revenue for each ad is determined by the product of its actual click-through rate (CTR) in the previous time slice, its expected conversion value per click, and its allocation ratio. Constraints include: Proportion constraint: The sum of the allocation proportions of all candidate ads is 1, and the allocation proportion of each ad is between 0 and 1. Budget constraint: The estimated cost of a single ad within the current time frame will not exceed the smoothed rate of the current remaining budget; Bid density constraint: Set bid percentile thresholds based on real-time bid density.
8. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The feedback signals include whether a click event has occurred, whether a preset conversion event has been completed, and the duration of the ad page stay. The training samples are concatenated as follows: the user's instantaneous intent vector and the creative feature vector are concatenated as the feature part, and the encoded value of the feedback signal is used as the label part, where clicks or conversions are marked as positive labels, no interaction is marked as a negative label, and the page dwell time is used as the sample weight coefficient.
9. The intelligent optimized advertising delivery method according to claim 1, characterized in that: The actual click-through rate (CTR) is calculated using the exponential moving average method, with the smoothing coefficient set to a fixed value. It is obtained by comparing the observed CTR of the ads in the current time slice with the old actual CTR, where the weight of the old actual CTR is a fixed value of the smoothing coefficient. The corrected actual CTR is synchronously written back to the ad database for use in the linear programming solution of the next time slice.
10. An intelligent optimized advertising delivery system, used to implement the intelligent optimized advertising delivery method according to any one of claims 1-9, characterized in that, include: Data acquisition module: Collects touch swipe trajectory sequences, segments them into independent swipe segments based on hand-raising events, maps them to a standard grid and generates a trajectory density matrix, calls the front camera to collect eye-tracking gaze heatmaps, maps effective gaze points with a dwell time greater than a preset threshold to the same grid, and calculates the heat value of each grid by multiplying the gaze duration by the inverse of the pupil jitter amplitude, and normalizes it to the grayscale range; The multimodal feature extraction module is configured to input the trajectory density matrix into the first 8 layers of a lightweight MobileNetV2 and output a path feature vector. It also inputs the eye-tracking gaze heatmap into the second convolutional neural network and outputs a gaze focus feature matrix. After flattening, the matrix is concatenated with the path feature vector and input into two fully connected layers to output a user instantaneous intent vector. The ad recall and planning module is configured to concatenate the user's instantaneous intent vector and behavioral feature vector, map them through a fully connected layer, calculate the cosine similarity with all ad feature vectors, and then filter them. It uses a simplex linear programming solver to maximize the sum of the products of the actual click-through rate, the expected conversion value per click, and the allocation ratio of each ad. It is subject to ratio constraints, budget constraints, and bid density constraints, and outputs the allocation ratio. Push and Feedback Update Module: Configured to push creative materials according to the allocated ratio. After the push, click events, conversion events and dwell time are collected as feedback signals. The user's instantaneous intent vector and creative feature vector are concatenated into features, feedback encoding is used as labels, and dwell time is used as weights. The concatenation is used as training samples and stored in the hot and cold partition sample pool. When the number of new samples in the hot zone reaches the preset threshold, the neural network parameters are asynchronously updated. Small batch stochastic gradient descent is used, and the actual click rate is corrected by the exponential moving average method. The data is written back to the database for the next time slice solution.