Advertisement placement optimization system and method for product promotion

Through space-time alignment of multi-source interactive data, cross-modal attention fusion and deep reinforcement learning, the advertising delivery strategy is optimized, and the problem that traditional advertising systems cannot adapt to changes in user interests and market environments is solved, and the advertising effect is improved and resource optimization is achieved.

CN120013607BActive Publication Date: 2025-08-12GUANGZHOU XINGCAN CULTURE MEDIA CO LTD

Patent Information

Application Number
CN202510490945.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-08-12
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Traditional advertising systems cannot accurately capture the dynamic changes in user interests, resulting in poor advertising delivery, lack of flexibility and dynamic adjustment capabilities, and cannot adapt to changes in the market environment.

Method used

Through space-time alignment of multi-source interactive data, cross-modal attention fusion, Hilbert spatial mapping and deep reinforcement learning, a user-advertising position benefit matrix is built, advertiser budget allocation and delivery strategies are optimized, indirect competition effect coefficients are introduced, and dynamic adjustment is achieved.

Benefits of technology

Accurately match user needs, improve ad click-through rate and conversion rate, optimize advertiser budget allocation, avoid resource waste, continuously adapt to market changes, and improve advertising delivery results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013607B_ABST
    Figure CN120013607B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of advertising delivery technology, specifically an advertising delivery optimization system and method for product promotion, comprising: a data alignment unit, the data alignment unit is used to obtain multi-source interaction data of users, wherein the multi-source interaction data includes click stream data, transaction data and social behavior data, and performs heterogeneous data spatiotemporal alignment on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data; a feature update unit, the feature update unit is used to perform cross-modal attention fusion on the aligned data to obtain attention features. The present invention performs spatiotemporal alignment on multi-source interaction data and utilizes cross-modal attention fusion technology to update user features, which can accurately capture the user's multi-dimensional interest changes and behavior patterns, thereby accurately matching user needs during the advertising delivery process and improving advertising effectiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of advertising placement, and in particular to an advertising placement optimization system and method for commodity promotion. Background Art

[0002] Traditional advertising systems often rely on static user profiles and a single data source (such as browsing history or purchasing behavior) for ad matching. This approach cannot accurately capture the dynamic changes and multi-dimensional characteristics of user interests, resulting in an inability to adapt to changes in user interests during ad delivery, which in turn affects the click-through rate and conversion rate of ads. Traditional systems also typically use simple tags or historical behavior-based data for interest matching. They lack in-depth feature mapping methods and do not utilize advanced technology to perform refined mathematical modeling of user interests. Therefore, the probability of mismatch between ads and user interests is high. Traditional systems are generally based on simple budget allocation strategies, often allocating budgets according to preset rules, and lack a real-time dynamic adjustment mechanism to optimize advertiser budgets. This approach can easily lead to resource waste and suboptimal advertising effectiveness, especially in rapidly changing market environments. The budget allocation of traditional systems may not be able to adapt to rapidly changing market demands. Traditional advertising systems also lack sufficient flexibility and dynamic adjustment capabilities. Advertising delivery strategies are generally difficult to adjust in a timely manner after they are initially set, and cannot be optimized based on real-time market and user feedback. As a result, advertising effectiveness may decline as the market environment changes. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to overcome the shortcomings of the above-mentioned prior art and provide an advertising delivery optimization system and method for product promotion.

[0004] The technical solution adopted to solve the above technical problems is: an advertising delivery optimization system for product promotion, comprising:

[0005] A data alignment unit, configured to obtain multi-source interaction data of a user, wherein the multi-source interaction data includes clickstream data, transaction data, and social behavior data, and perform spatiotemporal heterogeneous data alignment on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data;

[0006] a feature updating unit, configured to perform cross-modal attention fusion on the aligned data to obtain attention features, and perform feature update on the user using the attention features according to a dual-gating mechanism to update user features;

[0007] a benefit modeling unit configured to perform Hilbert space mapping on the user features to obtain orthogonal interest basis vectors, obtain ad slot feature vectors and environment feature vectors, and construct a user-ad slot benefit matrix based on the orthogonal interest basis vectors, the environment feature vectors, and the ad slot feature vectors;

[0008] A competition modeling unit is used to define an advertiser strategy space, wherein the advertiser strategy space includes ad slot budget allocation, creative aggressiveness coefficient and time period preference phase, and calculate the indirect competition effect coefficient based on the advertiser strategy space and the user-ad slot benefit matrix.

[0009] Preferably, the system further comprises:

[0010] a budget allocation unit, configured to construct a budget allocation objective function based on the indirect competition effect coefficient and the user-advertising space benefit matrix, and solve the budget allocation objective function according to a backward induction method to obtain an optimal budget allocation matrix for the advertiser;

[0011] An advertising delivery unit is used to obtain the spatiotemporal feature vector during advertising delivery, define a Q-function approximator based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize the advertising delivery strategy by maximizing the Q-function approximator based on deep reinforcement learning to obtain the best bidding strategy.

[0012] Preferably, performing spatiotemporal alignment of heterogeneous data on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data includes:

[0013] Performing cubic spline interpolation on the transaction data to obtain an interpolation result of the transaction data, wherein the calculation formula of the interpolation result of the transaction data is as follows:

[0014] ;

[0015] in, Indicates time The interpolation result of the transaction data at Indicates the total number of transaction data, represents the basis function based on Lagrange interpolation, and , Indicates time Number of transactions at

[0016] A social behavior heat decay coefficient is defined to perform heat decay processing on the social behavior data to obtain a heat decay result of the social behavior data. The calculation formula of the heat decay result of the social behavior data is as follows:

[0017] ;

[0018] in, Indicates time The heat decay results of social behavior data at Indicates time Social behavior data, represents the social behavior heat decay coefficient, and , Represents the seconds of the minute;

[0019] The interpolation results of the transaction data and the heat decay results of the social behavior data are aligned to a unified time axis according to the clickstream data to obtain aligned data corresponding to the multi-source interaction data. The calculation formula of the aligned data is as follows:

[0020] ;

[0021] in, Indicates alignment data, Indicates time Clickstream data from .

[0022] Preferably, cross-modal attention fusion is performed on the aligned data to obtain attention features, including:

[0023] The attention head of each modality in the aligned data is calculated to obtain the attention corresponding to each modality in the aligned data, wherein the calculation formula for the attention corresponding to each modality in the aligned data is as follows:

[0024] ;

[0025] in, Indicates the alignment data The attention corresponding to each modality, 、 and represents the query matrix, key matrix and value matrix, 、 and Indicates the alignment data The projection matrix corresponding to each mode is, Represents the dimension of the feature, represents the learnable modality mask matrix;

[0026] The attention corresponding to each modality in the aligned data is concatenated and linearly transformed to obtain the attention feature, where the calculation formula of the attention feature is as follows:

[0027] ;

[0028] in, Represents the attention feature, Represents a splicing operation, Represents a linear transformation matrix.

[0029] Preferably, updating the user's features by using the attention features according to the dual-gating mechanism to obtain user features includes:

[0030] The importance score is performed by combining the attention feature with the historical user features according to the first gating mechanism to obtain an importance score, wherein the calculation formula of the importance score is as follows:

[0031] ;

[0032] in, represents the importance score, represents the activation function, represents the first trainable weight matrix, represents the user features of the previous time step, Represents the attention feature of the current time step;

[0033] According to the second gating mechanism, the user's features are fused using the importance score to obtain fusion information, wherein the calculation formula of the fusion information is as follows:

[0034] ;

[0035] in, represents fusion information, represents the second trainable weight matrix;

[0036] The user characteristics are updated according to the importance score and the fusion information to update the user characteristics, wherein the updating formula of the user characteristics is as follows:

[0037] ;

[0038] in, represents the updated user features, represents the hyperbolic tangent function, represents the third trainable weight matrix;

[0039] When the importance score is greater than a preset importance score threshold, the updated user feature corresponding to the importance score is stored as an event to obtain a memory library, wherein the expression of the memory library is as follows:

[0040] ;

[0041] in, represents the updated user feature sequence stored in the memory bank, Indicates the size of the memory bank.

[0042] Preferably, the calculation formula of the user-advertising space benefit matrix is as follows:

[0043] ;

[0044] in, Represents a user and ad space The benefits between represents the orthogonal interest basis vector, represents the ad slot feature vector, represents the environmental feature vector, Represents the time decay factor, which is used to control the timeliness of the ad space. Indicates ad space Timeliness, Indicates the historical click-through rate mixing coefficient, balancing the effects of new and old ad positions. Indicates the historical click-through rate of the ad slot.

[0045] Preferably, the calculation formula of the indirect competition effect coefficient is as follows:

[0046] ;

[0047] in, Indicates advertiser With advertisers The indirect interference strength between Indicates advertiser In the advertising space Budget allocation on Represents a user and ad space The benefits between Represents the differences in advertiser strategy spaces;

[0048] The budget allocation objective function is as follows:

[0049] ;

[0050] in, Indicates ad space The benefits, represents the budget competition factor, represents the regularization coefficient, represents the indirect interference coefficient, Indicates advertiser total budget.

[0051] Preferably, the budget allocation objective function is solved according to the backward induction method to obtain the advertiser's optimal budget allocation matrix, including:

[0052] Initialize the budget allocation strategy matrix for each advertiser;

[0053] Iteratively update the budget allocation strategy of each advertiser, wherein the update formula of the advertiser's budget allocation strategy is as follows:

[0054] ;

[0055] in, Indicates at time Advertisers budget allocation strategy, Indicates that all advertisers at the time The sum of budget allocation strategies;

[0056] Define a convergence condition to control the accuracy of the budget allocation strategy update, where the convergence condition is as follows:

[0057] ;

[0058] in, Indicates at time Budget allocation strategy matrix, Indicates at time Budget allocation strategy matrix, Indicates the preset accuracy threshold;

[0059] When the update of the advertiser's budget allocation strategy satisfies a preset convergence condition, the iterative update of the advertiser's budget allocation strategy is stopped to obtain the advertiser's optimal budget allocation moment.

[0060] Preferably, a Q-function approximator is defined based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and an advertising delivery strategy is optimized by maximizing the Q-function approximator based on deep reinforcement learning to obtain an optimal bidding strategy, including:

[0061] A Q-function approximator is defined according to the advertiser's optimal budget allocation matrix and the spatiotemporal eigenvector, wherein the Q-function approximator is as follows:

[0062] ;

[0063] in, represents the Q-function myostat, represents the weight parameter of the Q function, and represents the network weight, Represents the spatiotemporal feature vector, which is obtained by extracting the geographic location information, time information and device information through MLP. Indicates the ad slot and bid embedding corresponding to the action, Indicates the user's status information;

[0064] Generate actions based on the state. That is, given a state, the policy network outputs the optimal ad placement selection and bidding combination;

[0065] The parameters of the Q-function approximator are optimized by back-propagation, and the policy network is updated so that the Q-function approximator can maximize the reward in the environment space.

[0066] The technical solution adopted to solve the above technical problems is: an advertising delivery optimization method for product promotion, which is applicable to the above-mentioned advertising delivery optimization system for product promotion, comprising:

[0067] Acquire multi-source interaction data of the user, wherein the multi-source interaction data includes clickstream data, transaction data, and social behavior data, and perform spatiotemporal alignment of heterogeneous data on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data;

[0068] Performing cross-modal attention fusion on the aligned data to obtain attention features, and updating the user features using the attention features according to a dual-gating mechanism to update the user features;

[0069] Performing Hilbert space mapping on the user features to obtain orthogonal interest basis vectors, obtaining ad slot feature vectors and environment feature vectors, and constructing a user-ad slot benefit matrix based on the orthogonal interest basis vectors, the environment feature vectors, and the ad slot feature vectors;

[0070] Defining an advertiser strategy space, wherein the advertiser strategy space includes ad space budget allocation, creative aggressiveness coefficient, and time period preference phase, and calculating an indirect competition effect coefficient based on the advertiser strategy space;

[0071] Constructing a budget allocation objective function based on the indirect competition effect coefficient and the user-advertising space benefit matrix, and solving the budget allocation objective function using a backward induction method to obtain an optimal budget allocation matrix for the advertiser;

[0072] Obtain the spatiotemporal feature vector when the advertisement is delivered, define a Q-function approximator based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize the advertisement delivery strategy by maximizing the Q-function approximator based on deep reinforcement learning to obtain the best bidding strategy.

[0073] The beneficial effects of the present invention are as follows: (1) The present invention performs spatiotemporal alignment through multi-source interactive data and uses cross-modal attention fusion technology to update user features, which can accurately capture the user's multi-dimensional interest changes and behavior patterns, thereby accurately matching user needs during the advertising process and improving advertising effects. By mapping user features to Hilbert space and obtaining orthogonal interest basis vectors, the different dimensions of user potential interests can be more accurately captured. This mapping method helps to accurately match user interests with advertising positions, reduce the mismatch between advertising and user interests, and thus improve advertising click-through rate and conversion rate; (2) The present invention further introduces the calculation of indirect competition effect coefficient through the user-advertising position benefit matrix and combines it with the advertiser's strategy space. This design can effectively optimize advertising. The budget allocation of advertisers can maximize the return on advertising investment, while avoiding resource waste or decline in advertising effect due to unreasonable budget allocation. Through deep reinforcement learning technology, the Q function approximator is used to continuously optimize the advertising delivery strategy. Reinforcement learning can adjust the strategy based on real-time feedback, so that advertising delivery can continuously adapt to market changes and obtain the best bidding strategy. This dynamic optimization process can help advertisers continuously improve the effectiveness of advertising delivery and adjust budget allocation in real time to cope with different market environments. (3) The present invention can optimize the advertising delivery strategy while considering market competition by introducing the indirect competition effect coefficient. This can not only help advertisers better allocate budgets, but also avoid excessive competition and market saturation, improve advertising effects, and ensure that each advertiser can obtain the best delivery return in a highly competitive environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] Figure 1 A schematic diagram of the system architecture of the overall system in an embodiment of the present invention;

[0075] Figure 2 The figure is a flowchart of the steps of the overall method in one embodiment of the present invention.

[0076] Figure numerals: 1. Data alignment unit; 2. Feature update unit; 3. Benefit modeling unit; 4. Competition modeling unit; 5. Budget allocation unit; 6. Advertising delivery unit. DETAILED DESCRIPTION

[0077] Example 1, as Figure 1 As shown, the present invention proposes an advertising delivery optimization system for product promotion, comprising:

[0078] Data alignment unit 1, which is used to obtain multi-source interaction data of users, wherein the multi-source interaction data includes click stream data, transaction data and social behavior data, and perform heterogeneous data spatiotemporal alignment on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data;

[0079] Feature update unit 2, which is used to perform cross-modal attention fusion on the aligned data to obtain attention features, and update the user features through the attention features according to the dual-gating mechanism to update the user features;

[0080] Benefit modeling unit 3, which is used to perform Hilbert space mapping on user characteristics to obtain orthogonal interest basis vectors, obtain ad slot feature vectors and environment feature vectors, and construct a user-ad slot benefit matrix based on the orthogonal interest basis vectors, environment feature vectors, and ad slot feature vectors;

[0081] Competition modeling unit 4, competition modeling unit 4 is used to define the advertiser strategy space, wherein the advertiser strategy space includes advertising budget allocation, creative aggressiveness coefficient and time period preference phase, and calculates the indirect competition effect coefficient based on the advertiser strategy space and the user-advertising benefit matrix.

[0082] In the present invention, multi-source interaction data includes interaction data between users and goods or advertisements. Usually, clickstream data reflects the user's click behavior on advertisements or goods, transaction data reflects whether the user has made a purchase, and social behavior data reflects the user's interaction on social platforms (such as likes, comments, etc.); heterogeneous data spatiotemporal alignment is due to the different spatiotemporal characteristics of these data sources, and they need to be aligned through timestamps so that they reflect the user's real behavior at the same time; cross-modal attention fusion is due to the different data sources (such as clickstream, transaction, social data) have different feature distributions and importance, and it is necessary to use attention mechanisms (especially cross-modal attention mechanisms) to perform weighted fusion on them to highlight the importance of each data in the current task; Hilbert space mapping refers to mapping user features into Hilbert space, which can better process high-dimensional data, and through the selection of orthogonal basis, each interest dimension can be made They are independent and unrelated to each other to avoid interference between different interests; the orthogonal interest basis vectors represent the user's independent interest points. Through orthogonalization, the model can more clearly distinguish the user's interest tendencies; the benefit matrix refers to the user-advertising benefit matrix constructed by combining the user's interests with the ad position and environmental characteristics. This matrix describes the matching degree between the user and each ad position and can be used to evaluate the effectiveness of advertising; the advertiser's strategy space includes budget allocation, creative aggressiveness coefficient (that is, the promotion strength of the advertisement) and time preference phase (that is, in what time period the advertiser wants to advertise). These factors affect the choice of ad position and the effectiveness of advertising; advertising is not only affected by direct competitors, but also by other indirect competitive factors. By analyzing the competitive relationship between ad positions and advertisers, the model can estimate these indirect effects and incorporate them into the delivery decision to avoid repeated or ineffective delivery.

[0083] In an optional embodiment, the system further includes:

[0084] Budget allocation unit 5, which is used to construct a budget allocation objective function based on the indirect competition effect coefficient and the user-advertising space benefit matrix, and solve the budget allocation objective function according to the backward induction method to obtain the advertiser's optimal budget allocation matrix;

[0085] Advertisement delivery unit 6, which is used to obtain the spatiotemporal feature vector during advertisement delivery, define a Q-function approximator based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize the advertisement delivery strategy by maximizing the Q-function approximator based on deep reinforcement learning to obtain the best bidding strategy.

[0086] It should be noted that the indirect competition effect coefficient represents the competition effect between ad spaces or advertisers during the advertising process. These effects include the impact of other advertisers' advertising, the supply and demand relationship of ad spaces and other factors. The introduction of indirect competition effect makes the budget allocation not only consider the direct competition between advertisers, but also consider how these external factors affect the effect of advertising. The construction of the objective function is usually to maximize the total benefit of the advertiser. The objective function will combine the indirect competition effect and the benefit matrix to determine the budget allocation of different ad spaces. The reverse induction method plays an important role in solving the budget allocation problem. Usually, this method starts from the final result and reversely calculates the optimal strategy in the decision-making process step by step. The reverse induction method can help Starting from the effectiveness of advertising, the budget allocation for each ad slot is gradually derived. The spatiotemporal feature vector in advertising reflects information such as the time period, location, and user activity of the ad. These features have a significant impact on advertising effectiveness, as different time periods, locations, and user behavior characteristics all affect performance indicators such as the click-through rate and conversion rate. The Q function is a core concept in reinforcement learning, representing the expected benefit of taking an action under a given state. In advertising, the Q function quantifies the benefits of advertising in a specific spatiotemporal environment. By defining a Q function approximator, the relationship between decision-making and benefits during the advertising process can be simulated, thereby providing advertisers with the optimal advertising strategy. Deep reinforcement learning is used here to optimize advertising strategies. By continuously interacting with the environment (i.e., continuously adjusting the advertising strategy and observing the results), the deep reinforcement learning model gradually learns how to maximize the Q function, thereby optimizing the advertising bidding strategy.

[0087] In the second embodiment, an advertisement placement optimization system for product promotion proposed by the present invention is provided. Compared with the first embodiment, this embodiment further includes: performing spatiotemporal alignment of heterogeneous data on multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data, including:

[0088] Perform cubic spline interpolation on the transaction data to obtain an interpolation result of the transaction data. The calculation formula for the interpolation result of the transaction data is as follows:

[0089] ;

[0090] in, Indicates time The interpolation result of the transaction data at Indicates the total number of transaction data, represents the basis function based on Lagrange interpolation, and , Indicates time Number of transactions at

[0091] A social behavior heat decay coefficient is defined to perform heat decay processing on the social behavior data to obtain a heat decay result of the social behavior data. The calculation formula for the heat decay result of the social behavior data is as follows:

[0092] ;

[0093] in, Indicates time The heat decay results of social behavior data at Indicates time Social behavior data, represents the social behavior heat decay coefficient, and , Represents the seconds of the minute;

[0094] Based on the clickstream data, the interpolation results of the transaction data and the heat decay results of the social behavior data are aligned to a unified timeline to obtain the aligned data corresponding to the multi-source interaction data. The calculation formula for the aligned data is as follows:

[0095] ;

[0096] in, Indicates alignment data, Indicates time Clickstream data from .

[0097] In this embodiment, cubic spline interpolation is a method of interpolation by connecting multiple segments of cubic polynomials (each segment of the polynomial has a continuous and smooth derivative at each interpolation point). It is very common in practical applications and is particularly suitable for problems that require maintaining data smoothness (such as curve fitting); Lagrange interpolation is a method of interpolation by calculating a polynomial for a set of given data points; social behavior data usually has a heat decay effect, which means that the influence of social behavior gradually weakens over time.

[0098] In an optional embodiment, cross-modal attention fusion is performed on the aligned data to obtain attention features, including:

[0099] The attention head of each modality in the aligned data is calculated to obtain the attention corresponding to each modality in the aligned data. The calculation formula for the attention corresponding to each modality in the aligned data is as follows:

[0100] ;

[0101] in, Indicates the alignment data The attention corresponding to each modality, 、 and represents the query matrix, key matrix and value matrix, 、 and Indicates the alignment data The projection matrix corresponding to each mode is, Represents the dimension of the feature, represents the learnable modality mask matrix;

[0102] The attention corresponding to each modality in the aligned data is concatenated and linearly transformed to obtain the attention feature. The calculation formula of the attention feature is as follows:

[0103] ;

[0104] in, Represents the attention feature, Represents a splicing operation, Represents a linear transformation matrix.

[0105] In an optional embodiment, user features are updated using attention features according to a dual-gating mechanism to obtain user features, including:

[0106] According to the first gating mechanism, historical user characteristics are combined with attention features to perform importance scoring to obtain an importance score. The calculation formula of the importance score is as follows:

[0107] ;

[0108] in, represents the importance score, represents the activation function, represents the first trainable weight matrix, represents the user features of the previous time step, Represents the attention feature of the current time step;

[0109] According to the second gating mechanism, user features are fused by importance score to obtain fusion information. The calculation formula of fusion information is as follows:

[0110] ;

[0111] in, represents fusion information, represents the second trainable weight matrix;

[0112] The user features are updated based on the importance score and fusion information to update the user features. The update formula of the user features is as follows:

[0113] ;

[0114] in, represents the updated user features, represents the hyperbolic tangent function, represents the third trainable weight matrix;

[0115] When the importance score is greater than the preset importance score threshold, the updated user features corresponding to the importance score are stored as events to obtain a memory library, where the expression of the memory library is as follows:

[0116] ;

[0117] in, represents the updated user feature sequence stored in the memory bank, Indicates the size of the memory bank.

[0118] It should be noted that the memory bank is used to store important user feature sequences. Whenever a user's behavior or features undergo significant changes at a certain time step, the model will store the feature sequence in the memory bank for future reference. The size of the memory bank may be limited, and when storing new features, there may be a strategy to decide whether to delete old features.

[0119] In an optional embodiment, the calculation formula of the user-advertising space benefit matrix is as follows:

[0120] ;

[0121] in, Represents a user and ad space The benefits between represents the orthogonal interest basis vector, represents the ad slot feature vector, represents the environmental feature vector, Represents the time decay factor, which is used to control the timeliness of the ad space. Indicates ad space Timeliness, Indicates the historical click-through rate mixing coefficient, balancing the effects of new and old ad positions. Indicates the historical click-through rate of the ad slot.

[0122] In an optional embodiment, the calculation formula of the indirect competition effect coefficient is as follows:

[0123] ;

[0124] in, Indicates advertiser With advertisers The indirect interference strength between Indicates advertiser In the advertising space Budget allocation on Represents a user and ad space The benefits between Represents the differences in advertiser strategy spaces;

[0125] The budget allocation objective function is as follows:

[0126] ;

[0127] in, Indicates ad space The benefits, represents the budget competition factor, represents the regularization coefficient, represents the indirect interference coefficient, Indicates advertiser total budget.

[0128] In an optional embodiment, the budget allocation objective function is solved according to the backward induction method to obtain the advertiser's optimal budget allocation matrix, including:

[0129] Initialize the budget allocation strategy matrix for each advertiser;

[0130] Iteratively update each advertiser's budget allocation strategy. The update formula for the advertiser's budget allocation strategy is as follows:

[0131] ;

[0132] in, Indicates at time Advertisers budget allocation strategy, Indicates that all advertisers at the time The sum of budget allocation strategies;

[0133] Define the convergence conditions to control the accuracy of the budget allocation strategy update, where the convergence conditions are as follows:

[0134] ;

[0135] in, Indicates at time Budget allocation strategy matrix, Indicates at time Budget allocation strategy matrix, Indicates the preset accuracy threshold;

[0136] When the update of the advertiser's budget allocation strategy meets the preset convergence condition, the iterative update of the advertiser's budget allocation strategy is stopped to obtain the advertiser's optimal budget allocation matrix.

[0137] It should be noted that each advertiser has a budget allocation strategy matrix, which is usually a two-dimensional matrix, where each row represents an advertiser's budget allocation strategy, and each column represents the budget allocation at different times (or different ad positions, etc.). The initialization matrix can be randomly initialized or set based on some prior information. Initialization is usually used to provide a starting point for solving nonlinear optimization problems.

[0138] In an optional embodiment, a Q-function approximator is defined based on the advertiser's optimal budget allocation matrix and spatiotemporal eigenvectors, and an advertising delivery strategy is optimized by maximizing the Q-function approximator based on deep reinforcement learning to obtain the optimal bidding strategy, including:

[0139] The Q-function approximator is defined based on the advertiser's optimal budget allocation matrix and the spatiotemporal eigenvector, where the Q-function approximator is as follows:

[0140] ;

[0141] in, represents the Q-function myostat, represents the weight parameter of the Q function, and represents the network weight, Represents the spatiotemporal feature vector, which is obtained by extracting the geographic location information, time information and device information through MLP. Indicates the ad slot and bid embedding corresponding to the action, Indicates the user's status information;

[0142] Generate actions based on the state. That is, given a state, the policy network outputs the optimal ad placement selection and bidding combination;

[0143] Through back propagation, the parameters of the Q-function approximator are optimized and the policy network is updated so that the Q-function approximator can maximize the reward in the environment space.

[0144] It should be noted that geographic location information usually indicates the user's location characteristics when the advertiser wants to display advertisements at a specific location. The geographic location can be represented by latitude and longitude, city information, etc.; time information indicates that the advertiser wants to place advertisements within a certain time period. Time information is usually represented by different granularities such as hours, days, weeks, etc., or directly encoded using timestamps; device information reflects the type of device used by the user, such as mobile phones, tablets, computers, etc.

[0145] Example 3, as Figure 2 As shown, the present invention proposes an advertising delivery optimization method for product promotion, which is applicable to the advertising delivery optimization system for product promotion, including:

[0146] S1. Acquire multi-source interaction data of users, where the multi-source interaction data includes clickstream data, transaction data, and social behavior data, and perform spatiotemporal alignment of heterogeneous data on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data;

[0147] S2. Perform cross-modal attention fusion on the aligned data to obtain attention features, and update the user features through the attention features according to the double-gating mechanism to update the user features;

[0148] S3. Perform Hilbert space mapping on the user features to obtain orthogonal interest basis vectors, obtain ad slot feature vectors and environment feature vectors, and construct a user-ad slot benefit matrix based on the orthogonal interest basis vectors, environment feature vectors, and ad slot feature vectors;

[0149] S4. Define an advertiser strategy space, where the advertiser strategy space includes ad placement budget allocation, creative aggressiveness coefficient, and time period preference phase, and calculate the indirect competition effect coefficient based on the advertiser strategy space;

[0150] S5. Construct a budget allocation objective function based on the indirect competition effect coefficient and the user-advertising space benefit matrix, and solve the budget allocation objective function using the backward induction method to obtain the advertiser's optimal budget allocation matrix;

[0151] S6. Obtain the spatiotemporal feature vector when the advertisement is delivered, define a Q-function approximator based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize the advertisement delivery strategy by maximizing the Q-function approximator based on deep reinforcement learning to obtain the best bidding strategy.

[0152] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.

Claims

1. An advertising optimization system for product promotion, characterized in that: include: A data alignment unit (1), the data alignment unit (1) is used to obtain multi-source interaction data of a user, wherein the multi-source interaction data includes click stream data, transaction data and social behavior data, and perform heterogeneous data spatiotemporal alignment on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data; A feature updating unit (2), the feature updating unit (2) is used to perform cross-modal attention fusion on the aligned data to obtain attention features, and perform feature update on the user through the attention features according to a dual-gating mechanism to update user features; A benefit modeling unit (3) is used to perform Hilbert space mapping on the user characteristics to obtain an orthogonal interest basis vector, obtain an ad slot feature vector and an environment feature vector, and construct a user-ad slot benefit matrix based on the orthogonal interest basis vector, the environment feature vector and the ad slot feature vector; A competition modeling unit (4), wherein the competition modeling unit (4) is used to define an advertiser strategy space, wherein the advertiser strategy space includes an advertisement budget allocation, a creative aggressiveness coefficient, and a time period preference phase, and calculate an indirect competition effect coefficient based on the advertiser strategy space and the user-advertising space benefit matrix; Performing spatiotemporal alignment of heterogeneous data on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data includes: Performing cubic spline interpolation on the transaction data to obtain an interpolation result of the transaction data, wherein the calculation formula of the interpolation result of the transaction data is as follows: ; in, Indicates time The interpolation result of the transaction data at Indicates the total number of transaction data, represents the basis function based on Lagrange interpolation, and , Indicates time Number of transactions at A social behavior heat decay coefficient is defined to perform heat decay processing on the social behavior data to obtain a heat decay result of the social behavior data. The calculation formula of the heat decay result of the social behavior data is as follows: ; in, Indicates time The heat decay results of social behavior data at Indicates time Social behavior data, represents the social behavior heat decay coefficient, and , Represents the seconds of the minute; The interpolation results of the transaction data and the heat decay results of the social behavior data are aligned to a unified time axis according to the clickstream data to obtain aligned data corresponding to the multi-source interaction data. The calculation formula of the aligned data is as follows: ; in, Indicates alignment data, Indicates time Clickstream data from .

2. The advertising optimization system for product promotion according to claim 1, characterized in that: The system further comprises: A budget allocation unit (5), the budget allocation unit (5) is used to construct a budget allocation objective function according to the indirect competition effect coefficient and the user-advertising space benefit matrix, and solve the budget allocation objective function according to the reverse induction method to obtain the advertiser's optimal budget allocation matrix; An advertisement delivery unit (6) is used to obtain a spatiotemporal feature vector when an advertisement is delivered, define a Q-function approximator according to the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize an advertisement delivery strategy by maximizing the Q-function approximator according to deep reinforcement learning to obtain an optimal bidding strategy.

3. The advertising optimization system for product promotion according to claim 2, characterized in that: Performing cross-modal attention fusion on the aligned data to obtain attention features, including: The attention head of each modality in the aligned data is calculated to obtain the attention corresponding to each modality in the aligned data, wherein the calculation formula for the attention corresponding to each modality in the aligned data is as follows: ; in, Indicates the alignment data The attention corresponding to each modality, 、 and represents the query matrix, key matrix and value matrix, 、 and Indicates the alignment data The projection matrix corresponding to each mode is, Represents the dimension of the feature, represents the learnable modality mask matrix; The attention corresponding to each modality in the aligned data is concatenated and linearly transformed to obtain the attention feature, where the calculation formula of the attention feature is as follows: ; in, Represents the attention feature, Represents a splicing operation, Represents a linear transformation matrix.

4. The advertising delivery optimization system for product promotion according to claim 3, characterized in that: The user feature is updated using the attention feature according to the dual-gating mechanism to obtain user features, including: The importance score is performed by combining the attention feature with the historical user features according to the first gating mechanism to obtain an importance score, wherein the calculation formula of the importance score is as follows: ; in, represents the importance score, represents the activation function, represents the first trainable weight matrix, represents the user features of the previous time step, Represents the attention feature of the current time step; According to the second gating mechanism, the user's features are fused using the importance score to obtain fusion information, wherein the calculation formula of the fusion information is as follows: ; in, represents fusion information, represents the second trainable weight matrix; The user characteristics are updated according to the importance score and the fusion information to update the user characteristics, wherein the updating formula of the user characteristics is as follows: ; in, represents the updated user features, represents the hyperbolic tangent function, represents the third trainable weight matrix; When the importance score is greater than a preset importance score threshold, the updated user feature corresponding to the importance score is stored as an event to obtain a memory library, wherein the expression of the memory library is as follows: ; in, represents the updated user feature sequence stored in the memory bank, Indicates the size of the memory bank.

5. The advertisement placement optimization system for product promotion according to claim 4, characterized in that: The calculation formula of the user-advertising space benefit matrix is as follows: ; in, Represents a user and ad space The benefits between represents the orthogonal interest basis vector, represents the ad slot feature vector, represents the environmental feature vector, Represents the time decay factor, which is used to control the timeliness of the ad space. Indicates ad space Timeliness, Indicates the historical click-through rate mixing coefficient, balancing the effects of new and old ad positions. Indicates the historical click-through rate of the ad slot.

6. The advertisement placement optimization system for product promotion according to claim 5, characterized in that: The calculation formula of the indirect competition effect coefficient is as follows: ; in, Indicates advertiser With advertisers The indirect interference strength between Indicates advertiser In the advertising space Budget allocation on Represents a user and ad space The benefits between Represents the differences in advertiser strategy spaces; The budget allocation objective function is as follows: ; in, Indicates ad space The benefits, represents the budget competition factor, represents the regularization coefficient, represents the indirect interference coefficient, Indicates advertiser total budget.

7. The advertisement placement optimization system for product promotion according to claim 6, characterized in that: The budget allocation objective function is solved according to the backward induction method to obtain the advertiser's optimal budget allocation matrix, including: Initialize the budget allocation strategy matrix for each advertiser; Iteratively update the budget allocation strategy of each advertiser, wherein the update formula of the advertiser's budget allocation strategy is as follows: ; in, Indicates at time Advertisers budget allocation strategy, Indicates that all advertisers at the time The sum of budget allocation strategies; Define a convergence condition to control the accuracy of the budget allocation strategy update, where the convergence condition is as follows: ; in, Indicates at time Budget allocation strategy matrix, Indicates at time Budget allocation strategy matrix, Indicates the preset accuracy threshold; When the update of the advertiser's budget allocation strategy satisfies a preset convergence condition, the iterative update of the advertiser's budget allocation strategy is stopped to obtain the advertiser's optimal budget allocation matrix.

8. The advertisement placement optimization system for product promotion according to claim 7, characterized in that: A Q-function approximator is defined based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and an advertising delivery strategy is optimized by maximizing the Q-function approximator based on deep reinforcement learning to obtain an optimal bidding strategy, including: A Q-function approximator is defined according to the advertiser's optimal budget allocation matrix and the spatiotemporal eigenvector, wherein the Q-function approximator is as follows: ; in, represents the Q-function myostat, represents the weight parameter of the Q function, and represents the network weight, Represents the spatiotemporal feature vector, which is obtained by extracting the geographic location information, time information and device information through MLP. Indicates the ad slot and bid embedding corresponding to the action, Indicates the user's status information; Generate actions based on the state. That is, given a state, the policy network outputs the optimal ad placement selection and bidding combination; The parameters of the Q-function approximator are optimized by back-propagation, and the policy network is updated so that the Q-function approximator can maximize the reward in the environment space.

9. An advertising delivery optimization method for product promotion, which is applicable to the advertising delivery optimization system for product promotion according to claim 8, characterized in that: include: Acquire multi-source interaction data of the user, wherein the multi-source interaction data includes clickstream data, transaction data, and social behavior data, and perform spatiotemporal alignment of heterogeneous data on the multi-source interaction data to obtain aligned data corresponding to the multi-source interaction data; Performing cross-modal attention fusion on the aligned data to obtain attention features, and updating the user features using the attention features according to a dual-gating mechanism to update the user features; Performing Hilbert space mapping on the user features to obtain orthogonal interest basis vectors, obtaining ad slot feature vectors and environment feature vectors, and constructing a user-ad slot benefit matrix based on the orthogonal interest basis vectors, the environment feature vectors, and the ad slot feature vectors; Defining an advertiser strategy space, wherein the advertiser strategy space includes ad space budget allocation, creative aggressiveness coefficient, and time period preference phase, and calculating an indirect competition effect coefficient based on the advertiser strategy space; Constructing a budget allocation objective function based on the indirect competition effect coefficient and the user-advertising space benefit matrix, and solving the budget allocation objective function using a backward induction method to obtain an optimal budget allocation matrix for the advertiser; Obtain the spatiotemporal feature vector when the advertisement is delivered, define a Q-function approximator based on the advertiser's optimal budget allocation matrix and the spatiotemporal feature vector, and optimize the advertisement delivery strategy by maximizing the Q-function approximator based on deep reinforcement learning to obtain the best bidding strategy.

Citation Information

Patent Citations

  • Video advertisement putting effect intelligent analysis and management system based on big data analysis

    CN119323441A

Cited By

  • Crowd portrait method and promotion system based on big data

    CN121352895A