Instant advertising strategy generation method based on transient gradient injection and entropy-directed search

By using Monte Carlo tree search of manifold space mapping and shadow simulation field, combined with transient gradient update of entropy-oriented objective function, the problem of targeted advertising strategy generation in long-tail user scenarios is solved, and high-return path capture and strategy optimization for long-tail users are achieved.

CN121981783BActive Publication Date: 2026-06-26SUZHOU PINWU INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610426036.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-02
Publication Date
2026-06-26
Estimated Expiration
2046-04-02

AI Technical Summary

Technical Problem

Existing advertising strategy generation methods struggle to target long-tail user scenarios that deviate from the main distribution of historical training data. This results in strategy generation results that fail to reflect the differentiated characteristics of the current request and underestimate potential high-value conversion opportunities.

Method used

By acquiring advertising request features and mapping them to a manifold space, a shadow simulation field is constructed and an extremum-oriented Monte Carlo tree search is performed. An entropy-oriented objective function is constructed using a set of trajectories, and transient gradient updates are performed on the low-rank adapter parameters to generate a highly targeted advertising strategy.

Benefits of technology

In the process of processing a single ad request, the targeting of ad strategy generation is improved, which can capture the potential high-return paths of long-tail users and improve the ability to make differentiated decisions on ad strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981783B_ABST
    Figure CN121981783B_ABST
Patent Text Reader

Abstract

The application relates to the field of advertisement delivery, and in particular to an instant advertisement strategy generation method based on transient gradient injection and entropy-oriented search, which comprises the following steps: performing manifold space mapping on user-side features, context features and candidate advertisement features in an advertisement request to obtain an initial state vector representing a current advertisement request state; loading a generative world model and introducing a strategy network, initializing low-rank adapter parameters in the strategy network, and constructing a shadow simulation field for the current advertisement request; performing extreme value-oriented Monte Carlo tree search in the shadow simulation field to generate a trajectory set corresponding to the current advertisement request; constructing an entropy-oriented objective function and performing transient gradient update on the low-rank adapter parameters to obtain an updated strategy network; and performing final forward reasoning on the initial state vector based on the updated strategy network to generate an advertisement strategy and resetting the low-rank adapter parameters after execution. The application can improve the pertinence of advertisement strategy generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of advertising technology, and in particular to a method for generating real-time advertising strategies based on transient gradient injection and entropy-guided search. Background Technology

[0002] In the process of internet advertising, platforms typically need to generate an advertising strategy corresponding to the current request based on user characteristics, access scenarios, candidate ad materials, and business constraints to determine specific decisions such as ad display, bid control, or material selection. As advertising scenarios become increasingly refined and user behavior becomes more complex, ad strategy generation not only needs to meet the timeliness requirements of online response but also needs to be as close as possible to the user's state and context of the current request to improve the actual effectiveness of ad delivery.

[0003] Existing advertising strategy generation methods typically employ a combination of offline training and online inference. Specifically, the system first builds training data based on historical advertising interaction logs, performing offline training on click-through rate prediction models, conversion rate prediction models, or joint decision-making models. Then, during the online phase, the model parameters are kept fixed. Upon receiving a new advertising request, user-side features, contextual features, and candidate advertising features are extracted and input into the already trained model for forward computation, generating the corresponding advertising strategy result. Since historical training data typically originates from highly active users with frequent interactions and relatively comprehensive behavioral characteristics, the offline-trained model primarily fits the existing historical distribution. The online phase mainly relies on the capabilities of this offline-trained model to generate real-time strategies.

[0004] However, based on the above processing methods, further analysis reveals that the training data relied upon for generating existing advertising strategies mainly comes from frequently interacting top-performing users, whose feature distribution exhibits a clear concentration. However, in actual online requests, compared to these high-frequency users, there exists a group of long-tail users with lower frequency of occurrence and rarer feature combinations. For these users, such as those speaking specific minority languages, those with atypical browsing paths, or potential high-net-worth individuals, their requests often fall outside the historical training data distribution. In this case, the offline-trained model, because it fits to an existing distribution and its parameters remain unchanged during the online phase, typically tends to output relatively conservative average predictions. This makes it difficult for the generated strategy to reflect the potentially differentiated characteristics within the current request, thus underestimating potential high-value conversions and making it difficult to capture high-return decision opportunities. Therefore, how to improve the targeting of advertising strategy generation for long-tail user scenarios that deviate from the main distribution of historical training data during the processing of a single advertising request has become a pressing problem in this field. Summary of the Invention

[0005] This application provides a real-time advertising strategy generation method based on transient gradient injection and entropy-guided search, which can improve the targeting of advertising strategy generation for long-tail user scenarios that deviate from the main distribution of historical training data during the processing of a single advertising request. This application provides the following technical solution:

[0006] In a first aspect, this application provides a method for generating instant advertising strategies based on transient gradient injection and entropy-guided search, the method comprising:

[0007] Obtain an ad request, and perform manifold space mapping on the user-side features, context features, and candidate ad features in the ad request to obtain an initial state vector representing the current ad request state;

[0008] Load the generative world model and introduce the policy network, initialize the low-rank adapter parameters in the policy network, and construct a shadow simulation field for the current ad request;

[0009] Based on the initial state vector, perform an extremum-oriented Monte Carlo tree search in the shadow simulation field to generate a trajectory set corresponding to the current ad request;

[0010] An entropy-oriented objective function is constructed based on the trajectory set, and transient gradient updates are performed on the low-rank adapter parameters to obtain a policy network for the current ad request.

[0011] The initial state vector is used to generate an advertising strategy based on the updated policy network, and the low-rank adapter parameters are reset after the advertising strategy is executed.

[0012] In a specific implementation, the manifold space mapping of the user-side features, context features, and candidate ad features in the ad request to obtain an initial state vector representing the current ad request state includes:

[0013] Extracting user-side features from ad requests Contextual features and candidate ad features After feature extraction, joint characterization processing is performed;

[0014] A pre-trained feature compression network is set up to perform manifold space mapping on the jointly represented features. The feature compression network adopts a multilayer perceptron structure to map the input high-dimensional sparse feature vector into a low-dimensional dense vector representation.

[0015] The mapping function of the feature compression network is expressed as: ;in, Represents a feature compression network; This represents the network parameters of the feature compression network; This represents the set of input features consisting of user-side features, contextual features, and candidate ad features; This represents the initial state vector of the current ad request in the low-dimensional manifold space; ,in, Represents the set of real numbers. Indicates by Composed of real numbers 3D real vector space, This represents the vector dimension of a low-dimensional manifold space.

[0016] In one specific implementation, the loading of the generative world model and the introduction of a policy network, the initialization of the low-rank adapter parameters in the policy network, and the construction of a shadow simulation field for the current ad request include:

[0017] Loading Generative World Model The input to the generative world model is the current state vector. With the proposed action The output is the predicted state for the next time step. and predicted rewards Its mathematical expression is: ;

[0018] Introducing policy networks The transient adaptation parameters in the policy network are initialized, where, Let represent the principal parameters of the policy network; a low-rank adapter is bypassed at the critical linear layer of the policy network, and the parameters of the low-rank adapter are denoted as . , As a transient adaptation parameter that can be partially updated during the current ad request processing and reset after the request ends;

[0019] low-rank adapter Includes two matrices and ,in, , ; Represents the set of real numbers. Indicates by A matrix space consisting of n real numbers Indicates by A matrix space consisting of n real numbers; The dimension of the state vector. Let represent the rank of the low-rank adapter, and The policy output after introducing a low-rank adapter can be expressed as:

[0020] ;

[0021] in, Indicates only by the main parameter The generated basic policy output; Represents the input state vector; This represents the amount of local correction applied by the low-rank adapter to the output of the basic policy;

[0022] By loading a generative world model And initialize transient adaptation parameters Together with the policy network, they form a shadow simulation field.

[0023] In one specific implementation, the step of performing an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate the trajectory set corresponding to the current advertising request includes:

[0024] With the initial state vector As the root node of the search tree, execute Each simulation cycle, in each simulation cycle, starts from the current state vector. Starting from the policy network in the state vector The generated action distribution determines the action. and the state vector With action Input Generative World Model The corresponding predicted state vector is obtained. and predicted rewards ; predict the state vector As the current state for the next step of the simulation, continue with action selection and state simulation until the preset maximum simulation depth is reached. ;

[0025] For any node Its child node actions The selection criteria are expressed as follows:

[0026] ;

[0027] in, This means selecting the action that maximizes the value of the node selection function from all candidate actions. This indicates that during the historical simulation process, the nodes... Select Action The maximum reward value observed later. Indicates the exploration coefficient; Indicates at node The action output by the policy network in the corresponding state The prior probability; Represents a node Next action Number of visits; Represents a node The total number of visits for all candidate actions, of which, Represents a node The candidate action index below;

[0028] After completion After each simulation cycle, the states, actions, and predicted reward sequences formed in each simulation cycle are collected to form a trajectory set.

[0029] In one specific implementation scheme, constructing the entropy-oriented objective function based on the trajectory set includes:

[0030] Based on trajectory set Constructing an entropy-oriented objective function Entropy-oriented objective function Represented as:

[0031] ;

[0032] in, This indicates the parameter to be updated, corresponding to the low-rank adapter parameter. ; Representing the trajectory From the set of trajectories Sampling; Represents the set of trajectories Calculate the expectation of the samples above; Representing the trajectory The corresponding weights; Indicates the state Below, the parameters are Policy network selects actions The logarithmic probability, Represents the entropy adjustment coefficient, and ; This represents the information entropy of the action distribution output by the policy network;

[0033] The trajectory weight The exponential weighting function is used to determine this, and its expression is as follows:

[0034] ;

[0035] in, Represents an exponential function; Representing the trajectory Cumulative rewards; Represents the set of trajectories The average cumulative reward for each trajectory; This indicates the reward scaling factor, representing the cumulative reward. It is obtained by summing the predicted rewards at each step in the trajectory, that is: ;in, The trajectory is in the th order. The predicted reward corresponding to each step; Indicates the sequence number of the deduction step in the trajectory. This indicates the maximum extrapolation depth of a single trajectory.

[0036] In one specific implementation, performing transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request includes:

[0037] Constructing an entropy-oriented objective function Then, the objective function with respect to the low-rank adapter parameters is calculated. gradient Its expression is:

[0038] ;

[0039] in, Indicates the parameter Find the partial derivative;

[0040] Obtain the gradient Then, a stochastic gradient descent optimizer or an Adam optimizer is used to optimize the low-rank adapter parameters. implement Step update, the updated low-rank adapter parameters are denoted as The update process is represented as follows: ;in, This represents the low-rank adapter parameters before the update. This represents the learning rate.

[0041] In one specific implementation, the step of generating an advertising strategy by performing final forward inference on the initial state vector based on the updated policy network, and resetting the low-rank adapter parameters after the advertising strategy is executed, includes:

[0042] After receiving the initial state vector, the updated policy network outputs the probability distribution of each candidate action in the current state. Based on the probability distribution, it selects the action with the highest probability among all candidate actions as the target action, and parses the action content corresponding to the target action into an advertising strategy to generate the final advertising strategy corresponding to the current advertising request.

[0043] After obtaining the final advertising strategy, the advertising strategy is sent to the advertising exchange system for execution, thus completing the strategy generation and placement decision for this advertising request;

[0044] After the advertising strategy is executed, the low-rank adapter parameters are reset to their initial state.

[0045] Secondly, this application provides a real-time advertising strategy generation system based on transient gradient injection and entropy-guided search, employing the following technical solution:

[0046] A real-time advertising strategy generation system based on transient gradient injection and entropy-guided search includes:

[0047] The request acquisition module is used to acquire advertising requests, and perform manifold space mapping on the user-side features, context features and candidate advertising features in the advertising requests to obtain an initial state vector representing the current advertising request state.

[0048] The simulation construction module is used to load the generative world model and introduce the policy network, initialize the low-rank adapter parameters in the policy network, and construct a shadow simulation field for the current advertising request.

[0049] The trajectory generation module is used to perform an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate a trajectory set corresponding to the current advertising request.

[0050] The network update module is used to construct an entropy-oriented objective function based on the trajectory set and perform transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request.

[0051] The strategy generation module is used to generate an advertising strategy by performing final forward inference on the initial state vector based on the updated policy network, and to reset the low-rank adapter parameters after the advertising strategy is executed.

[0052] Thirdly, this application provides an electronic device, the device including a processor and a memory; the memory stores a program, the program being loaded and executed by the processor to implement an instant advertising strategy generation method based on transient gradient injection and entropy-guided search as described in the first aspect.

[0053] Fourthly, this application provides a computer-readable storage medium storing a program that, when executed by a processor, is used to implement a real-time advertising strategy generation method based on transient gradient injection and entropy-guided search as described in the first aspect.

[0054] After obtaining the ad request, the user-side features, context features, and candidate ad features are mapped to a manifold space to obtain an initial state vector. Based on this, a shadow simulation field is constructed. The current ad request is virtually inferred through a generative world model and a policy network. Based on the initial state vector, an extremum-oriented Monte Carlo tree search is performed to generate a trajectory set. The trajectory set is then used to construct an entropy-oriented objective function, and transient gradient updates are performed on the low-rank adapter parameters. This allows the policy network to complete local adaptive adjustments within the state space corresponding to the current ad request. Finally, based on the updated policy network, a forward inference is performed on the initial state vector to generate an ad policy. After the policy is executed, the low-rank adapter parameters are reset. In the processing of a single ad request, this application does not directly rely on a fixed policy network for decision-making. Instead, it first generates a set of trajectories related to the current ad request in a shadow simulation field around the initial state vector corresponding to the current ad request. Then, it uses an entropy-driven objective function to transiently update the parameters of the low-rank adapter, enabling the policy network to form a local decision-making capability for the current request during the processing of the current request. Since this process is centered on the current ad request itself, it uses the trajectory set to characterize its potential high-return paths and uses parameter updates to make the action distribution concentrate on the actions corresponding to these paths. Therefore, it can reduce the dependence on the main distribution of historical training data. This allows the policy network to form differentiated decision results based on the characteristics of the current request when facing long-tail user scenarios that deviate from the main historical distribution, thereby improving the targeting of ad strategy generation for such scenarios.

[0055] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of this application are described in detail below with reference to the accompanying drawings. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating the instant advertising strategy generation method based on transient gradient injection and entropy-guided search in the embodiments of this application.

[0057] Figure 2 This is a schematic diagram of the overall process of the instant advertising strategy generation method based on transient gradient injection and entropy-guided search in the embodiments of this application.

[0058] Figure 3 This is a structural block diagram of the instant advertising strategy generation system based on transient gradient injection and entropy-guided search in the embodiments of this application.

[0059] Figure 4 This is a block diagram of an electronic device generated based on a real-time advertising strategy using transient gradient injection and entropy-guided search, as described in this application. Detailed Implementation

[0060] The specific embodiments of this application will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this application, but are not intended to limit the scope of this application.

[0061] Optionally, this application uses the instant advertising strategy generation method based on transient gradient injection and entropy-guided search provided in various embodiments as an example for application in an electronic device. The electronic device is a terminal or a server. The terminal can be a computer, tablet computer, etc. This embodiment does not limit the type of electronic device.

[0062] Reference Figure 1 This is a flowchart illustrating a real-time advertising strategy generation method based on transient gradient injection and entropy-guided search provided in an embodiment of this application. The method includes at least the following steps:

[0063] Step S101: Obtain the ad request, and perform manifold space mapping on the user-side features, context features and candidate ad features in the ad request to obtain the initial state vector representing the current ad request state.

[0064] In step S101, the advertising request initiated by the advertising exchange platform is obtained, and the user-side features, context features, and candidate advertising features in the advertising request are jointly modeled to construct a unified representation that can characterize the overall state of the current advertising request. Advertising strategy generation relies on an accurate depiction of the current advertising request state. However, the raw data in advertising requests typically comes from diverse sources, has high dimensionality, and differs in representation. Therefore, this step requires a unified representation of various features and compression into a low-dimensional space to reduce computational complexity and improve the stability of the state representation.

[0065] Specifically, it receives advertising request messages transmitted via the OpenRTB protocol and extracts three types of feature data from these messages. The first type is user-side features, denoted as... The first category includes device identifiers, long-term interest tags, and real-time browsing session sequences, used to characterize a user's historical behavioral preferences and current behavioral state. The second category is contextual features, denoted as... The first category includes application or webpage type, timestamp, geolocation information, and the current network environment, used to characterize the external environment in which the ad request occurs. The third category is candidate ad features, denoted as... Including those currently in the recall queue The identifier information, industry category, and preset bid cap of each candidate ad are used to represent the set of candidate ads participating in the generation of the current advertising strategy. Among them, This indicates the number of candidate ads in the current recall queue.

[0066] After extracting the three types of features mentioned above, joint representation processing is performed. Joint representation processing involves incorporating multiple features from different sources that collectively describe the same ad request into a unified encoding process. This ensures that user attribute information, environmental information, and candidate ad information are no longer isolated but form a holistic feature representation of the current ad request within the same feature representation framework. Due to user-side features... Contextual features and candidate ad features Significant differences exist in the sources, dimensions, and representations of features, and the original features are typically encoded using one-hot encoding or other sparse representations. Directly processing the original features can lead to excessive computational overhead and make it difficult to form a unified state representation. Therefore, a pre-trained feature compression network is set up to perform manifold space mapping on the jointly represented features. The feature compression network adopts a multilayer perceptron structure to map the input high-dimensional sparse feature vectors into low-dimensional dense vector representations.

[0067] In some implementations, the feature compression network includes an input layer, at least one hidden layer, and an output layer. The input layer receives an input feature set, the hidden layer performs nonlinear transformations and feature compression, and the output layer outputs a low-dimensional state vector. The activation function of the hidden layer may be a modified linear unit function.

[0068] Through the above structure, high-dimensional discrete features are transformed into continuous low-dimensional representations, allowing the original features to fall into a lower-dimensional, more compact representation space while preserving the main semantic relationships. This representation space is the low-dimensional manifold space. Here, manifold space mapping refers to using a feature compression network to map the original high-dimensional feature set to a low-dimensional continuous vector space that can preserve the main structural information of the current ad request.

[0069] The mapping function of the feature compression network is expressed as: ;in, Represents a feature compression network; This represents the network parameters of the feature compression network; This represents the set of input features consisting of user-side features, contextual features, and candidate ad features; This represents the initial state vector of the current ad request in the low-dimensional manifold space. ,in, Represents the set of real numbers. Indicates by Composed of real numbers 3D real vector space, The vector dimension representing the low-dimensional manifold space is used to characterize the dimension of the compressed state vector. In some implementations, 256 is a suitable value. A pre-trained feature compression network refers to a network trained before entering the instant advertising strategy generation process using historical advertising request samples. This enables the network to map multi-source heterogeneous features to a unified low-dimensional representation space. In some implementations, historical advertising request samples include user-side features, contextual features, and candidate advertising features corresponding to historical advertising requests. The training process can employ a supervised training method based on prediction tasks, allowing the feature compression network to learn parameters suitable for representing the advertising request state. During the processing of the current advertising request, the parameters... It remains fixed, only performing forward mapping computation, and does not participate in gradient updates under the current request.

[0070] Through the above processing, the multi-source heterogeneous features in the ad request are uniformly mapped to the initial state vector in the low-dimensional manifold space. Initial state vector This comprehensive representation, which integrates user attributes, context, and candidate ad information corresponding to the current ad request, can serve as a unified input representation of the current ad request state. This approach maintains the integrity of the state representation while avoiding the excessively high computational complexity associated with directly processing high-dimensional sparse features.

[0071] Step S102: Load the generative world model and introduce the policy network, initialize the low-rank adapter parameters in the policy network, and construct a shadow simulation field for the current ad request.

[0072] In step S102, the initial state vector obtained in step S101... This represents the overall state of the current ad request, but it only reflects the current request's representation in a low-dimensional manifold space and cannot be directly used for policy exploration and parameter updates in the absence of real feedback. Therefore, in this step, it is necessary to consider the initial state vector. A virtual environment capable of predicting the outcome of an action is constructed, and transient adaptation parameters that can be updated quickly are introduced into the policy network, enabling the current advertising request to complete state deduction and policy exploration in the virtual environment before the real feedback arrives.

[0073] Specifically, first, the generative world model is loaded, denoted as... Generative world models are pre-trained time-series prediction models used to describe the relationship between state transitions and reward changes during advertising decision-making. The input to a generative world model is the current state vector. With the proposed action The output is the predicted state for the next time step. and predicted rewards Its mathematical expression is:

[0074] ;

[0075] in, Represents a generative world model; Indicates the first The current state vector corresponding to the next deduction; Indicates the state The following actions are to be performed; Indicates action The predicted state vector after execution; Representation and Action The corresponding predicted reward. In advertising scenarios, actions... This may include bid amount, creative mix, or other advertising strategy actions, and predicted rewards. This can be characterized by both click probability and conversion value. In some implementations, generative world models can employ time-series prediction models based on the Transformer architecture. By training on historical ad request samples, action sequences, and interaction results, they learn state transition patterns and reward change patterns, thereby outputting the corresponding predicted state and predicted reward given a state and action. The reason for using a generative world model here is to enable multiple rounds of state deduction based on the prediction results even before the current ad request receives real feedback.

[0076] After loading the generative world model, it is also necessary to determine the action generation method under each inference state. Generative World Model Generative world models are used to describe the state transitions and reward changes corresponding to states and actions, but they do not directly provide the action to be chosen in the current state. Therefore, a policy network is introduced, and the transient adaptation parameters in the policy network are initialized. The policy network is denoted as... It is used to generate an action distribution based on a state vector, where, These represent the main parameters of the policy network. Considering the main parameters... The parameters are large, and directly updating them online would incur high computational costs. Therefore, in this step, a low-rank adapter is bypassed and mounted in the critical linear layer of the policy network, and the parameters of this low-rank adapter are denoted as... , As a transient adaptation parameter that can be partially updated during the current ad request processing and reset after the request ends, it is used to make local adjustments to the output of the policy network, while the main parameter Remain constant during the current ad request processing.

[0077] In some implementations, the low-rank adapter Includes two matrices and ,in, , . Represents the set of real numbers. Indicates by A matrix space consisting of n real numbers Indicates by A matrix space consisting of n real numbers; The dimension of the state vector. Let represent the rank of the low-rank adapter, and By representing the low-rank adapter as the product of two low-dimensional matrices, the number of parameters requiring updates can be significantly reduced while maintaining parameter tuning capabilities. At the initial request time, Initialize to zero or Gaussian noise, while the main parameters Keep frozen. The policy output after introducing the low-rank adapter can be expressed as:

[0078] ;

[0079] in, This represents the policy output after introducing transient adaptation parameters; Indicates only by the main parameter The generated basic policy output; Represents the input state vector; This represents the local correction amount applied by the low-rank adapter to the output of the basic policy. This configuration allows the policy network to maintain its original general decision-making capabilities while possessing a parameter base for rapid local adjustments based on the current ad request.

[0080] By loading a generative world model And initialize transient adaptation parameters Together with the policy network, they constitute a shadow simulation field. A shadow simulation field refers to a computational environment constructed outside the real advertising environment, with an initial state vector... Starting with a generative world model This section describes the state transition process, using a policy network to generate action distributions. This enables the deduction of states for different advertising policy actions and the generation of corresponding predicted rewards without relying on real user feedback. The output of this step includes a generative world model. Initialized transient adaptation parameters And by the main parameter With transient adaptation parameters The policy network is formed by these two entities. This configuration ensures that the initial state vector obtained in step S101... Instead of remaining at the static representation level, it enters the shadow simulation field where state evolution, action probing, and parameter adjustment are possible.

[0081] Step S103: Perform an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate the trajectory set corresponding to the current advertising request.

[0082] In step S103, within the shadow simulation field constructed in step S102, the initial state vector obtained in step S101 is used. As the root node of the search tree, an extremum-oriented Monte Carlo tree search is performed to discover the strategy path that may yield higher returns for the current ad request. Step S102 has already provided the generative world model. and includes low-rank adapter parameters The policy network enables prediction of state evolution and action effects even without providing real feedback. This step builds upon this foundation by generating correlation trajectories between states, actions, and predicted rewards through multiple simulations, thus characterizing the potential performance of different policy paths under the current ad request.

[0083] Specifically, with the initial state vector As the root node of the search tree, execute The simulation cycle is repeated several times, in which... This represents the total number of forward-looking deductions performed around the current ad request. In each simulation loop, the state vector is... Starting from the policy network in the state vector The generated action distribution determines the action. and the state vector With action Input Generative World Model The corresponding predicted state vector is obtained. and predicted rewards Then, the predicted state vector will be... As the current state for the next step of the simulation, continue with action selection and state simulation until the preset maximum simulation depth is reached. .in, Indicates the first The corresponding state vector is derived step by step. Indicates the first The selected action, Indicates the execution of an action The resulting predicted state vector This indicates the corresponding predicted reward. This indicates the maximum number of simulation layers in a single simulation cycle.

[0084] During the search process, for each node, the action to be prioritized for expansion needs to be determined among multiple candidate actions. To enable the search process to prioritize the action branch that may generate higher returns from the multiple strategy paths corresponding to the current ad request, an extreme value-oriented node selection method is used to determine the child node actions. For any node... Its child node actions The selection criteria are expressed as follows:

[0085] ;

[0086] in, Indicates at node The next selected action; This indicates that the action that maximizes the value of the node selection function is selected from all candidate actions. The value of the node selection function is determined by the maximum reward item and the exploration item corresponding to the action. This indicates that during the historical simulation process, the nodes... Select Action The maximum reward value observed later. This represents the exploration coefficient, used to adjust the weight of prior probability and number of visits in node selection; Indicates at node The action output by the policy network in the corresponding state The prior probability; Represents a node Next action Number of visits; Represents a node The total number of visits for all candidate actions, of which, Represents a node The candidate action index is used. Unlike node selection based on average reward, The maximum reward value of the action branch in the historical simulation is represented, so that the search process prioritizes the action branches that have had high reward results in the past, thus making it more likely to discover potential high-value strategy paths.

[0087] The above node selection method is used in the state vector corresponding to the current node. Next, determine the actions based on the action distribution generated by the policy network. Specifically, the policy network provides probability information for each candidate action in the current state. The node selection formula determines the action to be prioritized for expansion by combining the historical maximum reward value, prior probability, and access count of each candidate action. And determine the obtained action With state vector Input together into the generative world model Perform state simulation.

[0088] After completion After each simulation cycle, the states, actions, and predicted reward sequences formed in each simulation cycle are collected to form a trajectory set. The trajectory is recorded as , , is represented as: ;in, Indicates the first Trajectory generated by each simulation cycle; Indicates the first The state vector corresponding to the step; Represents the initial state vector The action selected is the first action determined by the policy network in the initial state corresponding to the current ad request; Represents the initial state vector Next action Subsequently, the corresponding reward predicted by the generative world model is used to characterize the expected return of this action in the current initial state of the ad request. Indicates the first The action chosen in the step; Indicates the first The predicted reward corresponding to each step; This indicates the maximum depth of trajectory projection. (From all) The trajectories generated by each simulation cycle constitute a trajectory set. , This represents the set of trajectory data generated in the shadow simulation field around the current ad request.

[0089] Through the above processing, the initial state vector is simulated in the shadow simulation field. A set of trajectories was generated. Trajectory set It contains multiple paths connecting states, actions, and predicted rewards, used to characterize the potential returns of different strategy paths under the current ad request.

[0090] Step S104: Construct an entropy-oriented objective function based on the trajectory set, and perform transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request.

[0091] In step S104, based on the trajectory set obtained in step S103 Construct an entropy-oriented objective function and apply it to the low-rank adapter parameters. Perform transient gradient updates to make the policy network form a more concentrated action selection tendency in the state space corresponding to the current ad request. The trajectory set obtained in step S103 The potential returns of different strategy paths under the current ad request have been characterized, but this trajectory set itself is still a deduction result in the shadow simulation field and has not yet been transformed into a strategy adjustment result at the parameter level. Therefore, in this step, it is necessary to utilize the trajectory set. The state, action, and reward information contained therein are used to construct an optimization objective, which is then applied to the low-rank adapter parameters. This allows the policy network to make local adaptive adjustments to the current ad request.

[0092] Specifically, firstly based on trajectory sets Constructing an entropy-oriented objective function Entropy-oriented objective function Represented as:

[0093] ;

[0094] in, This represents the entropy-oriented objective function; This indicates the parameter to be updated, which corresponds to the low-rank adapter parameter in this step. ; Representing the trajectory From the set of trajectories Sampling; Represents the set of trajectories Calculate the expectation of the samples above; Representing the trajectory The corresponding weights; Indicates the state Below, the parameters are Policy network selects actions The logarithmic probability, Represents the entropy adjustment coefficient, and ; The information entropy represents the distribution of output actions of the policy network. The objective function consists of two parts: a weighting term and an entropy term. The weighting term is used to enhance the influence of high-reward trajectories on parameter updates, while the entropy term is used to constrain the dispersion of the policy distribution, so that the output of the policy network gradually becomes more concentrated from a relatively dispersed state.

[0095] In the objective function described above, the trajectory weights The exponential weighting function is used to determine this, and its expression is as follows:

[0096] ;

[0097] in, Represents an exponential function; Representing the trajectory Cumulative rewards; Represents the set of trajectories The average cumulative reward for each trajectory; This represents the reward scaling factor, used to adjust the amplification of trajectory weights when the cumulative reward deviates from the mean. In some implementations, Based on the trajectory set The distribution of cumulative rewards is pre-set, or it can be determined based on the statistical dispersion of the cumulative rewards for each trajectory. For example, it can be a function value of the standard deviation or variance of the cumulative rewards, or a preset proportion, to ensure numerical stability during the exponential weighting process and prevent excessive amplification of the weight of high-reward trajectories. Through this weighting method, trajectories with cumulative rewards above the average level will receive greater weight, while trajectories with lower cumulative rewards will receive less weight, thus making the influence of high-reward trajectories in the objective function stronger. (Trajectory Cumulative Rewards) It can be obtained by summing the predicted rewards at each step in the trajectory, that is: ;in, The trajectory is in the th order. The predicted reward corresponding to each step; Indicates the sequence number of the deduction step in the trajectory. This represents the maximum extrapolation depth of a single trajectory. Entropy term. This is used to characterize the uncertainty of the action distribution output by the policy network. In this step, the entropy term is used as a penalty term in the objective function calculation, which reduces the information entropy of the action distribution during the parameter update process of the policy network. This causes the action probabilities to gradually concentrate on high-reward actions, rather than remaining scattered among multiple candidate actions.

[0098] Constructing an entropy-oriented objective function Then, the objective function is calculated with respect to the low-rank adapter parameters. The gradient of . This gradient is denoted as . Its expression is:

[0099] ;

[0100] in, Represents the entropy-oriented objective function Regarding low-rank adapter parameters The gradient; Indicates the parameter Find the partial derivative. The transient gradient update here refers to updating the gradient using the trajectory set only during the processing of the current ad request. For low-rank adapter parameters Perform partial updates, while the main parameters of the policy network... These parameters remain fixed and do not participate in gradient calculations or parameter updates in this step. Therefore, the low-rank adapter parameters... This constitutes the only parameter object that is updated under the current ad request, and the update is only valid within the lifecycle of the current request.

[0101] After obtaining the gradient Then, a stochastic gradient descent optimizer or an Adam optimizer is used to optimize the low-rank adapter parameters. implement Step update, the updated low-rank adapter parameters are denoted as The update process is represented as follows: ;in, This represents the low-rank adapter parameters before the update. This represents the learning rate. In some implementations, the learning rate... The value is set higher than the learning rate during offline training to enhance the speed of local adaptation under the current ad request, for example, 0.01. After gradient update, the low-rank adapter parameters The parameter values ​​were adjusted to better match the current distribution of ad request states, making the policy network more biased towards selecting the set of trajectories within the state space corresponding to the current request. Actions that correspond to high-return paths.

[0102] Through the above processing, the trajectory set obtained in step S103 is... Transform into low-rank adapter parameters The updated results transform the policy network from a general decision-making pattern in its initial state to a locally adaptive decision-making pattern specific to the current ad request. This step utilizes the trajectory set only during the processing of the current ad request. For low-rank adapter parameters Perform local gradient updates without changing the policy network's main parameters. Therefore, a transient gradient injection mechanism for the current ad request is formed, enabling the policy network to quickly adjust local parameters around the state space corresponding to the current ad request during the inference phase. On the other hand, through the optimization objective composed of a weighted method based on trajectory cumulative reward and entropy constraints, the policy network not only shifts towards the region corresponding to high-reward actions during parameter updates, but also reduces the information entropy of the action distribution, causing the action probability to concentrate on advantageous actions, thus forming an entropy-oriented search and update mechanism. This setting enables the policy network to more quickly capture potential high-reward actions in the context of the current ad request when faced with an ad request that deviates from the historical main sample distribution, and improves the targeting of the final ad strategy generation results.

[0103] Step S105: Based on the updated policy network, perform final forward inference on the initial state vector to generate an advertising policy, and reset the low-rank adapter parameters after the advertising policy is executed.

[0104] In step S105, the updated low-rank adapter parameters obtained in step S104 are... and policy network master parameters The updated policy network is invoked, and the initial state vector obtained in step S101 is used. Input the policy network and perform a final forward inference to obtain the action probability distribution corresponding to the current ad request. Step S104 has already enabled the policy network to complete local adaptive adjustment in the state space corresponding to the current ad request by updating the transient gradient of the low-rank adapter parameters. Therefore, in this step, multiple rounds of search or parameter updates are not performed. Instead, the updated policy network is used directly to output the final decision basis.

[0105] Specifically, the updated policy network receives the initial state vector. Then, the probability distribution of each candidate action in the current state is output. Based on this probability distribution, the action with the highest probability among all candidate actions is selected as the target action, and the action content corresponding to the target action is parsed into an advertising strategy, thereby generating the final advertising strategy corresponding to the current advertising request. Here, the target action represents the action selection that is most likely to bring higher returns in the current advertising request state, and the advertising strategy is the specific delivery decision content corresponding to the target action, including the bid amount, creative combination, and other control variables used to describe the advertising delivery method.

[0106] After obtaining the final advertising strategy, it is sent to the advertising exchange system for execution, thus completing the strategy generation and delivery decision for this advertising request. After the advertising strategy execution is complete, the low-rank adapter parameters are... A reset process is performed to restore it to its initial state. In step S104, the low-rank adapter parameters are only partially updated for the current ad request. Their values ​​are associated with the user-side features, context features, and candidate ad features corresponding to the current request. Therefore, if not reset, they will interfere with subsequent ad requests. In some implementations, the low-rank adapter parameters are allocated in temporary storage space, and the corresponding storage resources are released after the current ad request is processed, thereby ensuring that the parameters between different ad requests are independent.

[0107] Furthermore, this step records the actual interaction results corresponding to this ad request. Actual interaction results include user feedback on the ad, such as whether a click occurred, whether a conversion took place, and the corresponding conversion value. These actual interaction results are stored as real reward data and used in the offline phase to process the generative world model. and policy network master parameters The model is updated so that its state transition prediction and policy generation capabilities can be gradually improved during subsequent operation.

[0108] Through the above processing, the policy network after transient gradient update in step S104 is used to generate the final advertising policy. After execution, the low-rank adapter parameters are reset, and real interaction data is recorded to support offline updates. This setup ensures that each advertising request completes policy generation in an independent low-rank adapter parameter environment, guaranteeing the adaptability of the current request while avoiding interference with the main parameters of the policy network. This achieves an advertising policy generation mechanism that combines online local adaptation with offline global optimization.

[0109] In summary, combining Figure 2 This application proposes an instant advertising strategy generation method based on transient gradient injection and entropy-guided search. First, an advertising request is acquired. User-side features, contextual features, and candidate advertising features in the advertising request are jointly represented and mapped to a manifold space to obtain an initial state vector representing the current advertising request state. Then, a generative world model is loaded and introduced into a policy network. The low-rank adapter parameters in the policy network are initialized to form a shadow simulation field oriented towards the current advertising request. Next, an extremum-oriented Monte Carlo tree search is performed in the shadow simulation field with the initial state vector as the root node to generate a trajectory set corresponding to the current advertising request. Subsequently, an entropy-guided objective function is constructed based on the trajectory set, and transient gradient updates are performed only on the low-rank adapter parameters, enabling the policy network to form a locally adaptive decision-making capability within the state space corresponding to the current advertising request. Finally, based on the updated policy network, a final forward inference is performed on the initial state vector to determine the target action and generate an advertising strategy. After the advertising strategy is executed, the low-rank adapter parameters are reset, and the actual interaction results are recorded for offline updates of the generative world model and the main parameters of the policy network.

[0110] As can be seen from the above scheme, this application does not directly use a fixed policy network obtained through offline training to perform static inference on the current advertising request. Instead, it first performs state deduction, action probing, and trajectory generation in a shadow simulation field around the initial state vector corresponding to the current advertising request. Then, it uses transient gradient updates to enable the policy network to complete local adaptation around the advertising request. Since long-tail user scenarios that deviate from the main distribution of historical training data are difficult to accurately characterize using a fixed model based on existing high-frequency sample distributions, this application generates a separate trajectory set for the current advertising request and uses high-reward paths in the trajectory to adjust the low-rank adapter parameters in real time. This allows the policy network to no longer simply follow the average decision tendency formed under the main historical distribution during the processing of the current request. Instead, it can combine the feature combination corresponding to the current long-tail user, the context, and candidate advertising information to form an action selection result that is more in line with the request. At the same time, through entropy-guided search and updates, the action distribution is further shifted towards the action set that is more likely to generate high rewards under the current advertising request. Thus, in long-tail user scenarios that deviate from the main distribution of historical training data, the responsiveness and discriminative ability of the advertising strategy to the features of the current request can be improved, thereby enhancing the targeting of the advertising strategy generation.

[0111] Figure 3 This is a structural block diagram of an instant advertising strategy generation system based on transient gradient injection and entropy-guided search, provided in one embodiment of this application. The system includes at least the following modules:

[0112] The request acquisition module is used to acquire advertising requests, perform manifold space mapping on user-side features, context features and candidate advertising features in the advertising requests, and obtain an initial state vector representing the current advertising request state.

[0113] The simulation building module is used to load the generative world model and introduce the policy network, initialize the low-rank adapter parameters in the policy network, and build a shadow simulation field for the current ad request.

[0114] The trajectory generation module is used to perform an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate a set of trajectories corresponding to the current ad request.

[0115] The network update module is used to construct an entropy-oriented objective function based on the trajectory set and perform transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request.

[0116] The policy generation module is used to generate an advertising policy by performing final forward inference on the initial state vector based on the updated policy network, and to reset the low-rank adapter parameters after the advertising policy is executed.

[0117] For relevant details, please refer to the above method implementation examples.

[0118] Figure 4 This is a block diagram of an electronic device provided in one embodiment of this application. The device includes at least a processor 401 and a memory 402.

[0119] Processor 401 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0120] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the instant advertising strategy generation method based on transient gradient injection and entropy-guided search provided in the method embodiments of this application.

[0121] In some embodiments, the electronic device may also optionally include: a peripheral device interface and at least one peripheral device. The processor 401, memory 402, and peripheral device interface can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface via a bus, signal line, or circuit board. Indicatively, peripheral devices include, but are not limited to: radio frequency circuits, touch displays, audio circuits, and power supplies.

[0122] Of course, electronic devices may also include fewer or more components, and this embodiment does not limit this.

[0123] Optionally, this application also provides a computer-readable storage medium storing a program that is loaded and executed by a processor to implement the instant advertising strategy generation method based on transient gradient injection and entropy-guided search in the above method embodiments.

[0124] Optionally, this application also provides a computer product including a computer-readable storage medium storing a program, which is loaded and executed by a processor to implement the instant advertising strategy generation method based on transient gradient injection and entropy-guided search in the above method embodiments.

[0125] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0126] The above embodiments merely illustrate several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for generating real-time advertising strategies based on transient gradient injection and entropy-guided search, characterized in that, The method includes: Obtain an ad request, and perform manifold space mapping on the user-side features, context features, and candidate ad features in the ad request to obtain an initial state vector representing the current ad request state; Load the generative world model and introduce it into the policy network, initialize the low-rank adapter parameters in the policy network, and construct a shadow simulation field for the current ad request, including: Loading Generative World Model The input to the generative world model is the current state vector. With the proposed action The output is the predicted state for the next time step. and predicted rewards Its mathematical expression is: ; Introducing policy networks The transient adaptation parameters in the policy network are initialized, where, Let represent the principal parameters of the policy network; a low-rank adapter is bypassed at the critical linear layer of the policy network, and the parameters of the low-rank adapter are denoted as . , As a transient adaptation parameter that can be partially updated during the current ad request processing and reset after the request ends; low-rank adapter Includes two matrices and ,in, , ; Represents the set of real numbers. Indicates by A matrix space consisting of n real numbers Indicates by A matrix space consisting of n real numbers; The dimension of the state vector. Let represent the rank of the low-rank adapter, and The policy output after introducing a low-rank adapter can be expressed as: ; in, Indicates only by the main parameter The generated basic policy output; Represents the input state vector; This represents the amount of local correction applied by the low-rank adapter to the output of the basic policy; By loading a generative world model And initialize transient adaptation parameters Together with the policy network, they form a shadow simulation field; Based on the initial state vector, perform an extremum-oriented Monte Carlo tree search in the shadow simulation field to generate a trajectory set corresponding to the current ad request; An entropy-oriented objective function is constructed based on the trajectory set, and transient gradient updates are performed on the low-rank adapter parameters to obtain a policy network for the current ad request. The initial state vector is used to generate an advertising strategy based on the updated policy network, and the low-rank adapter parameters are reset after the advertising strategy is executed.

2. The instant advertising strategy generation method based on transient gradient injection and entropy-guided search according to claim 1, characterized in that, The process of mapping the user-side features, context features, and candidate ad features in the ad request to a manifold space to obtain an initial state vector representing the current ad request state includes: Extracting user-side features from ad requests Contextual features and candidate ad features After feature extraction, joint characterization processing is performed; A pre-trained feature compression network is set up to perform manifold space mapping on the jointly represented features. The feature compression network adopts a multilayer perceptron structure to map the input high-dimensional sparse feature vector into a low-dimensional dense vector representation. The mapping function of the feature compression network is expressed as: ;in, Represents a feature compression network; This represents the network parameters of the feature compression network; This represents the set of input features consisting of user-side features, contextual features, and candidate ad features; This represents the initial state vector of the current ad request in the low-dimensional manifold space; ,in, Represents the set of real numbers. Indicates by Composed of real numbers 3D real vector space, This represents the vector dimension of a low-dimensional manifold space.

3. The instant advertising strategy generation method based on transient gradient injection and entropy-guided search according to claim 1, characterized in that, The step of performing an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate the trajectory set corresponding to the current advertisement request includes: With the initial state vector As the root node of the search tree, execute Each simulation cycle, in each simulation cycle, starts from the current state vector. Starting from the policy network in the state vector The generated action distribution determines the action. and the state vector With action Input Generative World Model The corresponding predicted state vector is obtained. and predicted rewards ; predict the state vector As the current state for the next step of the simulation, continue with action selection and state simulation until the preset maximum simulation depth is reached. ; For any node Its child node actions The selection criteria are expressed as follows: ; in, This means selecting the action that maximizes the value of the node selection function from all candidate actions. This indicates that during the historical simulation process, the nodes... Select Action The maximum reward value observed later. Indicates the exploration coefficient; Indicates at node The action output by the policy network in the corresponding state The prior probability; Represents a node Next action Number of visits; Represents a node The total number of visits for all candidate actions, of which, Represents a node The candidate action index below; After completion After each simulation cycle, the states, actions, and predicted reward sequences formed in each simulation cycle are collected to form a trajectory set.

4. The instant advertising strategy generation method based on transient gradient injection and entropy-guided search according to claim 3, characterized in that, The construction of the entropy-oriented objective function based on the trajectory set includes: Based on trajectory set Constructing an entropy-oriented objective function Entropy-oriented objective function Represented as: ; in, This indicates the parameter to be updated, corresponding to the low-rank adapter parameter. ; Representing the trajectory From the set of trajectories Sampling; Represents the set of trajectories Calculate the expectation of the samples above; Representing the trajectory The corresponding weights; Indicates the state Below, the parameters are Policy network selects actions The logarithmic probability, Represents the entropy adjustment coefficient, and ; This represents the information entropy of the action distribution output by the policy network; The trajectory weight The exponential weighting function is used to determine this, and its expression is as follows: ; in, Represents an exponential function; Representing the trajectory Cumulative rewards; Represents the set of trajectories The average cumulative reward for each trajectory; This represents the reward scaling factor, or the cumulative reward. It is obtained by summing the predicted rewards at each step in the trajectory, that is: ;in, The trajectory is in the th order. The predicted reward corresponding to each step; Indicates the sequence number of the deduction step in the trajectory. This indicates the maximum extrapolation depth of a single trajectory.

5. The instant advertising strategy generation method based on transient gradient injection and entropy-guided search according to claim 4, characterized in that, The step of performing transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request includes: Constructing an entropy-oriented objective function Then, the objective function with respect to the low-rank adapter parameters is calculated. gradient Its expression is: ; in, Indicates the parameter Find the partial derivative; Obtain the gradient Then, a stochastic gradient descent optimizer or an Adam optimizer is used to optimize the low-rank adapter parameters. implement Step update, the updated low-rank adapter parameters are denoted as The update process is represented as follows: ;in, This represents the low-rank adapter parameters before the update. This represents the learning rate.

6. The instant advertising strategy generation method based on transient gradient injection and entropy-guided search according to claim 1, characterized in that, The step of generating an advertising strategy by performing final forward inference on the initial state vector based on the updated policy network, and resetting the low-rank adapter parameters after the advertising strategy is executed, includes: After receiving the initial state vector, the updated policy network outputs the probability distribution of each candidate action in the current state. Based on the probability distribution, it selects the action with the highest probability among all candidate actions as the target action, and parses the action content corresponding to the target action into an advertising strategy to generate the final advertising strategy corresponding to the current advertising request. After obtaining the final advertising strategy, the advertising strategy is sent to the advertising exchange system for execution, thus completing the strategy generation and placement decision for this advertising request; After the advertising strategy is executed, the low-rank adapter parameters are reset to their initial state.

7. A real-time advertising strategy generation system based on transient gradient injection and entropy-guided search, characterized in that, include: The request acquisition module is used to acquire advertising requests, and perform manifold space mapping on the user-side features, context features and candidate advertising features in the advertising requests to obtain an initial state vector representing the current advertising request state. The simulation building module is used to load the generative world model and introduce it into the policy network, initialize the low-rank adapter parameters in the policy network, and construct a shadow simulation field for the current ad request, including: Loading Generative World Model The input to the generative world model is the current state vector. With the proposed action The output is the predicted state for the next time step. and predicted rewards Its mathematical expression is: ; Introducing policy networks The transient adaptation parameters in the policy network are initialized, where, Let represent the principal parameters of the policy network; a low-rank adapter is bypassed at the critical linear layer of the policy network, and the parameters of the low-rank adapter are denoted as . , As a transient adaptation parameter that can be partially updated during the current ad request processing and reset after the request ends; low-rank adapter Includes two matrices and ,in, , ; Represents the set of real numbers. Indicates by A matrix space consisting of n real numbers Indicates by A matrix space consisting of n real numbers; The dimension of the state vector. Let represent the rank of the low-rank adapter, and The policy output after introducing a low-rank adapter can be expressed as: ; in, Indicates only by the main parameter The generated basic policy output; Represents the input state vector; This represents the amount of local correction applied by the low-rank adapter to the output of the basic policy; By loading a generative world model And initialize transient adaptation parameters Together with the policy network, they form a shadow simulation field; The trajectory generation module is used to perform an extremum-oriented Monte Carlo tree search in the shadow simulation field based on the initial state vector to generate a trajectory set corresponding to the current advertising request. The network update module is used to construct an entropy-oriented objective function based on the trajectory set and perform transient gradient updates on the low-rank adapter parameters to obtain a policy network for the current ad request. The strategy generation module is used to generate an advertising strategy by performing final forward inference on the initial state vector based on the updated policy network, and to reset the low-rank adapter parameters after the advertising strategy is executed.

8. An electronic device, characterized in that, The device includes a processor and a memory; the memory stores a program, which is loaded and executed by the processor to implement a real-time advertising strategy generation method based on transient gradient injection and entropy-guided search as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The storage medium stores a program that, when executed by a processor, is used to implement a real-time advertising strategy generation method based on transient gradient injection and entropy-guided search as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Intelligent recommendation method for optimizing advertisement keyword combination through cross validation

    CN120705408A

  • Advertisement effect evaluation method and system based on artificial intelligence

    CN120746652A