Simple method and system for an agent to access an advertisement real-time bidding system

CN122760176APending Publication Date: 2026-09-15SHANGHAI 吖吖 NETWORK TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610950225.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-09-15

AI Technical Summary

Benefits of technology

本发明通过广告领域语义映射模型完成异构参数字段的自动对齐与标准化转换,可有效降低不同平台接口的适配成本,缩短智能体接入周期;通过内置经多智能体自博弈预训练的出价策略网络,能够为接入智能体直接提供成熟的基础竞价能力,免去智能体侧的算法开发工作,缓解新接入智能体的冷启动效果波动与预算浪费问题;结合约束最优反应动态修正与双阶段约束管控机制,可自动适配动态竞争环境并实现合规的预算消耗控制,无需智能体自行开发博弈适配与风控逻辑,整体降低了智能体接入广告实时竞价系统的技术门槛与落地成本,提升了接入效率与运行稳定性,支撑智能体竞价技术的规模化推广应用。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122760176A_ABST
    Figure CN122760176A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of programmed advertising and multi-agent system, and particularly relates to a simple method and system for accessing an advertising real-time bidding system by an agent, comprising the following steps: obtaining an access request and unstructured configuration parameters, completing field alignment and standardized conversion through an advertising field semantic mapping model, and establishing a communication link; receiving a real-time bidding exposure request, extracting four types of state features, inputting a bidding strategy network pre-trained through multi-agent self-game, and generating an initial bidding coefficient; based on an opponent bidding distribution collected through a sliding window, performing a constraint optimal response dynamic correction on the initial coefficient; according to a budget consumption benchmark curve generated through dynamic programming, performing a two-stage constraint control on the corrected coefficient, generating a final bid, and sending the final bid to a platform, and simultaneously feeding back the bidding result to the accessing agent. The present application effectively reduces the agent access threshold and cost, and improves the bidding efficiency and budget control stability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of programmatic advertising technology and multi-agent systems, specifically to a simple method and system for agents to access a real-time advertising bidding system. Background Technology

[0002] With the continuous development of the programmatic advertising industry, real-time bidding has become the mainstream trading mechanism for display and feed ads. A single impression request can complete traffic distribution, multi-party bidding, and auction settlement within milliseconds, making it the core method for current internet advertising resource allocation. Meanwhile, reinforcement learning-based bidding agents are gradually becoming a core technology for advertisers to optimize campaign performance. All kinds of self-developed or third-party developed bidding agents need to connect to the real-time bidding system to complete transactions. Currently, the industry has conducted extensive research on optimizing bidding strategies themselves, resulting in various dynamic bidding and budget optimization solutions based on reinforcement learning.

[0003] However, the lack of standardized and integrated intelligent agent access solutions in existing technologies results in high barriers to entry and low adaptation efficiency for bidding intelligent agents when integrating them into real-time advertising bidding systems. Different advertising exchange platforms have different interface protocols and parameter field formats, requiring heterogeneous bidding intelligent agents to undergo customized interface development and field adaptation for each platform, leading to long integration cycles and high development costs. Furthermore, bidding intelligent agents need to independently complete the entire technical chain, including bidding algorithm construction, competitive environment adaptation, and budget control logic development. Most small and medium-sized advertisers lack the corresponding R&D capabilities, and newly integrated intelligent agents generally suffer from large fluctuations in performance and significant budget waste during the cold start phase. This significantly increases the difficulty of implementing intelligent agent bidding technology and hinders its large-scale promotion and application. Summary of the Invention

[0004] To address the technical deficiencies in the background technology, this invention proposes a simple method and system for intelligent agents to access a real-time advertising bidding system, which solves the aforementioned technical problems and meets practical needs. The specific technical solution is as follows: A simplified method for intelligent agents to access a real-time advertising bidding system includes the following steps: The system obtains the access requests and unstructured configuration parameters of the bidding agents to be accessed, and completes automatic field alignment and standardization transformation through a pre-trained advertising domain semantic mapping model to generate standard configuration parameters and establish a standardized communication link. Receive real-time bidding exposure requests from advertising exchange platforms, extract four types of state features, input them into a bidding strategy network pre-trained by multi-agent self-game, and generate initial bidding coefficients; Based on the opponent's bid distribution data collected by the sliding window, the initial bid coefficient is dynamically corrected by constrained optimal response to obtain the corrected bid coefficient; Based on the budget consumption baseline curve generated by dynamic programming, a two-stage constraint control is applied to the corrected bidding coefficient to generate the final bid and send it to the advertising exchange platform. At the same time, the bidding results are synchronously fed back to the connected bidding agent.

[0005] Furthermore, the advertising domain semantic mapping model completes training and field mapping through the following steps: Collect historical interface field descriptions, bidding parameter descriptions, and industry placement standard corpora from the advertising transaction field to construct a domain training corpus. Use a general lightweight pre-trained language model as the initial weights and employ a contrastive learning approach for domain fine-tuning training. Extract the core parameter fields of the entire real-time bidding process, input them into the trained advertising domain semantic mapping model to generate corresponding semantic embedding vectors, and build a standard field semantic vector library; Receive the parameter fields to be mapped uploaded by the bidding agent to be connected, input the semantic mapping model of the advertising domain to generate the semantic embedding vector to be mapped, calculate the cosine similarity between the semantic embedding vector to be mapped and each standard field vector in the standard field semantic vector library, and sort them in descending order of similarity value to obtain the matching candidate list. Select the first standard field in the matching candidate list and determine whether its corresponding similarity value is higher than the preset similarity threshold. If so, automatically complete the mapping relationship between the parameter field to be mapped and the standard field, and perform data format conversion to generate standard configuration parameters. If not, push the top three standard fields in the matching candidate list to the access party for manual confirmation. The system receives manually confirmed matching results and feeds the confirmed mapping pairs back into the incremental training set of the semantic mapping model in the advertising domain for subsequent iterative optimization of the model.

[0006] Furthermore, the four types of state features include estimated traffic value, remaining budget status, historical average winning bid, and bidding participation scale. The specific steps for training the bidding strategy network and generating the initial bidding coefficients are as follows: Construct an adversarial self-game multi-agent simulation environment to simulate multiple bidding agents participating in a two-price sealed auction. Generate global state data including flow value, budget state, and competition scale. Input the actor network corresponding to each bidding agent and output the bidding action of each agent. A centralized commentator network is set up to receive global state data and the joint bidding actions of all intelligent agents, output the global action value, and calculate the single-step reward value by combining the constraint perception composite reward rule; The backpropagation algorithm is executed based on the single-step reward value and the global action value to iteratively update the parameters of the actor network and the centralized critic network until the network converges; After training, one of the actor networks is fixed as the bidding strategy network for online inference. The four types of state features extracted from real-time bidding exposure requests are input into the bidding strategy network, and the initial bidding coefficient is output.

[0007] Furthermore, the specific steps for constraining the optimal response dynamic correction of the initial bid coefficient are as follows: A sliding window is used to collect opponent winning bid samples for a near-preset number of rounds to construct an opponent bid sample set. By fitting the opponent's bid sample set using the kernel density estimation algorithm, the probability density function and cumulative distribution function of the opponent's bids are obtained; Under the two-price sealed auction mechanism, with the goal of maximizing the expected return in a single step, and combining the constraints of budget consumption rate and return on investment, the constrained optimal bid anchor point is obtained. The initial bid coefficient is weighted and adjusted using a dynamic adjustment factor to obtain the adjusted bid coefficient. The adjustment formula is as follows:

[0008] in, This is the adjusted bid coefficient; This is the initial bid coefficient; To constrain the optimal bid anchor point; It is a dynamic correction coefficient, calculated by the sigmoid function, with a value range of (0,1). Its value increases as the difference between the current number of bidders and the historical average number of bidders increases, and also increases as the standard deviation of the recent round's winning bid increases.

[0009] Furthermore, the specific steps for implementing two-stage constraint control on the modified bid coefficient are as follows: Collect historical traffic quality distribution data and competition intensity time series data, and use dynamic programming algorithm to allocate the total budget according to the traffic value density of each time period to generate a dynamic budget consumption baseline curve. Real-time collection of current cumulative consumption budget data, and calculation of the deviation between the current actual consumption progress and the dynamic budget consumption baseline curve; When the deviation value is less than or equal to the preset deviation threshold, the controlled bid coefficient is obtained by linear multiplication. The adjustment formula is as follows:

[0010] in, This refers to the bid coefficient after control. This is the adjusted bid coefficient; This is the penalty gain coefficient; This represents the deviation between the current actual consumption progress and the baseline curve. When the deviation value exceeds the preset deviation threshold, an exponential penalty adjustment is used to obtain the controlled bid coefficient, and the penalty intensity increases exponentially with the increase of the deviation value.

[0011] Furthermore, the two-stage constraint control also includes a hard circuit breaker control mechanism, as detailed below: Two hard blocking rules are preset: a single bid cap and a budget safety threshold. The bid amount and remaining budget percentage data corresponding to the bid coefficient after control are obtained in real time. The first-level single bid limit check is performed. If the bid amount exceeds the preset single bid limit, the bid amount is automatically truncated to the single bid limit value, and a bid limit interception prompt is sent to the bidding agent to be connected; if the limit is not exceeded, the second-level check is performed. When performing the second-level budget circuit breaker check, if the remaining budget percentage is lower than the preset safety threshold, the budget circuit breaker state is triggered, blocking all subsequent bidding requests; The operation status is managed by a circuit breaker state machine. After the budget circuit breaker is triggered, the state machine switches to the circuit breaker pause state and synchronizes the budget exhaustion pause signal to the bidding agent waiting to be connected. When a budget increase or campaign period reset is detected, the state machine automatically switches back to normal campaign state, reinitializes the dynamic budget consumption baseline curve, and restores normal bidding functionality.

[0012] Furthermore, during the generation of the initial bid coefficients, the bidding strategy network possesses a lightweight online adaptive mechanism, specifically as follows: During the pre-training phase of the bidding strategy network, the basic weights are solidified based on the publicly available real-time bidding dataset and historical campaign data from multiple industries and advertisers, and a general initial bidding coefficient is output. Once the bidding agent is officially deployed, the returned bidding results will be accumulated in real time to construct a bidding feedback dataset. When the cumulative number of bidding feedback datasets reaches the preset number of sample rounds, the corresponding state features and the optimal bid label are extracted to construct a fine-tuning training sample set. A small learning rate is used to incrementally update the parameters of the top fully connected layer of the bidding strategy network, while the weights of the bottom feature extraction network remain fixed. The updated network output is adapted to the initial bidding coefficient of the current delivery scenario.

[0013] Furthermore, the unstructured configuration parameters are deployment descriptions in natural language form, and standard configuration parameters are generated through the following steps: A natural language configuration entry is set up in the access layer to receive the text of the delivery target and constraint requirements in natural language form input by the access party, which is then input into the advertising domain semantic mapping model as the parameter field to be mapped. The semantic mapping model in the advertising field performs semantic parsing and entity extraction on natural language text, and identifies four core parameter information contained in the text: budget size, bid cap, target cost per click, and delivery time period. The identified core parameter information is automatically mapped to the standard field system, and the corresponding constraint thresholds and reward weights are matched to generate initial standard configuration parameters; The parameter configuration confirmation message is pushed to the access party. After the access party confirms, the final standard configuration parameters are generated and take effect.

[0014] A simplified system for intelligent agents to access a real-time advertising bidding system includes a memory, a host computer, and a computer program stored in the memory and executable on the host computer. The computer program is configured to implement the steps of the simplified method for intelligent agents to access a real-time advertising bidding system as described above.

[0015] Compared with existing technologies, the simplified method and system for intelligent agents to access a real-time advertising bidding system provided by this invention have the following beneficial effects: This invention achieves automatic alignment and standardization of heterogeneous parameter fields through a semantic mapping model in the advertising domain, effectively reducing the adaptation costs of different platform interfaces and shortening the agent access cycle. By incorporating a bidding strategy network pre-trained through multi-agent self-game theory, it directly provides mature basic bidding capabilities to connected agents, eliminating the need for algorithm development on the agent side and mitigating the cold start effect fluctuations and budget waste issues for newly connected agents. Combined with a constraint-optimal response dynamic correction and a two-stage constraint control mechanism, it can automatically adapt to dynamic competitive environments and achieve compliant budget consumption control, eliminating the need for agents to develop their own game adaptation and risk control logic. Overall, it lowers the technical threshold and implementation costs for agents to access real-time advertising bidding systems, improves access efficiency and operational stability, and supports the large-scale promotion and application of agent bidding technology. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating a simplified method for an intelligent agent to access a real-time advertising bidding system according to the present invention. Detailed Implementation

[0017] In the description of this invention, it should be understood that the terms "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "middle," and "inner," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. Furthermore, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this invention, it should be noted that unless otherwise explicitly specified and limited, the terms "installed," "connected," and "joined" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention through specific circumstances.

[0018] The embodiments of the present invention will be described below with reference to the accompanying drawings and related examples. The embodiments of the present invention are not limited to the following examples, and the present invention relates to the relevant necessary components in this technical field, which should be regarded as well-known technology in this technical field and can be known and mastered by those skilled in this technical field.

[0019] See Figure 1 This invention provides a simple method for intelligent agents to access a real-time advertising bidding system, comprising the following steps: Step S100: Obtain the access request and unstructured configuration parameters of the bidding agent to be accessed, and complete the automatic alignment and standardization transformation of fields through the pre-trained advertising domain semantic mapping model to generate standard configuration parameters and establish a standardized communication link; Specifically, the access layer receives access requests initiated by the bidding agent to be accessed via the HTTPS protocol, parses the unstructured configuration parameters in the request body that are custom-named by the access party, calls the pre-trained advertising domain semantic mapping model to semantically encode each input parameter field, generates corresponding semantic vectors, and calculates the cosine similarity with the standard field semantic vector library one by one to complete field semantic matching. After successful matching, it automatically performs unified conversion of the parameter data type, unit of measurement, and numerical range to generate standard configuration parameters that conform to the platform data specifications. After the parameters pass the verification, it establishes a long-connection communication link with the bidding agent to be accessed based on the standardized bidding interaction primitives, completes permission verification and heartbeat mechanism initialization, and realizes bidirectional stable transmission of bidding data and control signals. The bidding agent to be accessed refers to a program entity with autonomous bidding decision-making capabilities that is to be accessed by the real-time advertising bidding system to participate in traffic bidding. It can be developed by advertisers or third-party service providers and is used to represent advertisers in executing bidding decisions and management of advertising placement. Unstructured configuration parameters refer to ad placement configuration data that does not follow unified standards and specifications, and whose naming and data formats are customized by the access party. Examples include parameter fields such as total budget, maximum bid, and ad placement target, which are named differently by different vendors. The advertising domain semantic mapping model refers to a semantic recognition model that is fine-tuned and trained on advertising transaction corpora based on a general pre-trained language model. It has the ability to accurately identify the semantics of parameter fields in the advertising bidding domain and can achieve semantic alignment of fields with different names but the same meaning. Automatic field alignment refers to the process of matching parameter fields with different names but the same business meaning through semantic similarity calculation, eliminating field naming differences and achieving data semantic uniformity. Standardization conversion refers to the process of uniformly converting semantically aligned parameters into standard data types, units of measurement, numerical precision, and value ranges to conform to the platform's unified data specifications. Standard configuration parameters refer to ad placement configuration data that fully conforms to the platform's unified specifications after field alignment and format conversion, including core ad placement parameters such as total budget, bid cap, and ad placement target. A standardized communication link refers to a bidirectional data transmission channel established based on a unified interaction protocol, message format, and interaction timing. It is used to carry out stable transmission of information such as bidding requests, bidding results, feedback data, and control signals.

[0020] Step S200: Receive real-time bidding exposure requests from the advertising exchange platform, extract four types of state features, input them into the bidding strategy network pre-trained by multi-agent self-game, and generate initial bidding coefficients. Specifically, the platform receives real-time bidding exposure request messages from the advertising exchange platform through a standardized interface, and parses the estimated traffic value data carried in the messages. It synchronously reads the remaining budget status of the current agent, the historical average winning bid over a fixed number of rounds, and the bidding participation scale corresponding to this bidding request from its local cache. It then performs numerical normalization on the four types of features—estimated traffic value, remaining budget status, historical average winning bid, and bidding participation scale—and concatenates them into a fixed-dimensional state feature vector. This state feature vector is input into a bidding strategy network pre-trained through multi-agent self-game theory, and after forward inference through a multi-layer fully connected network, it outputs a continuous initial bidding coefficient. The advertising exchange platform is an intermediary trading platform that aggregates media exposure traffic resources, organizes multiple bidding entities to participate in bidding, executes bidding settlement, and distributes results. It is the core hub of real-time bidding advertising transactions. A real-time bidding exposure request refers to a bidding request message sent by the advertising exchange platform to each bidding participant when an advertising exposure opportunity arises on the terminal. This message carries core information such as traffic attributes, user tags, and estimated traffic value. State features are the core input features of the bidding decision model. In this scheme, they specifically include four categories: estimated traffic value, remaining budget state, historical average winning bid, and bidding participation scale, which together characterize the core state of the current bidding scenario. Multi-agent self-game refers to constructing multiple equal bidding agents in a simulated bidding environment and training the strategy network through multiple rounds of adversarial bidding, enabling the network to learn the training paradigm of optimal bidding rules under dynamic competitive conditions. The bidding strategy network is a strategy model built based on a deep neural network. It receives state feature inputs and outputs corresponding bidding coefficients, serving as the core computational unit for generating basic bidding decisions. The initial bidding coefficient is the basic multiplier parameter output by the bidding strategy network; the final bidding amount is calculated by multiplying this coefficient by the estimated traffic value.

[0021] Step S300: Based on the opponent's bid distribution data collected by the sliding window, the initial bid coefficient is dynamically corrected by constrained optimal response to obtain the corrected bid coefficient; Specifically, the platform maintains a fixed-length sliding window cache. After each round of bidding, the returned competitor winning bid data is updated in the sliding window, forming a timely matching competitor bid sample set. An independent background thread asynchronously fits the sample set using a kernel density estimation algorithm to obtain the probability density function and cumulative distribution function of the current market competitor bids. Combining the settlement rules of a two-price sealed auction, with the goal of maximizing the expected return in a single step, and superimposed with budget consumption rate constraints and return on investment constraints, the optimal bid anchor point under the current environment is obtained. Based on the difference between the current number of bidders and the historical average, and the standard deviation of the recent round winning bids, a dynamic correction coefficient is calculated to weight and correct the initial bid coefficient, outputting a corrected bid coefficient adapted to the current competitive environment. The sliding window is a time-sensitive data caching mechanism that only retains historical data from the most recent preset number of rounds to fit the market competition status of the current period, ensuring the data's matching degree with the current environment. Competitor bid distribution data refers to the statistical distribution data of bids from other competing entities participating in the same auction, used to quantitatively characterize the intensity of competition and bidding patterns in the current market. Constrained optimal response dynamic correction refers to an optimization mechanism that, based on the characteristics of competitors' bid distribution, solves for the optimal bid anchor point under constraints such as budget consumption rate and return on investment, and dynamically adjusts the initial bid to better adapt it to the current competitive environment. The corrected bid coefficient refers to the bid multiplier parameter after being adjusted to adapt to the competitive environment, which is more in line with the current market competition state compared to the initial bid coefficient.

[0022] Step S400: Based on the budget consumption baseline curve generated by dynamic programming, implement two-stage constraint control on the corrected bidding coefficient, generate the final bid and send it to the advertising exchange platform, and simultaneously feed back the bidding results to the connected bidding agent.

[0023] Specifically, based on historical traffic quality distribution and competition intensity time-series data, a dynamic programming algorithm is used to allocate the total budget according to the traffic value density of each time period, generating a dynamic budget consumption baseline curve and storing it locally. The current cumulative budget consumption deviation from the baseline curve is calculated in real time. First, soft constraint rules are used to smooth the adjusted bidding coefficient, and then hard interception rules are used to verify the bid cap and budget balance, resulting in a controlled bidding coefficient. The controlled bidding coefficient is multiplied by the estimated traffic value to calculate the final bid amount. The final bid is encapsulated into a standard message and sent to the ad exchange platform for bidding. After receiving the bidding success / failure status and actual deduction amount from the ad exchange platform, the results are simultaneously fed back to the bidding agent to be integrated, and the budget balance and historical winning bids in the local cache are updated for subsequent rounds of decision-making. Dynamic programming is an algorithm that breaks down a complex problem into multiple sub-problems and solves the globally optimal solution in stages; here, it is used to optimally allocate the total budget according to the traffic value density of each time period. The budget consumption baseline curve refers to the planned consumption progress curve of the budget in each period of the complete campaign cycle, serving as a benchmark for budget consumption rhythm control. Dual-stage constraint control refers to a budget control mechanism that includes two stages: soft constraint smooth adjustment and hard constraint circuit breaker interception. The soft constraint stage enables smooth adjustment of the consumption rhythm, while the hard constraint stage ensures the compliance baseline of the budget and bid. The final bid refers to the actual bid amount submitted to the ad exchange platform for bidding after the entire process of basic generation, competition correction, and constraint control. The bidding result refers to the transaction result data returned by the ad exchange platform after the bidding process ends, including core information such as bidding success / failure status, actual deduction amount, and the second highest bid.

[0024] In one embodiment of the present invention, the advertising domain semantic mapping model completes training and field mapping through the following steps: Step S101: Collect historical interface field descriptions, bidding parameter descriptions, and industry placement standard corpora in the advertising transaction field, construct a domain training corpus, use a general lightweight pre-trained language model as the initial weights, and use a contrastive learning approach for domain fine-tuning training. Specifically, a domain training corpus is constructed by collecting multi-source advertising transaction corpora. The corpus sources include public interface documents from mainstream advertising transaction platforms, field mapping logs from historical projects, parameter specification documents for programmatic advertising, and parameter description texts from advertising delivery systems. This covers descriptive corpora of all parameter fields across the entire chain, including budget, bidding, performance, and constraint categories. The raw corpus undergoes cleaning and preprocessing to remove redundant symbols and invalid text, and synonym and anisotropic field pairs are labeled as training samples. Using a general lightweight pre-trained language model as initial weights, domain fine-tuning is performed using a contrastive learning approach: synonymous parameter field pairs are used as positive samples, and semantically unrelated field pairs are used as negative samples. After inputting into the model, the network parameters are optimized using a contrastive loss function, ensuring that the semantic embedding vectors output by the model satisfy the characteristic of "synonymous field vectors being close in distance and anisotropic field vectors being far in distance." After training, a semantic mapping model adapted to the advertising bidding domain is obtained, balancing low inference latency and high domain recognition accuracy, and adapting to high-concurrency agent access scenarios.

[0025] Step S102: Extract the core parameter fields of the real-time bidding full link, input them into the trained advertising domain semantic mapping model to generate corresponding semantic embedding vectors, and build a standard field semantic vector library; Specifically, the core parameter fields of the entire real-time bidding process are identified, covering all core configuration fields such as total budget, daily budget, single bid cap, target cost per click, target ROI, and ad placement time. Each field is accompanied by standardized business definition description text. These standardized description texts are then sequentially input into a trained advertising domain semantic mapping model to generate corresponding fixed-dimensional semantic embedding vectors. All standard fields and their corresponding semantic vectors are stored in a vector database to construct a standard field semantic vector library. A vector index is also established to support millisecond-level similarity retrieval. The vector library supports dynamic expansion; when new parameter fields are added, corresponding semantic vectors are generated and updated in the vector library, ensuring that the field coverage matches the business iteration needs.

[0026] Step S103: Receive the parameter fields to be mapped uploaded by the bidding agent to be accessed, input the semantic mapping model of the advertising domain to generate the semantic embedding vector to be mapped, calculate the cosine similarity between the semantic embedding vector to be mapped and each standard field vector in the standard field semantic vector library, and sort them in descending order of similarity value to obtain the matching candidate list. Specifically, after receiving the parameter fields to be mapped uploaded by the bidding agent to be connected, the access layer first performs text preprocessing, including removing special characters, unifying text case, and performing Chinese word segmentation and normalization to eliminate the interference of format differences on semantic recognition. The preprocessed parameter fields to be mapped are then input into the advertising domain semantic mapping model to generate a semantic embedding vector of the same dimension as the standard fields. The similarity retrieval interface of the vector database is called to calculate the cosine similarity between the semantic embedding vector to be mapped and all vectors in the standard field semantic vector library. The similarity value is calculated by the cosine of the angle between the two vectors, and the value ranges from 0 to 1. The higher the value, the higher the semantic overlap between the two fields. All standard fields are sorted from high to low according to the similarity value to generate a matching candidate list. The list includes the similarity value and standard business definition of each candidate field.

[0027] Step S104: Select the first standard field in the matching candidate list and determine whether its corresponding similarity value is higher than the preset similarity threshold. If yes, automatically complete the mapping relationship binding between the parameter field to be mapped and the standard field, and perform data format conversion to generate standard configuration parameters. If no, push the top three standard fields in the matching candidate list to the access party for manual confirmation. Specifically, the system extracts the top-ranked standard field from the candidate matching list and its corresponding similarity value. This value is then compared to a preset similarity threshold, which is based on historical mapping accuracy and used to balance the efficiency and accuracy of automatic mapping. If the highest similarity value is higher than the preset threshold, the semantic matching is deemed reliable, and the system automatically establishes a mapping binding relationship between the parameter field to be mapped and the standard field. After binding, the system automatically performs data format conversion, including unifying the parameter's unit of measurement, data type, numerical precision, and value range. For example, it maps the access party's custom "daily maximum consumption" field to the standard field "daily budget," while unifying the numerical unit to yuan and the data type to floating-point, ultimately generating standard configuration parameters that conform to the platform specifications. If the highest similarity value is lower than the preset threshold, the automatic matching confidence is deemed insufficient, and the system pushes the top three standard fields from the candidate matching list to the access party's console interface, along with business descriptions and similarity values ​​for each field, for the access party to manually confirm the correct mapping fields.

[0028] Step S105: Receive the manually confirmed matching results and feed the confirmed mapping pairs back to the incremental training set of the semantic mapping model in the advertising domain for subsequent iterative optimization of the model.

[0029] Specifically, the system receives the manual confirmation result returned by the access party. If the access party selects the correct field from the candidate list or adds a field already existing in the standard field library, the mapping pair of "parameter field to be mapped - standard field" is marked as a positive sample and added to the incremental training set of the semantic mapping model in the advertising domain. When the cumulative number of samples in the incremental training set reaches a preset number, the model incremental fine-tuning process is triggered. The top-level parameters of the model are updated with a very small learning rate, so that the model gradually adapts to the semantic recognition of more custom fields, improving the accuracy and coverage of subsequent automatic mapping. During the incremental training process, the underlying feature extraction parameters of the model remain unchanged, which not only ensures the stability of the model's basic semantic capabilities, but also enables continuous iterative optimization, forming a closed-loop mechanism of "mapping-feedback-optimization".

[0030] In one embodiment of the present invention, the four types of state features include estimated traffic value, remaining budget status, historical average winning bid, and bidding participation scale. The specific steps for training the bidding strategy network and generating initial bidding coefficients are as follows: Step S201: Construct an adversarial self-game multi-agent simulation environment, simulate multiple bidding agents participating in a two-price sealed auction, generate global state data including flow value, budget state, and competition scale, input the actor network corresponding to each bidding agent, and output the bidding action of each agent. Specifically, the simulation environment constructs a real traffic sample pool based on a publicly available real-time bidding dataset, reconstructing the value distribution, arrival rate, and fluctuation patterns of traffic in a time-series manner to ensure the similarity between the simulation scenario and the real-world advertising environment. The environment pre-defines multiple equivalent bidding agent roles, with all agents sharing the same action and state spaces, simulating a game scenario in real-world advertising where multiple advertisers compete for traffic. During each simulation iteration, the environment samples a single exposure traffic sample from the traffic pool, generating global state data containing the estimated value of a single traffic item, the remaining budget percentage for each agent, the total number of agents participating in the current round of bidding, and historical winning bid statistics. The local observation portion of the global state data corresponding to each agent is then input into an independent actor network for each agent. The actor network employs a multi-layer fully connected network structure, using a linear rectified function as the activation function. The output layer uses a hyperbolic tangent activation function to map the output to a preset continuous value range, ultimately outputting the bidding coefficient corresponding to each agent. Combined with the estimated traffic value, this yields the bidding action for each agent. The environment executes bidding and settlement according to the two-price sealed auction rules, determines the winning intelligent agent and the actual deduction amount, and completes a single round of bidding interaction.

[0031] Step S202: Set up a centralized commentator network to receive global state data and the joint bidding actions of all agents, output the global action value, and calculate the single-step reward value by combining the constraint-aware composite reward rule; Specifically, a multi-agent reinforcement learning architecture with centralized training and decentralized execution is adopted, setting up a single centralized reviewer network. This network is a fully connected value network, with the input being a joint vector formed by concatenating the global state vector and the bidding actions of all agents, and the output being the action value corresponding to the global joint action. This is used to provide a global perspective on value assessment during the training phase, solving the training instability problem caused by the non-stationarity of the multi-agent environment. Simultaneously, a constraint-aware composite reward rule is used to calculate the single-step reward value of each agent. The reward rule consists of four parts: a revenue component, a cost component, a budget constraint penalty component, and an ROI constraint penalty component. Successful agents receive click-through conversion revenue, deducting actual costs; unsuccessful agents receive neither revenue nor costs. The budget constraint penalty component is calculated based on the deviation between the agent's current budget consumption progress and the time schedule, applying an increasing penalty when consumption exceeds the target, guiding the network to smoothly consume the budget. The ROI constraint penalty component is calculated based on the difference between the cumulative actual ROI and the target ROI, applying a penalty proportional to the cost when the ROI is lower than the target, guiding the strategy to balance revenue and cost. The reward value for each agent is used to update the corresponding actor network, while the global action value is used to update the critic network.

[0032] Step S203: Execute the backpropagation algorithm based on the single-step reward value and the global action value, and iteratively update the parameters of the actor network and the centralized critic network until the network converges; Specifically, the training process employs an experience replay mechanism, storing the global state, joint actions, reward values, and next-time global state generated in each round of interaction into an experience replay pool. During training, batch samples are randomly sampled to update the network, breaking the temporal correlation of samples and improving training stability. Simultaneously, a dual-objective network architecture is adopted, with separate target actor and target critic networks. The parameters of the main network are periodically synchronized to the target networks to mitigate value estimation bias during training. When updating parameters, the parameters of the centralized critic network are first updated based on temporal difference error to minimize the mean squared error of value prediction. Then, based on the action-value gradient output by the critic network, the parameters of each actor network are updated using a policy gradient method to maximize the agent's long-term cumulative reward. During training, four core indicators—cumulative reward, win rate, budget utilization, and return on investment—are continuously monitored. When the fluctuation range of these core indicators is below a preset threshold and the overall trend is stable across multiple iterations, the network is considered converged, and training is terminated.

[0033] Step S204: After training, one of the actor networks is fixed as the bidding strategy network for online inference. The four types of state features extracted from the real-time bidding exposure request are input into the bidding strategy network, and the initial bidding coefficient is output.

[0034] Specifically, because all agents have the same actor network structure and a symmetrical training environment during training, the learned strategy is universal. Therefore, after training, only the weight parameters of one actor network are extracted and solidified as the bidding strategy network deployed online. During the online inference stage, a centralized critic network is not invoked, nor is it necessary to obtain the internal states of other bidders. Inference can be completed based on four locally observable state features, meeting the information constraints and real-time requirements of real-world scenarios. When generating the initial bid coefficients online, the extracted traffic prediction value, remaining budget state, historical average winning bid, and bidding participation scale are first numerically normalized to eliminate the impact of differences in feature dimensions on network inference. The normalized feature vectors are input into the bidding strategy network, and after forward propagation, continuous initial bid coefficients are output. The output results are directly passed to the subsequent game correction module to complete the basic bid generation stage.

[0035] In one embodiment of the present invention, the specific steps for constraining the optimal response dynamic correction of the initial bid coefficient are as follows: Step S301: Use a sliding window to collect opponent winning bid samples from nearly a preset number of rounds to construct an opponent bid sample set; Specifically, the system maintains a fixed-length first-in-first-out (FIFO) sliding window queue in memory as a storage medium for competitor bidding samples. After each round of bidding is completed and the bidding results are received from the ad exchange platform, the winning bid data (i.e., the second-highest bid in a two-price auction, corresponding to the actual deduction price for the winning bidder) is extracted from the result message and written as a competitor bidding sample to the end of the queue. If the queue length reaches the preset window size, the oldest sample at the head of the queue is automatically removed, ensuring that the window always contains bidding samples from the most recent preset rounds, ultimately forming a time-matched competitor bidding sample set. The window size is set according to the frequency of traffic fluctuations, typically ranging from 100 to 500 rounds of bidding samples, balancing the statistical significance of distribution fitting with the tracking sensitivity of environmental changes; the window is reduced in high-fluctuation scenarios to improve tracking speed, and expanded in stable scenarios to improve fitting accuracy.

[0036] Step S302: Fit the opponent's bid sample set using the kernel density estimation algorithm to obtain the probability density function and cumulative distribution function of the opponent's bid; Specifically, kernel density estimation, a non-parametric probability density estimation method, is used to fit the distribution of competitor bid samples within a sliding window. This method requires no pre-defined distribution shape and accurately recreates the bidding distribution characteristics of the real market. During execution, a Gaussian kernel is selected as the kernel function, and the optimal bandwidth parameter is determined through cross-validation. All bid samples within the window are substituted into the kernel density estimation formula to calculate the continuous competitor bid probability density function, and the corresponding cumulative distribution function is obtained through integration. The probability density function describes the probability density of competitor bids falling near a certain price point; the cumulative distribution function describes the probability of competitor bids falling below a certain price, i.e., the winning probability of the corresponding bid, and is the core computational basis for subsequently solving for the optimal bid anchor point. This fitting process is executed asynchronously by an independent background computation thread, refitting every fixed number of rounds, and updating the fitting results to a shared cache. Each bid decision directly reads the latest distribution function from the cache, avoiding the fitting calculation from occupying the main decision-making process time and ensuring that the single bid response latency is controlled within milliseconds.

[0037] Step S303: Under the two-price sealed auction mechanism, with the goal of maximizing the expected return in a single step, and combining the budget consumption rate constraint and the return on investment constraint, the constrained optimal bid anchor point is obtained. Specifically, based on the revenue mechanism of a two-price sealed-bid auction, with the optimization objective of maximizing the expected revenue per step, and superimposed with constraints on budget consumption rate and return on investment (ROI), the optimal bidding coefficient under these constraints is solved using a numerical grid search method, i.e., the optimal bid anchor point. In a two-price auction, when one's bid is the product of the bidding coefficient and the estimated value of the traffic, the winning condition is that one's bid is higher than the highest bid of all competitors; the actual revenue after winning is the estimated value of the traffic minus the highest bid of the competitors. Therefore, the expected revenue per step can be calculated in integral form, with the integral interval from 0 to one's bid, and the integrand being the product of "traffic value minus bid" and the probability density of the highest bid of the competitors. The first term is the budget consumption rate constraint, requiring that the expected deduction corresponding to the current bid must not exceed the upper limit of the budget consumption allowed in the current period to avoid budget overspending; the second term is the ROI constraint, requiring that the expected ROI corresponding to the current bid must not be lower than the preset target value to ensure that the campaign effect meets the target. Within the feasible range of bid coefficient values, a grid traversal is performed at a preset step size, calculating the expected return value corresponding to each candidate bid coefficient one by one, and verifying whether it meets two constraints. Among all candidate coefficients that meet the constraints, the coefficient with the largest expected return is selected as the optimal bid anchor point. This numerical solution method has low computational cost, high stability, and is suitable for asynchronous computing engineering implementations.

[0038] Step S304: The initial bid coefficient is weighted and adjusted using a dynamic adjustment factor to obtain the adjusted bid coefficient. The adjustment formula is as follows:

[0039] in, This is the adjusted bid coefficient; This is the initial bid coefficient; To constrain the optimal bid anchor point, it represents the theoretical optimal bid coefficient under the combined effects of the current competitive environment, budget constraints, and return on investment constraints; It is a dynamic correction coefficient, calculated by the sigmoid function, with a value range of (0,1). Its value increases as the difference between the current number of bidders and the historical average number of bidders increases, and also increases as the standard deviation of the recent round's winning bid increases.

[0040] Its calculation logic is as follows:

[0041] The current bidding participation scale, i.e. the number of bidders participating in this bidding, is directly extracted from the real-time bidding exposure request message issued by the advertising exchange platform; The historical average number of bidders is obtained by statistically analyzing the average bidding participation scale of the same period in the past 7 days. It is pre-stored in the local configuration and updated daily. The near-round winning bid volatility is the standard deviation of all counterparty winning bid samples within the sliding window, calculated in real time based on the counterparty bid sample set. To pre-determine the weighting coefficients, simulation experiments were conducted to calibrate the weights of the impact of bidding size deviation and price fluctuation on the correction strength. In a typical information flow advertising scenario where the winning bid unit is yuan and the historical average number of bidders is approximately 5, the calibration value is... =0.3, =2; The two coefficients control the weight of the impact of bidding scale deviation and price fluctuation on the correction strength, respectively, and can be adjusted adaptively according to the different industry deployment scenarios; The sigmoid function is an S-shaped activation function that can map input values ​​to the (0,1) interval, ensuring that the correction coefficient is always within a reasonable range and avoiding overcorrection that could lead to price fluctuations.

[0042] When the number of bidders is significantly higher than the historical average, or the win price volatility is high, it indicates that the current market competition state deviates significantly from the pre-training environment, and the probability of deviation in the base bid is higher. In this case, the value of α automatically increases to enhance the correction range and quickly adapt to environmental changes. When the market state is stable, the value of α automatically decreases to maintain the stability of the basic strategy and avoid excessive correction leading to frequent bid fluctuations. Taking the calibration parameters as an example: when the number of bidders is 3 more than the historical average and the win price volatility is 0.4 yuan, the calculated input value is 0.3. 3+2.0 0.4 = 1.7, corresponding to α ≈ 0.846, indicating a strong correction. When the number of bidders is 2 fewer than the historical average and the win price volatility is 0.1 yuan, the calculated input value is 0.3. (﹣2)+2.0 0.1 = -0.4, corresponding to α≈0.401, the correction is weakened, which is in line with the design logic of prioritizing the stability of the strategy when the competition is mild.

[0043] The fitting of the opponent's bid distribution and the calculation of the optimal bid anchor point are both executed asynchronously by independent threads, and the calculation results are stored in a shared cache. When making a single bid decision, the latest anchor point and correction coefficient in the cache are read directly. The calculation time for a single step correction is less than 0.2ms, which fully meets the response requirements of real-time bidding at the level of hundreds of milliseconds.

[0044] This formula is a first-order linear interpolation correction formula. It uses the initial bid coefficients output by the pre-trained network as a basis to perform a weighted approximation towards the constrained optimal bid anchor point in the current environment; the correction magnitude is determined by the dynamic correction coefficients. control: The larger the value, the closer the adjusted bid is to the optimal anchor point of the current environment, and the stronger the environment adaptability. The smaller the value, the more the bidding inertia of the basic strategy is preserved, resulting in stronger operational stability. Through... The dynamic adjustment can balance the strategy's environmental adaptability and operational stability.

[0045] In one embodiment of the present invention, the specific steps for implementing two-stage constraint control on the modified bid coefficient are as follows: Step S401: Collect historical traffic quality distribution data and competition intensity time series data, and use dynamic programming algorithm to allocate the total budget according to the traffic value density of each time period to generate a dynamic budget consumption baseline curve. Specifically, historical campaign data from the past 7-14 days for the target advertiser's industry is retrieved, and the campaign is divided into time periods by hourly granularity. The average estimated traffic value, average win price, click-through rate, and number of bidding participants are calculated for each time period to construct a traffic value density dataset for each time period. Traffic value density is defined as the expected conversion revenue per unit of exposure, used to quantify the return on investment for different time periods. With maximizing the total expected revenue over the entire campaign cycle as the optimization objective and the total budget amount as a global hard constraint, the daily budget allocation problem is decomposed into sub-allocation problems for each time period. The optimal budget allocation for each time period is solved recursively through dynamic programming. Higher budget amounts are allocated to high-value-density time periods, and budgets are reduced accordingly for low-value-density time periods, replacing the traditional uniform consumption model of average allocation. Based on the budget allocation results for each time period, a continuous cumulative budget consumption baseline curve is generated through linear interpolation. Each time point on the curve corresponds to the planned cumulative consumption amount at the current moment, serving as a benchmark for subsequent consumption progress control. The curve is initialized daily and can be slightly updated on the day of campaign based on real-time traffic quality fluctuations to ensure a good match between the baseline and actual traffic characteristics.

[0046] Step S402: Collect the current cumulative consumption budget data in real time and calculate the deviation between the current actual consumption progress and the dynamic budget consumption baseline curve; Specifically, the system uses deduction logs to statistically analyze the cumulative actual deduction amount for the day for each bidding agent to be connected, and calculates the actual consumption progress (cumulative consumption amount / total budget amount). Simultaneously, it reads the planned consumption progress corresponding to the current time point in the dynamic budget consumption baseline curve. The actual consumption progress is subtracted from the planned consumption progress to obtain the deviation value. A positive deviation value indicates that the budget consumption is ahead of the baseline plan, while a negative deviation value indicates that the budget consumption is behind the baseline plan. The larger the absolute value of the deviation, the greater the degree of deviation between the actual consumption and the baseline, and the stronger the control measures required.

[0047] Step S403: When the deviation value is less than or equal to the preset deviation threshold, the controlled bid coefficient is obtained by linear multiplication. The adjustment formula is as follows:

[0048] in, This refers to the bid coefficient after control. This is the adjusted bid coefficient; is the penalty gain coefficient; is the preset linear adjustment gain parameter used to control the degree of influence of the deviation magnitude on the bid; based on the iPinYou public RTB dataset, it is calibrated in a multi-agent bidding simulation environment through grid search, and the calibration value is 1.2 in the general information flow advertising scenario; this coefficient can be adaptively adjusted according to the strictness of budget control in different industries, and the stricter the control requirements, the larger the coefficient value.

[0049] This represents the deviation between the current actual consumption progress and the baseline curve; the value can be positive or negative, with positive values ​​indicating consumption ahead of schedule and negative values ​​indicating consumption lagging behind. It is calculated as the ratio of actual cumulative consumption to the total budget, minus the planned consumption ratio at the corresponding point in time on the baseline curve. The preset deviation threshold is the critical value used to distinguish between linear adjustment and exponential penalty phases, and is preset to 0.05 (i.e., a 5% deviation tolerance) based on the stability requirements of the campaign. When the absolute value of the deviation is within 5%, a gentle linear adjustment is adopted to avoid frequent price fluctuations interfering with the bidding strategy's effectiveness.

[0050] Specifically, this formula is a linear soft-constraint adjustment formula. Based on the corrected bid coefficient, it adjusts the bid coefficient linearly according to the consumption progress deviation: when consumption is ahead of schedule (Δb>0), the bid coefficient is lowered proportionally to slow down the consumption pace; when consumption is behind schedule (Δb<0), the bid coefficient is raised accordingly to accelerate the consumption pace. The formula uses a max function to set a lower limit of 0.5 to prevent the bid coefficient from being too low and thus failing to win bids, ensuring a basic delivery volume. This stage is suitable for typical scenarios with small deviations, with a gentle adjustment to avoid frequent bid fluctuations affecting campaign stability.

[0051] Step S404: When the deviation value is greater than the preset deviation threshold, an exponential penalty adjustment is used to obtain the controlled bid coefficient. The penalty intensity increases exponentially with the increase of the deviation value.

[0052] Specifically, when the consumption progress deviation exceeds the preset deviation threshold, it indicates that the consumption rhythm deviates significantly from the baseline, and the linear adjustment is insufficient to quickly bring the rhythm back on track. Therefore, it switches to exponential penalty adjustment, with the adjustment formula as follows:

[0053] Where δ is the preset deviation threshold (valued at 0.05), and exp is the natural exponential function. Under this mechanism, the greater the deviation value exceeds the threshold, the more rapidly the penalty increases exponentially, and the larger the reduction in the bid coefficient becomes. This can slow down the consumption pace in a short period of time, preventing the budget from being consumed far ahead of schedule; at the same time, a bid floor of 0.5 is still maintained to ensure basic exposure acquisition capabilities. If the deviation value is negative and the absolute value exceeds the threshold (significantly delayed consumption), a symmetrical exponential gain adjustment is adopted to appropriately increase the bid coefficient, accelerate budget consumption, and avoid excessive budget remaining at the end of the campaign period. Through two-stage hierarchical control, both the stability of campaigns in normal scenarios and the effectiveness of control under abnormal deviations are taken into account.

[0054] It should be noted that the two-stage constraint control also includes a hard circuit breaker control mechanism, as detailed below: Step S405: Preset two hard blocking rules: single bid limit and budget safety threshold; obtain in real time the bid amount and remaining budget percentage data corresponding to the bid coefficient after control. Specifically, the system pre-configures two hard-blocking rule thresholds as inviolable compliance bottom lines. The single bid cap is derived from the standard configuration parameters of the bidding agent to be integrated; it is a core constraint set by the agent during the integration phase, representing the highest acceptable bid amount for a single exposure. The budget safety threshold is a system-preset fallback ratio threshold, generally set at 1% of the total budget, but can be adjusted adaptively according to the strictness of control in the campaign scenario. The real-time data required for verification is synchronously obtained through internal system links: first, the post-control bid amount, calculated by multiplying the output post-control bid coefficient by the estimated value of the current traffic; second, the remaining budget percentage, calculated as (total budget amount - cumulative actual deduction amount for the day) / total budget amount. This data comes from the system's real-time deduction statistics module, automatically updated after each round of bidding and actual deduction, with a data latency of less than 100ms, ensuring timely verification.

[0055] Step S406: Perform the first-level single bid limit verification. If the bid amount exceeds the preset single bid limit, the bid amount will be automatically truncated to the single bid limit value, and a bid limit interception prompt will be sent to the bidding agent to be connected; if the limit is not exceeded, proceed to the second-level verification. Specifically, this layer of verification targets the compliance of individual bids and is performed before the bid is submitted to the advertising exchange platform. The system compares the calculated bid amount with the preset single bid limit: If the bid amount exceeds the single bid limit, the final bid amount will be automatically truncated to the single bid limit value. At the same time, a bid limit interception prompt will be generated and synchronized to the bidding agent to be connected through a standardized communication link. The prompt content includes the original bid, the bid after truncation, and the triggering reason. The truncated bid will still participate in the bidding normally, preserving the opportunity to place orders as much as possible while adhering to the bid limit.

[0056] If the bid amount does not exceed the upper limit, the verification passes, and the bid data enters the second-level budget circuit breaker verification stage. This layer of verification can effectively avoid abnormally high bids in extreme scenarios such as bidding strategy network anomalies and excessive game correction, ensuring that the cost of a single bid always remains within the preset range.

[0057] Step S407: Perform the second-level budget circuit breaker verification. When the remaining budget percentage is lower than the preset safety threshold, the budget circuit breaker state is triggered, and all subsequent bidding requests are blocked. Specifically, this layer of verification targets the compliance of the total budget and serves as the ultimate safety net mechanism for budget control. The system uses the remaining budget percentage calculated from the current cumulative deductions as the verification basis, comparing it with a budget safety threshold. When the remaining budget percentage falls below the preset safety threshold, a budget circuit breaker is immediately triggered, and the system sets a bidding interception flag at the access adaptation layer. After the flag takes effect, all subsequent real-time bidding exposure requests will no longer enter the bidding calculation process and will be directly intercepted and discarded, blocking subsequent budget consumption from the source and ensuring that the total consumption throughout the entire cycle does not exceed the budget. Compared with soft constraint adjustments, this layer of verification is a mandatory safety net, not relying on the gradual adjustment of the bidding coefficient. It can immediately stop losses in extreme scenarios and completely eliminate the risk of budget overruns.

[0058] Step S408: Manage the operating state through the circuit breaker state machine. After triggering the budget circuit breaker, the state machine switches to the circuit breaker pause state and synchronizes the budget exhaustion pause signal to the bidding agent to be connected. Specifically, the system uses a finite state machine to uniformly manage the operational status of budget control. The state machine includes three stable states: normal deployment, early warning, and circuit breaker suspension. State transitions are triggered only by the remaining budget percentage, with a defined flow logic to avoid state oscillations. The normal deployment state is the default state when the remaining budget is sufficient; the entire bidding process executes normally, and soft constraints are smoothly adjusted. The early warning state is triggered when the remaining budget percentage falls below the early warning threshold (generally set at 3% of the total budget). The state machine switches from normal deployment to early warning, sending a budget warning notification to the bidding agents to be connected via a standardized communication link, informing them in advance that the budget is about to run out, allowing the agents to prepare for deployment termination. The circuit breaker suspension state is triggered when the remaining budget percentage falls below the safety threshold. The state machine switches from early warning to circuit breaker suspension, simultaneously executing bid interception operations and pushing a budget exhaustion suspension signal to the bidding agents to be connected. This signal includes information such as the remaining budget amount, trigger time, and recovery conditions. All state changes are logged, allowing for traceability of the entire control execution process.

[0059] Step S409: When a budget increase or campaign period reset is detected, the state machine automatically switches back to the normal campaign state, reinitializes the dynamic budget consumption baseline curve, and restores the normal bidding function.

[0060] Specifically, when the daily campaign cycle is reset, the accumulated consumption is automatically cleared, the daily dynamic budget consumption baseline curve is regenerated, and the campaign is switched back to normal. After the agent submits a budget increase request and updates the total budget parameters, if the remaining budget percentage rises above the safety threshold, the blocking is automatically lifted, the baseline curve is reset, and bidding is resumed. No manual intervention is required on the system side throughout the process. The agent is notified synchronously after the recovery is completed, forming a complete control loop from abnormal blocking to automatic recovery.

[0061] In one embodiment of the present invention, during the process of generating the initial bid coefficient, the bidding strategy network has a lightweight online adaptive mechanism, as follows: Step a: In the pre-training stage of the bidding strategy network, the basic weights are solidified based on the publicly available real-time bidding dataset and historical campaign data from multiple industries and advertisers, and a general initial bidding coefficient is output. Specifically, using publicly available real-time bidding datasets and historical campaign data from multiple industries and advertisers as training foundations, the bidding strategy network is fully trained in an adversarial self-game multi-agent simulation environment. This allows the network to learn bidding decision logic and budget smoothing control rules under general scenarios. After training, the weights of the network's bottom feature extraction layer and intermediate strategy layer are solidified, while only the parameters of the top fully connected output layer are retained as updatable. The solidified network can directly output a general initial bidding coefficient that is compatible with most campaign environments, providing stable bidding support for newly connected agents during the cold start phase, allowing them to go online and run without starting training from scratch.

[0062] Step b: After the bidding agent is officially launched, the returned bidding results will be accumulated in real time to construct a bidding feedback dataset. Specifically, after the bidding intelligence is officially launched, after each round of bidding, the feedback data such as the win / loss status, actual cost, and winning bid level returned by the advertising exchange platform and transmitted through the bidding result synchronization link will be written to the local feedback data cache queue in real time. The cache adopts a fixed-length first-in-first-out mechanism, continuously accumulating full-link data under the real bidding scenario as the launch process progresses, without the need for additional manual collection and labeling, providing a native data source that fits the current environment for subsequent online adaptive fine-tuning.

[0063] Step c: When the accumulated number of bidding feedback datasets reaches the preset number of sample rounds, extract the corresponding state features and the optimal bid label to construct a fine-tuning training sample set; Specifically, when the cumulative number of bidding feedback samples in the cache reaches a preset round threshold, the fine-tuning sample construction process is automatically triggered. The four types of state features corresponding to each sample are extracted as model input features: estimated traffic value, remaining budget status, historical average winning bid, and bidding participation scale. At the same time, the constraint optimal bid anchor point calculated by the constraint optimal response dynamic correction module for the corresponding round is matched as a supervision label. The fine-tuning training sample set is constructed according to the one-to-one correspondence between input and label, and the training batches are divided according to a fixed batch size.

[0064] Step d: Incrementally update the parameters of the top fully connected layer of the bidding strategy network using a small learning rate, while keeping the weights of the bottom feature extraction network fixed. The updated network output is adapted to the initial bidding coefficient of the current delivery scenario.

[0065] Specifically, a learning rate much lower than that used in the pre-training phase is employed. Incremental updates are performed on the bidding strategy network based on the constructed fine-tuned sample set. During the update process, only the parameters of the top fully connected output layer are adjusted, while the weights of the bottom feature extraction network remain fixed throughout. This update method consumes less than 5% of the computational cost of full training. It avoids the catastrophic forgetting and basic strategy drift problems caused by full updates, and allows the network to quickly adapt to the current advertiser's industry attributes, traffic characteristics, and competitive environment. After the updated weights are smoothly replaced and put online, the output initial bidding coefficient will be more in line with the current advertising scenario, achieving continuous optimization of advertising performance with low overhead.

[0066] In one embodiment of the present invention, the unstructured configuration parameters are a delivery description in natural language form, and standard configuration parameters are generated through the following steps: Step S105: Set up a natural language configuration entry in the access layer to receive the natural language text of the delivery target and constraint requirements input by the access party, and input it as the parameter field to be mapped into the advertising domain semantic mapping model. Specifically, a natural language configuration entry is set up in the intelligent agent access console of the access layer, providing a text input box to receive the natural language text of the target and constraint requirements for the ad placement. It supports descriptions of ad placement needs in everyday language, such as "Daily budget of 1,000 yuan, maximum bid of 2 yuan per bid, try to keep the cost per click within 1 yuan, ad placement from 9 am to 6 pm on weekdays". After the input is completed, the system performs basic preprocessing on the text, including removing redundant punctuation, unifying the number format, and completing semantic omissions. The preprocessed text is then input as a parameter field to be mapped into the semantic mapping model of the advertising domain.

[0067] Step S106: The advertising domain semantic mapping model performs semantic parsing and entity extraction on natural language text to identify four core parameter information contained in the text: budget size, bid cap, target cost per click, and delivery time period. Specifically, the semantic mapping model in the advertising domain combines domain-fine-tuned named entity recognition capabilities with a semantic slot filling mechanism to perform sentence-by-sentence semantic parsing and entity extraction on the input natural language text. It automatically identifies four types of core parameter information contained in the text: budget size, bid cap, target cost per click, and ad placement period. The model can identify parameters with the same meaning in different expressions. For example, it identifies "daily budget" and "daily spending cap" as daily budget parameters, and "maximum bid" and "no more than one bid per time" as single bid cap parameters. At the same time, it extracts the corresponding numerical values, units, and constraints of the parameters and outputs a structured set of parameter entities.

[0068] Step S107: Automatically map the identified core parameter information to the standard field system, match the corresponding constraint thresholds and reward weights, and generate initial standard configuration parameters; Specifically, the four types of parameter entities identified are semantically matched with standard fields in the standard field semantic vector library to complete the automatic mapping of parameter information to the standard field system. At the same time, according to the numerical range of the parameters and the type of the target, the corresponding constraint thresholds and control parameters are automatically matched. For example, the budget safety threshold and circuit breaker warning threshold are automatically matched according to the total budget value, and the default value of the corresponding penalty gain coefficient is automatically matched according to the target click cost. Finally, the initial standard configuration parameters are generated. The parameter format, data type and unit are completely consistent with the platform standard specifications.

[0069] Step S108: Generate parameter configuration confirmation information and push it to the access party. After the access party confirms, the final standard configuration parameters are generated and take effect.

[0070] Specifically, the system organizes the generated initial standard configuration parameters into a structured configuration confirmation form, clearly displaying the standard name, value, unit, and business meaning of each parameter, and pushes it to the access party's console interface for verification. The access party can manually modify any incorrectly parsed parameters, and submit a confirmation command after confirming that everything is correct. After receiving the confirmation command, the system generates the final standard configuration parameters and writes them to the configuration database. Once officially effective, these parameters serve as the binding basis for the entire subsequent bidding process. If the access party does not confirm, the parameters will not take effect, avoiding configuration errors caused by semantic parsing biases and balancing access convenience with configuration accuracy.

[0071] This invention completes the entire process—field adaptation, basic bid generation, competitive environment correction, and budget constraint management—sequentially on the access side, eliminating the need for bidding agents to independently develop corresponding algorithms and management logic. Specifically, it utilizes a semantic mapping model fine-tuned for the advertising domain to automatically align and standardize heterogeneous parameter fields, coupled with a natural language parameter parsing entry point, significantly reducing the adaptation costs and access barriers for different platform interfaces, effectively shortening the agent access cycle. The built-in bidding strategy network, pre-trained through multi-agent self-game theory, directly provides mature basic bidding capabilities to accessing agents, eliminating algorithm development work on the agent side and effectively mitigating the cold start effect fluctuations and budget waste issues for newly accessed agents. This is further enhanced by a constraint-optimal response dynamic correction mechanism. It can dynamically adjust bids in real time by tracking changes in market competition, improving the adaptability of pre-trained strategies to dynamic competitive environments and ensuring the stability of bidding results. It employs a dynamically programmed budget consumption baseline curve combined with a two-stage soft constraint adjustment and hard circuit breaker control mechanism, which can both allocate budgets to high-value traffic periods to improve campaign revenue and achieve smooth budget consumption and compliance protection throughout the entire cycle. A lightweight online adaptive mechanism is also included, which can incrementally fine-tune the top-level network parameters based on real bidding feedback, achieving continuous optimization of campaign results with low computing power. Overall, it significantly reduces the technical threshold and implementation cost of intelligent agents accessing real-time advertising bidding systems, improves access efficiency and operational stability, and supports the large-scale promotion and application of intelligent agent bidding technology.

[0072] The present invention also provides a simplified system for intelligent agents to access a real-time advertising bidding system, including a memory, a host computer, and a computer program stored in the memory and executable on the host computer, wherein the computer program is configured to implement the steps of the simplified method for intelligent agents to access a real-time advertising bidding system as described above.

[0073] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A simplified method for an intelligent agent to access a real-time advertising bidding system, characterized in that, Includes the following steps: The system obtains the access requests and unstructured configuration parameters of the bidding agents to be accessed, and completes automatic field alignment and standardization transformation through a pre-trained advertising domain semantic mapping model to generate standard configuration parameters and establish a standardized communication link. Receive real-time bidding exposure requests from advertising exchange platforms, extract four types of state features, input them into a bidding strategy network pre-trained by multi-agent self-game, and generate initial bidding coefficients; Based on the opponent's bid distribution data collected by the sliding window, the initial bid coefficient is dynamically corrected by constrained optimal response to obtain the corrected bid coefficient; Based on the budget consumption baseline curve generated by dynamic programming, a two-stage constraint control is applied to the corrected bidding coefficient to generate the final bid and send it to the advertising exchange platform. At the same time, the bidding results are synchronously fed back to the connected bidding agent.

2. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, The advertising domain semantic mapping model completes training and field mapping through the following steps: Collect historical interface field descriptions, bidding parameter descriptions, and industry placement standard corpora from the advertising transaction field to construct a domain training corpus. Use a general lightweight pre-trained language model as the initial weights and employ a contrastive learning approach for domain fine-tuning training. Extract the core parameter fields of the entire real-time bidding process, input them into the trained advertising domain semantic mapping model to generate corresponding semantic embedding vectors, and build a standard field semantic vector library; Receive the parameter fields to be mapped uploaded by the bidding agent to be connected, input the semantic mapping model of the advertising domain to generate the semantic embedding vector to be mapped, calculate the cosine similarity between the semantic embedding vector to be mapped and each standard field vector in the standard field semantic vector library, and sort them in descending order of similarity value to obtain the matching candidate list. Select the first standard field in the matching candidate list and determine whether its corresponding similarity value is higher than the preset similarity threshold. If so, automatically complete the mapping relationship between the parameter field to be mapped and the standard field, and perform data format conversion to generate standard configuration parameters. If not, push the top three standard fields in the matching candidate list to the access party for manual confirmation. The system receives manually confirmed matching results and feeds the confirmed mapping pairs back into the incremental training set of the semantic mapping model in the advertising domain for subsequent iterative optimization of the model.

3. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, The four types of state features include estimated traffic value, remaining budget status, historical average winning bid, and bidding participation scale. The specific steps for training the bidding strategy network and generating the initial bidding coefficients are as follows: Construct an adversarial self-game multi-agent simulation environment to simulate multiple bidding agents participating in a two-price sealed auction. Generate global state data including flow value, budget state, and competition scale. Input the actor network corresponding to each bidding agent and output the bidding action of each agent. A centralized commentator network is set up to receive global state data and the joint bidding actions of all intelligent agents, output the global action value, and calculate the single-step reward value by combining the constraint perception composite reward rule; The backpropagation algorithm is executed based on the single-step reward value and the global action value to iteratively update the parameters of the actor network and the centralized critic network until the network converges; After training, one of the actor networks is fixed as the bidding strategy network for online inference. The four types of state features extracted from real-time bidding exposure requests are input into the bidding strategy network, and the initial bidding coefficient is output.

4. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, The specific steps for constraining the optimal response dynamic correction of the initial bid coefficient are as follows: A sliding window is used to collect opponent winning bid samples for a near-preset number of rounds to construct an opponent bid sample set. By fitting the opponent's bid sample set using the kernel density estimation algorithm, the probability density function and cumulative distribution function of the opponent's bids are obtained; Under the two-price sealed auction mechanism, with the goal of maximizing the expected return in a single step, and combining the constraints of budget consumption rate and return on investment, the constrained optimal bid anchor point is obtained. The initial bid coefficient is weighted and adjusted using a dynamic adjustment factor to obtain the adjusted bid coefficient. The adjustment formula is as follows: , in, This is the adjusted bid coefficient; This is the initial bid coefficient; To constrain the optimal bid anchor point; It is a dynamic correction coefficient, calculated by the sigmoid function, with a value range of (0,1). Its value increases as the difference between the current number of bidders and the historical average number of bidders increases, and also increases as the standard deviation of the recent round's winning bid increases.

5. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, The specific steps for implementing two-stage constraint control on the corrected bid coefficient are as follows: Collect historical traffic quality distribution data and competition intensity time series data, and use dynamic programming algorithm to allocate the total budget according to the traffic value density of each time period to generate a dynamic budget consumption baseline curve. Real-time collection of current cumulative consumption budget data, and calculation of the deviation between the current actual consumption progress and the dynamic budget consumption baseline curve; When the deviation value is less than or equal to the preset deviation threshold, the controlled bid coefficient is obtained by linear multiplication. The adjustment formula is as follows: , in, This refers to the bid coefficient after control. This is the adjusted bid coefficient; This is the penalty gain coefficient; This represents the deviation between the current actual consumption progress and the baseline curve. When the deviation value exceeds the preset deviation threshold, an exponential penalty adjustment is used to obtain the controlled bid coefficient, and the penalty intensity increases exponentially with the increase of the deviation value.

6. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 5, characterized in that, The two-stage constraint control also includes a hard circuit breaker control mechanism, as detailed below: Two hard blocking rules are preset: a single bid cap and a budget safety threshold. The bid amount and remaining budget percentage data corresponding to the bid coefficient after control are obtained in real time. Perform the first-level single bid limit check. If the bid amount exceeds the preset single bid limit, automatically truncate the bid amount to the single bid limit value and send a bid limit interception prompt to the bidding agent to be connected. If the limit is not exceeded, proceed to the second level of verification; When performing the second-level budget circuit breaker check, if the remaining budget percentage is lower than the preset safety threshold, the budget circuit breaker state is triggered, blocking all subsequent bidding requests; The operation status is managed by a circuit breaker state machine. After the budget circuit breaker is triggered, the state machine switches to the circuit breaker pause state and synchronizes the budget exhaustion pause signal to the bidding agent waiting to be connected. When a budget increase or campaign period reset is detected, the state machine automatically switches back to normal campaign state, reinitializes the dynamic budget consumption baseline curve, and restores normal bidding functionality.

7. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, During the generation of the initial bid coefficients, the bidding strategy network possesses a lightweight online adaptive mechanism, as detailed below: During the pre-training phase of the bidding strategy network, the basic weights are solidified based on the publicly available real-time bidding dataset and historical campaign data from multiple industries and advertisers, and a general initial bidding coefficient is output. Once the bidding agent is officially deployed, the returned bidding results will be accumulated in real time to construct a bidding feedback dataset. When the cumulative number of bidding feedback datasets reaches the preset number of sample rounds, the corresponding state features and the optimal bid label are extracted to construct a fine-tuning training sample set. A small learning rate is used to incrementally update the parameters of the top fully connected layer of the bidding strategy network, while the weights of the bottom feature extraction network remain fixed. The updated network output is adapted to the initial bidding coefficient of the current delivery scenario.

8. A simplified method for an intelligent agent to access a real-time advertising bidding system according to claim 1, characterized in that, The unstructured configuration parameters are descriptions in natural language, and standard configuration parameters are generated through the following steps: A natural language configuration entry is set up in the access layer to receive the text of the delivery target and constraint requirements in natural language form input by the access party, which is then input into the advertising domain semantic mapping model as the parameter field to be mapped. The semantic mapping model in the advertising field performs semantic parsing and entity extraction on natural language text, and identifies four core parameter information contained in the text: budget size, bid cap, target cost per click, and delivery time period. The identified core parameter information is automatically mapped to the standard field system, and the corresponding constraint thresholds and reward weights are matched to generate initial standard configuration parameters; The parameter configuration confirmation message is pushed to the access party. After the access party confirms, the final standard configuration parameters are generated and take effect.

9. A simplified system for intelligent agents to access a real-time advertising bidding system, characterized in that, The system includes a memory, a host computer, and a computer program stored in the memory and executable on the host computer, the computer program being configured to implement the steps of a simplified method for an agent to access a real-time advertising bidding system as described in any one of claims 1 to 8.