Marketing Segmentation Intelligent Identification and Positioning Methods and Systems
By constructing a track positioning weight set and auction feedback factors, the problems of insufficient positioning accuracy and strategy oscillation in e-commerce advertising are solved. This enables accurate identification of tracks with multiple attributes and low-cost, high-efficiency customer acquisition, improving the stability and robustness of the advertising strategy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU YUNZHIDACHUANG TECH CO LTD
- Filing Date
- 2026-04-22
- Publication Date
- 2026-05-26
AI Technical Summary
Existing e-commerce advertising technologies suffer from insufficient targeting accuracy when identifying niche markets, an inability to distinguish the reasons for rising costs, resulting in strategy oscillations and audience fatigue, and a lack of ability to identify the endogenous feedback mechanism of the auction environment.
The market data aggregation unit generates a track positioning weight set, which is then combined with the cost sensitivity analysis unit and the strategy optimization unit. The strategy optimization model is trained using calibration return values to construct continuous track positioning weights and auction feedback factors. This quantifies the endogenous impact of placement actions on the auction environment, removes negative feedback, and achieves accurate identification and low-cost, high-efficiency customer acquisition.
It improves the stability and robustness of e-commerce advertising strategies, avoids strategy oscillations caused by auction bias, and achieves accurate identification of multi-attribute tracks and low-cost, high-efficiency customer acquisition.
Smart Images

Figure CN122089401A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of Internet data processing technology, and more specifically, to a method and system for intelligent identification and positioning of marketing segments. Background Technology
[0002] In current digital marketing, especially e-commerce advertising, advertisers typically utilize automated systems to identify high-potential niche markets (such as specific categories, price points, or scenario combinations) and adjust their advertising strategies based on monitored performance data. Existing mainstream solutions generally employ an observation-execution logic: the system collects public opinion or search trends from outside the platform as market signals, combines this with on-platform conversion costs or ROI as feedback rewards, and uses reinforcement learning or rule engines to dynamically adjust ad frequency or budget. However, existing automated advertising technologies have significant limitations in practical applications in e-commerce retail media, primarily in the following aspects:
[0003] First, existing ad placement control models typically treat ad frequency as a linear variable, assuming that increasing frequency linearly increases exposure. However, actual e-commerce advertising systems generally employ a hybrid mechanism combining real-time bidding and quality score ranking. Under this mechanism, ad impressions depend not only on the bid but also heavily on quality signals such as click-through rate (CTR). When the system repeatedly exposes the same audience to the same ad frequently within a short period, audience fatigue or response decay occurs, leading to a decrease in CTR. Under the bidding ranking mechanism, a decrease in CTR directly results in a lower quality score. To maintain the original exposure level, the system is often forced to accept higher cost-per-click (CPC), thus driving up the final conversion cost.
[0004] Secondly, existing technical solutions lack the ability to identify the aforementioned endogenous feedback mechanisms. Traditional models only focus on the final observed changes in conversion costs, failing to distinguish whether the cost increase stems from intensified external market competition (a decline in the value of the track itself) or from quality score penalties caused by high-frequency ad placements (auction bias). When cost increases occur due to the user's own actions, existing technologies are prone to misjudging them as a less popular track or a deteriorating environment, thus incorrectly reducing weights or stopping ad placements, causing the strategy to oscillate between aggressive ad placements and erroneous stop-loss orders.
[0005] Furthermore, in identifying niche markets, traditional methods often rely on rigid label classifications, which are ill-suited to the reality that e-commerce products frequently span multiple scenarios simultaneously (e.g., belonging to both the camping and coffee niche markets), resulting in insufficient positioning accuracy. Additionally, conversion data typically exhibits a significant time lag; directly utilizing real-time data for strategy learning can introduce noise due to immature data, further impacting the stability of decision-making. Summary of the Invention
[0006] This invention provides a method and system for intelligent identification and positioning of marketing segments, which solves the technical problems mentioned in the background.
[0007] The first aspect is the intelligent identification and positioning system for marketing segmentation, including:
[0008] The market data aggregation unit is configured to collect public data from multiple sources and map it to preset sub-sectors, and to perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, wherein the track positioning weight set represents the degree of belonging of the target object in different sub-sectors.
[0009] The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs.
[0010] The strategy optimization unit is configured to construct an enhanced state input including the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor, and to train the strategy optimization model using the calibration reward value; the calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data, and the auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor;
[0011] The task distribution unit is configured to use the trained strategy optimization model to output the suggested delivery frequency for the current period, and translate the suggested delivery frequency into multi-channel execution tasks based on the principle of total strength conservation.
[0012] Secondly, the intelligent identification and positioning method for marketing segments, applied to any of the aforementioned intelligent identification and positioning systems for marketing segments, includes:
[0013] Collect publicly available data from multiple sources and map it to preset sub-tracks. Perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, which represents the degree of belonging of the target object to different sub-tracks.
[0014] The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs.
[0015] An enhanced state input is constructed, which includes the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor. The model is optimized by training a strategy using a calibration reward value. The calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data. The auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor.
[0016] The trained strategy optimization model outputs the suggested delivery frequency for the current period, and the suggested delivery frequency is translated into multi-channel execution tasks based on the principle of total strength conservation.
[0017] The beneficial effects of this invention are as follows: By constructing a continuous set of track positioning weights and auction feedback factors, it effectively solves the problem of cost misjudgment caused by audience fatigue due to high-frequency exposure and auction mechanism penalties in e-commerce advertising. The system can quantify the endogenous impact of advertising actions on the auction environment and uses calibration return values to remove negative feedback caused by the actions themselves, thereby forcing the strategy model to focus on the real value changes of the track rather than the system's penalty mechanism. This approach significantly improves the stability and robustness of advertising strategies in complex bidding environments, avoids strategy oscillations caused by auction bias, and achieves accurate identification of tracks with multiple attributes and low-cost, high-efficiency customer acquisition. Attached Figure Description
[0018] Figure 1 This is a flowchart of the marketing segmentation intelligent identification and positioning system of the present invention;
[0019] Figure 2 This is a schematic diagram of a specific implementation scenario of the present invention. Detailed Implementation
[0020] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0021] Example 1: As Figure 1 As shown, the intelligent identification and positioning system for marketing segments includes:
[0022] The market data aggregation unit is configured to collect public data from multiple sources and map it to preset sub-sectors, and to perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, wherein the track positioning weight set represents the degree of belonging of the target object in different sub-sectors.
[0023] The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs.
[0024] The strategy optimization unit is configured to construct an enhanced state input including the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor, and to train the strategy optimization model using the calibration reward value; the calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data, and the auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor;
[0025] The task distribution unit is configured to use the trained strategy optimization model to output the suggested delivery frequency for the current period, and translate the suggested delivery frequency into multi-channel execution tasks based on the principle of total strength conservation.
[0026] Preferably, collecting publicly available data from multiple sources and mapping it to preset sub-tracks includes:
[0027] The preset sub-tracks Defined as a leaf category of products Fixed price range and application scenario words The combination of three-dimensional Cartesian products formed by these three elements is represented as follows: ,in Belongs to category collection , Belongs to the price range set , Belongs to the set of scene words ;
[0028] For each piece of collected multi-source publicly available content data Perform the following mapping operations:
[0029] Configure each of the product leaf categories Category feature keyword set For the aforementioned multi-source publicly available content data text content The matching calculation is performed using the following formula:
[0030]
[0031] in, As keywords, For the indicator function; select The product leaf category to which this content belongs ;
[0032] Analyzing the multi-source publicly available content data The price information in the data is used to determine the fixed price range to which it belongs. If no explicit price information is available, it will be mapped to the default price range for that category.
[0033] Using a pre-built scene dictionary to analyze the multi-source publicly available content data Perform a longest string match and determine the longest matching term as the application scenario term. ;
[0034] Comprehensively determined , and The multi-source publicly available content data Uniquely mapped to the corresponding preset sub-track .
[0035] The segmented market segment 'k' is a three-dimensional Cartesian product of product leaf categories, fixed price ranges, and application scenario keywords. It's used to precisely define specific sub-sectors in e-commerce marketing. Ideally, it should encompass all combinations of the product leaf category set multiplied by the fixed price range set multiplied by the application scenario keyword set, covering real-world scenarios where multiple attributes of e-commerce products intersect, ensuring no omissions and clear boundaries. Product leaf category 'g' is one of the three dimensions of the segmented market segment, referring to the specific product category at the lowest level of the e-commerce platform's category tree. Ideally, it should be an existing list of leaf categories within the e-commerce site, such as coffee pots under outdoor equipment or dresses under clothing. Fixed price range 'p' is another three-dimensional dimension of the segmented market segment, referring to pre-defined price range segments. Ideally, it should have five ranges: 0-50 yuan, 50-150 yuan, 150-300 yuan, 300-800 yuan, and 800 yuan and above, referencing mainstream price distribution patterns in the e-commerce industry to cover low, medium, and high price ranges. Application scenario keywords (s) are one of the three dimensions of a niche market, referring to core words describing the usage scenarios or purposes of a product. They are preferably high-frequency scenario keywords such as commuting, camping, gift-giving, oil control, sensitive skin, and weddings, covering e-commerce user search habits and common product marketing scenarios. Category set G is the set of all product leaf categories, preferably a complete list of leaf categories from the e-commerce platform, including all underlying categories under primary categories such as 3C digital products, clothing and footwear, and home furnishings, thus comprehensively covering the product categories operated in e-commerce and avoiding omissions of potential niche markets. Price range set P is the set of all fixed price ranges, preferably including sets containing 0 to 50 yuan, 50 to 150 yuan, 150 to 300 yuan, 300 to 800 yuan, and 800 yuan and above, to conform to the long-tail characteristic of e-commerce product price distribution, distinguishing between low-priced traffic-driving products and high-priced profit products without complicating operations due to too many ranges. The scenario keyword set is a collection of all application scenario keywords, preferably including 20 to 50 high-frequency scenario keywords such as commuting, camping, gift-giving, oil control, sensitive skin, weddings, maternal and infant products, and outdoor activities. This is achieved by analyzing product titles, user reviews, and search keywords on e-commerce platforms to filter out the most frequently used scenario keywords, ensuring coverage of mainstream needs. Multi-source public content data (d) consists of single content data entries collected from multiple public platforms, including product descriptions, user reviews, and topic discussions. This can be determined based on public content from channels such as Xiaohongshu, Douyin, public comment sections of e-commerce platforms, and industry news websites. The category feature keyword set is a dedicated set of keywords configured for each product's leaf category, used to match the category to which the content data belongs. It is preferably composed of 10 to 30 high-frequency core keywords for each category. For example, keywords for the outdoor coffee pot category include camping coffee pots, portable coffee pots, and outdoor coffee brewing equipment. This is to extract high-frequency words from product titles, attribute descriptions, and user search terms within this category, ensuring accurate category matching.The text content refers to the text information contained in the multi-source publicly available content data d, including titles, body text, tags, etc. It can be obtained by extracting the text portion of the content data using data collection tools, removing non-text elements such as images and videos. Text in images can be extracted using optical character recognition (OCR) technology. The keyword w is a single specific word in the category feature keyword set, preferably any word from the set. For example, if the category feature keyword set is for "outdoor coffee pot," w could be "camping coffee pot." This word accurately represents the product characteristics of the corresponding category, aiding in category matching of the content data. The indicator function is used to determine whether the keyword w belongs to the text content; it returns 1 if it does, and 0 otherwise. Standard indicator function logic is preferred. The score is the statistical score of the multi-source publicly available content data d belonging to a certain product leaf category. It is used to determine the most suitable category for the content. The calculation method is to count the number of keywords in the category feature keyword set that belong to the text content. Each matched keyword scores 1 point, and each unmatched keyword scores 0 points. The total score is the category score. The category to which the content data 'd' ultimately belongs is the product leaf category, i.e., the category with the highest score. It is determined by comparing the scores of the content across all product leaf categories and selecting the category with the highest score. If multiple categories have the same score, the category with the smallest number is selected. Price information refers to the price-related information contained in the multi-source public content data 'd, including the listed price, discounted price, and promotional price. This can be obtained by parsing the text of the content data, price tags in product links, and price descriptions in images. Vague descriptions such as "high cost-performance ratio" can be considered as having no explicit price information. The preset default price range is the default fixed price range corresponding to the content's category when the content data has no explicit price information. It is preferably the mainstream price range of the corresponding product leaf category. For example, the default price range for the outdoor coffee pot category is 50 to 150 yuan. Based on the statistical distribution of historical transaction prices for products in this category, the price range with the highest transaction volume is taken as the default value to ensure the reasonableness of the default value. The scenario dictionary is a pre-built dictionary used to match application scenario terms. It contains all application scenario terms and related variations, preferably a dictionary that includes all words and common variations in the scenario term set. For example, variations of camping include wilderness camping, outdoor camping, etc., to cover different ways users describe scenarios and improve the comprehensiveness of scenario matching. The longest term is the term matched by the scenario dictionary after performing the longest string match on the text content. It is determined by comparing the text content with all terms in the scenario dictionary and selecting the longest term that is a perfect match. If no perfect match is found, it is considered a no-match. Application scenario terms are the application scenario terms finally determined from the multi-source publicly available content data d, i.e., the terms matched by the longest string match. They are determined by using the term if the longest string match matches, otherwise a general scenario is used as the application scenario term, ensuring that each piece of content data has a corresponding scenario term.This approach uses a three-dimensional Cartesian product of product leaf categories, fixed price ranges, and application scenario keywords to define segmented markets. Traditional market segmentation often uses a single category or a two-dimensional approach of category plus price, failing to reflect the scenario attributes of the product. For example, an outdoor coffee maker belongs to the outdoor equipment category, falls within the 50-150 yuan price range, and corresponds to the camping scenario. This three-dimensional combination accurately positions the segment, allowing marketing to better align with user needs and scenarios, solving the positioning ambiguity caused by the lack of scenario attributes in traditional segmentation methods. The method also determines the category by matching category feature keyword sets and ranking based on statistical scores. Traditional methods often rely on single keyword matching, which is prone to misjudgment. For instance, if content mentions that a portable coffee maker is suitable for camping, the keyword set for the outdoor coffee maker category includes words like "portable coffee maker" and "camping," resulting in a higher score due to the greater number of matches. However, the ordinary coffee maker category only matches the word "coffee maker," resulting in a lower score. Therefore, it is classified as an outdoor coffee maker. This method, through multi-keyword statistical scoring, improves the accuracy of category matching and avoids the accidental misjudgments caused by single-keyword matching. For content data without explicit price information, a rule is designed to map it to the corresponding category's preset default price range. Traditional methods often simply discard or classify it into a uniform range, leading to data waste or insufficient precision. For example, content describing the use of an outdoor coffee maker without mentioning the price is classified as follows: the default price range for this category is 50 to 150 yuan. Therefore, it is classified into this range, preserving valid data while ensuring the rationality of price dimension division through the category's mainstream price range, thus solving the problem of mapping content without price information to specific price categories. The longest string matching rule is used to extract application scenario words from the scenario dictionary. Traditional matching methods often prioritize short terms or perform fuzzy matching, easily matching non-core scenario words. For example, if the text content is "coffee maker for outdoor camping," the scenario dictionary includes terms like "camping," "outdoor camping," and "outdoor." The longest match result would be "outdoor camping," not the shorter "camping." This method accurately captures the specific scenario described in the text, avoiding scenario generalization problems caused by matching short terms, and improving the accuracy of scenario matching. By integrating three dimensions, a unique mapping of content data to specific sub-categories is achieved, forming a mapping mechanism that combines multi-dimensional verification with deterministic attribution. Traditional methods often involve single-dimensional mapping followed by manual correction, which is inefficient and prone to errors. For example, if content is categorized as an outdoor coffee pot, priced between 50 and 150 yuan, and set in a camping setting, the corresponding sub-category is this three-dimensional combination. Unique attribution can be determined without manual intervention, ensuring both mapping accuracy and processing efficiency, and adapting to the needs of large-scale data processing.
[0036] The specific composition of category set G: The product leaf category list is determined based on the category tree of mainstream e-commerce platforms, such as the official category system of Taobao and JD.com. The category hierarchy is divided into second-level categories under first-level categories, third-level categories under second-level categories, and so on down to the leaf categories at the bottom level. It does not include more detailed custom categories. For example, the first-level category is outdoor equipment, the second-level category is outdoor cookware, and the third-level category is coffee pot. Coffee pot is a leaf category. Category set G contains all such third-level and lower leaf categories to ensure comprehensive coverage and clear hierarchy.
[0037] The specific range of the price band set P: The fixed price range is divided into five levels, namely 0 to 50 yuan, 50 to 150 yuan, 150 to 300 yuan, 300 to 800 yuan, and 800 yuan and above. The boundary of the range is set based on the price distribution data of e-commerce products. 0 to 50 yuan covers low-priced traffic-driving products, 50 to 150 yuan covers mainstream best-selling products, 150 to 300 yuan covers mid-to-high-end products, 300 to 800 yuan covers high-end products, and 800 yuan and above covers luxury or high-end customized products. The span of each range gradually increases, which conforms to the long-tail characteristic of price distribution.
[0038] The specific content of the scenario keyword set includes 30 high-frequency scenario keywords such as commuting, camping, gift-giving, oil control, sensitive skin, wedding, maternal and infant, outdoor, office, student, fitness, travel, home, kitchen, car, beauty, skincare, digital products, etc. The keyword list is updated every quarter. By analyzing the search keywords, user reviews and product titles of e-commerce platforms in the previous quarter, scenario keywords with increased usage frequency are added and scenario keywords with extremely low usage frequency are deleted to ensure the timeliness and practicality of the keyword list.
[0039] The rules for constructing category feature keyword sets are as follows: Keyword sources are high-frequency words from the titles, attribute descriptions, user search terms, and industry news of products in the corresponding category. The number of keywords should be controlled between 10 and 30. The selection criteria are the top 30 words with the highest frequency of occurrence and those that are strongly related to the category. General and ambiguous words are excluded. For example, for the keyword selection of the outdoor coffee pot category, the most frequently occurring words are extracted from the product titles of this category. General words such as coffee pot are removed, and strongly related words such as camping coffee pot, portable coffee pot, and outdoor coffee brewing equipment are retained.
[0040] The specific settings for the default price range are as follows: The default price range for each product leaf category is the median price range of the transaction prices of products in that category over the past six months. By statistically analyzing the transaction price data of products in that category on e-commerce platforms, the median is calculated, and the default value is determined based on the fixed price range in which the median falls. For example, the median transaction price of outdoor coffee pots over the past six months is 98 yuan, which falls into the range of 50 to 150 yuan. Therefore, the default price range is 50 to 150 yuan, ensuring that the default value conforms to the actual price distribution of the category.
[0041] The specific content of the scenario dictionary: The scenario dictionary contains all the words in the scenario word set and their corresponding common variations. The number of entries is 2 to 3 times the number of scenario words. The entry selection logic is to include common expressions, synonyms and related derivatives of scenario words. They are stored in categories according to scenario type. For example, the travel scenario includes entries such as commuting, travel, outdoor, camping, etc. The beauty scenario includes entries such as oil control, sensitive skin, moisturizing, whitening, etc. For example, the camping scenario includes entries such as camping, wilderness camping, outdoor camping, suburban camping, etc., to ensure that it covers the expression habits of different users.
[0042] Execution details of longest string matching: When multiple terms of equal length exist, they are sorted by priority in the scene dictionary, and the term with higher priority is selected. The priority is set according to the frequency of term usage. Case sensitivity and Chinese and English punctuation are ignored during matching, and only the text content is compared. For example, if the text content is "camping coffee pot", and the scene dictionary includes "camping" and "wilderness camping", the longest term "wilderness camping" is selected during matching; if the text content is "Camping coffee pot", the term "camping" is matched after ignoring case sensitivity.
[0043] The calculation boundary of the statistical score: When the keyword hit score of all product leaf categories is 0, the content data is classified into other categories. Other categories are preset unified categories used to include content data that cannot be matched with a specific leaf category. For example, if a certain content only mentions coffee but does not involve specific category feature keywords, the score of all categories is 0, so it is classified into other categories to avoid the situation where the data has no affiliation.
[0044] Preferably, multi-dimensional business indicators are fused to generate a track positioning weight set, which represents the degree of belonging of the target object to different sub-tracks, including:
[0045] For each of the preset sub-tracks The topic search index is calculated based on the data mapped to this track. Competitor interaction popularity and intensity of the period when the population is active These constitute the multidimensional business indicators;
[0046] Detect the current cycle If any of the multidimensional business metrics are missing, then fill them in using the following formula:
[0047]
[0048] in, Represents any one of the aforementioned multidimensional business metrics. The preset attenuation coefficient;
[0049] Using the preset length Calculate the mean of historical data using a scrolling window. with standard deviation The multidimensional business indicators are standardized to obtain normalized business indicators. , and :
[0050]
[0051] in, To prevent extremely small constants from being divided by zero, the standard deviation calculation formula includes this constant;
[0052] Utilize preset business weights The normalized business metrics are linearly weighted and summed to obtain each of the preset sub-sectors. External opportunity score :
[0053]
[0054] By introducing a temperature coefficient The exponential normalization function for the external opportunity score Calculations are performed to generate a normalized set of track positioning weights. , of which Weight of each track for:
[0055]
[0056] in, The total number of the preset sub-tracks.
[0057] The topic search index is calculated based on publicly available data mapped to specific market segments, reflecting the degree of user search interest in topics within that segment. Competitor interaction heat is calculated based on publicly available data mapped to specific market segments, reflecting the level of user interaction with competitor content within that segment. Audience activity intensity during specific time periods is calculated based on publicly available data mapped to specific market segments, reflecting the intensity of attention or participation from the target audience in content related to that segment during a specific time period. Multidimensional business metrics are a collection of topic search index, competitor interaction heat, and audience activity intensity, used to comprehensively evaluate market opportunities within specific market segments. The period is the time cycle for data processing, used to unify the time unit for data collection, calculation, and updates, preferably 1 hour, to meet the real-time requirements of e-commerce advertising, enabling timely capture of market changes while ensuring sufficient data volume to support calculations. The decay coefficient is a preset coefficient used to fill in missing data, preserving data correlation through historical data decay, preferably 0.9, to balance the reference value of historical data with the time decay effect, avoiding drastic interference from missing value filling in subsequent calculations. The rolling window length is a preset window length for calculating the historical data mean and standard deviation, used to dynamically adapt to data changes. A preferred length is 168 hours (7 days), as the e-commerce market exhibits weekly fluctuations, and a 7-day window covers the entire fluctuation cycle, ensuring the reasonableness of the statistical results. The mean is the historical average of the multi-dimensional business indicators within the rolling window, used as a benchmark for standardization. The standard deviation is the historical standard deviation of the multi-dimensional business indicators within the rolling window, used to measure data dispersion and support standardization calculations. The division-to-zero constant is a preset constant used to avoid division by zero in standard deviation calculations, preferably 10 to the power of -6. This value is extremely small and will not affect the actual calculation result of the standard deviation, while effectively avoiding division-to-zero errors. The normalized business indicator is the result of standardizing the topic search index, eliminating the influence of data dimensions and facilitating the fusion calculation of multiple indicators. The normalized business indicator is the result of standardizing the competitor interaction heat, possessing a unified scale and allowing direct comparison with other indicators. The normalized business indicator is the result of standardizing the intensity of active user periods, ensuring fair weighting of different dimensional indicators during fusion. The business weight is a preset weighting coefficient corresponding to the topic search index, used to adjust the importance of this indicator in opportunity assessment. A value of 0.4 is preferred because the topic search index reflects user demand and has a relatively high weighting in market opportunity assessment. The business weight is a preset weighting coefficient corresponding to the competitor interaction heat, used to set the weight of this indicator in the integrated calculation. A value of 0.4 is preferred because competitor interaction heat reflects the level of competition in the market. The business weight is a preset weighting coefficient corresponding to the intensity of user activity during peak hours, used to allocate the integrated weight of this indicator. A value of 0.2 is preferred because the intensity of user activity during peak hours is a secondary assessment indicator with a relatively minor impact on opportunity judgment.The external opportunity score is a weighted sum of the sub-segments' scores in the current period, calculated from normalized business indicators according to preset weights, used to quantify the market opportunity size of the segment. The temperature coefficient is a preset parameter in the exponential normalization function used to control the sharpness of the weight distribution, preferably 0.7. This value balances the concentration and dispersion of weights, preventing any segment from having an excessively high weight or an overly even distribution. The segment positioning weight set is a collection containing the normalized weights of all sub-segments, used to present the relative importance of each segment as a whole. The sub-segment weight is the positioning weight of a sub-segment in the current period, representing its relative priority. The total number of sub-segments is a preset total number, determined by the number of combinations of categories, price ranges, and scenario keywords, preferably between 1000 and 5000, thus covering mainstream e-commerce business scenarios, ensuring segmentation while avoiding excessive computational complexity. We select a three-dimensional combination of topic search index, competitor interaction intensity, and audience activity intensity during peak hours as the core business indicator. Traditional evaluations often use a single two-dimensional indicator, such as popularity or popularity plus conversion, ignoring the impact of audience activity times on campaign timing. For example, the camping sector has high topic search index and competitor interaction intensity, but the target audience is mainly active on weekends. The three-dimensional indicator can simultaneously capture demand, competition, and time-specific characteristics, making market opportunity assessment more comprehensive and suitable for the actual needs of e-commerce campaigns. When data is missing, we use a decay coefficient multiplied by the filling rule of the previous period's data. Traditional methods often use zero filling or mean filling. Zero filling loses historical correlation, and mean filling cannot reflect the time-series changes of data. For example, if a sub-sector is missing data on Monday due to data collection failure, we use 0.9 of Sunday's data to fill it, which preserves the historical trend characteristics of the sector and reflects the impact of time through decay, avoiding excessive deviation between the filled value and the actual situation. Standardization is achieved by calculating the mean and standard deviation using a fixed-length rolling window. Traditional standardization often uses a global fixed threshold, which cannot adapt to the dynamic changes in the e-commerce market. For example, before the Spring Festival, e-commerce consumption is constantly rising. The rolling window can incorporate the latest data in real time to update the mean and standard deviation, allowing the standardization results to be dynamically adjusted according to market changes. However, the global threshold can lower the benchmark due to low-intensity data in the early stages, leading to distortion in the standardized data later. External opportunity scores are generated by linearly weighting and summing based on preset business weights. Multi-dimensional heterogeneous indicators are difficult to compare directly due to their different dimensions and meanings. This method transforms them into a unified quantitative standard. For example, the numerical range of topic search index may be 0 to 10000, and the intensity of active period of the population may be 0 to 100. After standardization, weighted calculation can integrate indicators of different scales into a single score, solving the problem of collaborative evaluation of multiple indicators.An exponential normalization function incorporating a temperature coefficient is used to generate a continuous positioning weight set. Traditional methods often employ hard classification or proportional normalization. Hard classification fails to reflect the cross-segment attributes, while proportional normalization leads to overly even weight distribution. For example, if a product fits both the camping and outdoor coffee segments, a continuous weight set can allocate weights of 0.6 and 0.4 to the two segments, achieving flexible allocation and adapting to the actual cross-segment attributes of e-commerce products. Hard classification, on the other hand, forces it into a single segment, resulting in positioning bias. The specific calculation logic for the topic search index is as follows: the search indices of all topics or trending items mapped to the segment are summed. The search index for a single topic uses the official index provided by the data source platform. If the platform does not provide one, the number of searches is multiplied by a coefficient of 1 to convert it into an index. The summation is the topic search index for that segment. For example, three topics mapped to the camping coffee pot segment have search indices of 2000, 1500, and 1000 respectively. After summing, the topic search index for that segment is 4500. The specific calculation rules for competitor interaction popularity are as follows: Interaction data includes likes, comments, favorites, and reposts, with weighting coefficients set to 1, 3, 2, and 4 respectively. The calculation method is that the interaction popularity of a single competitor's content is equal to the number of likes multiplied by 1, the number of comments multiplied by 3, the number of favorites multiplied by 2, and the number of reposts multiplied by 4. The interaction popularity of all competitor content mapped to this track is summed to obtain the competitor interaction popularity of this track. For example, if a competitor's content has 50 likes, 20 comments, 15 favorites, and 10 reposts, its interaction popularity is 50×1+20×3+15×2+10×4=50+60+30+40=180. Definition and calculation of audience activity intensity during peak hours: Peak hours are divided into hours, with each period containing 24-hour segments. The number of content posts mapped to the track in each hour segment is counted (the number of posts is used as a substitute when there is no exposure data). The number of posts in a given hour segment is the intensity value for that time period. The audience activity intensity during the current period is the maximum of the intensity values of all hour segments. For example, if the number of content posts in the camping track at 10 AM on Saturday is 80, which is the highest for the day, the audience activity intensity for that period is 80. The attenuation coefficient is specifically set to 0.9. This value ensures that after filling in missing data, the error between the subsequent calculation results and the complete data scenario is controlled within 5%, balancing accuracy and practicality. The rolling window length is specifically set to 168 hours, or 7 days. The logic behind this is based on the analysis of e-commerce market data over the past year, which revealed that the popularity changes of most tracks exhibit a 7-day cyclical fluctuation. A 7-day window can fully cover one fluctuation, ensuring that the mean and standard deviation reflect the true data characteristics. The specific value of the division-to-zero minimum constant is 10 to the power of -6. This value is much smaller than the actual standard deviation of the business indicators and will not have a substantial impact on the standardization results. It is only used to avoid division-to-zero errors in extreme cases.The specific values for the business weights are: topic search index 0.4, competitor interaction popularity 0.4, and audience activity intensity during peak hours 0.2. These values are set based on regression analysis of historical campaign data. This weight combination maximizes the correlation between external opportunity scores and actual campaign performance, resulting in optimal evaluation accuracy. The temperature coefficient is set to 0.7. Its impact on the weight distribution is as follows: when the temperature coefficient is greater than 1, the weight distribution tends to be even; when it is less than 0.5, the weight is concentrated in a few high-scoring tracks; and at 0.7, a balance is achieved between concentration and dispersion, making it suitable for most e-commerce business scenarios.
[0058] The total number of sub-segments is determined by multiplying the number of leaf categories, the number of fixed price ranges, and the number of scenario keywords. For example, with 100 leaf categories, 5 price ranges, and 20 scenario keywords, the total number is 100 × 5 × 20 = 10,000. The number of scenario keywords can be adjusted according to the scale of the business. Small businesses can reduce the number of scenario keywords to 10, and the total number should be controlled within 5,000.
[0059] Boundary handling for standardized calculations: When there is insufficient historical data in the scrolling window, if the data volume is greater than or equal to 3 periods, the mean and standard deviation are calculated using the existing data; if there are less than 3 periods, the mean is taken as the initial value of the data in the current period, and the standard deviation is taken as the default standard deviation of the historical statistics of the indicator (such as the default standard deviation of the topic search index is 1000). For example, if only 2 days of data are collected in the initial period, the mean is taken as the average of the data in these 2 days, and the standard deviation adopts the default value.
[0060] External opportunity score range and anomaly handling: The normal score range is from -5 to 5. When the score is greater than 5, it is truncated to 5. When it is less than -5, it is truncated to -5. This is to avoid extreme outliers affecting subsequent weight calculations. For example, if a certain track's score reaches 8 due to a sudden hot topic, it is truncated to 5 according to the rules to prevent the weight of this track from being too high and squeezing out the advertising resources of other tracks.
[0061] Preferably, an auction feedback factor is calculated based on the ratio of the change in unit conversion cost between adjacent periods to the recommended frequency of placement. This auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs, including:
[0062] Set a preset ripening waiting window For the current cycle Get the consumption amount for this period. and the number of maturity conversions confirmed within the maturity waiting window. ;
[0063] Dividing the consumption amount by the mature conversion number yields the unit conversion cost for the current period. And in the next cycle After the aforementioned maturity waiting window Then, the unit conversion cost for the next cycle is calculated similarly. The calculation formula is as follows:
[0064]
[0065] The difference between the natural logarithm of the unit conversion cost in the next period and the natural logarithm of the unit conversion cost in the current period is calculated as the change in the unit conversion cost.
[0066] Divide the magnitude of the change by the recommended delivery frequency for the current period. With the preset minimum constant The sum of these factors yields the auction feedback factor. The calculation formula is as follows:
[0067]
[0068] The maturity waiting window is a preset time window used to confirm the number of mature conversions, ensuring complete data collection. 24 hours is preferred, as conversion lag in the e-commerce industry is typically between 12 and 24 hours, covering over 80% of effective conversions and balancing data accuracy and timeliness. Cost expenditure is the total advertising expenditure for the current period, i.e., the total actual advertising fees paid. This can be obtained through the advertising platform's financial data interface, bill export function, or third-party data analytics tools, directly extracting the expenditure records for the period. The number of mature conversions is the number of clicks or attribution points that occurred within the maturity waiting window, reflecting truly effective conversion results. Cost per conversion (CPC) is the mature unit conversion cost for the current period, calculated by dividing the cost expenditure by the number of mature conversions, and is a core metric for measuring campaign performance. The cost per conversion after the maturity waiting window in the next period is the mature unit conversion cost, forming a benchmark for comparison with the current period's cost. Cost change margin is the natural logarithmic difference between the current and next period's CPC, used to quantify the relative degree of cost change. The recommended placement frequency is the placement frequency output by the current cycle strategy optimization model and is the core action variable driving cost changes. The minimum constant is a preset constant to avoid division by zero when the recommended placement frequency is 0. Ideally, it should be 10 to the power of -6. This value is extremely small and will not affect the calculation result of the ratio of cost change magnitude to placement frequency; it is only used to avoid mathematical calculation errors. The auction feedback factor is a core factor quantifying the endogenous amplification effect of placement actions on auction environment costs, directly reflecting the intensity of the impact of frequency adjustments on costs. A maturity waiting window is introduced to address conversion delay issues. Existing technologies often directly use real-time conversion data to calculate costs, which can easily lead to cost miscalculations due to incomplete conversions. For example, if a user clicks on an ad within a cycle but completes an order conversion 20 hours later, real-time data may miss this conversion, leading to an overestimation of costs. However, by setting a 24-hour maturity waiting window, this conversion will be included in the statistics, making the unit conversion cost calculation more accurate and closely reflecting the actual time distribution characteristics of e-commerce conversions. By employing natural logarithmic transformation, the multiplicative fluctuations of unit conversion costs are converted into additive changes. In e-commerce auction environments, costs are affected by bidding competition and quality signals, often exhibiting multiplicative fluctuation characteristics. For example, if the unit conversion cost in a certain track increases from 100 yuan to 200 yuan (doubling), and then to 400 yuan (doubling again), the linear changes are 100 and 200 respectively, failing to reflect a consistent fluctuation intensity. However, after natural logarithmic transformation, the change is consistently 0.693, accurately reflecting the relative fluctuation pattern and solving the problem that traditional linear calculations cannot adapt to multiplicative fluctuations.By constructing a ratio between cost change and recommended campaign frequency, the system directly links campaign actions with environmental feedback. Existing technologies often focus on either cost or frequency in isolation, failing to quantify the causal relationship between the two. For example, if the recommended campaign frequency increases by 10 times in a certain period, and the cost change in the next period is 0.5, the ratio is 0.05, meaning that for every additional campaign, the cost increases by 5%. This quantification method clearly presents the intrinsic impact of campaign actions on costs, providing a precise basis for subsequent strategy adjustments. Introducing a minimal constant into the denominator solves the division-by-zero problem when the recommended campaign frequency is 0, without altering the core calculation logic. For instance, if no campaigns are conducted in a certain period due to track adjustments, the recommended campaign frequency is 0. In this case, the denominator is the minimal constant 10 to the power of -6. The cost change can be normally divided by this constant to obtain the auction feedback factor, avoiding computational interruption. Furthermore, this constant is extremely small and will not substantially interfere with the calculation results in non-zero frequency scenarios, achieving robust design in extreme scenarios. The preferred duration of the maturity waiting window is 24 hours. This is based on analysis of e-commerce industry advertising data over the past year, which found that 85% of conversions occur within 24 hours of a click. This duration covers the vast majority of effective conversions. It can be fine-tuned based on product type. For fast-moving consumer goods (FMCG) users with short decision-making cycles, it can be set to 12 hours; for durable goods users with longer decision-making cycles, it can be set to 48 hours, ensuring data integrity across different scenarios. The statistical scope for expenditure is the direct advertising cost recorded by the advertising platform, excluding coupon discounts, platform service fees, and taxes. The statistical boundary is based on the platform's financial statements. For example, if the actual advertising cost for a certain period is 5000 yuan, coupon discounts are 1000 yuan, and platform service fees are 300 yuan, the statistical expenditure will be 5000 yuan. The attribution rule for mature conversions uses click attribution, and the attribution validity period is consistent with the maturity waiting window. That is, conversions completed within the maturity waiting window after a user clicks on an ad are counted in the mature conversion count for the corresponding period. If multi-touchpoint attribution is used, the period corresponding to the last valid click is used for statistics. For example, if a user clicks on an ad in two different periods and completes a conversion 10 hours after the second click, this conversion is counted in the period of the second click. The specific value of the minimum constant is 10 to the power of -6. The reasonableness of this value is that this value is much smaller than the minimum actual value of the recommended ad frequency (usually 1). When the ad frequency is non-zero, the denominator is approximately equal to the ad frequency, and the calculation result is not affected. When the ad frequency is zero, the division operation can be completed normally, avoiding system errors.
[0069] The abnormal handling rule for unit conversion cost is that when the number of mature conversions is 0, the unit conversion cost is set to a preset maximum value of 10,000 yuan. This value is higher than the normal cost of most tracks in the e-commerce industry. This can not only avoid the mathematical problem of infinite cost when there is no conversion, but also indicate that the current campaign performance of the track is extremely poor through high cost signals, guiding the strategy model to be adjusted.
[0070] The criteria for defining adjacent periods are that the time boundaries are continuous, and the periods are divided by hours. Period t is the tth hour, and period t+1 is the t+1th hour. The period division is not adjusted across natural days or holidays. For example, the 23rd hour (the previous day) and the 24th hour (the current day) are still considered adjacent periods to ensure the continuity of the time series and the consistency of data comparison.
[0071] The rule for truncating outliers in cost changes is that when the absolute value of the difference in the natural logarithm is greater than 2, it is truncated to ±2. For example, if a certain track experiences a sudden increase in costs due to unexpected competition, the difference is 3, and it is truncated to 2 according to the rule; if a sudden decrease in costs is due to favorable policies, the difference is -3, and it is truncated to -2. This is to avoid extreme outliers from distorting the auction feedback factor and affecting the stability of subsequent strategy training.
[0072] Preferably, constructing an enhanced state input that includes the track positioning weight set, the suggested delivery frequency for historical periods, and the auction feedback factor includes:
[0073] Construct the enhanced state input The calculation formula is as follows:
[0074]
[0075] in, The normalized business metrics for the current period, including those for all the aforementioned preset sub-sectors. to Normalized topic search index Competitor interaction popularity and intensity of active periods of the population ;
[0076] The weight set for locating the track in the current period;
[0077] This refers to the auction feedback factor mentioned in the previous cycle;
[0078] The recommended delivery frequency for the previous period;
[0079] symbol This represents a vector concatenation operation that uses the values from the previous cycle to satisfy the causal constraints of real-time inference.
[0080] The enhanced state input is a model input generated by vector concatenation of a normalized business metric set, a track positioning weight set, and historical feedback information. It provides comprehensive state information support for the strategy optimization model. The normalized business metric set includes normalized topic search indices, normalized competitor interaction heat, and normalized intensity of active user periods for all preset sub-tracks; it is a core data combination reflecting the external market state. Historical feedback information is a collective term for the auction feedback factors and suggested placement frequency of the previous period, used to record the correlation data between past placement actions and corresponding environmental feedback. This approach integrates exogenous publicly available data-derived indicators, track positioning results, and endogenous feedback variables to construct an enhanced state input. Traditional models often rely solely on exogenous data or a single state dimension, failing to adapt to the closed-loop characteristics of dynamic auctions and action feedback in e-commerce advertising. For example, in the "Camping Coffee Pot" track, exogenous data-derived indicators include the track's normalized search index, track positioning results have a positioning weight of 0.6, and endogenous feedback variables are the auction feedback factor of 0.05 from the previous period and the frequency of 50 placements. By integrating these three, the model can simultaneously perceive the market environment, track priority, and feedback from past actions, avoiding biased decision-making caused by a single state dimension. It explicitly uses the auction feedback factor and suggested placement frequency from the previous period as historical feedback information, rather than current period data. Traditional designs often overlook temporal causal constraints, easily introducing future information leakage. For instance, if the current period is the 5th hour, misusing the auction feedback factor from the 5th hour (which has not yet been calculated) would allow the model to prematurely obtain feedback that has not yet occurred, leading to training bias. Using historical data from the 4th hour, however, strictly adheres to the causal logic of real-time inference, ensuring the authenticity of the training process. By using vector concatenation to achieve unified integration of heterogeneous indicators, traditional methods struggle to handle the synergistic effects of high-dimensional and heterogeneous data. For example, normalized business indicators are 3000-dimensional high-dimensional data (1000 tracks × 3 indicators), track positioning weight sets are 1000-dimensional continuous vectors, and historical feedback variables are two scalars. Vector concatenation integrates these into a 4002-dimensional structured input, allowing data of different types and dimensions to collaboratively provide decision support for the model, thus solving the technical challenge of integrating heterogeneous data.
[0081] The specific structure of the normalized business metric set is as follows: it is assembled according to the metric type. First, the normalized topic search index of all sub-tracks is arranged in order, then the normalized competitor interaction heat of all tracks is arranged, and finally the normalized audience activity intensity of all tracks is arranged. The dimension matching rule is that the number of sub-tracks is fixed. If a new track is added, the three indicators of the track are added to the end of the set in the order of normalized topic search index, normalized competitor interaction heat, and normalized audience activity intensity, to ensure that the set dimension is always 3 times the total number of sub-tracks.
[0082] Technical details of vector concatenation: A horizontal concatenation method is adopted, in which the normalized business indicator set, track positioning weight set, and historical feedback information are connected horizontally in sequence. The total dimension of the concatenated vector is 3K+K+2=4K+2 (K is the total number of sub-tracks). The dimension verification rule is that the total dimension needs to be calculated after concatenation. If it is not equal to 4K+2, it is judged as a concatenation failure. The data of each part is re-acquired and concatenated again until the dimension meets the requirements.
[0083] The rules for handling missing historical feedback information are as follows: When there is no t-1 period data in the initial period, the auction feedback factor should be filled with 0.000001 first, and the placement frequency should be filled with 10 first. This value is close to the average feedback level and basic frequency of the initial placement in the e-commerce industry. This can avoid the model training oscillation caused by the initial value of 0, and will not cause excessive interference to the initial training results.
[0084] Enhanced anomaly handling for status input: When the absolute value of a normalized business indicator is greater than 5, the sum of the track positioning weight set deviates from 1 by more than 0.1, the absolute value of the auction feedback factor in historical feedback information is greater than 2, or the suggested placement frequency exceeds the range of 0 to 300, it is judged as an anomaly. The handling method is to replace the anomaly with the historical 7-day average of the corresponding part. For example, if a normalized indicator is 6, it is replaced with the average of the indicator over the past 7 days, which is 0.3.
[0085] Scale adaptation of data from different dimensions: No additional secondary normalization is required. The normalized business indicators have been standardized through a rolling window (mean 0, standard deviation 1). The track positioning weight set has been normalized through exponential normalization (sum 1, value 0 to 1). The auction feedback factor in the historical feedback information is naturally in a small range of -2 to 2. The recommended placement frequency has been constrained to 0 to 300 through upper and lower limits. The scales of each part of the data are mutually adapted and can be directly spliced into vectors.
[0086] Preferably, a calibration reward value training strategy is used to optimize the model; the calibration reward value is configured to be obtained by subtracting an auction endogenous penalty term from the observed raw reward data, the auction endogenous penalty term being generated based on the correlation between the suggested placement frequency and the auction feedback factor, including:
[0087] For the current cycle Get the next cycle Unit conversion cost Calculate the original return data The formula is as follows:
[0088]
[0089] Calculate the recommended delivery frequency for the current period. The auction feedback factor for the current period The product of these terms is used as an estimate of the auction endogenous penalty term;
[0090] Calculate the calibration report value The formula is as follows:
[0091]
[0092] Construct the policy optimization model, which includes the following parameters: Policy network and parameters are value network ;
[0093] Calculate the temporal difference objective of the value network and loss function The formula is as follows:
[0094]
[0095]
[0096] in, The preset discount factor, This serves as the input for the enhanced state in the next cycle. The prediction frequency for the next cycle;
[0097] Calculate the deterministic policy gradient of the policy network. And update parameters The formula is as follows:
[0098]
[0099]
[0100] in, This is the preset learning rate.
[0101] The raw return data is the negative natural logarithm of the unit conversion cost for the next cycle, used to transform cost indicators into positive incentive return signals. The auction-endogenous penalty term is the product of the suggested placement frequency for the current cycle and the auction feedback factor, used to quantify the intensity of negative feedback in the auction environment caused by the placement action. The calibrated return value is the true market return after removing negative feedback from the auction environment, obtained by adding the raw return data and the auction-endogenous penalty term, providing unbiased feedback for model training. The strategy optimization model is a dual-network model containing a policy network and a value network, used to output placement frequency and evaluate state value. The policy network is a neural network that outputs the suggested placement frequency, with parameters in the theta format, receiving augmented state input and outputting continuous frequencies. The value network is a neural network that evaluates state value, with parameters in the omega format, outputting the expected return in the current state. The policy network parameters are learnable parameters of the policy network, updated through policy gradients to optimize the frequency output. The value network parameters are learnable parameters of the value network, updated through temporal difference errors to improve the accuracy of value evaluation. The temporal difference objective is the discounted sum of the calibration return value and the estimated value of the next cycle state, used as the training objective of the value network. The discount factor is a preset coefficient used to balance current and future returns, preferably 0.98, because short-term returns account for a high proportion in the e-commerce sector; this value balances immediate effects and long-term benefits, avoiding excessive short-sightedness. The next cycle prediction frequency is the suggested deployment frequency for the next cycle output by the strategy network, used in calculating the temporal difference objective. The value network loss function is the squared loss based on the temporal difference error, used to measure the deviation between the value network's predicted value and the target value. The temporal difference error is the difference between the value network's predicted value and the temporal difference objective, and is the core basis for updating the value network parameters. The policy gradient is a chain product of the value network gradient and the policy network's output gradient, used to guide the direction of policy network parameter updates. The policy network learning rate is a preset step size for updating the policy network parameters, preferably 10 to the power of -4. This step size balances update speed and stability, avoiding parameter oscillations or excessively slow convergence. This paper defines an auction-inherent penalty term, linking campaign actions with auction environment feedback. Existing technologies often overlook this relationship. For example, if a suggested campaign frequency of 50 times is used in a given period, with an auction feedback factor of 0.05 and a penalty term of 2.5, this value accurately quantifies the intensity of the auction penalty caused by increased frequency, solving the problem of not being able to distinguish between internal and external causes of cost changes. A calibration logic using raw return data plus the auction-inherent penalty term is employed to remove data bias introduced by the auction mechanism. For instance, if a certain track experiences an increase in unit conversion cost due to high-frequency campaigning, the raw return data may decrease. The calibration return value, by adding the penalty term, restores the true value of the track, preventing the model from misjudging a track as unpopular and incorrectly reducing campaigns, thus solving the strategy oscillation problem.The original return data is designed as the negative natural logarithm of the unit conversion cost to adapt to the multiplicative fluctuations in costs in the e-commerce auction environment. For example, if the unit conversion cost increases from 100 yuan to 200 yuan and then to 400 yuan, the linear changes are 100 and 200 respectively. After the conversion to the negative natural logarithm, the return changes are consistent, accurately reflecting the relative cost changes. This differs from the limitation of traditional linear return functions, which cannot adapt to multiplicative fluctuations. A dual-network optimization model consisting of a policy network and a value network is constructed to form an evaluation and optimization closed loop. The policy network focuses on output frequency, while the value network evaluates the value of the state. For example, if the value network evaluates the expected return in the current enhanced state as 3.2, the policy network adjusts the frequency of its output accordingly. Compared to a single network directly outputting actions, this improves the robustness and convergence speed of the strategy and avoids the overfitting problem of a single network. The strategy gradient calculation employs a chain rule of strategy network output versus value network gradient, ensuring that the strategy update direction aligns with the value enhancement goal. For example, the value network gradient indicates that increasing frequency in the current state will improve returns, while the strategy gradient guides the strategy network to increase the frequency output, ensuring that the deployment actions are always adjusted towards maximizing the true value of the track, rather than blindly pursuing short-term exposure. The discount factor is specifically set at 0.98, a value that ensures the model maintains a return weight above 0.8 for the next 10 cycles, emphasizing both current performance and long-term benefits, adapting to the 1-2 week short-term return cycle of the e-commerce track. The strategy network is a fully connected neural network with 3 hidden layers, each with 64 neurons. The activation function except for the output layer uses the ReLU function, and the output layer uses the Sigmoid function to constrain the output between 0 and 1. Scaling is then used to obtain the suggested deployment frequency. The network input dimension is the total dimension of the enhanced state inputs: 4K + 2, and the output dimension is 1. The value network is structured as a fully connected neural network with three hidden layers, each containing 64 neurons. The ReLU activation function is used for all layers. The input dimension is 4K + 2, consistent with the policy network, while the output dimension is 1. The input-output dimension matching rule is that the input dimension equals the total dimension of the augmented state inputs, and the output dimension is a single value assessment. The policy network's learning rate is 10 to the power of -4. A fixed learning rate is used for adjustment. Because the system updates data in real-time each cycle, a fixed learning rate ensures training stability, avoids parameter oscillations caused by adaptive learning rates, and simplifies implementation.
[0102] The rule for handling anomalies in temporal difference errors is that when the absolute value of the error is greater than 2, it is truncated to ±2. For example, when the error is 3, it is truncated to 2, and when the error is -3, it is truncated to -2. This is to avoid extreme errors causing the value network parameters to update too much, which would affect the training stability.
[0103] The model parameter initialization rule is that the weight parameters of the policy network and the value network are initialized using Xavier, and the bias parameters are initialized to 0. This initialization method can make the variance of the input and output of each layer of the network consistent, and avoid gradient vanishing or exploding caused by the initial values being too large or too small.
[0104] The model is updated once per cycle, using a single-sample update method. Since the system processes data on an hourly basis, real-time requirements are high. Single-sample updates can respond to market changes in a timely manner without waiting for batch sample accumulation, making it suitable for dynamic bidding environments.
[0105] The rule for truncating outliers in calibration reward values is that the value range is constrained to between -5 and 5. When the value is greater than 5, it is truncated to 5, and when it is less than -5, it is truncated to -5. This avoids extreme outliers from causing model training oscillations and ensures a smooth training process.
[0106] The numerical stability of policy gradient calculation is handled by gradient pruning, which prunes the L2 norm of the gradient to 1.0. When the gradient norm exceeds 1.0, the gradient is scaled proportionally to avoid gradient explosion. At the same time, the BatchNorm technique is introduced into the network layer to normalize the layer input and alleviate the gradient vanishing problem.
[0107] The optimizer for the loss function of the value network is the Adam optimizer, with hyperparameters set to a learning rate of 10^-3, a momentum coefficient of 0.9, and a weight decay coefficient of 10^-5. This optimizer can adaptively adjust the learning rate, balance convergence speed and stability, and adapt to the loss optimization requirements of the value network.
[0108] Preferably, the trained strategy optimization model outputs the suggested delivery frequency for the current period, and based on the principle of total intensity conservation, the suggested delivery frequency is translated into a multi-channel execution task, including:
[0109] Utilizing the policy network in the policy optimization model Receive the enhanced state input for the current period via the Sigmoid activation function The suggested delivery frequency is processed and output in continuous numerical form. The formula is as follows:
[0110]
[0111] in, and These are the preset lower and upper limits of frequency, respectively;
[0112] The recommended frequency of deployment With the preset period duration constant Multiply the products and round the result to the nearest integer to obtain the total target deployment amount. To satisfy the principle of total strength conservation:
[0113]
[0114] in, This indicates the rounding operation;
[0115] Based on preset channel weights Calculate each execution channel Channel allocation ratio The total target delivery volume Multiply by the aforementioned channel allocation ratio Then, rounding is performed to obtain the initial number of tasks allocated to each execution channel. :
[0116]
[0117] in, Total number of channels;
[0118] Calculate the total target delivery volume The difference between the sum of the initial assigned tasks and the sum of the initial assigned tasks across all execution channels yields the rounding remainder error. :
[0119]
[0120] The rounding remainder error Superimposed on the execution channel with the smallest preset number ( The final number of tasks for that channel is generated based on the initial number of tasks allocated. The initial number of tasks allocated to other execution channels is used as the final number of tasks, and these are combined to form the multi-channel execution task:
[0121]
[0122] The activation function is used to normalize the output of the strategy network. Here, the Sigmund function is preferred, as it constrains the network output between 0 and 1. The Sigmund function is ideally chosen because it smoothly maps the output range, adapts to frequency scaling requirements, and avoids interference from extreme values. The lower frequency limit is a preset minimum suggested delivery frequency, used to limit the minimum threshold for delivery frequency. It is preferably 0 to support scenarios where delivery is paused and to accommodate situations where the track has no delivery value. The upper frequency limit is a preset maximum suggested delivery frequency, used to constrain the highest threshold for delivery frequency. It is preferably 300, referencing the maximum hourly capacity of a single e-commerce channel to avoid resource waste or platform limitations caused by high-frequency delivery. The cycle duration constant is a preset cycle duration used to calculate the total target delivery volume, consistent with the system data processing cycle, preferably 1 hour. The total target delivery volume is the integer result of the product of the suggested delivery frequency and the cycle duration constant, serving as the total base for multi-channel task allocation. Channel weight is a preset weight assigned to each execution channel, used to set the task allocation priority for each channel. Ideally, it's a weight combination based on the conversion rate of each channel over the past three months, such as 2 for Xiaohongshu, 1 for Douyin, and 1 for Kuaishou, thus allowing high-conversion channels to receive more advertising resources. Execution channel is the specific channel for ad placement, the vehicle for task implementation. Ideally, it includes mainstream e-commerce marketing channels such as Xiaohongshu, Douyin, Kuaishou, and Taobao Express, to cover platforms where target users are concentrated and improve reach efficiency. Channel allocation ratio is the task allocation ratio for each channel calculated based on the channel weight, used to split the total target ad volume. Total number of channels is the preset total number of ad execution channels, ideally 3 to 5, to balance ad coverage and management complexity, avoiding resource dispersion due to too many channels. Initial number of allocated tasks is the initial number of tasks calculated for each channel according to the allocation ratio, rounded to the nearest integer. Rounding remainder error is the difference between the total target ad volume and the sum of the initial number of tasks for all execution channels, generated by the rounding operation. The minimum-numbered execution channel is the preset remainder error receiving channel, fixed at channel number 1, preferably channel number 1, to establish deterministic processing rules and avoid chaotic error allocation. The final task number is the number of tasks after adding the remainder error to the minimum-numbered channel, ensuring that the total task volume is consistent with the total target delivery volume. The final task number of other channels is the final task number of channels with numbers greater than 1, retaining the initial allocated task number unchanged. Multi-channel execution tasks are the set of the final task numbers of all execution channels, which is a list of tasks that the delivery platform can directly execute. A combination of frequency upper and lower limits plus a Sigmund activation function is used. Traditional designs often directly output discrete frequencies or unconstrained continuous frequencies, which are prone to frequency anomalies. For example, if the strategy network outputs 0.8, after calculation with the activation function and upper and lower limits, the suggested delivery frequency is 0 + (300 - 0) × 0.8 = 240, which is within a reasonable range. If there is no constraint, the network output may be 2, resulting in a frequency of 600, exceeding the channel's capacity. This combination ensures both strategy flexibility and avoids frequency anomalies.The principle of total intensity conservation is proposed. Traditional multi-channel allocation often ignores the consistency between the intensity of the campaign and the strategy output, leading to deviations between actual and expected deployments. For example, if the suggested frequency is 200, the cycle duration is a constant of 1 hour, and the total target deployment volume is 200, the total number of tasks after multi-channel allocation will still be 200. This ensures that the intensity of the campaign is consistent with the strategy output and avoids under- or over-deployment due to allocation splitting. A three-level allocation logic is designed, from channel weight to allocation ratio to initial task number. Traditional methods often use simple average allocation or direct allocation based on weight, which cannot balance fairness and priority. For example, if the channel weights are 2, 1, and 1, the allocation ratios are 0.5, 0.25, and 0.25, and the total target deployment volume is 100, the initial task numbers are 50, 25, and 25. This not only aligns with weight priority but also achieves reasonable resource allocation, solving the multi-channel allocation problem. To address the remainder error caused by rounding, a deterministic processing rule is adopted to aggregate the results to the channel with the smallest preset number. Traditional methods often involve cyclical allocation or discarding remainders, which are inefficient and prone to errors. For example, if the total target delivery volume is 101, the initial task counts are 50, 25, and 25, with a remainder of 1. After being aggregated to channel number 1, the final task counts are 51, 25, and 25, while the total task count remains 101. This approach ensures the accuracy of the total count while avoiding complex logic, improving execution efficiency and reproducibility. A complete translation process is implemented, from continuous frequency to total delivery volume and then to multi-channel tasks. Traditional technologies involve multiple strategy outputs and lightweight execution translation, leading to implementation deviations. For example, if the strategy outputs a continuous frequency of 200.5 and a total target delivery volume of 201, after channel allocation and remainder processing, the final task count for each channel is directly generated without additional branch judgments. This adapts to the actual scenario of multi-channel collaborative delivery in e-commerce, ensuring the accurate implementation of strategy results. The minimum and maximum frequency values are 0 and 300, respectively. These values are based on experience in e-commerce advertising, where exceeding 300 times per hour on a single channel significantly reduces marginal returns and easily triggers the platform's anti-spam mechanism. A value of 0 allows for a complete pause in advertising. The cycle duration constant is set to 1 hour, perfectly matching the system's data processing cycle to ensure the recommended frequency matches the time granularity of task allocation and avoids data misalignment across cycles.
[0123] The channel weight setting rules are as follows: based on the conversion volume ratio of each channel in the past 3 months, the conversion volume ratio = conversion volume of a certain channel / total conversion volume of all channels, and the weight = conversion volume ratio × 10 (rounded to the nearest integer). For example, Xiaohongshu's conversion volume ratio is 20%, and the weight is 2; Douyin's conversion volume ratio is 10%, and the weight is 1. The calibration method is to recalculate the conversion volume ratio and update the weight every month.
[0124] The total number of channels is 3 to 5, and channels can be dynamically added or removed. When adding a new channel, the weight is allocated according to the initial conversion rate of the new channel, and the weight of the original channel is reduced proportionally. When deleting a channel, its weight is evenly distributed to the remaining channels to ensure that the total weight is a fixed value.
[0125] The rule for handling negative integer remainder errors is as follows: When the total target delivery volume is less than the sum of the initial task numbers, the error is still deducted from the channel with the smallest number. For example, if the total target delivery volume is 99 and the initial task numbers are 50, 25, and 25, the remainder is -1, and the final task number for channel number 1 is 49. Other channels remain unchanged to ensure that the total number of tasks is consistent with the total target delivery volume.
[0126] The available activation functions include Sigmoid function and ReLU function. The replacement condition is: if ReLU function is used, a truncation operation must be added after the output to constrain the result to between 0 and 1 to avoid exceeding the frequency scaling range.
[0127] The task distribution order for multi-channel execution is as follows: tasks are distributed in ascending order of channel number. The retry rules after a failed distribution are as follows: after the first failure, a retry will be made within 10 minutes, with a maximum of 3 retries. If all 3 retries fail, the task for that channel will be temporarily assigned to the channel with the next smallest number to ensure that no task is missed.
[0128] The upper limit constraint for the number of tasks in a channel is 500 per cycle for a single channel. When the initial number of tasks exceeds this limit, the initial number of tasks in all channels is reduced proportionally according to the weight ratio of each channel until the total number of tasks in all channels does not exceed the limit. For example, if the initial number of tasks in channel 1 is 550 and the upper limit is 500, then the number of tasks in all channels will be reduced by a ratio of 500 / 550.
[0129] Example 2: A method for intelligent identification and positioning of marketing segments, applied to any of the marketing segment intelligent identification and positioning systems described herein, including:
[0130] Collect publicly available data from multiple sources and map it to preset sub-tracks. Perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, which represents the degree of belonging of the target object to different sub-tracks.
[0131] The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs.
[0132] An enhanced state input is constructed, which includes the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor. The model is optimized by training a strategy using a calibration reward value. The calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data. The auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor.
[0133] The trained strategy optimization model outputs the suggested delivery frequency for the current period, and the suggested delivery frequency is translated into multi-channel execution tasks based on the principle of total strength conservation.
[0134] Example 3: This example mainly optimizes the selection and fusion processing logic of multi-dimensional business indicators in the market data aggregation unit.
[0135] In the current competitive market environment, considering only user-side search and interaction activity (i.e., market demand) can easily lead the system to mistakenly enter an extremely competitive market. To more accurately identify high-potential markets with high demand but low supply, this embodiment introduces market supply-side data as a key evaluation dimension.
[0136] The specific implementation steps are as follows:
[0137] Step 1: Construct a multi-dimensional business indicator system incorporating supply-side data. The market data aggregation unit is configured to simultaneously collect the number of ad placements mapped to the specific sub-sector, in addition to collecting topic search indexes, competitor interaction popularity, and the intensity of user activity during peak hours. The number of ad placements represents the market congestion and competition saturation of the specific sub-sector within the current statistical period. Specific data sources include, but are not limited to: the number of newly added commercial notes, the number of short videos promoting sales, and the number of in-feed ads placed within the specific sub-sector.
[0138] Step Two: Standardization of Metric Data. For the collected content delivery data, standardization is performed using the same logic as for the topic search index. First, it checks for missing data in the current period's content delivery count. If missing data is found, it is filled by multiplying the previous period's value by a preset decay coefficient. Second, a preset rolling window is used to calculate the mean and standard deviation of the content delivery count over historical periods, transforming the current period's value into a dimensionless normalized value, allowing it to be compared on the same scale as topic search index, competitor interaction intensity, and the intensity of peak user activity.
[0139] Step 3: Calculating Track Opportunity Scores Based on Supply-Demand Game Theory. When generating external opportunity scores for sub-tracks, a calculation logic of positive and negative weighting is adopted:
[0140] Set positive gain metrics: Set topic search index, competitor interaction popularity, and intensity of user activity during peak hours as positive metrics. The higher these three values are, the stronger the market demand, and the more positively they contribute to the opportunity score of the track.
[0141] Set a negative penalty indicator: Set the number of content submissions as a negative indicator. The higher this value, the more the market supply is excessive and the competition is fierce, resulting in a negative deduction (i.e., penalty) on the opportunity score of the track.
[0142] Calculate the net opportunity score: Using preset business weights, the normalized values of the above positive indicators are weighted and summed to obtain the market demand score; simultaneously, using preset congestion weights, the weighted result of the normalized values of the number of content placements is calculated. Subtracting the result calculated by the congestion weight from the market demand score yields the supply and demand net opportunity score for this sub-segment.
[0143] Step 4: Generating Track Positioning Weight Sets. Using an exponential normalization function incorporating a temperature coefficient, the supply and demand net opportunity scores for all sub-tracks are processed. During this process, for tracks with excessively high content submissions resulting in negative net opportunity scores, their normalized weights will approach zero, automatically reducing the system's focus on overly competitive tracks. Conversely, for tracks with high search popularity and relatively low content submission volumes, their weights will be significantly increased. The resulting track positioning weight sets will precisely target high-quality sub-tracks that combine high market demand with low competitive pressure.
[0144] like Figure 2 As shown, Figure 2 The diagram illustrates the implementation scenario of a marketing segment intelligent identification and positioning system: First, publicly available data is obtained from Xiaohongshu and Douyin and fed into the system's market data aggregation unit. After processing by the cost sensitivity analysis unit and strategy optimization unit, the task distribution unit outputs suggested placement frequencies to the unified advertising placement management platform. This platform pushes the information to e-commerce store product pages, influencer content publishing, and other channels. After a user places an order, conversion and consumption data are fed back to the system, completing the entire process of data collection → analysis and optimization → placement execution → effect feedback.
[0145] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A marketing segmentation intelligent identification and positioning system, characterized in that, include: The market data aggregation unit is configured to collect public data from multiple sources and map it to preset sub-sectors, and to perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, wherein the track positioning weight set represents the degree of belonging of the target object in different sub-sectors. The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs. The strategy optimization unit is configured to construct an enhanced state input including the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor, and to train the strategy optimization model using the calibration reward value; the calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data, and the auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor; The task distribution unit is configured to use the trained strategy optimization model to output the suggested delivery frequency for the current period, and translate the suggested delivery frequency into multi-channel execution tasks based on the principle of total strength conservation.
2. The marketing segmentation intelligent identification and positioning system according to claim 1, characterized in that, Collect publicly available data from multiple sources and map it to preset sub-tracks, including: The pre-defined segmented tracks include: a three-dimensional Cartesian product combination consisting of product leaf categories, fixed price ranges, and application scenario keywords; For each piece of collected multi-source public content data, perform the following mapping operation: The text content of the multi-source public content data is matched and calculated using a pre-set set of category feature keywords. Based on the statistical score of keyword hits, the multi-source public content data is classified into the product leaf category with the highest score. The price information in the multi-source publicly available content data is analyzed to determine the fixed price range to which it belongs; The longest string match is performed on the multi-source public content data using a pre-built scenario dictionary, and the longest matched term is taken as the application scenario term. By combining the product leaf categories, the fixed price range, and the application scenario terms, the multi-source public content data is mapped to the corresponding preset sub-segments.
3. The intelligent identification and positioning system for marketing segmentation according to claim 2, characterized in that, Multi-dimensional business indicators are fused to generate a track positioning weight set, which represents the degree of belonging of the target object to different sub-tracks, including: For each of the preset sub-tracks, the topic search index, competitor interaction popularity, and intensity of active user time periods are calculated based on the data mapped to that sub-track, forming the multi-dimensional business indicators; Fill in missing data for the multidimensional business metrics in the current period; The mean and standard deviation of historical multidimensional business indicators are calculated using a preset rolling window, and the multidimensional business indicators of the current period are standardized to obtain normalized business indicators. The normalized business indicators are linearly weighted and summed using preset business weights to obtain the external opportunity score for each preset sub-segment. The external opportunity score is calculated using an exponential normalization function that includes a temperature coefficient, generating a normalized track positioning weight set.
4. The intelligent identification and positioning system for marketing segmentation according to claim 3, characterized in that, Based on the ratio of the change in unit conversion cost between adjacent periods to the recommended frequency of placement, an auction feedback factor is calculated. This auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs, including: Set a maturity waiting window, and for the current period, obtain the consumption amount of the current period and the number of maturity conversions within the maturity waiting window; Divide the consumption amount by the number of mature conversions to obtain the unit conversion cost for the current period, and calculate the unit conversion cost for the next period after the maturity waiting window has passed in the next period. The difference between the natural logarithm of the unit conversion cost in the next period and the natural logarithm of the unit conversion cost in the current period is calculated as the change in the unit conversion cost. The auction feedback factor is obtained by dividing the magnitude of the change by the sum of the suggested delivery frequency in the current period and the preset minimum constant.
5. The intelligent identification and positioning system for marketing segmentation according to claim 4, characterized in that, Construct an enhanced state input that includes the track positioning weight set, the suggested delivery frequency for historical periods, and the auction feedback factor, including: Obtain the normalized business metrics for the current period, which include standardized values for topic search index, competitor interaction popularity, and intensity of active user periods for all the preset sub-sectors. Obtain the track positioning weight set for the current period; Obtain the auction feedback factor and the suggested placement frequency of the previous period as historical feedback information; The normalized business metrics, the track positioning weight set, the auction feedback factor from the previous period, and the suggested delivery frequency from the previous period are vectorized and concatenated to generate the enhanced state input.
6. The intelligent identification and positioning system for marketing segmentation according to claim 5, characterized in that, The model is optimized using a training strategy based on calibration reward values. These calibration reward values are configured to be obtained by subtracting an auction-endogenous penalty term from the observed raw reward data. The auction-endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor, and includes: For the current period, obtain the unit conversion cost for the next period, calculate the negative value of its natural logarithm, and use it as the observed raw return data; Calculate the product of the suggested delivery frequency in the current period and the auction feedback factor in the current period, and use it as the auction endogenous penalty term; The original return data is added to the auction endogenous penalty term to obtain the calibrated return value, which serves as the true market return after removing the negative feedback of the auction environment. Construct the policy optimization model, which includes a policy network for outputting actions and a value network for evaluating state values; Using the discounted sum of the calibration reward value and the estimated value of the next cycle state as the target value, the temporal difference error of the value network is calculated, and the parameters of the value network are updated based on the temporal difference error. The policy gradient is calculated using the chain rule based on the gradient of the value network with respect to the output of the policy network, and the parameters of the policy network are updated based on the policy gradient to maximize the evaluation value of the value network.
7. The intelligent identification and positioning system for marketing segmentation according to claim 6, characterized in that, The trained strategy is used to optimize the model, which outputs the suggested delivery frequency for the current period. Based on the principle of total intensity conservation, the suggested delivery frequency is translated into multi-channel execution tasks, including: The policy network in the policy optimization model receives the enhanced state input for the current period, and outputs the suggested delivery frequency in continuous numerical form via an activation function; Multiply the suggested delivery frequency by the preset cycle duration constant, and round the product to the nearest integer to obtain the total target delivery amount, so as to satisfy the principle of total intensity conservation. The channel allocation ratio for each execution channel is calculated based on the preset channel weight. The total target delivery volume is multiplied by the channel allocation ratio and rounded to the nearest integer to obtain the initial number of tasks allocated to each execution channel. Calculate the difference between the total target delivery volume and the sum of the initial allocated tasks for all execution channels to obtain the integer remainder error; The rounding remainder error is added to the initial task number of the execution channel with the smallest preset number to generate the final task number of that execution channel. The initial task numbers of the other execution channels are used as their final task numbers, and the combination constitutes the multi-channel execution task.
8. A method for intelligent identification and positioning of marketing segments, applied to the intelligent identification and positioning system for marketing segments as described in any one of claims 1-7, characterized in that, include: Collect publicly available data from multiple sources and map it to preset sub-tracks. Perform fusion processing on multi-dimensional business indicators to generate a track positioning weight set, which represents the degree of belonging of the target object to different sub-tracks. The cost sensitivity analysis unit is configured to calculate the auction feedback factor based on the ratio of the change in unit conversion cost between adjacent periods to the recommended placement frequency. The auction feedback factor is used to quantify the endogenous amplification effect of placement actions on auction environment costs. An enhanced state input is constructed, which includes the track positioning weight set, the suggested delivery frequency of the historical period, and the auction feedback factor. The model is optimized by training a strategy using a calibration reward value. The calibration reward value is configured to be obtained by subtracting the auction endogenous penalty term from the observed raw reward data. The auction endogenous penalty term is generated based on the correlation between the suggested delivery frequency and the auction feedback factor. The trained strategy optimization model outputs the frequency of the proposed delivery in the current period, and the frequency of the proposed delivery is translated into multi-channel execution tasks based on the principle of total strength conservation.