Information flow putting operation compass construction and optimization method based on artificial intelligence

By integrating multi-source data and using deep learning to identify user status, the system dynamically generates tailored advertising content and optimizes delivery strategies, solving the problems of data fragmentation and rigid strategies in information flow delivery systems, and enabling real-time personalized decision-making and improved user experience.

CN122089404APending Publication Date: 2026-05-26XIAOKAN TECHNOLOGY (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAOKAN TECHNOLOGY (BEIJING) CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

In existing information flow delivery systems, fragmented data utilization, crude user intent recognition, and rigid content generation and delivery strategies make it impossible to achieve a balance between real-time personalized decision-making and a healthy user experience.

Method used

By collecting and fusing multi-source heterogeneous data in real time, combining lightweight deep learning models to identify user status and avoidance tendencies, dynamically generating appropriate advertising content, optimizing the delivery strategy based on reinforcement learning, and building a visual decision dashboard to support real-time adjustments.

Benefits of technology

It achieves precise matching between ad content and user intent, improves ad relevance and user experience, and dynamically adjusts the delivery strategy to maximize long-term user value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122089404A_ABST
    Figure CN122089404A_ABST
Patent Text Reader

Abstract

The invention discloses an artificial intelligence-based information flow putting operation compass construction and optimization method. The method comprises the following steps: S1, carrying out real-time acquisition and fusion processing on multi-source heterogeneous data; s2, user dynamic state perception and intention recognition; s3, advertisement content intelligent generation and multi-mode adaptation are carried out; s4, putting strategy dynamic adjustment and real-time bidding optimization are carried out; s5, performing effect multidimensional evaluation and compass parameter closed-loop optimization; and S6, a compass visualization and manual cooperative intervention interface. According to the method, multi-dimensional data such as user behaviors, environments, contents, portraits and the like are synchronized and fused through a streaming computing framework and a time window alignment technology, a dynamically updated structured feature matrix is constructed, the problem of data islands is solved, a high-quality and high-timeliness unified data base is provided for subsequent precise modeling, and the modeling efficiency is improved. And the completeness and the real-time performance of model input information are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information processing technology, specifically involving a method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence. Background Technology

[0002] In-feed advertising has become a core form of digital marketing. However, current in-feed advertising delivery and operation systems generally face the following technical bottlenecks:

[0003] Data utilization is fragmented. Multi-source heterogeneous data, such as real-time user behavior, environmental context, advertising content features, and static profiles, are often collected and processed independently, lacking low-latency, high-consistency real-time fusion capabilities. This results in delayed and one-sided judgments of users' instantaneous intentions and states, making it difficult to support truly real-time personalized decision-making.

[0004] User intent identification is crude. Existing technologies mostly rely on lagging indicators such as historical click-through rates or simple rule thresholds (such as dwell time) to infer user interests. They cannot effectively distinguish between different information processing modes (convergent and divergent) such as "deep browsing" and "rapid scrolling," nor can they accurately predict users' immediate avoidance tendencies towards advertisements (such as blocking or quickly skipping), resulting in inaccurate timing of ad delivery and content selection.

[0005] Content generation and delivery strategies are rigid. Advertising content production often relies on pre-built material libraries, making it impossible to dynamically generate highly adapted creatives based on users' real-time status. At the same time, delivery strategies (such as bidding and frequency control) are mostly based on static rules or short-term conversion goal optimization, lacking adaptive decision-making models that are centered on long-term user value and can balance immediate revenue with a healthy user experience.

[0006] Therefore, the industry urgently needs a systematic approach that can integrate end-to-end data, deeply understand users' real-time intentions, dynamically generate adapted content, and make intelligent decisions and continuously optimize in a closed loop with long-term value as the goal, in order to break through existing technological bottlenecks and achieve a dual improvement in the effectiveness of information flow advertising and user experience. Summary of the Invention

[0007] This application provides a method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence, aiming to solve the problems of fragmented data utilization, crude user intent recognition, and rigid content generation and delivery strategies in existing technologies.

[0008] A method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence, the method comprising:

[0009] S1. Real-time acquisition and fusion processing of multi-source heterogeneous data: Integrate real-time user behavior stream data, contextual environment data, advertising content feature data and user static tag data, perform real-time cleaning, normalization and feature extraction, and construct a dynamically updated user-environment-content interaction feature matrix.

[0010] S2. User dynamic state perception and intent recognition: Based on the user-environment-content interaction feature matrix output by S1, a lightweight deep learning model is used to identify the user's current convergent or divergent information processing state and predict the user's real-time avoidance tendency of advertisements.

[0011] S3. Intelligent generation and multimodal adaptation of advertising content: Based on the user status recognition results and avoidance tendency prediction results output by S2, the generative AI module is dynamically invoked to generate differentiated advertising content, and the generated advertising content is multimodal adapted.

[0012] S4. Dynamic Adjustment of Placement Strategy and Real-time Bidding Optimization: Based on the user status and avoidance tendency output by S2, the advertising content generated by S3, and historical data, a real-time bidding and placement strategy model based on reinforcement learning is constructed and run. With the goal of maximizing long-term user value, the model dynamically decides on the timing of ad placement, bidding strategy, and frequency control. The historical data is a specific set of data from the past of users and the platform, which is necessary for the reinforcement learning model to make real-time decisions and is used to support long-term value assessment and current strategy calculation.

[0013] S5. Multidimensional effect evaluation and closed-loop optimization of compass parameters: Define and calculate a multidimensional effect evaluation index system that includes short-term indicators, avoidance indicators, experience indicators and long-term indicators; based on the evaluation results, automatically optimize the model hyperparameters and decision thresholds in S2 to S4 through an asynchronous parallel Bayesian optimization framework.

[0014] S6, Compass Visualization and Human Collaborative Intervention Interface: Construct a visual decision dashboard to display in real time user status distribution, content adaptability, strategy effectiveness and effect trend data generated from S1 to S5; and provide an intervention interface to allow the injection of expert rules or fine-tuning of strategies in specific scenarios, with the intervention instructions fed back to the execution process of S3 or S4.

[0015] Optionally, in S1, constructing a dynamically updated user-environment-content interaction feature matrix includes: using a streaming computing framework to process multi-source heterogeneous data in real time;

[0016] On a user-by-user basis, within a scrollable time window, events from different data sources are time-series aligned and sessions reconstructed based on timestamps.

[0017] The processed data undergoes heterogeneous feature unified encoding and feature interaction calculation to generate a structured feature matrix for each user-advertisement-context triplet to be decided.

[0018] Optionally, in step S2, identifying the user's current information processing state specifically involves inputting the content consumption feature sequence and macro-behavioral statistical features from the feature matrix into a two-stream neural network;

[0019] The content consumption feature sequence is processed by the temporal feature extraction layer and the multi-head self-attention layer in the dual-stream neural network, and the macro-behavioral statistical features are processed by the fully connected layer.

[0020] The processed features are concatenated and output through a classification layer as a binary probability distribution of the user's state in convergent and divergent states, as well as a state focus score.

[0021] Optionally, in S2, predicting the user's real-time avoidance tendency towards the advertisement specifically involves: constructing a multi-task learning model based on the user's micro-behavioral sequence with ultra-high temporal accuracy;

[0022] The multi-task learning model simultaneously predicts the probability of a user's tendency to block, skip, or be indifferent to advertisements, and outputs a tendency probability vector.

[0023] Optionally, in S3, dynamically calling the generative AI module to generate differentiated advertising content based on the output of S2 includes: generating differentiated content generation instructions through an intelligent strategy controller based on the user's state type, state intensity, and avoidance tendency probability;

[0024] The content generation instruction is sent to a generative AI engine group that includes a structured copywriting generation engine and a creative content generation engine to generate corresponding advertising copy and creative materials.

[0025] The dynamic assembler automatically typesets and composites the generated materials according to the platform specifications of the target ad format.

[0026] Optionally, in S4, the real-time bidding and placement strategy model based on reinforcement learning models the advertising placement decision as a sequential decision process. Its state space integrates the user's real-time state vector, user session context, advertising inventory and environmental information, and user long-term value indicators output by S2. Its action space includes a composite decision on whether to place an ad, which ad content to select, the bid amount, and the placement channel. Its reward function is a weighted function that combines short-term gains, long-term user satisfaction, costs, and penalties for avoidance behavior.

[0027] Optionally, S4 also includes a cold start processing mechanism for new users or users with sparse behavioral data: based on the user's data sufficiency index, it is determined whether the user is in the cold start stage, data accumulation period or mature stage;

[0028] For users in the cold start phase, based on their initial limited interaction data, the basic strategy network parameters are quickly fine-tuned using a meta-learner to generate an initial personalized delivery strategy.

[0029] During the data accumulation period, a weighted fusion decision is made using the strategy generated by the meta-learner, the group strategy obtained from real-time dynamic clustering of users, and the output of the gradually growing fully personalized model. The weights are dynamically adjusted according to the data sufficiency index.

[0030] Optionally, in S5, automatically tuning the model parameters using the asynchronous parallel Bayesian optimization framework includes: encoding the hyperparameters to be optimized and the decision thresholds in S2 to S4 into parameter vectors;

[0031] With the goal of maximizing the "compass comprehensive health score" and on the premise of satisfying various constraints, Gaussian process regression is used as a surrogate model for modeling.

[0032] By optimizing the acquisition function that comprehensively considers the expected improvement of the objective function and the probability of constraint satisfaction, multiple sets of parameter combinations are selected for parallel online experiments.

[0033] The surrogate model is updated based on the experimental feedback, and new parameter combinations are iteratively selected for evaluation until the optimization process converges.

[0034] Optionally, S5 further includes constructing a high-fidelity digital twin simulation environment: the simulation environment includes a user behavior simulator capable of simulating the response behavior of virtual users to a given advertising strategy;

[0035] The new parameter combinations or new policy logic generated through Bayesian optimization will first be subjected to large-scale parallel simulation testing and security verification in the simulation environment.

[0036] Only when the simulation test results meet the preset security and performance gating conditions will the parameter combination or strategy be approved to enter the online low-traffic test phase.

[0037] Optionally, the intervention interface in S6 includes: a strategy rule injection engine, used to receive and compile business rules defined by operations personnel, and to correct the AI ​​strategy output of S3 or S4 in real time with high priority;

[0038] The parameter dynamic adjustment panel allows operators to dynamically adjust the weight combination in multi-objective optimization and trigger the system to automatically select or optimize strategies based on the new weights.

[0039] The specific traffic takeover mode allows operators to make manual decisions regarding designated high-value traffic, while the AI ​​system only provides predictions and suggestions;

[0040] All human interventions must be certified and fully recorded, and important interventions must be rehearsed and their effects predicted in a strategy simulation environment before taking effect.

[0041] Compared with the prior art, this application has at least the following beneficial effects:

[0042] This application uses a streaming computing framework and time window alignment technology to synchronize and integrate multi-dimensional data such as user behavior, environment, content, and profiles, and construct a dynamically updated structured feature matrix. This solves the problem of data silos and provides a high-quality, high-time unified data foundation for subsequent accurate modeling, improving the completeness and real-time performance of model input information.

[0043] This application effectively distinguishes between users' convergent and divergent information processing states by designing a lightweight dual-stream neural network and combining temporal patterns with macroscopic statistical features. At the same time, it uses a multi-task learning model to analyze micro-behavioral sequences and accurately predicts specific avoidance tendencies such as blocking, skipping, and indifference. This enables the system to respond to users' psychological changes and provides unprecedentedly accurate evidence for the dynamic adaptation of content and strategies.

[0044] Based on accurate user status and intent recognition results, this application dynamically drives a generative AI engine through a strategy controller to achieve real-time generation of advertising copy and creative materials that are "personalized for each user and time". Combined with dynamic assembly and device network adaptation technology, it ensures that the generated content not only matches the user's intent in terms of information dimension, but also seamlessly integrates with the presentation environment in terms of form, which greatly improves the relevance, nativeness and reception friendliness of the advertisement. Attached Figure Description

[0045] Figure 1 A flowchart illustrating an embodiment of this application of a method for constructing and optimizing an AI-based information flow delivery and operation compass. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments.

[0047] The method for constructing and optimizing an AI-based information flow delivery and operation compass provided in this application includes the following steps:

[0048] S1. Real-time acquisition and fusion processing of multi-source heterogeneous data: Integrate real-time user behavior stream data, contextual environment data, advertising content feature data and user static tag data; use a streaming computing framework to perform real-time cleaning, normalization and feature extraction, and construct a dynamically updated user-environment-content interaction feature matrix;

[0049] S1 is specifically implemented through the collaborative efforts of a data acquisition layer, a real-time processing layer, and a feature fusion layer, as detailed below:

[0050] Data Acquisition Layer: This layer asynchronously acquires multi-source heterogeneous data from multiple independent data sources. Preferably, this invention employs an event-driven architecture, achieving low-latency data capture through lightweight proxies deployed on different data source ends or by directly calling the platform's open interfaces.

[0051] Real-time user behavior stream data: This data is captured in real-time using event tracking technology, tracking the sequence of micro-interactions within the information feed interface. This sequence includes not only discrete events such as clicks, triggering of a block button (explicit rejection), and swiping to skip, but also continuous variables such as the duration of dwell time on a single piece of content and the instantaneous speed and acceleration during swiping. This data is linked using user session IDs and timestamps.

[0052] Contextual environment data: Collect contextual information at the moment the ad request is sent, including timestamps accurate to the second (used to determine time periods such as weekdays, holidays, and morning / evening peak hours), the geographical location of the user's device's GPS or IP resolution, the current network type and signal strength (such as 4G / 5G / Wi-Fi and signal strength level), and environmental factors that may affect the user's patience for interaction, such as the device's battery status.

[0053] Content feature data: Preprocessing and feature extraction of candidate advertising creatives. For text and image creatives, extracting text topic word vectors, sentiment polarity, and readability index; for visual creatives, extracting deep feature vectors extracted through a pre-trained convolutional neural network, as well as statistical features such as color distribution and complexity; for video creatives, additionally extracting keyframe sequence features and audio features.

[0054] User static tag data: Long-term stable features related to the current user are asynchronously retrieved from the user profiling system. These features are not completely "static" but are updated at a low frequency (e.g., daily), including demographic attributes, long-term interest tags based on historical behavior (generated through clustering algorithms), historical average click-through rates, past ad blocking and skipping category preferences, etc.

[0055] Real-time processing layer: This layer performs stream cleaning, alignment, and windowed aggregation. The raw data collected is pushed to a message queue (such as Kafka). The real-time processing layer consumes queue data based on a streaming computing engine (such as Apache Flink or Spark Streaming) and performs the following core operations;

[0056] Cleaning and Validation: The behavior flow data is deduplicated, and invalid or abnormal sessions are filtered out (e.g., robot behaviors with dwell times exceeding a reasonable threshold). Missing location information in the context data is imputed using a prediction model based on historical trajectories.

[0057] Time window alignment and session reconstruction: Define a scrollable time window (e.g., the most recent 30 minutes) for each user. Within this window, asynchronous events from different data sources are time-series aligned and merged based on high-precision timestamps to reconstruct a complete recent user interaction session. This step resolves the issue of inconsistent transmission latency between different data sources.

[0058] Real-time feature calculation: Within the session window, real-time feature derivation is performed. For example, the distribution of non-ad content topics viewed by the user in the current session is calculated to infer their immediate interests; the ratio of recent (e.g., within the past hour) ad exposure frequency to avoidance behavior (blocking / skipping) is calculated as an indicator of current ad fatigue; and the estimated content loading latency for the current network environment is calculated.

[0059] Feature Fusion Layer: Constructs a dynamic interactive feature matrix. This layer receives multi-source features that have been processed and aligned in real time, fuses and structures them, and outputs a feature matrix that can be directly consumed by the model.

[0060] Heterogeneous Feature Unified Encoding: All features are uniformly encoded as numerical vectors. For categorical features (such as city or advertising category), a dynamic weighted embedding technique is used to map them into low-dimensional dense vectors. The weights of this embedding layer can be jointly trained in subsequent models. For numerical features, dynamic normalization is performed, and the normalization parameters (such as mean and variance) are updated online based on historical data within a sliding time window to adapt to possible shifts in data distribution.

[0061] Feature Interactions and Associations: Explicitly calculate the interaction terms between key feature pairs. For example, calculate the cosine similarity between a user's long-term interest tags and the features of the current candidate ad content; calculate the match between the user's real-time swiping speed and the visual complexity of the ad creative (assuming that highly complex creatives are more easily ignored during rapid swiping);

[0062] Finally, a structured feature matrix is ​​generated for each "user-advertisement-context" triple to be decided. The rows of the matrix can represent different feature groups (behavior, content, environment, profile), and the columns can represent aggregated values ​​at different time granularities (such as immediate, the last 5 minutes, the last 1 hour) from the current moment. Each cell in the matrix is ​​an encoded and normalized feature value.

[0063] S2. User dynamic state perception and intent recognition: Based on the real-time features extracted in S1, a lightweight deep learning model is used to identify the user's current information processing state, which includes at least a convergent state and a divergent state; at the same time, the user's real-time avoidance tendency towards advertisements is predicted by combining micro-behavioral sequences, which includes at least a blocking tendency, a skipping tendency, and an indifferent tendency.

[0064] Specifically, the lightweight deep learning model is a lightweight two-stream neural network, designed to distinguish between convergent and divergent states. The specific process is as follows:

[0065] Receive two specific subsets from the S1 feature matrix. The first subset is the content consumption feature sequence, including the user's recent time window. Within the first 10 viewed items, the sequence of dwell time, clicks, and scrolling speed for non-advertising information is analyzed. The second subset represents the macro-level behavioral statistics of the current session, including session duration, the topic dispersion of viewed content (calculated based on information entropy), and the presence and specificity of search query terms.

[0066] For content consumption feature sequences, a temporal feature extraction layer (which can employ one-dimensional convolution or gated recurrent units, GRU) is first used to extract local and long-term dependency patterns from the sequence. Subsequently, a multi-head self-attention layer is introduced, which automatically evaluates the importance of behaviors at different times in the sequence and captures patterns of user attention shifts. Macro-level behavioral statistical features are directly encoded through a fully connected layer. Finally, the feature vectors output from the two processing branches are concatenated.

[0067] The concatenated feature vector is passed through a softmax classification layer, which outputs a binary probability distribution indicating whether the user is in a convergent or divergent state. , In addition, the network outputs a continuous state focus score, which integrates indicators such as the stability of dwell time and topic concentration to quantify the intensity of the current state;

[0068] The micro-behavioral sequence prediction of users' real-time avoidance tendency towards advertisements is performed in parallel with a lightweight deep learning model, focusing on predicting users' immediate avoidance responses when they are next exposed to advertisements, as detailed below:

[0069] Three main avoidance tendencies are defined: blocking tendency (actively clicking to close), skipping tendency (swiping abnormally quickly, such as swiping speed exceeding its historical average speed by 2 standard deviations), and insensitivity tendency (not skipping, but with extremely short dwell time and no interaction). The input to this module is a sequence of micro-behaviors with ultra-high temporal precision, including: a sequence of interaction events based on ultra-high temporal precision (such as click coordinates, swipe trajectory, speed, and acceleration) and inferences about dwell time in specific areas of the screen (which can be estimated by analyzing touch events and interface layout).

[0070] Predictive Model and Feature Engineering: A multi-task learning model is constructed to simultaneously predict the probability of three avoidance tendencies. The basic input to the model is a smoothed micro-behavioral time series. Key feature engineering includes: calculating the abrupt change points of swipe acceleration, statistically analyzing the cumulative gaze dwell time in a specific area of ​​the current screen (usually the expected location of the ad), and analyzing the overall interaction rhythm after entering the current information feed page (such as the regularity of click intervals). These features are then fused with the user's historical avoidance rate and the inherent intrusiveness score of the ad creative obtained in S1.

[0071] Propensity output: The model ultimately outputs a three-dimensional vector. These represent the estimated probabilities of triggering blocking, quick skipping, and imperceptible ignoring, respectively. These probability values ​​are independent of, yet potentially correlated with, the output of the state-aware module, together forming a complete profile of the user's intent.

[0072] Furthermore, to ensure that state awareness and intent recognition can continuously adapt to dynamically changing user behavior patterns and platform ecosystems, a closed-loop online learning and adaptive process was constructed. Through streaming data processing and incremental learning technology, the model parameters are continuously, stably, and securely evolved.

[0073] The core of online learning is to establish a low-latency, high-fidelity feedback data loop. This loop operates through the following stages:

[0074] Feedback Signal Definition and Collection: Two types of feedback signals are clearly defined. The first type is state verification signals: After determining that a user is in a certain state, this is verified by monitoring the user's behavior patterns in the following short period (e.g., the next 30 seconds). For example, if it is determined to be a "converging state," but the user immediately begins to quickly and randomly swipe through different topics, this determination can be considered a "negative feedback"; conversely, if the user subsequently clicks deeply or stays on the same topic for a long time, it is "positive feedback." The second type is genuine avoidance behavior signals: These are the user's actions of blocking, skipping, or being indifferent (extremely short stay) to the actually displayed advertisements. This is the most direct monitoring signal for the intent prediction module.

[0075] Feedback Data Queue and Labeling: All generated feedback signals, along with the original features that initially triggered the prediction, the model prediction results, and the corresponding context information, are stored in real-time in a dedicated feedback data queue. Each data entry is automatically labeled with a timestamp, session ID, and associated with a specific model version. For state verification signals, "soft labels" (i.e., confidence weights) are automatically generated using a predefined rule base; for example, the consistency strength of subsequent actions is used as the weight for the feedback sample.

[0076] An online learning framework based on mini-batch stochastic gradient descent is adopted, and the idea of ​​elastic weight consolidation is incorporated to prevent catastrophic forgetting.

[0077] Incremental training sample stream: From the feedback data queue, a small batch of the latest feedback samples is dynamically extracted in time windows (e.g., every minute) to form the training batch. To ensure the timeliness and representativeness of the samples, feedback samples from slightly earlier times (e.g., within the past 24 hours) are also extracted in a certain proportion to form a mixed training set;

[0078] Continuous parameter optimization: Model parameter updates are not performed immediately upon each sample arrival, but rather triggered by mini-batch gradient descent updates at a fixed frequency (e.g., after every N new feedback samples). The loss function, based on the original cross-entropy loss, adds a regularization term. This term penalizes drastic changes to important parameters of the old task (defined by historical data), thus striking a balance between adapting to new data and preserving historical knowledge. This importance is evaluated through online approximation of the Fisher information matrix.

[0079] Personalized Fine-tuning: For high-value or highly active users, lightweight personalized fine-tuning is supported on top of the global model. This involves maintaining a set of "personal offset" parameters for a specific user. When that user generates enough feedback data, their personal data is used locally (or after encrypted upload) to fine-tune only this set of offset parameters, making the predictions more aligned with the user's unique habits, without altering the massive global model parameters.

[0080] Meanwhile, online learning must include rigorous version control and security measures, specifically including:

[0081] Shadow Mode and A / B Testing: Any new parameter version generated through online learning will first run in "shadow mode." This means the new model processes online traffic in parallel, but its predictions are only used for recording and evaluation and do not actually influence advertising decisions. Simultaneously, the new model will undergo small-scale A / B testing against the current stable online version, rigorously comparing its performance on core metrics (such as state judgment accuracy and avoidance behavior prediction AUC).

[0082] Performance monitoring and anomaly detection: Real-time monitoring of the distribution changes of the new model's prediction results. If a sharp shift in the prediction distribution is detected in a short period of time (such as a sudden change in the mean of the "blocking tendency" probability exceeding the threshold), or a significant drop in key indicators during low-traffic testing, an alert will be automatically triggered and the online deployment of that version's parameters will be paused;

[0083] Versioned storage and fast rollback: All model parameters, corresponding feature engineering code versions, and training data snapshots are stored completely and in a versioned manner. Once a problem is confirmed in a new version, it can be automatically rolled back to the previous stable version within one minute to ensure the continuity of online services.

[0084] Through the above mechanism, state perception and intent recognition have the ability to continuously learn and evolve from real interactions, track the slow migration of group trends, and respond quickly to sudden hot events or new user behavior patterns, thereby maintaining the accuracy and timeliness of predictions in the long term.

[0085] S3. Intelligent Generation and Multimodal Adaptation of Ad Content: Based on the user status and avoidance tendency identified in S2, the ad content generation module based on generative AI is dynamically invoked to generate differentiated ad content; at the same time, the generated ad content is multimodally adapted to ensure that the ad format is seamlessly integrated with the current user device, network environment and platform interface;

[0086] Specifically, an intelligent policy controller is established, whose core inputs are the {state type, state strength, [avoidance tendency probability]} output from step two. The controller internally contains a pre-defined "policy-content" mapping matrix, which is driven by both business rules and machine learning.

[0087] when When the user is highly focused and in a high-level state, the controller triggers the "rational persuasion" mode. The generated instructions explicitly require that the core content revolve around concrete elements such as a key attribute comparison table, a list of technical parameters, a summary of third-party testing reports, and step-by-step solution diagrams. The language style must be rigorous, precise, and information-dense. Simultaneously, if the probability of "blocking tendency" is high, the instructions will specifically require avoiding the use of subjective or exaggerated modifiers to reduce resistance caused by perceptual interference.

[0088] when When the threshold is high, the controller triggers the "emotional resonance" mode. The generated instructions require that the core content construct a scenario-based narrative, highlight aesthetic design, utilize metaphors and symbols, and reinforce abstract elements such as brand value propositions. Visually, it requires vibrant colors, dynamic composition, or an artistic feel. If the "skip tendency" probability is high, the instructions will require setting a strong attraction element (such as suspense or stunning visuals) in the visual focus or opening sentence within the first 3 seconds to counteract the inertia of quickly scrolling through the text.

[0089] When the state probability distribution is similar or the avoidance tendency is complex, the controller adopts a weighted synthesis strategy. For example, it generates a hybrid structure that begins with an emotional scenario (to address divergent thinking) but embeds clear product function anchors in the middle (while also considering convergence). The proportion of each part is dynamically determined by the weighting of the state probability and tendency probability.

[0090] The generation instructions are sent to the backend generative AI engine group. This engine group adopts a microservice architecture and contains multiple specialized sub-engines, specifically including:

[0091] The structured copy generation engine receives "rational persuasion" instructions. First, it retrieves relevant entities (such as competitor models, technical parameters, and certification codes) from the product knowledge graph. Then, it calls a finely tuned large-scale language model, whose training data includes a large amount of product manuals, evaluation reports, and question-and-answer pairs. Based on the instructions, the model fills the retrieved structured information into predefined logical argument templates (such as "advantage comparison - principle explanation - evidence presentation" templates) to generate accurate and objective advertising copy. The copy length and numerical presentation methods (such as whether to use charts) are fine-tuned according to the intensity of the situation.

[0092] Creative Content Generation Engine: Receiving "Emotional Resonance" Commands. This engine integrates text generation and image / video generation capabilities. Its text generation component is based on a language model trained on brand stories, slogans, and social media copy, excelling at generating concise and impactful statements. Its image generation component, based on a diffusion model, can generate compliant creative images or keyframes in real-time, based on text descriptions and brand visual guidelines (primary color scheme, logo placement, font). For videos, the engine intelligently retrieves clips matching the theme from a library of high-quality video clips already labeled with emotional tags, and quickly edits and splices them according to the emotional curve, adding generated text and music.

[0093] Dynamic assembler: Responsible for automatically laying out, compositing, and rendering the text, visual, and audio elements output by each engine, according to the platform's UI specifications for the target ad format (large images, small images, videos, interactive cards). The assembly process strictly adheres to the platform's design guidelines, ensuring that the generated content is visually highly consistent with the native news feed, reducing any sense of rejection caused by abrupt formatting.

[0094] The generated content needs to pass through an adaptation filtering and optimization layer before final deployment;

[0095] Device and network adaptation: Real-time reading of user device model, screen resolution, operating system version, and current network bandwidth and latency. For high-performance devices and high-speed networks, high-definition videos or lightweight interactive ads (such as swipeable cards) are prioritized. For low-end devices or weak network environments, an automatic content degradation strategy is triggered: videos are converted into keyframe GIFs or high-definition still images, complex animations are removed, and even long text summaries are converted into bullet points to ensure that ads can load smoothly and present core information within a limited time.

[0096] Real-time quality scoring and alternatives: Each generated content set is quickly evaluated by an independent content quality scoring model. This model scores content across multiple dimensions, including creativity, relevance, clarity, visual appeal, and brand safety. If the generated main content score is below a threshold, the highest-scoring alternative will be immediately selected from several alternative options generated for that request (typically 2-3 variations are generated in parallel during the process). If all options fail to meet the requirements, the system will revert to a pre-prepared library of premium content matching the user's tags for distribution, and this failed generation case will be recorded for analysis.

[0097] A / B testing traffic injection: A small portion of traffic is reserved for real-time A / B testing of newly generated content against the current best online solution. Test results (click-through rate, engagement rate, conversion rate) are fed back to the strategy controller and generation engine in real time as part of online learning, continuously optimizing generation strategies and model parameters;

[0098] Furthermore, to ensure the legality, compliance, and brand consistency of intelligently generated advertising content, and to avoid legal risks and reputational damage caused by inappropriate content, a real-time compliance review filter based on multimodal deep learning needs to be designed and deployed. This filter serves as an essential quality checkpoint in the content generation process, completing the synchronous detection and judgment of text, visual, and audio elements.

[0099] Specifically, a parallel multi-branch neural network architecture is adopted to perform independent and collaborative analysis of each modality of the advertising content;

[0100] The text compliance detection branch takes the text, title, and tags of generated advertisements as input. At its core is a pre-trained language model (such as BERT or a similar structure) fine-tuned with a large-scale corpus of prohibited terms. Its training corpus not only includes publicly available sensitive word lists but also deeply integrates historical rejection case texts from advertising review platforms, covering dozens of subcategories such as false advertising, disparaging competitors, absolute terms, promotion of prohibited categories, and inappropriate mentions of privacy data. The model output is a multi-dimensional vector, with each dimension representing the confidence score for the corresponding violation category. Furthermore, this branch integrates a semantic consistency detection module to identify misleading content such as discrepancies between text and images and clickbait titles. This is achieved by calculating the cross-modal similarity between the text description and the image feature vector.

[0101] Visual Compliance Inspection Branch: Processes advertising images or video keyframes. It employs a deep convolutional neural network as the basic feature extractor, with multiple proprietary classification heads connected to the backend. Its core capabilities include: 1) Sensitive Scene and Object Recognition: Accurately identifies visual elements involving violence, gore, pornography, politically sensitive symbols, and prohibited items (such as tobacco and firearms). 2) Brand Identity and Copyright Element Detection: Prevents unauthorized misuse or infringement of trademarks, mascots, and specific IP characters by comparing them with a database of registered trademarks and copyrighted images. 3) Visual Quality and Authenticity Assessment: Detects whether images are excessively blurry, distorted, or contain obvious, deceptive synthetic traces (such as Photoshop forgery).

[0102] Audio compliance detection branch (e.g., ads containing audio): For video or audio ads, an audio feature extraction network is used to detect whether the background music or voice-over contains unauthorized copyrighted music clips, or whether the voice content contains illegal information;

[0103] The confidence score output by the detection model will be input into a dynamic strategy engine, which will perform corresponding processing actions based on the violation category, confidence score, and preset business risk level threshold.

[0104] Tiered Judgment and Handling: The strategy engine pre-sets three levels of judgment: "Pass," "Review," and "Block." For serious violations with high confidence (>threshold A) (such as pornography, political content, or absolute false advertising), a "hard block" is automatically executed, and the content will not be delivered, while a detailed review log is generated. For potentially risky content with medium confidence (threshold B < confidence < threshold A) (such as mildly violent images that may cause discomfort or borderline language), it is marked as "requires review" and routed to the manual review queue for final decision by reviewers; during this period, the content will not be delivered. For minor issues or misjudgments with low confidence (<threshold B), the content may be allowed to "pass," but relevant tags will be recorded.

[0105] Content Correction and Replacement Suggestions: For certain non-serious violation categories (such as the use of inappropriate exaggerated language), a correction suggestion generation module can be triggered simultaneously with the "review" or "blocking" decision. This module, based on constraint text generation technology, attempts to rewrite the violating text fragments in a synonymous and compliant manner, or suggests replacing a specific area in an image. The corrected content will then re-enter the review process.

[0106] To adapt to the dynamic changes in review requirements and ensure efficiency, the following mechanisms are necessary:

[0107] Model lightweighting and acceleration: The detection models deployed in the production environment have been optimized by pruning, quantization and other methods, and may use dedicated hardware for inference acceleration to ensure that a single review is completed within tens of milliseconds, meeting the latency requirements of the real-time generation process.

[0108] Incremental learning and hot rule updates: A continuous feedback loop is in place. The final rulings of human reviewers on "reviewed" content, as well as subsequent content reports from users, are used as new labeled data for periodic incremental fine-tuning of the detection model. Simultaneously, the latest plaintext sensitive word lists, policy and regulatory keywords, and other rules can be injected into the strategy engine in real time through the hot update system without requiring a service restart.

[0109] End-to-end audit trail: Every step of content generation, review, and deployment, including original generation instructions, content for each modality, detailed scores from model detection, the judgment results and basis of the strategy engine, and even manual review records, is immutably recorded in the audit log. This provides complete data support for post-event accountability, model performance evaluation, and compliance verification.

[0110] Through the aforementioned multi-layered, multi-modal, and dynamic strategy-driven review and filtering mechanism, this invention fully leverages the creative capabilities of generative AI while constructing a solid and reliable technical defense line to ensure that the generated advertising content achieves a balance between creativity, commercial viability, and security compliance.

[0111] S4. Dynamic Adjustment of Advertising Strategies and Real-Time Bidding Optimization: Construct a real-time bidding and advertising strategy model based on reinforcement learning, with the goal of maximizing long-term user value, and dynamically decide on the timing of advertising, bidding strategies, frequency control, and channel selection;

[0112] The model models advertising placement decisions as a continuous sequence of decision-making processes;

[0113] State space (S) definition: states It is a complete snapshot at time step t, which integrates: 1) the user's real-time state vector (from step two, including state type, intensity, and probability of each avoidance tendency); 2) the user's current session context (topics of viewed content, duration of this session, and historical exposure frequency); 3) advertising inventory and environmental information (candidate ad set and its characteristics, estimated real-time market competition intensity, and the platform's given traffic floor price); and 4) the user's long-term value indicator (the user's future LTV potential predicted based on historical data).

[0114] Action Space (A) Definition: Action It is a composite decision vector, which includes: whether to display the ad at the current moment (binary decision), if so, which specific ad content to choose (from S3), the bid amount for this display, and the suggested ad display channels (such as app feed, video overlay, lock screen banner, etc.).

[0115] Reward Function (R) Design: Reward The design of this function is crucial for guiding long-term value in the model. It is a weighted synthesis function: Among them, "user satisfaction" can be implicitly measured by whether the user's subsequent normal browsing time of non-ad content decreases after the campaign is launched; "penalty items" are directly related to whether the avoidance behavior predicted in step two actually occurs and its severity.

[0116] The decision model approximates the mapping from state to optimal action using a deep Q-network or actor-critic network. Its core output decision logic is as follows:

[0117] Decision-making on ad delivery timing: The model internally learns the user's psychological acceptance curve. Its input is a sequence of user state changes (such as nodes transitioning from "divergent" to "convergent"), and the output is the "optimal delivery probability," which is combined with the current quality of the ad inventory and the competitive environment to make a comprehensive decision. For example, even if the acceptance rate is high, if only low-relevance ads are currently available, the model may choose to forgo the opportunity to protect the user experience and look forward to the next opportunity with a higher matching degree.

[0118] Real-time bidding strategy: The bidding action is directly output by the model as a specific amount. This amount is generated based on a value estimation subnetwork, which predicts the combined value of this impression for the advertiser (click-to-conversion value) and the platform (long-term ecosystem health). The bidding model simulates competitor behavior in real time, employing a variant of a contextual gambling algorithm to strike a balance between "exploration" (trying different bids to learn market reactions) and "exploitation" (using known optimal bids) to win high-value exposure opportunities at the optimal cost.

[0119] Intelligent Frequency Control: The user fatigue level is not a fixed threshold but a dynamic model. This model is calculated in real time based on the user's historical exposure interval decay curve, recent avoidance behavior frequency, and the repeated exposure effect of the same ad creative. When the "dislike probability" predicted by the model exceeds the safety threshold, even if there are other reasons for delivery, this delivery will be forcibly suppressed, and a "rest period" timer will be started for this user;

[0120] Multi-channel Collaborative Optimization: For platforms with multiple user touchpoints, the state space of the model will be extended to include the real-time context of each channel. "Channel selection" in the action space is a multi-armed bandit problem. The model maintains an independent value estimator for each channel and considers the synergy effect between channels (e.g., seeing an ad in the information flow and then receiving a reminder on another channel may result in a higher conversion rate) and the intrusion effect (multiple exposures across channels in a short period may cause annoyance). Ultimately, the channel or sequence of channel combinations with the highest comprehensive expected value is selected;

[0121] The specific process of model training and online deployment strategy includes:

[0122] Offline Training and Simulation: The model is first pre-trained in an offline environment composed of historical log data through an offline reinforcement learning algorithm (such as conservative Q-learning) to safely learn strategies from historical behaviors and avoid the costs of early online exploration;

[0123] Online Learning and Safe Exploration: After deployment, the model is fine-tuned online through a safe sandbox. Only a small portion of the traffic (such as 1%) is used to execute an "exploratory" strategy with slight random perturbations to learn the latest dynamics of the market. The data generated by the exploratory strategy, together with the data of the regular strategy, is used to update the model parameters regularly. All policy updates need to pass the policy performance evaluation in the offline simulator, and can only take effect after confirming that the key metrics (such as long-term value, user avoidance rate) have not decreased;

[0124] Dynamic Adjustment of Multi-objective Trade-off: The weight parameters ( ) in the reward function are not fixed; the platform operator can dynamically adjust these weights through a high-level control panel according to business goals (such as emphasizing user acquisition or activation recently), and the model will automatically adapt and learn the optimal strategy under the new goals;

[0125] Furthermore, to address the problem of model decision failure caused by feature缺失 for new users or users with sparse behavior data, and to ensure the decision coherence and effect stability throughout the user's life cycle, a hierarchical progressive cold start processing mechanism is designed. This mechanism provides reliable basic decisions in the data-scarce stage and can achieve a smooth and automatic transition to a fully personalized model during the data accumulation process;

[0126] The specific process is as follows: Maintain a dynamic data sufficiency metric for each user. This metric is not a simple count of behavior occurrences, but a weighted score that comprehensively considers: 1) the number of unique sessions the user has on the current platform; 2) the diversity of recorded valid behavioral events (such as clicks, pauses, and swipes); and 3) the number of stable feature dimensions that can be extracted in step one. One or more progressive thresholds are preset (e.g., ...). , When user metrics fall below the minimum threshold When the indicator is between, it is judged as a "cold start user"; when the indicator is between and In the interim, it is in the "data accumulation period"; beyond that, it is in the "data accumulation period". Then they are considered "mature users";

[0127] For cold-start users, the main reinforcement learning model will invoke a dedicated meta-policy model. The core objective of this model is to leverage experience learned from a vast number of other users on "how to quickly learn from a new user," and the specific process is as follows:

[0128] Meta-model training: In the offline phase, training is conducted by simulating a cold start scenario with a large number of "new users." Specifically, the historical data of all established users is truncated on the timeline, retaining only the first N interactions to simulate the cold start period, with subsequent data used for validation. A meta-learner (e.g., based on MAML or its variants) is trained to quickly adjust the parameters of a base policy network based on the support set formed by the user's initial limited interaction data, generating a task-specific policy adapted to the user's initial characteristics.

[0129] Rapid online adaptation: When encountering users experiencing a cold start online, their initial few interactions (even just 2-3) are used as a support set and input into the meta-learner. The meta-learner rapidly fine-tunes a set of pre-trained, general-purpose basic policy network parameters, producing an initial personalized policy. This policy already possesses a certain degree of specificity, far superior to random or averaged policies, thus enabling relatively reasonable deployment decisions in the early stages of data collection. Furthermore, these decisions themselves aim to more efficiently explore user preferences and accelerate data accumulation.

[0130] In addition to the meta-strategy, a strategy mapping mechanism based on user profiles for rapid segmentation is run in parallel as a supplement or alternative.

[0131] Real-time dynamic clustering: Maintaining an efficient clustering model based on static user attributes (such as device model, registration channel, and region) and the initial small amount of dynamic behavior available (such as the category of the first click). Once a new user generates any recordable behavior, its similarity to existing dynamic clusters is calculated in real time, and it is assigned to the most similar cluster. The strategy preferences of these clusters (such as which ad content users in that cluster click on most frequently, during which time periods they are more active, and the frequency of ads they can tolerate) are continuously updated.

[0132] Group policy inheritance and fine-tuning: Each dynamic cluster is associated with a group-optimal policy, which is derived by analyzing the historical optimal decisions of all mature users within the cluster. When a cold-start user is assigned to a cluster, they immediately inherit that cluster's group policy as their initial policy. Similar to the meta-policy, as the user accumulates their own data, the group policy is progressively updated using Bayesian methods or fine-tuned online based on their personal data, gradually shifting it towards their true personal preferences.

[0133] A weighted fusion approach is used to achieve a seamless transition from a cold start strategy to a mature personalized strategy. The specific process is as follows:

[0134] Adaptive Weight Fusion: During the user data accumulation phase, the final decision is jointly determined by the output of the meta-policy model, the group policy of the user's group, and the gradually evolving fully personalized model. The weights of these three factors are dynamically controlled by the aforementioned data sufficiency indicators. Initially, the group policy and meta-policy dominate; as the indicators improve, the weight of the personalized model increases linearly or non-linearly; when the indicators exceed... At the threshold, the weights are fully transferred to the personalized model, and the cold start strategy is completely stripped away.

[0135] Knowledge distillation and model warm-up: During the transition phase, the fully personalized model is not trained from scratch. It utilizes knowledge distillation technology, using the decisions from the fusion strategy during the cold start phase as "soft labels," while simultaneously combining this with real-world user behavior data gradually accumulated for collaborative training. This allows the personalized model to achieve a relatively high-quality initialization early on, avoiding a steep performance slump in the early stages and achieving a smooth rise in the learning curve.

[0136] Through the aforementioned composite cold start processing mechanism, which provides an intelligent starting point through meta-learning, reliable prior knowledge through group profiling, and ensures a smooth experience through fusion transition, the inherent "cold start" problem in recommendation systems and advertising is effectively solved, enabling new users to obtain a consistent and rapidly optimized service experience from the very beginning of the platform's lifecycle.

[0137] S5. Multidimensional effect evaluation and closed-loop optimization of compass parameters: Define and calculate a multidimensional effect evaluation index system that includes short-term indicators, avoidance indicators, experience indicators and long-term indicators; based on the evaluation results, establish an automatic adjustment mechanism for operation compass parameters, and regularly adjust the hyperparameters and decision thresholds of each model in S2 to S4.

[0138] The multidimensional performance evaluation index system specifically includes:

[0139] Short-term effect layer: Directly calculate traditional metrics such as click-through rate and conversion rate, but place them at the user granularity for time-series smoothing calculation to reduce noise;

[0140] User avoidance layer: Accurately calculates ad blocking rate and abnormal skip rate as defined by S2. Abnormal skips are determined based on the user's individual historical scrolling baseline, rather than a fixed speed threshold.

[0141] Experience Health Layer: Calculate the rate of change in the total dwell time and number of interactions of users on N consecutive non-ad content sessions after ad exposure. A significant decrease is considered a deterioration in experience. Simultaneously, monitor whether the probability of users actively exiting the current news feed session increases due to ad exposure;

[0142] Long-term value layer: By embedding long-term tracking code, a proxy model for user lifetime value is constructed. This model can make weekly or monthly predictions based on leading indicators such as short-term user activity, retention tendency, and willingness to upgrade consumption. For brand advertising, models of changes in brand awareness and favorability are fitted using regular, small-scale brand enhancement survey data.

[0143] Indicator Integration and Health Score: The above four levels of indicators are aggregated into a "Compass Overall Health Score" through a dynamic weighted comprehensive algorithm. The weights of each level are not fixed and can be configured and adjusted by operations personnel according to the stage strategy (such as the new user acquisition period, brand building period, and revenue generation period). A dashboard is provided to display the changing trends of indicators at each level and the overall score in real time.

[0144] The multiple models in S2 to S4 contain numerous hyperparameters and decision thresholds, making manual tuning inefficient. Therefore, an asynchronous parallel Bayesian optimization framework is used for automated optimization. The specific process is as follows:

[0145] Parameter Types and Ranges: Each parameter is defined in the registry with its name, type (continuous, discrete, or ordered categorical), mathematical boundary, or enumerated value. For example, the classification confidence threshold for the state classification model in S2 is defined as continuous, with a range of [0.5, 0.95]; the exploration rate for the reinforcement learning model in S4... Defined as continuous, ranging from [0.01, 0.3]; the long-term value discount factor in the bidding strategy. Defined as continuous, ranging from [0.8, 0.99]. Furthermore, decision thresholds, such as the sliding speed multiplier threshold for determining "quick skip," are also included as continuous parameters.

[0146] Before optimization, all parameters were uniformly encoded into a high-dimensional vector. ,in This represents the entire parameter space. For categorical parameters, one-hot encoding is used; for continuous and discrete parameters, normalization is performed so that each dimension falls within the [0,1] interval, facilitating subsequent modeling.

[0147] The optimization problem is defined as: satisfying a series of inequality constraints Under the premise of maximizing the objective function This refers to the "Compass Overall Health Score";

[0148] Gaussian process regression is used as a surrogate model to model the unknown objective function. and each constraint function A Gaussian process provides a prediction of the mean and variance (uncertainty) for each function to be modeled. In the initial stage, the surrogate model is initialized based on historical experimental data or prior knowledge;

[0149] The acquisition function employs feasibility constraints. At each step, when selecting the next evaluation point, the optimizer considers not only the expected improvement of the objective function but also the probability of each constraint being satisfied. Specifically, the probability of satisfying each constraint is defined. This probability can be calculated using a Gaussian process model of the constraint function. Finally, the next evaluation point... It is determined by optimizing an improved expectation acquisition function that comprehensively considers the expected improvement of the objective function and the probability of constraint satisfaction;

[0150] Parallelized Asynchronous Evaluation: To accelerate the optimization process, parallel evaluation of multiple sets of parameters is supported. A batch Bayesian optimization strategy is employed, optimizing a batch acquisition function to select a set of parameter combinations that are diverse in the parameter space and have potential in both objective and constraint aspects. They were simultaneously assigned to multiple independent online experimental groups for parallel testing.

[0151] The entire tuning process is driven by an optimization controller module, forming a closed loop of "proposal-experiment-observation-update". The specific process is as follows:

[0152] Experimental grouping and traffic isolation: The controller, through the experimental platform, provides each group of parameter combinations to be evaluated. Create independent, non-interfering online experiment groups. Allocate a small, but statistically representative, portion of user traffic to each group (e.g., 0.5% of total traffic). Group creation, parameter configuration, and code deployment are all automated using scripts.

[0153] Metrics Collection and Target Calculation: The experimental group runs for a predetermined observation period (e.g., 24 hours). During this period, all relevant behavioral logs of users within the group are collected, and short-term metrics, avoidance metrics, experience metrics, and long-term proxy metrics are calculated according to the multidimensional metric system defined in step five. At the end of the period, the comprehensive health score corresponding to the parameters of the group is calculated based on the preset fusion formula and dynamic weights. And check each constraint. Does it meet the requirements?

[0154] Agent Model Update and Next Round Proposal: Experimental Observations The data is fed back to the optimization controller in real time. The controller asynchronously updates all Gaussian process surrogate models (objective function and constraint functions) using the new data points. After the model is updated, the controller immediately runs the next optimization iteration: based on the latest surrogate model, it optimizes the acquisition function from the parameter space. The system intelligently proposes the next batch (or the next) of parameter combinations to be evaluated. For parameter combinations that have already been evaluated, regardless of their performance, their data are incorporated into the surrogate model to continuously improve the understanding of the parameter space.

[0155] The optimization process continues until one of the following conditions is met: 1) In N consecutive iterations, the increase in the target score is less than a preset threshold. ;2) The number of experimental rounds has reached the maximum budget M;3) The optimal parameter point found has been confirmed through statistical testing. Its performance has consistently outperformed historical baselines;

[0156] Real-time monitoring and circuit breaking: During the operation of each experimental group, key guardrail indicators (such as shielding rate) are monitored in real time. If the shielding rate of an experimental group significantly exceeds the safety limit within a short period of time (such as the first 2 hours), the experimental circuit breaking will be automatically triggered, the traffic allocation of that group will be stopped immediately, and the parameters of that group will be marked as "constraint violation". Some of its valid data can still be used to update the proxy model, but the target score will be penalized.

[0157] Optimal parameter deployment and rollback: When the optimization process converges and the optimal parameter combination is determined. Then, it will undergo final verification in a simulation environment. Once successful, The new parameters will be pushed to the production environment, replacing the original parameters. Simultaneously, a snapshot of the previous stable parameter set will be retained. After the new parameters are deployed, an enhanced monitoring period will begin. If key metrics unexpectedly decline, a one-click rollback can be performed.

[0158] To improve the stability of responding to abnormal or malicious behavior, adversarial training will be further introduced;

[0159] Adversarial Example Generation: Train an adversarial generative network with the goal of generating virtual user behavior sequences that can "fool" the core state recognition and intent prediction models. The generator receives normal behavior sequences and random noise, and outputs adversarial sequences that retain basic temporal features but may cause the model to misjudge (e.g., causing the model to misjudge "quick skip" as "unnoticeable", or "divergent state" as "convergent state").

[0160] Adversarial training: These generated adversarial examples are periodically mixed with real data to retrain or fine-tune the core model from step two. This process is equivalent to allowing the model to continuously evolve in an adversarial process of "spear and shield," improving its ability to discriminate disturbances, noise, and unconventional behavior patterns, thereby making it more robust in actual operation;

[0161] To avoid deploying risky strategies directly, a high-fidelity digital twin simulation environment was built. The specific process is as follows:

[0162] Environmental modeling: The core of the simulation environment is a user behavior simulator. By learning from massive amounts of historical log data, it can simulate the feedback behaviors of virtual users with different attributes under a given advertising strategy, including clicks, conversions, and avoidances. The simulator not only simulates individuals but also a simplified market competition environment.

[0163] Rapid strategy validation: Any new parameter combination generated through Bayesian optimization, or any new strategy logic implemented through manual intervention, must first undergo large-scale, rapid parallel simulation testing in a simulation environment. Testing can simulate operational effects over days or even weeks, completed within minutes.

[0164] Security Gating: The simulation environment outputs the estimated overall health score and various level indicators for the new strategy. Only new strategies that simultaneously meet the following conditions will be approved to proceed to the next stage of online small-scale testing: 1) the overall score is significantly higher than the current online baseline; 2) all constraints are met; and 3) the strategy performs stably across multiple simulated user groups. Otherwise, the strategy will be automatically rejected and the reason recorded.

[0165] S6. Compass Visualization and Human-Assisted Intervention Interface: Build a visual decision dashboard to real-time display the user status distribution, content adaptation, strategy effectiveness, and trend of effects; Provide key intervention interfaces for operation staff, allowing the injection of expert rules or fine-tuning of strategies in specific scenarios;

[0166] The visual decision dashboard is a collection of hierarchical views designed according to operation roles and attention granularity, specifically including:

[0167] Macro Operations Room View: This view is presented in the form of a large data screen and is targeted at strategy decision-makers. The core shows the trend curve of the "Compass Comprehensive Health Score" summarized in real-time across the entire platform and decomposes it into a stacked chart of the contribution degrees of the three major pillars of effects, experience, and long-term value. At the same time, it shows the real-time status distribution (convergence vs. divergence ratio) of different user groups (such as new customers, active customers, dormant customers) and the coverage rate of the current optimal strategy in the form of a geographical heat map or topological map. Key alerts (such as an abnormal soar in the blocking rate in a certain area) pop up in a prominent manner;

[0168] Mid-level Strategy Diagnosis View: This view is targeted at operation experts and provides drill-down analysis capabilities. The core of the interface is a strategy effectiveness matrix. The vertical axis can list different user segmentation groups or content categories, and the horizontal axis can list different delivery strategies (such as "rational persuasion - high-value bidding", "emotional resonance - frequency control"). In each matrix cell, the traffic scale is represented by dots of different sizes, and the comprehensive effectiveness score of the strategy for this group is represented by color. Operation staff can click on any cell to view the detailed index breakdown, user status transition diagram, and representative delivery case samples (including the generated advertisement content and the subsequent behavior sequence of users) under this strategy;

[0169] Micro-level Real-time Tracking View: This view supports post-mortem review or real-time diagnosis of single delivery decisions. Operation staff can input a specific session ID or user ID to reproduce the complete timeline of that session. The timeline is clearly marked with: the predicted status and intentions of the user at each moment, the key decisions made (such as which content to select and when to bid), and the subsequent real feedback (whether clicked, blocked, skipped). By comparing the prediction with the fact, the accuracy of AI decisions can be intuitively evaluated;

[0170] The intervention interface follows the principles of "clear permissions, controllable scope, and traceable impact". In key scenarios, operation staff have the ability to calibrate or override AI strategies, specifically including:

[0171] Strategy rule injection engine: Operations personnel can define temporary or permanent business rules through a graphical or DSL (Domain-Specific Language) editor. For example, during "Brand Safety Month," the following rule could be injected into all luxury brand advertisements: "Mandate the use of 'elegant narrative' style templates and prohibit the use of price comparison-related concrete elements." After the rules are compiled, they are deployed to the online system in real time through the strategy rule injection engine. The engine ensures that the rules are executed with high priority; the AI's original strategy output needs to be corrected through rule filters, and the correction log is recorded.

[0172] Dynamic parameter adjustment panel: For the Pareto optimal solution set after multi-objective optimization, operators are not limited to selecting only one. A dynamic weight adjustment slider is provided. For example, facing a Pareto front composed of three-dimensional objectives—"effectiveness-experience-cost"—operators can drag the sliders on the three dimensions to change the weight combination of the objectives in real time. The backend will automatically select the best-matching solution from the pre-calculated Pareto solution set based on the new weights, or quickly launch a new round of targeted optimization and deploy the new strategy to the specified traffic range. This feature enables one-click dynamic adjustment of business objectives;

[0173] God Mode: For extremely high-value user groups (such as VIP customers) or extremely important brand events, a "manual takeover" mode is provided. In this mode, the AI ​​system plays a supporting role, only providing data predictions and suggestions (such as "This user is currently in a divergent state, emotional content is recommended"). The final decisions on content selection, bidding, and timing of delivery are manually specified and confirmed by operations personnel in a simulated environment, and then directly applied to the target traffic. This mode ensures absolute human control at critical junctures.

[0174] All human intervention actions are strictly recorded and controlled, including:

[0175] Operational audit trail: All rule injections, parameter adjustments, and manual takeover operations must undergo two-factor authentication and record the operator, time, modified content, expected impact scope, and reasons, generating a complete audit log so that any intervention behavior can be traced.

[0176] Sandbox simulation: Any intervention rules or parameter adjustments can be run in the strategy simulation environment before officially taking effect. The simulation simulates all traffic that the rule may affect in the next 24 hours and estimates its impact on key performance indicators, generating a simulation report for final confirmation by operations personnel, thereby avoiding online incidents caused by incorrect rule definitions.

[0177] Intervention effectiveness decay and automatic exit: Most intervention rules are designed with time-limited validity or effectiveness decay mechanisms. For example, an aggressive bidding rule set for a promotional campaign will automatically expire after the campaign ends. For long-term rules, the system will periodically evaluate their continued effectiveness. If it finds that the performance of the AI's self-learning strategy has consistently outperformed the rule's effect in the scenarios covered by the rule, the system will prompt the operations staff, suggesting that they review and possibly exit the manual rule, returning control to a better AI strategy.

[0178] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence, characterized in that, The method includes: S1. Real-time acquisition and fusion processing of multi-source heterogeneous data: Integrate real-time user behavior stream data, contextual environment data, advertising content feature data and user static tag data, perform real-time cleaning, normalization and feature extraction, and construct a dynamically updated user-environment-content interaction feature matrix. S2. User dynamic state perception and intent recognition: Based on the user-environment-content interaction feature matrix output by S1, a lightweight deep learning model is used to identify the user's current convergent or divergent information processing state and predict the user's real-time avoidance tendency of advertisements. S3. Intelligent generation and multimodal adaptation of advertising content: Based on the user information processing status recognition results and avoidance tendency prediction results output by S2, the generative AI module is dynamically invoked to generate differentiated advertising content, and the generated advertising content is multimodal adapted. S4. Dynamic Adjustment of Placement Strategy and Real-time Bidding Optimization: Based on the user information processing status and avoidance tendency output by S2, the advertising content generated by S3, and historical data, a real-time bidding and placement strategy model based on reinforcement learning is constructed and run. With the goal of maximizing long-term user value, the model dynamically decides the timing of ad placement, bidding strategy, and frequency control. The historical data is a specific set of data from the past of users and the platform, which is necessary for the reinforcement learning model to make real-time decisions and supports long-term value assessment and current strategy calculation. S5. Multidimensional effect evaluation and closed-loop optimization of compass parameters: Define and calculate a multidimensional effect evaluation index system that includes short-term indicators, avoidance indicators, experience indicators and long-term indicators; based on the evaluation results, automatically optimize the model hyperparameters and decision thresholds in S2 to S4 through an asynchronous parallel Bayesian optimization framework. S6, Compass Visualization and Human Collaborative Intervention Interface: Construct a visual decision dashboard to display in real time user status distribution, content adaptability, strategy effectiveness and effect trend data generated from S1 to S5; and provide an intervention interface to allow the injection of expert rules or fine-tuning of strategies in specific scenarios, with the intervention instructions fed back to the execution process of S3 or S4.

2. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In S1, constructing a dynamically updated user-environment-content interaction feature matrix includes: using a streaming computing framework to process multi-source heterogeneous data in real time. On a user-by-user basis, within a scrollable time window, events from different data sources are time-series aligned and sessions reconstructed based on timestamps. The processed data undergoes heterogeneous feature unified encoding and feature interaction calculation to generate a structured feature matrix for each user-advertisement-context triplet to be decided.

3. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In step S2, identifying the user's current information processing state specifically involves inputting the content consumption feature sequence and macro-behavioral statistical features from the feature matrix into a two-stream neural network. The content consumption feature sequence is processed by the temporal feature extraction layer and the multi-head self-attention layer in the dual-stream neural network, and the macro-behavioral statistical features are processed by the fully connected layer. The processed features are concatenated and output through a classification layer as a binary probability distribution of the user's state in convergent and divergent states, as well as a single-state focus score.

4. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In S2, predicting the user's real-time avoidance tendency towards advertisements specifically involves: constructing a multi-task learning model based on the user's micro-behavioral sequence with ultra-high temporal accuracy. The multi-task learning model simultaneously predicts the probability of a user's tendency to block, skip, or be indifferent to advertisements, and outputs a tendency probability vector.

5. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In S3, dynamically calling the generative AI module to generate differentiated advertising content based on the output of S2 includes: generating differentiated content generation instructions through an intelligent strategy controller based on the user's state type, state intensity, and avoidance tendency probability. The content generation instruction is sent to a generative AI engine group that includes a structured copywriting generation engine and a creative content generation engine to generate corresponding advertising copy and creative materials. The dynamic assembler automatically typesets and composites the generated materials according to the platform specifications of the target ad format.

6. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In S4, the real-time bidding and placement strategy model based on reinforcement learning models the advertising placement decision as a sequential decision process. Its state space integrates the user's real-time state vector, user session context, advertising inventory and environmental information, and user long-term value indicators output by S2. Its action space includes a composite decision of whether to place an ad, which ad content to select, the bid amount, and the placement channel. Its reward function is a weighted function that combines short-term gains, long-term user satisfaction, costs, and penalties for avoidance behavior.

7. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, The S4 also includes a cold start processing mechanism for new users or users with sparse behavioral data: based on the user's data sufficiency index, it is determined whether the user is in the cold start stage, data accumulation period or mature stage; For users in the cold start phase, based on their initial limited interaction data, the basic strategy network parameters are quickly fine-tuned using a meta-learner to generate an initial personalized delivery strategy. During the data accumulation period, a weighted fusion decision is made using the strategy generated by the meta-learner, the group strategy obtained from real-time dynamic clustering of users, and the output of the gradually growing fully personalized model. The weights are dynamically adjusted according to the data sufficiency index.

8. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, In S5, the automatic tuning of model parameters through the asynchronous parallel Bayesian optimization framework includes: encoding the hyperparameters to be optimized and the decision thresholds in S2 to S4 into parameter vectors; With the goal of maximizing the "compass comprehensive health score" and on the premise of satisfying various constraints, Gaussian process regression is used as a surrogate model for modeling. By optimizing the acquisition function that comprehensively considers the expected improvement of the objective function and the probability of constraint satisfaction, multiple sets of parameter combinations are selected for parallel online experiments. The surrogate model is updated based on the experimental feedback, and new parameter combinations are iteratively selected for evaluation until the optimization process converges.

9. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, The S5 also includes constructing a high-fidelity digital twin simulation environment: the simulation environment includes a user behavior simulator that can simulate the feedback behavior of virtual users to a given advertising strategy; The new parameter combinations or new strategy logic generated through Bayesian optimization will first be subjected to large-scale parallel simulation testing and security verification in the simulation environment. Only when the simulation test results meet the preset security and performance gating conditions will the parameter combination or strategy be approved to enter the online low-traffic test phase.

10. The method for constructing and optimizing an information flow delivery and operation compass based on artificial intelligence according to claim 1, characterized in that, The intervention interface in S6 includes: a strategy rule injection engine, which is used to receive and compile business rules defined by operations personnel, and correct the AI ​​strategy output of S3 or S4 in real time with high priority; The parameter dynamic adjustment panel allows operators to dynamically adjust the weight combination in multi-objective optimization and trigger the system to automatically select or optimize strategies based on the new weights. The specific traffic takeover mode allows operators to make manual decisions regarding designated high-value traffic, while the AI ​​system only provides predictions and suggestions; All human interventions must be certified and fully recorded, and important interventions must be rehearsed and their effects predicted in a strategy simulation environment before taking effect.