An ai investment flow precise matching method fusing multi-modal features

CN122615218APending Publication Date: 2026-08-21SHAOXING MAIKE CULTURE MEDIA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610596387.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

首先,现有投流模型多基于单一维度特征建模,仅依托简单的用户画像或广告文本特征完成匹配预估,无法整合广告视觉、用户时序行为、投放场景等多模态异构数据,特征表征维度片面,难以精准匹配用户真实偏好与投放场景,流量错配率高,预估精准度较差

Benefits of technology

[0071]与现有技术相比,本发明的优点在于:通过同步采集并标准化预处理用户行为、广告素材及场景环境三类多模态异构数据,弥补了传统投流特征维度单一的缺陷,同时依托BERT、CNN+ViT、Bi-GRU多分支网络独立提取各模态深度特征,并结合跨模态注意力交互融合算法实现多类异构特征自适应加权融合,有效保留有效特征、过滤冗余噪声;其次,通过搭建搭载多专家共享结构的MMoE双任务预估模型,同步完成广告点击与转化概率预测,配合定制化联合损失函数大幅提升模型预估精度与泛化能力;此外,本发明还基于实时流量竞争度实现匹配权重动态自适应调整,结合多目标效用打分与贪心调度策略,精准平衡投放效果与投放成本,同时搭配孤立森林异常流量检测与滑动窗口增量迭代机制,构建完整的线上反馈优化闭环,有效规避虚假作弊流量干扰,持续迭代优化模型参数与模态权重,显著提升广告与用户流量的匹配精度、投流稳定性与资源利用率,适配各类复杂多变的广告投放场景,满足精细化、智能化、低成本的商业化AI投流需求。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122615218A_ABST
    Figure CN122615218A_ABST
Patent Text Reader

Abstract

The application relates to an AI investment flow precise matching method fusing multi-modal features. It solves the problems of single feature dimension, weak cross-modal fusion capability and low model prediction accuracy of the prior art AI investment flow matching technology. It comprises the following steps: extracting multi-dimensional features through a multi-branch network, combining a cross-modal attention mechanism to complete feature fusion, building an MMoE double-task prediction model to realize precise prediction, adaptively adjusting weights according to flow competition degree, combining multi-target scoring to complete intelligent investment flow scheduling, and additionally adding an abnormal flow filtering and incremental iteration mechanism. The application has the advantages that high-precision, low-cost and high-stability AI intelligent investment flow matching is realized under a complex flow scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital media technology, and specifically to an AI-based precise matching method for streaming that integrates multimodal features. Background Technology

[0002] With the rapid development of internet feed ads and e-commerce live streaming ads, AI-powered intelligent ad delivery technology has become a core technology for ad traffic distribution, ad budget optimization, and precise user reach, and is widely used in commercial scenarios such as short video promotion, product sales, and online traffic generation. Traditional manual ad delivery relies on the personal experience of operations personnel, which has many shortcomings, such as strong subjectivity in traffic judgment, uneven budget allocation, low accuracy in audience matching, and inability to adapt to traffic fluctuations in real time. Therefore, the industry generally adopts artificial intelligence algorithms to replace manual processes to complete automated ad delivery decisions, thereby improving ad conversion rates and reducing ad delivery costs.

[0003] Currently, traditional AI-powered ad matching technology still suffers from numerous technical shortcomings. First, existing ad matching models are mostly based on single-dimensional feature modeling, relying solely on simple user profiles or ad text features for matching and prediction. They fail to integrate multimodal heterogeneous data such as ad visuals, user temporal behavior, and ad placement scenarios. This results in a one-sided feature representation, making it difficult to accurately match users' true preferences with the ad placement scenario, leading to a high traffic mismatch rate and poor prediction accuracy. Second, traditional algorithms lack cross-modal interaction and fusion mechanisms. Various heterogeneous features operate independently, failing to achieve adaptive weighted complementarity, easily losing effective features and introducing noise interference. Furthermore, most are single-task prediction models, unable to simultaneously consider ad click and conversion predictions, resulting in weak model generalization ability and insufficient adaptability to complex traffic scenarios. Finally, existing ad matching strategies have fixed weights, unable to adaptively adjust based on real-time traffic competition and placement costs, making it difficult to balance ad revenue and costs. They also lack the ability to identify abnormal traffic and a robust online incremental iteration loop, failing to optimize model parameters and feature weights in real time, leading to poor long-term ad stability and significant waste of traffic resources.

[0004] In summary, current AI-powered ad delivery matching technologies generally suffer from technical defects such as limited feature dimensions, weak cross-modal fusion capabilities, low model prediction accuracy, and rigid traffic matching strategies, making it difficult to meet the current demand for refined, intelligent, and cost-effective ad delivery. Summary of the Invention

[0005] The purpose of this invention is to address the above-mentioned technical problems by providing an AI-based precise matching method that integrates multimodal features.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: 1. A method for accurate AI streaming matching that integrates multimodal features, the method comprising the following steps:

[0007] S1. Multimodal source data acquisition and standardized preprocessing: Real-time acquisition of three types of heterogeneous data: user behavior modality U, advertising material modality A, and scene environment modality S. After missing value imputation, normalization and time-series alignment, a unified input dataset D={U,A,S} is constructed, where U is the user feature set, A is the advertising feature set, and S is the scene feature set.

[0008] S2. Construct a multimodal branch feature extraction network: Encode the three modalities of text, vision, and behavioral time series independently, and output modality-specific low-dimensional dense feature vectors;

[0009] S3. Based on the cross-modal attention interaction fusion algorithm, multi-branch features are weighted and fused to generate a globally unified multi-modal fusion feature vector F;

[0010] S4. Build a dual-task joint prediction model, and simultaneously output the ad click prediction probability p_CTR and conversion prediction probability p_CVR based on the fusion feature F. Complete the end-to-end training of the model through the joint loss function.

[0011] S5. Calculate the traffic matching score based on the multi-objective utility ranking formula, and sort the scores from high to low to complete the precise matching of the three elements of advertisement-traffic-scenario and real-time traffic scheduling.

[0012] S6. Construct an online feedback loop, collect real clicks, conversions and negative feedback data after traffic is generated, and incrementally update and iterate the feature weights and model parameters.

[0013] In step S1, the user behavior modal data U includes basic user attributes, behavior sequences, and interest tags; the advertising material modal data A includes text information, visual information, and audio information; the scene environment modal data S includes the placement scene, device information, and time information; missing value imputation adopts a submodal imputation strategy, with missing user behavior modal values ​​imputed by the median of users with the same interest tags, missing advertising visual modal values ​​imputed by the mean of visual features of similar advertisements, and missing scene modal values ​​imputed by a pattern imputation method; normalization processing includes: logarithmic features are normalized using Min-Max, mapping feature values ​​to the [0,1] interval, and the calculation formula is:

[0014] x_norm=(x-x_min) / (x_max-x_min),

[0015] Where: x is the original feature value.

[0016] x_min is the minimum value of this feature.

[0017] x_max is the maximum value of this feature;

[0018] Categorical features are converted into vector form using one-hot encoding;

[0019] The time-series alignment steps include: aligning user behavior time-series data and scene dynamic data with a time granularity of 100ms, timestamping static data of advertising creatives, and constructing a unified input dataset.

[0020] D={U,A,S}

[0021] Where: U∈R^(m×n1),

[0022] A∈R^(k×n2), S∈R^(t×n3),

[0023] m represents the number of users, k represents the number of advertisements, t represents the number of scenarios, and n1, n2, and n3 represent the feature dimensions of the three modalities.

[0024] In step S2, the text modality branch construction and encoding include:

[0025] S2-11. Using a BERT-based pre-trained encoder, input advertising text, user comment text, and keywords, and perform word segmentation and masking processing on the text.

[0026] S2-12. Semantic encoding is performed using a 12-layer Transformer encoder.

[0027] S2-13. Take the output corresponding to [CLS] token as the text feature vector F_text, F_text∈R^768, where: the hidden layer dimension of the BERT-base encoder is 768, the number of attention heads is 12, and the activation function is GELU;

[0028] Visual modality branch construction and coding include:

[0029] S2-21. A CNN+ViT hybrid encoder is adopted, which extracts local visual features of advertising images / video frames through 4 layers of CNN and outputs local feature vectors with a dimension of 256.

[0030] S2-22. Input the local features into the ViT encoder to extract global visual features, and finally output the visual feature vector F_vis, where F_vis∈R^512.

[0031] Behavioral timing branch construction and coding include:

[0032] S2-31. A Bi-GRU encoder is used as the input, with a time-series sequence of user behavior. The sequence length is fixed at 30. The Bi-GRU includes both forward GRU and backward GRU, and the hidden layer dimension is 256 for each.

[0033] S2-32. After concatenating the outputs of the forward and backward hidden layers, map them through a fully connected layer to form a behavior feature vector F_behavior, where F_behavior∈R^512;

[0034] Feature dimension unification involves adjusting the dimensions of F_text, F_vis, and F_behavior through fully connected layers, mapping them to the same dimension d=512, and obtaining modality-specific low-dimensional dense feature vectors with consistent dimensions, denoted as F_text'∈R^512, F_vis'∈R^512, and F_behavior'∈R^512, respectively.

[0035] In the above-mentioned AI-based precise matching method that integrates multimodal features, step S3 specifically includes the following steps:

[0036] S31. Cross-modal attention weight calculation: Construct a cross-modal attention matrix, calculate the correlation between features of each modality, and the attention weight calculation formula is: W_att=softmax((F_text'·W_q)·(F_vis'·W_k)^T+(F_text'·W_q)·(F_behavior'·W_k)^T+(F_vis'·W_q)·(F_behavior'·W_k)^T);

[0037] Where W_q and W_k are both learnable weight matrices, W_q∈R^(512×256) and W_k∈R^(512×256). The softmax function is used to normalize the weights to ensure that the sum of the weights is 1.

[0038] S32. Modal Feature Weighting Enhancement: Attention weights are applied to the corresponding modal features to obtain weighted modal features.

[0039] F_text''=W_att[0]·F_text', F_vis''=W_att[1]·F_vis', F_behavior''=W_att[2]·F_behavior';

[0040] Where: W_att[0], W_att[1], and W_att[2] are the attention weights corresponding to text, visual, and behavioral modalities, respectively;

[0041] S33. Feature Fusion and Nonlinear Mapping: A globally unified multimodal fusion feature vector F is generated by combining feature concatenation and activation functions. The calculation formula is: F=σ(F_text''·W_text⊕F_vis''·W_vis⊕F_behavior''·W_behavior); where σ is the LeakyReLU activation function, ⊕ is the feature concatenation operation, W_text, W_vis, and W_behavior are learnable modal weight matrices, and satisfy ||W_text||2+||W_vis||2+||W_behavior||2=1, and W_text, W_vis, and W_behavior are all ∈ R^(512×512). The final output fusion feature vector F ∈ R^512.

[0042] In the above-mentioned AI-based precise matching method that integrates multimodal features, step S4 includes the following steps:

[0043] S41. The underlying shared structure of MMoE (Multi-gate Mixture-of-Experts) is adopted. The bottom layer is set up with 4 expert networks, each of which is a 3-layer fully connected layer with GELU activation function. The upper layer is set up with two independent output heads, namely CTR prediction tower and CVR prediction tower. Each output head is a 2-layer fully connected layer, and the final output dimension is 1.

[0044] S42. Input the multimodal fusion feature vector F obtained in S3 into the MMoE bottom shared structure, extract deep features through the expert network, and output the ad click prediction probability p_CTR through the CTR tower.

[0045] S43. Construction of Joint Loss Function: Construct a joint loss function L_total that includes CTR loss, CVR loss, and regularization term;

[0046] S44. Model Training: The Adam optimizer is used for model training. The initial learning rate η = 1e-4, and the learning rate decays using a cosine annealing strategy. The batch size is 256, and the number of training epochs is 50. An early stopping strategy is adopted: training is stopped and the optimal model parameters are saved when the validation set loss does not decrease for 5 consecutive epochs. 6. According to claim 3, a method for accurate AI streaming matching that integrates multimodal features, step S5 includes the following steps:

[0047] S51. Construction of Multi-Objective Utility Ranking Formula: Taking into account click probability, conversion probability, similarity between user and ad features, and campaign cost, a formula for calculating the traffic matching score is constructed:

[0048] Score = w1·p_CTR + w2·p_CVR + w3·Sim(F_user,F_ad) - w4·Cost; where w1, w2, w3, and w4 are dynamic scheduling weights.

[0049] Sim(F_user, F_ad) represents the cosine similarity between user and ad features.

[0050] Cost refers to the unit cost per exposure.

[0051] S52, Precise Matching: For the current set of ads to be delivered, user traffic set, and scenario set, calculate the matching score for each group and sort them from high to low according to the score;

[0052] S53. Real-time traffic scheduling: Millisecond-level real-time decision-making is adopted, with a preset score threshold θ=0.6. When a certain group's score ≥ θ, the delivery is triggered; only the top-3 high-scoring ads are matched in the same traffic position, and traffic allocation is completed according to a greedy strategy.

[0053] In the above-mentioned AI-based precise matching method that integrates multimodal features, step S6 specifically includes the following steps:

[0054] S61. Feedback Data Collection: Real-time collection of actual feedback data after the flow is delivered;

[0055] S62. Abnormal Data Removal: The isolated forest unsupervised learning algorithm is used to detect anomalies in the collected feedback data and remove abnormal traffic and cheating samples.

[0056] S63, Incremental Learning Update: An incremental learning mechanism is adopted, with a sliding window of N=5000 real-time valid samples as one update cycle. New samples are input into the trained model, and the Adam optimizer is used to incrementally update the model parameters. The learning rate η=1e-5.

[0057] S64. Feature weight adjustment: Based on the effect of the feedback data, dynamically adjust the weight values ​​of the modal weight matrices W_text, W_vis, and W_behavior in S3;

[0058] S65. Closed-loop iteration: Repeat steps S61-S64 to achieve continuous iterative optimization of model parameters and feature weights, ensuring that the accuracy of flow matching improves dynamically with real-time data.

[0059] In step S1, user behavior modality data is collected using a tracking method, which includes page tracking, behavior tracking, and conversion tracking.

[0060] In step S21, the CNN+ViT hybrid encoder of the visual modality branch includes an image preprocessing step: normalizing the size of the advertisement picture / video frame (uniformly adjusted to 224×224 pixels), converting the color gamut, and denoising with Gaussian blur; for video frame acquisition, a key frame extraction algorithm is used, and 1 key frame is extracted every 5 frames.

[0061] In step S5, the adaptive adjustment of the dynamic scheduling weights w1, w2, w3, and w4 is calculated based on the real-time traffic competition degree. The calculation formula for the traffic competition degree Comp is:

[0062] Comp=(Ad_num / Traffic_num)×(Bid_avg / Bid_min);

[0063] where Ad_num is the number of advertisements to be placed currently,

[0064] Traffic_num is the number of available user traffic currently,

[0065] Bid_avg is the average bid of all current advertisers,

[0066] Bid_min is the minimum bid of all current advertisers;

[0067] The weight adjustment rule is:

[0068] When Comp≥1.2, w1 = 0.2, w2 = 0.4, w3 = 0.3, w4 = 0.1;

[0069] When 0.8 < Comp < 1.2, w1 = 0.3, w2 = 0.3, w3 = 0.2, w4 = 0.2;

[0070] When Comp≤0.8, w1 = 0.4, w2 = 0.2, w3 = 0.1, w4 = 0.3.

[0071] Compared with existing technologies, the advantages of this invention are as follows: By simultaneously collecting and standardizing the preprocessing of three types of multimodal heterogeneous data—user behavior, advertising materials, and scene environment—it overcomes the deficiency of single-dimensional features in traditional ad delivery. Simultaneously, it independently extracts deep features of each modality using BERT, CNN+ViT, and Bi-GRU multi-branch networks, and combines a cross-modal attention interaction fusion algorithm to achieve adaptive weighted fusion of multiple heterogeneous features, effectively preserving valid features and filtering redundant noise. Secondly, by building an MMoE dual-task prediction model with a multi-expert shared structure, it simultaneously completes the prediction of ad click and conversion probabilities, coupled with a customized joint loss function. This invention significantly improves the model's prediction accuracy and generalization ability. Furthermore, it dynamically and adaptively adjusts matching weights based on real-time traffic competitiveness, combining multi-objective utility scoring and a greedy scheduling strategy to accurately balance campaign performance and costs. Simultaneously, it incorporates an isolated forest abnormal traffic detection and a sliding window incremental iteration mechanism to construct a complete online feedback optimization loop, effectively avoiding interference from fraudulent traffic. Continuous iteration and optimization of model parameters and modal weights significantly improve the matching accuracy between ads and user traffic, campaign stability, and resource utilization. This adapts to various complex and ever-changing advertising scenarios, meeting the refined, intelligent, and low-cost commercial AI campaign requirements. Attached Figure Description

[0072] Figure 1 This is a flowchart of the method of the present invention;

[0073] Figure 2 This is the operational logic diagram of the joint prediction model of the present invention; Detailed Implementation

[0074] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0075] like Figure 1-2 As shown, an AI-based precise matching method integrating multimodal features is proposed. This method includes the following steps:

[0076] S1. Multimodal source data acquisition and standardized preprocessing: Real-time acquisition of three types of heterogeneous data: user behavior modality U, advertising material modality A, and scene environment modality S. After missing value imputation, normalization and time-series alignment, a unified input dataset D={U,A,S} is constructed, where U is the user feature set, A is the advertising feature set, and S is the scene feature set.

[0077] S2. Construct a multimodal branch feature extraction network: Encode the three modalities of text, vision, and behavioral time series independently, and output modality-specific low-dimensional dense feature vectors;

[0078] S3. Based on the cross-modal attention interaction fusion algorithm, multi-branch features are weighted and fused to generate a globally unified multi-modal fusion feature vector F;

[0079] S4. Build a dual-task joint prediction model, and simultaneously output the ad click prediction probability p_CTR and conversion prediction probability p_CVR based on the fusion feature F. Complete the end-to-end training of the model through the joint loss function.

[0080] S5. Calculate the traffic matching score based on the multi-objective utility ranking formula, and sort the scores from high to low to complete the precise matching of the three elements of advertisement-traffic-scenario and real-time traffic scheduling.

[0081] S6. Construct an online feedback loop, collect real clicks, conversions and negative feedback data after traffic is generated, and incrementally update and iterate the feature weights and model parameters.

[0082] In step S1, the user behavior modal data U includes basic user attributes (age, gender, region), behavior sequences (clicks, dwell time, favorites, conversions, jumps), and interest tags; the advertising material modal data A includes text information (ad title, product description, keywords), visual information (ad image frames, video keyframes, cover images), and audio information (ad voiceovers, background sounds); the scene environment modal data S includes the delivery scene (information feed, search page, live stream pop-up), device information (phone model, system version, network type), and time information (delivery time period, weekday, holidays); missing value imputation adopts a modal imputation strategy, missing values ​​of user behavior modal are imputed using the median of users with the same interest tags, missing values ​​of advertising visual modal are imputed using the mean of visual features of similar advertisements, and missing values ​​of scene modal are imputed using the pattern imputation method (taking the value with the highest frequency); normalization processing includes: logarithmic features are normalized using Min-Max, mapping feature values ​​to the [0,1] interval, and the calculation formula is:

[0083] x_norm=(x-x_min) / (x_max-x_min),

[0084] Where: x is the original feature value.

[0085] x_min is the minimum value of this feature.

[0086] x_max is the maximum value of this feature;

[0087] Categorical features are converted into vector form using one-hot encoding;

[0088] The time-series alignment steps include: aligning user behavior time-series data and scene dynamic data with a time granularity of 100ms; timestamping advertising creative static data to ensure consistency in the time dimension of the three types of data; and finally constructing a unified input dataset.

[0089] D={U,A,S}

[0090] Where: U∈R^(m×n1),

[0091] A∈R^(k×n2), S∈R^(t×n3),

[0092] m represents the number of users, k represents the number of advertisements, t represents the number of scenarios, and n1, n2, and n3 represent the feature dimensions of the three modalities.

[0093] In step S2, the text modality branch construction and encoding include:

[0094] S2-11. Using a BERT-based pre-trained encoder, input advertising text, user comment text, and keywords, and perform word segmentation and masking processing on the text.

[0095] S2-12. Semantic encoding is performed using a 12-layer Transformer encoder.

[0096] S2-13. Take the output corresponding to [CLS] token as the text feature vector F_text, F_text∈R^768, where: the hidden layer dimension of the BERT-base encoder is 768, the number of attention heads is 12, and the activation function is GELU;

[0097] Visual modality branch construction and coding include:

[0098] S2-21. A CNN+ViT hybrid encoder is used to extract local visual features of advertising images / video frames through 4 layers of CNN (with convolutional kernel sizes of 3×3, 3×3, 5×5, and 5×5, stride of 1, and padding of 1), and output a local feature vector with a dimension of 256.

[0099] S2-22. Input the local features into the ViT encoder (12 layers, 8 attention heads) to extract global visual features, and finally output the visual feature vector F_vis, F_vis∈R^512;

[0100] Behavioral timing branch construction and coding include:

[0101] S2-31. A Bi-GRU encoder is used as the input, which is a time-series sequence of user behavior (clicks, dwell times, conversions, etc., arranged in chronological order). The sequence length is fixed at 30 (padded with zeros if insufficient, truncated if too long). The Bi-GRU includes both forward GRU and backward GRU, and the hidden layer dimension is 256 for each.

[0102] S2-32. After concatenating the outputs of the forward and backward hidden layers, map them through a fully connected layer to form a behavior feature vector F_behavior, where F_behavior∈R^512;

[0103] Feature dimension unification involves adjusting the dimensions of F_text, F_vis, and F_behavior through fully connected layers, mapping them to the same dimension d=512, and obtaining modality-specific low-dimensional dense feature vectors with consistent dimensions, denoted as F_text'∈R^512, F_vis'∈R^512, and F_behavior'∈R^512, respectively.

[0104] Step S3 specifically includes the following steps:

[0105] S31. Cross-modal attention weight calculation: Construct a cross-modal attention matrix, calculate the correlation between features of each modality, and the attention weight calculation formula is: W_att=softmax((F_text'·W_q)·(F_vis'·W_k)^T+(F_text'·W_q)·(F_behavior'·W_k)^T+(F_vis'·W_q)·(F_behavior'·W_k)^T);

[0106] Where W_q and W_k are both learnable weight matrices, W_q∈R^(512×256) and W_k∈R^(512×256). The softmax function is used to normalize the weights to ensure that the sum of the weights is 1.

[0107] S32. Modal Feature Weighting Enhancement: Attention weights are applied to the corresponding modal features to obtain weighted modal features.

[0108] F_text''=W_att[0]·F_text', F_vis''=W_att[1]·F_vis', F_behavior''=W_att[2]·F_behavior'

[0109] Where: W_att[0], W_att[1], and W_att[2] are the attention weights corresponding to text, visual, and behavioral modalities, respectively;

[0110] S33. Feature Fusion and Nonlinear Mapping: A globally unified multimodal fusion feature vector F is generated by combining feature concatenation and activation functions. The calculation formula is: F=σ(F_text''·W_text⊕F_vis''·W_vis⊕F_behavior''·W_behavior);

[0111] Where: σ is the LeakyReLU activation function (slope is 0.01), ⊕ is the feature concatenation operation, W_text, W_vis, and W_behavior are learnable modality weight matrices, and satisfy ||W_text||2+||W_vis||2+||W_behavior||2=1, W_text, W_vis, and W_behavior are all ∈ R^(512×512), and the final output fused feature vector F ∈ R^512.

[0112] Step S4 includes the following steps:

[0113] S41. The underlying shared structure of MMoE (Multi-gate Mixture-of-Experts) is adopted. The bottom layer is set up with 4 expert networks, each of which is a 3-layer fully connected layer (the hidden layer dimensions are 512, 256, and 128 respectively), and the activation function is GELU. The upper layer is set up with two independent output heads, namely the CTR prediction tower (used to output the click probability p_CTR) and the CVR prediction tower (used to output the conversion probability p_CVR). Each output head is a 2-layer fully connected layer, and the final output dimension is 1.

[0114] S42. Input the multimodal fusion feature vector F obtained in S3 into the MMoE bottom-level shared structure. After extracting deep features through the expert network, output the ad click-through rate prediction probability p_CTR (p_CTR∈[0,1]) through the CTR tower and the ad conversion prediction probability p_CVR (p_CVR∈[0,1]) through the CVR tower. S43. Construction of joint loss function: Construct a joint loss function L_total containing CTR loss, CVR loss and regularization term for end-to-end model training. The calculation formula is as follows:

[0115] L_total=α·L_CTR+β·L_CVR+γ·L_reg;

[0116] Where: L_CTR is the cross-entropy loss, used to optimize the accuracy of CTR prediction.

[0117] L_CTR=-[y_ctr·ln(p_CTR)+(1-y_ctr)·ln(1-p_CTR)],

[0118] Where: y_ctr represents the clicked label (1 for clicked, 0 for unclicked);

[0119] L_CVR is the weighted cross-entropy loss, used to address the imbalance of positive and negative samples in CVR.

[0120] L_CVR=-[y_cvr·ln(p_CVR)+λ(1-y_cvr)·ln(1-p_CVR)],

[0121] Where: y_cvr is the conversion label (converted to 1, not converted to 0).

[0122] λ is the positive and negative sample balance coefficient, λ=3.5;

[0123] L_reg is the L2 regularization term, used to suppress model overfitting. L_reg = ∑(W_i²), where W_i is all the learnable parameters of the model.

[0124] α and β are task balancing coefficients, α+β=1, where α=0.4 and β=0.6, used to balance the training weights of the two tasks CTR and CVR;

[0125] γ is the regularity coefficient, γ = 1e-5;

[0126] S44. Model Training: The Adam optimizer is used for model training. The initial learning rate η = 1e-4. The learning rate decays using a cosine annealing strategy. The batch size is 256, the number of training epochs is 50, and an early stopping strategy is adopted. When the validation set loss does not decrease for 5 consecutive epochs, training is stopped and the optimal model parameters are saved.

[0127] Step S5 includes the following steps:

[0128] S51. Construction of Multi-Objective Utility Ranking Formula: Taking into account click probability, conversion probability, similarity between user and ad features, and campaign cost, a formula for calculating the traffic matching score is constructed:

[0129] Score=w1·p_CTR+w2·p_CVR+w3·Sim(F_user,F_ad)-w4·Cost;

[0130] Among them, w1, w2, w3, and w4 are dynamic scheduling weights, which are adaptively adjusted by the real-time traffic competition level, and satisfy w1+w2+w3+w4=1. When the traffic competition level is high, w2 (conversion weight) and w3 (similarity weight) increase, and w4 (cost weight) decreases; when the traffic competition level is low, w1 (click weight) increases, and w4 (cost weight) increases.

[0131] Sim(F_user, F_ad) is the cosine similarity between user and ad features, used to measure the matching degree between users and ads. Sim=cosθ=(F_user·F_ad) / (|F_user|·|F_ad|), where F_user is the user multimodal fusion feature (obtained from user behavior, interest and other modal features through steps S2 and S3), F_ad is the ad multimodal fusion feature (obtained from steps S2 and S3), and |F_user| and |F_ad| are the L2 norms of the user and ad fusion features, respectively.

[0132] Cost is the cost per impression, which is the cost of each ad impression. It is determined by the advertiser's bid and the platform's pricing. Cost = Bid ​​ / Impression, where Bid is the advertiser's bid and Impression is the estimated number of impressions.

[0133] S52, Precise Matching: For the current set of ads, user traffic, and scenarios to be delivered, calculate the matching score (Score) for each group (ad-user-scenario) and sort them from high to low according to the score;

[0134] S53 Real-time Traffic Scheduling: Employs millisecond-level real-time decision-making (response time ≤ 50ms), with a preset score threshold θ = 0.6. When the Score of a certain group (ad-user-scenario) is ≥ θ, the delivery is triggered. Only the top-3 high-scoring ads are matched for the same traffic position (same user, same scenario). Traffic allocation is completed according to a greedy strategy, that is, the ad with the highest Score is given priority, and the remaining traffic is allocated to the ads with the second and third highest Scores in turn, to avoid duplicate delivery and traffic waste.

[0135] Step S6 specifically includes the following steps:

[0136] S61. Feedback Data Collection: Collect real-time feedback data after ad placement, including positive feedback data (clicks, conversions, favorites, repeat purchases) and negative feedback data (skip, block, report, dwell time <1s), and collect abnormal data during ad placement (brushing, fake clicks, malicious conversions).

[0137] S62. Abnormal Data Removal: The Isolation Forest unsupervised learning algorithm is used to detect anomalies in the collected feedback data and remove abnormal traffic and cheating samples. The anomaly detection threshold is set to 0.8. When the abnormal score of a sample is ≥0.8, it is judged as an abnormal sample and removed.

[0138] S63, Incremental Learning Update: An incremental learning mechanism is adopted, with a sliding window of N=5000 real-time valid samples as one update cycle. New samples are input into the trained model, and the Adam optimizer is used to incrementally update the model parameters. The learning rate η=1e-5 (lower than the initial training learning rate to avoid model oscillation) is used. Only the fully connected layers and attention weight matrix of the model are updated, and the pre-trained encoder parameters are not updated.

[0139] S64. Feature Weight Adjustment: Based on the effect of the feedback data, dynamically adjust the weight values ​​of the modality weight matrices W_text, W_vis, and W_behavior in S3. For modalities with high positive feedback rates, increase the corresponding weight values ​​by 5%-10%; for modalities with high negative feedback rates, decrease the corresponding weight values ​​by 5%-10% to ensure the effectiveness of the fused features.

[0140] S65. Closed-loop iteration: Repeat steps S61-S64 to achieve continuous iterative optimization of model parameters and feature weights, ensuring that the accuracy of flow matching improves dynamically with real-time data.

[0141] In step S1, user behavior modality data is collected using a tracking method, which includes page tracking, behavior tracking, and conversion tracking.

[0142] Page tracking is used to collect data on the advertising pages visited by users, behavioral tracking is used to collect user actions on the advertising pages (clicks, swipes, pauses), and conversion tracking is used to collect user's final conversion actions (placing an order, making a payment, leaving contact information). The collection frequency is 10ms / time to ensure the real-time nature and completeness of the data.

[0143] In step S21, the CNN+ViT hybrid encoder of the visual modality branch includes image preprocessing steps: normalizing the size of the advertising images / video frames (uniformly adjusting them to 224×224 pixels), converting the color gamut (from RGB to YUV), and Gaussian blur denoising (Gaussian kernel size is 3×3, standard deviation is 0.5) to reduce the impact of visual noise on feature extraction; the video frame acquisition adopts a key frame extraction algorithm, extracting 1 key frame every 5 frames, balancing feature integrity and computational efficiency.

[0144] The adaptive adjustment of dynamic scheduling weights w1, w2, w3, and w4 in step S5 is based on real-time traffic contention calculation. The formula for calculating traffic contention Comp is:

[0145] Comp=(Ad_num / Traffic_num)×(Bid_avg / Bid_min);

[0146] Where Ad_num represents the number of ads to be delivered.

[0147] Traffic_num is the current available user traffic volume,

[0148] Bid_avg is the average bid of all current advertisers,

[0149] Bid_min is the minimum bid of all current advertisers;

[0150] The weight adjustment rule is as follows:

[0151] When Comp ≥ 1.2 (high competition), w1 = 0.2, w2 = 0.4, w3 = 0.3, w4 = 0.1;

[0152] When 0.8 < Comp < 1.2 (medium competition), w1 = 0.3, w2 = 0.3, w3 = 0.2, w4 = 0.2;

[0153] When Comp ≤ 0.8 (low competition), w1 = 0.4, w2 = 0.2, w3 = 0.1, w4 = 0.3.

[0154] In summary, the principle of this embodiment is as follows: By collecting three types of modal data of user behavior, advertisement materials, and scene environment and completing standardization and temporal alignment processing, using multi-branch networks of BERT, CNN+ViT, and Bi-GRU to extract deep features of text, vision, and temporal behavior respectively and unify the feature dimensions, and then through the cross-modal attention mechanism to achieve interactive weighting and fusion of multi-heterogeneous features, constructing a global fusion feature that can comprehensively represent user preferences, advertisement quality, and placement environment, relying on the MMoE multi-expert network architecture to build a CTR and CVR dual-task joint prediction model, combining a custom joint loss function to complete model training and convergence, and based on the real-time traffic competition degree to adaptively adjust dynamic weights, combining user-advertisement similarity and placement cost to construct a multi-objective matching scoring mechanism, completing the three-way precise matching and real-time traffic scheduling of advertisements, users, and scenes, and at the same time combining the Isolation Forest algorithm to filter abnormal cheating traffic, and continuously updating model parameters and modal weights through the sliding window incremental learning method, forming a complete traffic prediction, matching scheduling, and online iteration closed-loop, thereby achieving high-precision, low-cost, and high-stability AI intelligent traffic matching in complex traffic scenarios.

[0155] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Those skilled in the art of the present invention can make various modifications or supplements to the described specific embodiments or use similar methods to replace them, but will not deviate from the spirit of the present invention or exceed the scope defined by the appended claims.

Claims

1. A method for accurate AI streaming matching that integrates multimodal features, characterized in that, This method includes the following steps: S1. Multimodal source data acquisition and standardized preprocessing: Real-time acquisition of three types of heterogeneous data: user behavior modality U, advertising material modality A, and scene environment modality S. After missing value imputation, normalization and time-series alignment, a unified input dataset D={U,A,S} is constructed, where U is the user feature set, A is the advertising feature set, and S is the scene feature set. S2. Construct a multimodal branch feature extraction network: Encode the three modalities of text, vision, and behavioral time series independently, and output modality-specific low-dimensional dense feature vectors; S3. Based on the cross-modal attention interaction fusion algorithm, multi-branch features are weighted and fused to generate a globally unified multi-modal fusion feature vector F; S4. Build a dual-task joint prediction model, and simultaneously output the ad click prediction probability p_CTR and conversion prediction probability p_CVR based on the fusion feature F. Complete the end-to-end training of the model through the joint loss function. S5. Calculate the traffic matching score based on the multi-objective utility ranking formula, and sort the scores from high to low to complete the precise matching of the three elements of advertisement-traffic-scenario and real-time traffic scheduling. S6. Construct an online feedback loop, collect real clicks, conversions and negative feedback data after traffic is generated, and incrementally update and iterate the feature weights and model parameters.

2. The AI ​​streaming precise matching method integrating multimodal features according to claim 1, characterized in that, In step S1, the user behavior modal data U includes user basic attributes, behavior sequences, and interest tags; the advertising material modal data A includes text information, visual information, and audio information; the scene environment modal data S includes the placement scene, device information, and time information; the missing value imputation adopts a submodal imputation strategy, the missing values ​​of user behavior modal data are imputed by the median of users with the same interest tags, the missing values ​​of advertising visual modal data are imputed by the mean of visual features of similar advertisements, and the missing values ​​of scene modal data are imputed by the pattern imputation method; The normalization process includes: applying Min-Max normalization to numerical features, mapping feature values ​​to the [0,1] interval, and the calculation formula is as follows: x_norm=(x-x_min) / (x_max-x_min), Where: x is the original feature value. x_min is the minimum value of this feature. x_max is the maximum value of this feature; Categorical features are converted into vector form using one-hot encoding; The time-series alignment steps include: aligning user behavior time-series data and scene dynamic data with a time granularity of 100ms, timestamping advertising creative static data, and constructing a unified input dataset. D={U,A,S} Where: U∈R^(m×n1), A∈R^(k×n2), S∈R^(t×n3), m represents the number of users, k represents the number of advertisements, t represents the number of scenarios, and n1, n2, and n3 represent the feature dimensions of the three modalities.

3. The AI ​​streaming precise matching method integrating multimodal features according to claim 2, characterized in that, In step S2, the text modality branch construction and encoding includes: S2-11. Using a BERT-based pre-trained encoder, input advertising text, user comment text, and keywords, and perform word segmentation and masking processing on the text. S2-12. Semantic encoding is performed using a 12-layer Transformer encoder. S2-13. Take the output corresponding to [CLS] token as the text feature vector F_text, F_text∈R^768, where: the hidden layer dimension of the BERT-base encoder is 768, the number of attention heads is 12, and the activation function is GELU; The aforementioned visual modality branch construction and encoding includes: S2-21. A CNN+ViT hybrid encoder is adopted, which extracts local visual features of advertising images / video frames through 4 layers of CNN and outputs local feature vectors with a dimension of 256. S2-22. Input the local features into the ViT encoder to extract global visual features, and finally output the visual feature vector F_vis, where F_vis∈R^512. The construction and encoding of the behavioral temporal branching includes: S2-31. A Bi-GRU encoder is used as the input, with a time-series sequence of user behavior. The sequence length is fixed at 30. The Bi-GRU includes both forward GRU and backward GRU, and the hidden layer dimension is 256 for each. S2-32. After concatenating the outputs of the forward and backward hidden layers, map them through a fully connected layer to form a behavior feature vector F_behavior, where F_behavior∈R^512; The aforementioned feature dimension unification involves adjusting the dimensions of F_text, F_vis, and F_behavior through a fully connected layer, mapping them to the same dimension d=512, and obtaining a mode-specific low-dimensional dense feature vector with consistent dimensions, denoted as F_text'∈R^512, F_vis'∈R^512, and F_behavior'∈R^512, respectively.

4. The AI ​​streaming precise matching method integrating multimodal features according to claim 3, characterized in that, Step S3 specifically includes the following steps: S31. Cross-modal attention weight calculation: Construct a cross-modal attention matrix, calculate the correlation between features of each modality, and the attention weight calculation formula is: W_att=softmax((F_text'·W_q)·(F_vis'·W_k)^T+(F_text'·W_q)·(F_behavior'·W_k)^T+(F_vis'·W_q)·(F_behavior'·W_k)^T); Where W_q and W_k are both learnable weight matrices, W_q∈R^(512×256) and W_k∈R^(512×256). The softmax function is used to normalize the weights to ensure that the sum of the weights is 1. S32. Modal Feature Weighting Enhancement: Attention weights are applied to the corresponding modal features to obtain weighted modal features. F_text''=W_att[0]·F_text', F_vis''=W_att[1]·F_vis', F_behavior''=W_att[2]·F_behavior'; Where: W_att[0], W_att[1], and W_att[2] are the attention weights corresponding to text, visual, and behavioral modalities, respectively; S33. Feature Fusion and Nonlinear Mapping: A globally unified multimodal fusion feature vector F is generated by combining feature concatenation and activation functions. The calculation formula is: F=σ(F_text''·W_text⊕F_vis''·W_vis⊕F_behavior''·W_behavior); where σ is the LeakyReLU activation function, ⊕ is the feature concatenation operation, W_text, W_vis, and W_behavior are learnable modal weight matrices, and satisfy ||W_text||2+||W_vis||2+||W_behavior||2=1, and W_text, W_vis, and W_behavior are all ∈ R^(512×512). The final output fusion feature vector F ∈ R^512.

5. The AI ​​streaming precise matching method integrating multimodal features according to claim 3, characterized in that, Step S4 includes the following steps: S41. The underlying shared structure of MMoE (Multi-gate Mixture-of-Experts) is adopted. The bottom layer is set up with 4 expert networks, each of which is a 3-layer fully connected layer with GELU activation function. The upper layer is set up with two independent output heads, namely CTR prediction tower and CVR prediction tower. Each output head is a 2-layer fully connected layer, and the final output dimension is 1. S42. Input the multimodal fusion feature vector F obtained in S3 into the MMoE bottom shared structure, extract deep features through the expert network, and output the ad click prediction probability p_CTR through the CTR tower. S43. Construction of Joint Loss Function: Construct a joint loss function L_total that includes CTR loss, CVR loss, and regularization term; S44. Model Training: The Adam optimizer is used for model training. The initial learning rate η = 1e-4. The learning rate decays using a cosine annealing strategy. The batch size is 256, the number of training epochs is 50, and an early stopping strategy is adopted. When the validation set loss does not decrease for 5 consecutive epochs, training is stopped and the optimal model parameters are saved.

6. The AI ​​streaming precise matching method integrating multimodal features according to claim 5, characterized in that, Step S5 includes the following steps: S51. Construction of Multi-Objective Utility Ranking Formula: Taking into account click probability, conversion probability, similarity between user and ad features, and campaign cost, a formula for calculating the traffic matching score is constructed: Score = w1·p_CTR + w2·p_CVR + w3·Sim(F_user,F_ad) - w4·Cost; where w1, w2, w3, and w4 are dynamic scheduling weights. Sim(F_user, F_ad) represents the cosine similarity between user and ad features. Cost refers to the unit cost per exposure. S52, Precise Matching: For the current set of ads to be delivered, user traffic set, and scenario set, calculate the matching score for each group and sort them from high to low according to the score; S53. Real-time traffic investment scheduling: Adopt millisecond-level real-time decision-making, preset the score threshold θ = 0.6, and trigger the placement when the Score of a certain group ≥ θ; Only match the Top-3 high-score ads for the same traffic position, and complete the traffic allocation according to the greedy strategy.

7. The AI ​​streaming precise matching method integrating multimodal features according to claim 1, characterized in that, Step S6 specifically includes the following steps: S61. Feedback data collection: Real-time collect the real feedback data after traffic investment. S62. Abnormal data elimination: Use the Isolation Forest unsupervised learning algorithm to perform anomaly detection on the collected feedback data, and eliminate abnormal traffic and cheating samples. S63. Incremental learning update: Adopt an incremental learning mechanism, take a sliding window of N = 5000 real-time valid samples as one round of update cycle, input the new samples into the trained model, and use the Adam optimizer to perform incremental updates on the model parameters, with the learning rate η = 1e-5. S64. Feature weight adjustment: Dynamically adjust the weight values of the modality weight matrices W_text, W_vis, and W_behavior in S3 according to the effect of the feedback data. S65. Closed-loop iteration: Repeat steps S61 - S64 to achieve continuous iterative optimization of the model parameters and feature weights, and ensure that the traffic investment matching accuracy dynamically improves with real-time data.

8. The method according to claim 1, characterized in that, In step S1, the user behavior modality data is collected using the data embedding method, and the data embedding is divided into page data embedding, behavior data embedding, and conversion data embedding.

9. The AI ​​streaming precise matching method integrating multimodal features according to claim 2, characterized in that, The CNN + ViT hybrid encoder of the visual modality branch in step S21 includes an image preprocessing step: Normalize the size of the ad image / video frame (uniformly adjust to 224×224 pixels), perform color gamut conversion, and Gaussian blur denoising; The video frame collection uses the key frame extraction algorithm, and 1 key frame is extracted every 5 frames.

10. The method according to claim 5, characterized in that, In step S5, the adaptive adjustment of the dynamic scheduling weights w1, w2, w3, and w4 is calculated based on the real-time traffic competition degree. The calculation formula for the traffic competition degree Comp is: Comp = (Ad_num / Traffic_num) × (Bid_avg / Bid_min); where Ad_num is the number of ads to be placed currently, Traffic_num is the number of available user traffic currently, Bid_avg is the average bid of all current advertisers, Bid_min is the minimum bid of all current advertisers; The weight adjustment rule is: When Comp ≥ 1.2, w1 = 0.2, w2 = 0.4, w3 = 0.3, w4 = 0.1; When 0.8 < Comp < 1.2, w1 = 0.3, w2 = 0.3, w3 = 0.2, w4 = 0.2; When Comp ≤ 0.8, w1 = 0.4, w2 = 0.2, w3 = 0.1, w4 = 0.3.