AI-based marketing advertisement accurate placement method

Through the optimization of multimodal sensor data and dynamic delivery strategies, the problem of insufficient perception of users' emotional states in advertising delivery has been solved, and accurate matching of advertising content with user emotions has been achieved, thereby improving the personalization of advertising and user experience.

CN120598617BActive Publication Date: 2025-10-17SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511106414.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-17
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing advertising delivery technologies lack the ability to perceive users' real-time emotional states and attention changes in a fine-grained manner, resulting in insufficient timing and matching of advertising displays, an inability to adaptively adjust delivery strategies, content duplication, and a degraded user experience.

Method used

By collecting multimodal sensory data from users and utilizing pre-trained emotion classifiers and matching models, we dynamically evaluate the emotional adaptability of advertising materials, select the best display timing based on the user attention decay curve, adjust the delivery strategy in real time, and record negative feedback data to optimize the model.

Benefits of technology

It achieves a high degree of fit between advertising content and user emotions, improves the personalization level and exposure efficiency of recommendation results, enhances user emotional acceptance and advertising influence, and improves the system's adaptability and long-term strategy stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120598617B_ABST
    Figure CN120598617B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of advertisement delivery, in particular to an AI-based marketing advertisement precise delivery method, which comprises the following steps: collecting multi-modal sensor data of a user terminal, extracting facial micro-expression, touch screen pressure, device rotation angular velocity and other behavior information, generating an emotion feature vector through a pre-trained emotion classifier; matching the emotion feature with a sentiment annotation matrix of advertisement materials, selecting advertisement materials with an emotion adaptation degree higher than a threshold value to construct a candidate advertisement set; combining a user attention decay curve with an advertisement material display duration to perform duration matching, so that delivery can be completed before an attention critical point; when a user skips an advertisement, generating a negative feedback feature matrix by extracting pupil focus position, finger trajectory and acceleration information, dynamically updating the weight of the emotion classifier and optimizing the matching threshold value. The application realizes an emotion-driven, attention-aware and negative feedback self-learning precise delivery mechanism, and significantly improves the personalized matching degree and delivery effect of advertisements.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of advertisement delivery, and in particular to an AI-based marketing advertisement accurate delivery method. BACKGROUND

[0002] With the development of mobile Internet and the popularity of smart terminals, advertisement delivery gradually evolves from the traditional content matching mode to the user portrait-driven personalized recommendation mode. In particular, in high-frequency interactive scenarios such as social platforms, short videos, and news clients, user browsing behavior, emotional response, and content acceptance become key factors affecting the conversion effect of advertisements. To improve the accuracy of advertisement delivery and user response rate, more and more systems begin to introduce artificial intelligence and sensor fusion technology, collect user behavior data, build multi-dimensional feature models, and dynamically adapt to the content characteristics of advertisement materials. Emotion perception, as a core direction of user state recognition, has important value in personalized advertisement recommendation, which can effectively improve the matching degree of advertisement content and user psychology, thereby driving accurate conversion.

[0003] However, the existing advertisement delivery technology mainly relies on static user portraits or coarse-grained behavior tags, lacks fine-grained perception of user real-time emotional state and attention changes, and cannot dynamically determine the display timing and matching degree of advertisements. In addition, there is a lack of in-depth modeling of user negative feedback behavior in the advertisement delivery process, making it difficult for the system to adaptively adjust the delivery strategy in multiple skipping behaviors, causing problems such as repeated advertisement content, mismatched delivery time, and decreased user experience. At the same time, the emotional labeling method of advertisement materials is relatively rough, lacking alignment mechanism in structure and dimension with user emotional features, resulting in inaccurate emotion adaptation degree evaluation. SUMMARY

[0004] The application provides an AI-based marketing advertisement accurate delivery method, which realizes closed-loop adaptive control from perception-matching-delivery-optimization based on multi-modal perception, dynamic attention modeling, and negative feedback learning mechanism.

[0005] The AI-based marketing advertisement accurate delivery method comprises the following steps:

[0006] S1: Collecting multi-modal sensor data of the user terminal, including facial micro-expression change rate, touch screen pressure value, and device rotation angular velocity, and outputting an emotional feature vector through a pre-trained emotion classifier;

[0007] S2: Inputting the emotional feature vector generated in S1 into a matching degree model, calculating the emotional adaptation degree of each material in the advertisement material library, and selecting materials with an adaptation degree greater than a threshold alpha to generate a candidate advertisement set;

[0008] S3: Select an advertisement material with a matching display duration from the candidate advertisement set based on the current attention decay curve of the user, and display it before the user's attention decays to the critical point;

[0009] S4: When the advertisement is skipped, record the user's pupil focus position and finger movement trajectory 0.5 seconds before skipping, generate a negative feedback feature matrix, and update the emotion classifier weight of S1 in reverse.

[0010] Optionally, the pre-trained emotion classifier includes a multi-modal input layer, a time series feature extraction module, a dynamic attention fusion layer, and a classification output layer.

[0011] Optionally, S1 includes:

[0012] S11: Capture the facial video stream in real time through the front camera of the user terminal, calculate the facial micro-expression change rate between consecutive frames, simultaneously obtain the touch screen pressure value of the screen contact area through the pressure sensor, and collect the device rotation angular velocity in three-dimensional space through the gyroscope, timestamp align and noise filtering on the above three kinds of original data, and generate standardized multi-modal sensing data.

[0013] S12: Input the standardized multi-modal sensing data into the time series feature extraction module of the pre-trained emotion classifier, wherein the facial micro-expression change rate is output as a visual time series feature vector through a first time series feature extraction channel, and the touch screen pressure value and the device rotation angular velocity are output as an inertial time series feature vector through a second time series feature extraction channel.

[0014] S13: Input the visual time series feature vector and the inertial time series feature vector output by S12 into the dynamic attention fusion layer of the pre-trained emotion classifier, generate a fusion feature matrix through cross-modal attention weight distribution, and finally map it to a 128-dimensional emotion feature vector through the classification output layer.

[0015] Optionally, the matching degree model includes a dynamic weight generation module, an adaptation degree calculation engine, and a sorting filter.

[0016] Optionally, S2 includes:

[0017] S21: Extract the visual emotional features and audio emotional features of each advertisement material in the advertisement material library, output the material pleasantness and material arousal through the pre-trained material emotional analysis model, and generate an emotional labeling matrix containing two-dimensional emotional indicators.

[0018] S22: Input the emotion feature vector generated by S1 into the dynamic weight generation module of the matching degree model, output the weight coefficient W1 and the weight coefficient W2, combine the emotional labeling matrix output by S21, calculate the emotional adaptation degree of each material, and form a dynamic adaptation degree set.

[0019] S23: screening the emotion adaptation degree > threshold value from the dynamic adaptation degree set , generating a candidate advertisement set in descending order of adaptation degree, wherein the threshold value is dynamically adjusted according to the historical advertisement conversion rate of the user.

[0020] Optionally, the pre-trained material emotion analysis model comprises a visual feature extraction branch, an audio feature extraction branch, and a bimodal attention fusion layer, wherein;

[0021] The visual feature extraction branch converts the RGB frame of the advertisement material into the HSV color domain through the HSV conversion layer, and calculates the weighted product of the dominant color saturation and brightness in the color emotion mapping layer to generate visual emotion features;

[0022] The audio feature extraction branch extracts 128-dimensional mel-frequency cepstrum coefficients in the range of 0-8 kHz through the mel-frequency spectrum coding layer, and generates a beat-emotion intensity curve through the beat emotion analysis layer;

[0023] The bimodal attention fusion layer calculates the correlation matrix of visual and audio features, outputs the fusion feature vector to the fully connected classification layer, and outputs the material pleasantness and material arousal to form the emotion labeling matrix.

[0024] Optionally, S3 comprises:

[0025] S31: Real-time monitoring of the page stay duration and scrolling speed of the user on the current browsing page, acquiring time sequence data at a sampling frequency of 10 times per second, and synchronously recording the screen coordinate range of the advertisement display area.

[0026] S32: Based on the page stay duration and scrolling speed collected in S31, fitting the current attention decay curve through an exponential decay function.

[0027] S33: Extracting the best display duration of each advertisement material from the candidate advertisement set generated in S2, calculating the overlap degree of the critical attention threshold corresponding to the time window, and selecting the advertisement material with an overlap degree > overlap degree threshold to generate a duration adaptation advertisement subset.

[0028] S34: If the current attention decay curve falls below the critical attention threshold , select the advertisement material with the highest priority from the duration adaptation advertisement subset, and display the advertisement material in the corresponding channel window according to the screen coordinate range recorded in S31.

[0029] Optionally, the exponential decay function is expressed as:

[0030] ;

[0031] wherein, represents the attention level of the user at the time , is an initial attention value, is an attention decay coefficient, reflecting the activity level of the user's scrolling behavior, is the duration of the user's stay on the current page.

[0032] Optionally, the S4 comprises:

[0033] S41: When the user skips the advertisement, capture the sequence of user pupil focus positions 0.5 seconds before skipping through the front camera at a sampling rate of 120fps, while recording the coordinate point set of the finger movement trajectory through the touch screen sensor, and synchronously collecting the peak data of the three-axis accelerometer.

[0034] S42: Calculate the average Euclidean distance of the sequence of user pupil focus positions obtained in S41 from the boundary of the advertisement area, extract the path curvature of the starting point to the advertisement close button for the finger movement trajectory, and generate a three-dimensional negative feedback feature matrix in combination with the acceleration peak value.

[0035] S43: Input the negative feedback feature matrix into the pre-trained emotion classifier of S1, calculate the cosine similarity deviation of the output vector and the original emotion feature vector, and when the deviation value is greater than the threshold , update the convolution kernel weight of the emotion classifier through the back propagation algorithm.

[0036] S44: Count the number N of times the user continuously skips the same type of advertisement, and when N>3, adaptively reduce the threshold of S33 and the overlap threshold of S33.

[0037] Advantages of the present application:

[0038] The present application, by constructing a multi-modal sensing input system containing visual, tactile and motion information, uses a pre-trained emotion classifier to model the current real emotional state of the user, and further uses a matching degree model to fuse and analyze the emotional state with the emotional dimension of the advertisement material, to realize dynamic evaluation of the emotional adaptation degree of the advertisement. This mechanism ensures that the advertisement content is highly contextually matched with the user's emotions before being delivered, significantly improving the personalization level of the recommendation results and the user's emotional acceptance, avoiding user's aversion or skipping behavior caused by emotional mismatch, and enhancing the actual influence of the advertisement content.

[0039] The application uses user page staying time and scrolling speed data to fit the attention decay curve in real time, matches the dynamic curve with the optimal display time of the candidate advertising materials, and realizes the advertising adaptation and selection in the attention dimension. The system only selects the materials suitable for the current attention level change trend for delivery, and completes the display before the critical decay point of attention is reached, and further combines the recorded screen coordinate range for cross-channel delivery. This mechanism can accurately grasp the advertising display window, avoid skipping or ignoring the materials due to long display or misplaced delivery, and improve the advertising exposure efficiency and the coordination of the user interface.

[0040] The application collects high-frequency behavior data such as pupil focus position, finger movement trajectory and acceleration signal in real time after the user skips the advertisement, constructs a three-dimensional negative feedback feature matrix, and adjusts the emotion classifier based on the feedback information. At the same time, the system dynamically adjusts the emotion matching threshold and the display time overlap threshold according to the type and frequency of the skipping behavior, forming a delivery strategy optimization mechanism with self-learning ability. This feedback correction process enables the system to have behavior adaptive ability for individual users, effectively avoids the problem of performance degradation of the recommendation caused by model aging or deviation, and significantly improves the stability of the long-term delivery strategy and the overall intelligence level of the system. BRIEF DESCRIPTION OF DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only a part of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0042] Fig. 1 The method flowchart of the embodiment of the present application is shown in the figure.

[0043] Fig. 2 The S2 flowchart of the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0044] The present application will be described in detail below in combination with the drawings and specific embodiments. It should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and other alternative ways can also be used by those skilled in the art to implement them; and the drawings are only used to describe the embodiments more specifically, and are not intended to specifically limit the present application.

[0045] It is noted that the recitations "one embodiment," "an embodiment,” “one example embodiment,” “some embodiments,” etc. indicate the described features, structures, or characteristics can be included in one or more embodiments, but each as such an embodiment can not necessarily include than the particular feature, structure, or characteristic. In addition, when describing features, structures, or characteristics in the foregoing description, it should be understood that such features, structures, or characteristics can be implemented in other embodiments, whether or not explicit description of such an embodiment is provided. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it should be understood that explicit description of such feature, structure, or characteristic can also appear in the description or claims of another embodiment, as such an embodiment is within the knowledge of one of ordinary skill in the art.

[0046] Generally, the terminology can be understood at least in part from usage in context. For example, the term “one or more” as used herein, depending at least in part upon context, can be used to describe any feature, structure, or characteristic in a singular sense or can be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “based on,” can be understood as not necessarily delimiting a set of exclusive factors and, instead, can allow for additional factors not necessarily expressly described.

[0047] As shown in Figs. 1-2 the AI-based marketing advertisement accurate placement method includes the following steps:

[0048] S1: Collecting multi-modal sensor data of the user terminal, including facial micro-expression change rate, touch screen pressure value, and device rotation angular velocity, outputting an emotional feature vector through a pre-trained emotion classifier, specifically:

[0049] S11, multi-modal sensor data acquisition and preprocessing: first, the front camera of the user terminal is used to collect the video stream of the user's face in real time. The video stream is processed by face key point detection and action unit coding algorithm, and the facial micro-expression change rate between consecutive image frames is extracted to depict the visual emotional cues of the user in the current use context. At the same time, the touch screen pressure value from the touch screen is collected to record the user's operation strength during the interface process. The built-in gyroscope of the device is called synchronously to obtain the rotation angular velocity of the device in three-dimensional space to reflect the user's dynamic operation state of the device. The above three types of data are bound to a unified time label, and noise filtering processing is performed respectively, including removing the light disturbance of the facial video, smoothing the discrete value of the touch screen data, and stabilizing the fluctuation of the rotation data. After time alignment and filtering processing, the standardized multi-modal sensor data of unified structure is finally formed as the input data source of the emotion recognition model.

[0050] S12, time sequence feature vector extraction: the standardized multi-modal sensor data is input into the time sequence feature extraction module of the pre-trained emotion classifier. The module consists of two channels, which process information of different modalities respectively. Among them, the facial micro-expression change rate extracts visual-related time sequence features through the first channel, which can identify the expression change trend in a short time and form a stable visual time sequence feature vector; the touch screen pressure value and the device rotation angular velocity are modeled by the second channel. The channel can extract the dynamic behavior pattern of the user in operation and output the inertial time sequence feature vector. The results extracted by the two channels will be used as the basic features for multi-modal fusion.

[0051] S13, feature fusion and emotion feature generation: the visual time sequence feature vector and the inertial time sequence feature vector are input into the dynamic attention fusion layer of the pre-trained emotion classifier. In this fusion layer, the system weights the input features according to the weight distribution strategy of different modalities, realizes the dynamic fusion of visual information and action information at the decision level. The fused feature vector is input into the classification output layer to generate an emotion feature vector representing the current user's emotional state. The feature vector is the key input for the subsequent advertisement emotion matching and intelligent delivery.

[0052] S2: input the emotion feature vector generated in S1 into the matching degree model, calculate the emotion adaptation degree of each material in the advertisement material library, select the materials with adaptation degree > threshold a to generate a candidate advertisement set, specifically:

[0053] S21, generation of advertisement material emotion labeling matrix: first, perform emotion feature extraction process on each advertisement material in the advertisement material library. The process is completed by a pre-trained material emotion analysis model, which includes a visual feature extraction branch, an audio feature extraction branch, and a dual-modal attention fusion layer.

[0054] In the visual feature extraction branch, the image frames of the advertisement material are subjected to color space conversion, converting the RGB color representation of the original image into HSV color domain representation, from which the dominant hue information is extracted, combined with the saturation and brightness indicators of the image, to generate visual emotion features representing the overall visual emotional tendency of the advertisement.

[0055] In the audio feature extraction branch, the background sound or dubbing content in the advertisement material is analyzed, the structural features of the audio in the specified frequency range are extracted through spectrum transformation, and the relationship between rhythm and emotion is identified to obtain the dynamic mapping between beat and emotion intensity, forming audio emotion features.

[0056] The visual and audio emotion features output by the two branches are sent to a dual-modal attention fusion layer, in which the system dynamically allocates fusion weights according to the relevance between the two types of features, thereby forming a unified fusion feature vector as input for emotion analysis. The fusion feature vector is then input into a fully connected classification layer, which outputs two emotion indicators, namely, the material pleasantness and the material arousal. Finally, the emotion analysis results of all the advertising materials are organized into an emotion labeling matrix, which serves as the basis for the subsequent mood adaptation degree calculation.

[0057] S22, mood adaptation degree calculation: the mood feature vector generated in S1 is input into the dynamic weight generation module of the matching degree model. This module dynamically generates weight coefficients W1 and W2 according to the user state reflected in the current mood feature, which are used to adjust the proportion of pleasantness and arousal in the matching process. Subsequently, the system combines these two weight coefficients with the emotion labeling matrix output in step S21 to calculate the matching degree between each advertising material and the current user mood state. This matching process is completed by the adaptation degree calculation engine in the matching degree model, which can integrate the material pleasantness and arousal by weighting according to the weight coefficients, and take the integration result as the mood adaptation degree value of the material.

[0058] Mood adaptation degree = W1 x material pleasantness + W2 x material arousal

[0059] The final output result is a set of dynamic adaptation degrees, each element of which represents the adaptation degree of an advertising material to the current user mood state.

[0060] The dynamic generation of weight coefficients W1 and W2 includes:

[0061] 1. Emotion factor projection of mood feature vector: First, input the output mood feature vector into the emotion factor mapping layer in the dynamic weight generation module. This layer has an emotion factor dictionary trained based on clustering labels, which has divided the training samples into high-dimensional features according to the pleasantness dominant type, arousal dominant type, and neutral state region. Use the projection function to normalize and map the input mood feature vector, and calculate its relative offset position in the "pleasantness-arousal" two-dimensional emotion plane.

[0062] 2. Weight offset coefficient calculation: according to the projection result, evaluate the degree of dependence of the user's current mood state on pleasantness or arousal. For example:

[0063] If the mood feature falls in the high-pleasantness and low-arousal region (such as relaxation, satisfaction), it will tend to give a higher weight to pleasantness;

[0064] If the mood feature falls in the high-arousal and low-pleasantness region (such as anxiety, tension), it will tend to give a higher weight to arousal;

[0065] If it is in the emotional center area, the weighting of the two is close to equal.

[0066] This strategy is implemented by the bidirectional attention normalization network in the module, and the output is two floating-point weight coefficients.

[0067] 3. Weight coefficient normalization output: The two preliminary coefficients obtained above are Softmax normalized or linearly scaled to satisfy the constraint that the total weight is 1, forming the final:

[0068] Weight coefficient W1: used to adjust the proportion of the material's pleasantness;

[0069] Weight coefficient W2: used to adjust the proportion of material awakening degree.

[0070] S23, candidate ad set generation: After obtaining the dynamic fitness set, the set is input into the sorting filter in the fitness model. First, all the ads with a fitness higher than the current threshold are selected. creatives, the threshold This is a reference standard dynamically adjusted based on historical user ad interaction data, ensuring recommendations have personalized conversion potential. Selected creatives are sorted by suitability from high to low and organized into a candidate ad set. This candidate ad set serves as input for subsequent ad display, enabling precise delivery before user attention reaches critical levels.

[0071] S3: Based on the user's current attention decay curve (fitted by page dwell time and scrolling speed), select ad creatives with matching display durations from the candidate ad set and deliver them before the user's attention decays to the critical point. Specifically:

[0072] S31, collection of attention decay parameters: After the user enters a specific content page, a real-time monitoring mechanism for attention behavior parameters is started. This monitoring includes two key dimensions: one is the page dwell time, that is, the continuous browsing time since the user entered the current page; the other is the scrolling speed, that is, the user's sliding rate in the vertical direction. These two parameters are recorded at a fixed sampling frequency of ten times per second to form time series data for subsequent attention modeling. At the same time, the coordinate range of the current advertising display area on the screen is synchronously recorded to determine the display position of the delivery material on the terminal interface. The continuous collection of the above parameters runs through the entire process of user page browsing and serves as important input data for subsequent attention decay modeling.

[0073] S32, attenuation curve modeling: The page dwell time and scrolling speed data collected in S31 are input into the attenuation modeling module. This module establishes a dynamic attenuation model of attention over time based on the user's browsing rhythm and scrolling behavior pattern in the current page.

[0074] First, the system analyzes the complexity of the text, images, and interactive elements on the page's first screen, and calculates the information entropy based on the structured content density, which is used to estimate the initial state of user attention when the page is first loaded. This initial state is denoted as the initial attention value .

[0075] Next, the system performs statistical analysis on the scrolling speed data to obtain the fluctuation degree in the time series, and dynamically adjusts the decay speed control factor based on the scrolling speed variance to construct the decay coefficient , which is specifically:

[0076] 1. Scrolling speed time series extraction: Record the user's vertical scrolling speed sequence in the current page at a fixed sampling frequency (such as 10Hz). This speed sequence reflects the user's browsing rhythm for the page content, and frequent value changes mean fast browsing, while stable changes represent in-depth reading.

[0077] 2. Scrolling speed variance calculation: Calculate the variance of the scrolling speed within a sliding time window (such as the last 5 seconds) to obtain the volatility indicator of the scrolling behavior. The larger the variance value, the more intense the user's scrolling and the less stable the attention; otherwise, the attention is concentrated.

[0078] 3. Mapping to construct decay coefficient: To map the scrolling speed fluctuation to the attention decay speed, the system uses linear normalization and interval mapping mechanism to construct the decay coefficient . The specific process is as follows:

[0079] Set the upper and lower limits of the scrolling speed variance (such as the minimum is 0.0 and the maximum is an empirical value), and normalize the currently calculated variance to the range of 0~1;

[0080] According to the normalized fluctuation value, perform linear interpolation in the preset decay coefficient interval (for example: 0.1 to 1.0) to output the decay coefficient under the current session;

[0081] If the user continuously exhibits high fluctuation behavior in a short period of time, the system can weight the previous round value to form a smooth transition and avoid jitter.

[0082] Through this mechanism, it can flexibly reflect different user behavior characteristics and achieve personalized fitting of the attention model.

[0083] For example, when the scrolling speed sequence is relatively stable (such as variance close to 0.05), the mapped is about 0.2, and the attention decay curve decreases slowly; while when the user quickly scrolls the page (variance reaches 0.8), the system maps the 0.9, corresponding to fast attention loss. This dynamic adjustment process enables the exponential decay function to accurately depict the user's current interaction state.

[0084] Based on the above two parameters, the system uses an exponential decay function to fit the user's attention, and the function form used is: ;

[0085] wherein, represents the attention level of the user at time , is the initial attention value, is the attention decay coefficient, reflecting the activity level of the user's scrolling behavior, is the duration of the user's stay on the current page.

[0086] This decay function can dynamically reflect the user's attention trend in the page. If the user quickly slides the page, the scrolling speed fluctuates greatly, and the value increases, and the decay speed is faster; if the user browses smoothly, the value decreases, and the attention time is longer.

[0087] Finally, the system generates a complete attention decay curve based on the function, which serves as the basis for subsequent advertisement material length matching and display timing determination, providing mathematical support for user attention-driven precise advertisement placement.

[0088] S33, advertisement material length matching: obtain the candidate advertisement set generated in S2, and extract the best display length of each advertisement material one by one. This display length can be preset or dynamically evaluated according to factors such as material content density, recommendation strategy, historical placement feedback, etc. Then, the system compares the display length of each advertisement material with the key time window in the attention decay curve, which is the effective display interval before the attention value drops to the critical attention threshold . Calculate the matching degree of the advertisement material display length in this time window, i.e. the overlap degree, select the advertisement material with overlap degree > overlap degree threshold to generate the length-adapted advertisement subset, which only contains advertisement materials that closely match the current attention period, ensuring effective display within the remaining attention window.

[0089] S34, cross-channel placement triggering: during the user's browsing process, when the attention decay curve drops to the critical attention threshold When the time is up, the system immediately triggers the advertisement placement decision mechanism. From the length-adapted advertisement subset, the material priority strategy is sorted according to the pre-set material priority strategy. The priority can be calculated based on the material quality score, brand budget, user interest preference and other factors. Then, the highest priority advertisement material is selected, combined with the advertisement display area screen coordinate range recorded in S31, and the selected advertisement material is automatically placed in the current content page or associated application window. If there are multiple parallel display channels (such as the main page, embedded video player, recommendation card, etc.), the system can also determine the most suitable placement entrance combined with the terminal state, realize intelligent switching and placement across channels, and improve the visibility of the advertisement and the user response probability.

[0090] S4: When the advertisement is skipped, record the user's pupil focus position and finger movement trajectory 0.5 seconds before skipping, generate a negative feedback feature matrix, and update the emotion classifier weight of S1 in the reverse direction, specifically:

[0091] S41, negative feedback signal capture: when the system detects that the user skips the playing advertisement material, immediately start the high-frequency behavior data acquisition mechanism. This mechanism mainly works through the cooperation of three sensing subsystems to capture the negative feedback signal. First, call the front camera to continuously obtain the user's face image within the time period of 0.5 seconds before the advertisement is skipped at a sampling frequency of 120 frames per second, extract the user's pupil focus position sequence in this time period, and analyze the shift direction of the user's visual attention. Secondly, the touch screen sensor synchronously records the finger movement trajectory generated in the user's skipping action, including the spatial coordinates of all contact points. At the same time, the system also activates the three-axis accelerometer to collect three-axis directional acceleration data during the skipping operation process and extract its peak information to represent the force characteristics of the user's action. These three types of data constitute a complete negative feedback behavior perception input, which provides support for subsequent feature modeling.

[0092] S42, feature matrix construction: structure the multi-source perception data obtained in S41 to construct a three-dimensional negative feedback feature matrix reflecting the user's skipping behavior intention. First, the system calculates the average distance from each position point in the user's pupil focus position sequence to the boundary coordinates of the advertisement display area in the interface, thereby quantifying whether the user is in a visual transfer state outside the advertisement area before skipping. Secondly, the path of the finger movement trajectory is geometrically modeled to extract the path curvature from the starting point of the trajectory to the advertisement close button, which is used to measure the operation accuracy and initiative of the skipping behavior. Finally, combined with the three-axis peak data collected by the accelerometer, the system aligns the above visual, tactile and motion signals in time sequence and fuses the features, finally outputs a negative feedback feature matrix containing three dimensions of indicators as a structured representation of the system's skipping behavior.

[0093] S43, classifier weight correction: the negative feedback feature matrix constructed in S42 is input into the pre-trained emotion classifier in S1 for reanalysis. By comparing the output vector of the input in the emotion classifier with the original emotion feature vector generated for the first time in the user's current browsing process, the similarity deviation between the two is calculated. The deviation is measured based on the angle of the vector space to identify the degree of deviation between the user's actual emotional state and the classifier prediction. When the deviation value exceeds the set error threshold , the back propagation mechanism of the model is automatically called to locally update the convolution kernel weight in the emotion classifier, thereby improving the recognition accuracy of the model for similar emotional states and enhancing its sensitivity and adaptability to negative behavior signals.

[0094] S44, dynamic threshold optimization: in order to further improve the response speed of the system to the changes in user interest, the system introduces a dynamic threshold adjustment mechanism. When it is continuously detected that the user skips more than three times of the same type of advertising materials, the system judges that there is a problem of insufficient adaptability in the current advertising screening strategy, and the key matching parameters need to be adjusted. Specifically, the threshold used in the matching degree model is adjusted for convergence, so that more materials close to the user's emotional state but slightly lower than the original threshold can enter the candidate advertising set. At the same time, the overlap threshold used in the length matching of the advertising materials is contracted to expand the coverage of the length adaptive advertising subset.

[0095] The threshold is updated to × 0.8, and the overlap threshold is updated to γ × 0.7.

[0096] The present application encompasses any substitutions, modifications, equivalent methods and solutions made on the essence and scope of the present application. In order for the public to have a thorough understanding of the present application, specific details are described in the following preferred embodiments of the present application, and the present application can also be fully understood without the description of these details to those skilled in the art. In addition, in order to avoid unnecessary confusion to the essence of the present application, well-known methods, processes, procedures, elements and circuits, etc. are not described in detail.

[0097] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. The AI-based marketing advertising precision delivery method is characterized by: The following steps are involved: S1: Collects multimodal sensor data from the user terminal, including facial micro-expression change rate, touch screen pressure value, and device rotation angular velocity, and outputs emotion feature vectors through a pre-trained emotion classifier; S2: Input the emotion feature vector generated by S1 into the matching model, calculate the emotion compatibility with each material in the advertising material library, and select materials with a compatibility greater than a threshold α to generate a candidate advertising set; S3: Based on the user's current attention decay curve, select ad creatives with matching display durations from the candidate ad set and deliver them before the user's attention decays to the critical point, including: S31: monitor the user's page dwell time and scrolling speed on the current page in real time, obtain time series data at a sampling frequency of 10 times per second, and simultaneously record the screen coordinate range of the advertisement display area; S32: Based on the page dwell time and scrolling speed collected in S31, the current attention decay curve is fitted using an exponential decay function; S33: Extract the optimal display duration of each ad creative from the candidate ad set generated in S2, and calculate its optimal display duration relative to the current attention decay curve at the critical attention threshold. The overlap of the corresponding time window, select overlap > overlap threshold The duration of ad creative generation is adapted to the ad subset; S34: If the current attention decay curve drops to the critical attention threshold When the ad is displayed, the ad creative with the highest priority is selected from the duration-adapted ad subset, and the ad creative is placed in the corresponding channel window according to the screen coordinate range recorded by S31; S4: When an ad is skipped, the user's pupil focus position and finger movement trajectory are recorded 0.5 seconds before the skip, and a negative feedback feature matrix is ​​generated to reversely update the emotion classifier weights of S1, including: S41: When a user skips an ad, the front-facing camera captures the user's pupil focus position sequence for the first 0.5 seconds of the skip at a sampling rate of 120 fps. The touchscreen sensor also records the coordinate points of the finger's movement trajectory, and simultaneously collects peak data from the three-axis accelerometer. S42: Calculate the average Euclidean distance between the user pupil focus position sequence obtained in S41 and the border of the advertising area, extract the path curvature from the starting point to the ad close button of the finger movement trajectory, and generate a three-dimensional negative feedback feature matrix based on the acceleration peak value; S43: Input the negative feedback feature matrix into the pre-trained emotion classifier of S1, calculate the cosine similarity deviation between its output vector and the original emotion feature vector, and when the deviation value is greater than the threshold When , the convolution kernel weights of the emotion classifier are updated through the back propagation algorithm; S44: Count the number of times N that users skip the same type of advertisements in a row. When N>3, the threshold of S2 is adaptively lowered. Overlap threshold with S33 .

2. The AI-based marketing advertising precision delivery method according to claim 1, characterized in that: The pre-trained emotion classifier includes a multimodal input layer, a temporal feature extraction module, a dynamic attention fusion layer and a classification output layer.

3. The AI-based marketing advertising precision delivery method according to claim 2, characterized in that: Said S1 comprises: S11: The front-facing camera of the user terminal captures facial video streams in real time and calculates the rate of change of facial micro-expressions between consecutive frames. Simultaneously, a pressure sensor acquires the touch screen pressure value of the screen contact area, and a gyroscope collects the device's angular velocity in three-dimensional space. Timestamps are aligned and noise filtered on these three types of raw data to generate standardized multimodal sensor data. S12: Inputting the standardized multimodal sensor data into a temporal feature extraction module of a pre-trained emotion classifier, wherein the facial micro-expression change rate is output as a visual temporal feature vector via a first temporal feature extraction channel, and the touch screen pressure value and the device rotation angular velocity are output as an inertial temporal feature vector via a second temporal feature extraction channel; S13: The visual temporal feature vector and the inertial temporal feature vector output by S12 are input into the dynamic attention fusion layer of the pre-trained emotion classifier, and a fusion feature matrix is ​​generated through cross-modal attention weight distribution. Finally, it is mapped into a 128-dimensional emotion feature vector through the classification output layer.

4. The AI-based marketing advertising precision delivery method according to claim 3, characterized in that: The matching model includes a dynamic weight generation module, a fitness calculation engine and a sorting filter.

5. The AI-based marketing advertising precision delivery method according to claim 4, characterized in that: The S2 includes: S21: For each advertising creative in the advertising creative library, extract its visual emotional features and audio emotional features, output the creative pleasure and creative arousal through a pre-trained creative emotional analysis model, and generate an emotional annotation matrix containing two-dimensional emotional indicators; S22: Input the emotion feature vector generated by S1 into the dynamic weight generation module of the matching model, output the weight coefficient W1 and the weight coefficient W2, and combine it with the emotion labeling matrix output by S21 to calculate the emotion adaptability of each material to form a dynamic adaptability set; S23: Filter the emotional fitness from the dynamic fitness set to be greater than the threshold The advertising materials are sorted in descending order of suitability to generate candidate advertising sets, where the threshold Dynamically adjust based on the user's historical advertising conversion rate.

6. The AI-based marketing advertising precision delivery method according to claim 5, characterized in that: The pre-trained material sentiment analysis model includes a visual feature extraction branch, an audio feature extraction branch and a bimodal attention fusion layer, wherein; The visual feature extraction branch converts the RGB frame of the advertising material into the HSV color gamut through the HSV conversion layer, and calculates the weighted product of the saturation and brightness of the main color in the color emotion mapping layer to generate visual emotion features; The audio feature extraction branch extracts 128-dimensional Mel-frequency cepstral coefficients in the 0-8kHz range through the Mel-frequency spectrum encoding layer, and generates a beat-emotion intensity curve through the beat emotion analysis layer; The bimodal attention fusion layer calculates the correlation matrix of visual and audio features, outputs the fused feature vector to the fully connected classification layer, outputs the material pleasantness and material arousal, and forms an emotion labeling matrix.

7. The AI-based marketing advertising precision delivery method according to claim 1, characterized in that: The exponential decay function is expressed as: ; in, Indicates that the user is at time The level of attention, is the initial attention value, is the attention decay coefficient, which reflects the activity level of the user's scrolling behavior. The duration that the user stays on the current page.

Citation Information

Patent Citations

  • Dynamic advertisement content intelligent delivery method based on user emotion recognition

    CN119887307A

  • Cognitive training emotion interaction method and system based on large model

    CN120148767A