A dynamic visual material intelligent putting management method and system for video advertisements
By separating video ad layers and combining them with real-time scene parameters for filtering and bitrate adaptation, the problem of low efficiency in video ad creative delivery is solved, achieving adaptive optimization and efficient delivery of the creative library.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 三明医学科技职业学院
- Filing Date
- 2026-06-22
- Publication Date
- 2026-07-21
AI Technical Summary
Existing methods for delivering video ad creatives cannot adapt to real-time scene changes that cater to individual users, resulting in image cropping or stretching distortions, visual abruptness, frequent stuttering or playback failures, and a lack of dynamic optimization capabilities based on positive responses, leading to low efficiency in delivering creative materials.
By separating the foreground, background, and overlay layers of video ads, combining real-time scene parameters for three-layer filtering, layered rendering and bitrate adaptation, constructing feedback feature records, and dynamically adjusting the material matching set, adaptive optimization of the material library is achieved.
It significantly improved the image quality and playback smoothness of ad frames, and achieved an overall improvement in the delivery efficiency of the creative library. Through self-optimization driven by real feedback, it ensures that the creatives are delivered in the optimal scenarios.
Smart Images

Figure CN122434604A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent marketing technology, and in particular to a method and system for intelligent delivery management of dynamic visual materials for video advertising. Background Technology
[0002] Current video ad creative delivery methods typically employ a fixed creative library matched with fixed ad placements. This means that ad creatives are pre-bound to ad placements of specific sizes, without dynamic filtering based on real-time terminal environment parameters during delivery. This approach often results in cropped or stretched images of the same creative at different viewport ratios, creating visual jarring when the main color tone of the content differs significantly from the ad creative's hue. Furthermore, insufficient bitrate can lead to frequent stuttering or playback failures due to the lack of bitrate prediction. Overall, the delivery effectiveness heavily relies on manually preset creative compatibility ranges and cannot adapt to the ever-changing real-time scenarios of diverse user experiences.
[0003] Current advertising delivery feedback mechanisms are mostly limited to macro-level metrics such as impressions and click-through rates. They fail to perform fine-grained correlation analysis between delivery scenario parameters, visual characteristics of creative materials, and user viewing completion, viewing behavior, and interaction behavior. Furthermore, they lack the closed-loop optimization capability to dynamically adjust the applicable scenario range of creative materials based on the positive response concentration trend. This results in the repeated delivery of the same creative material in low-response scenarios, leading to wasted traffic, while in high-response scenarios, delivery opportunities are missed due to overly strict constraints. The overall delivery efficiency of the creative material library remains suboptimal for a long time, failing to achieve the evolution from static matching to adaptive matching. Summary of the Invention
[0004] This invention provides a method and system for intelligent delivery management of dynamic visual materials for video advertising, the main purpose of which is to solve the problem of low efficiency in intelligent delivery management of dynamic visual materials for video advertising.
[0005] To achieve the above objectives, the present invention provides a method for intelligent delivery management of dynamic visual materials for video advertising, comprising: The basic visual attribute parameters of the target process are compiled into a material feature description set, and the delivery environment indicators of the advertising delivery request in the target process are combined into the real-time scene parameters of the target process. The candidate material group for the target process is selected based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set; The materials in the candidate material group are rendered layer by layer according to the image hierarchy to obtain the bitrate-adapted frame sequence of the target process; The current response level of the target process is determined based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and the delivery characteristics, visual fingerprints and current response level of the target process are organized into the feedback feature record of the target process. The applicable scenario range of the visual fingerprint is adjusted according to the positive response concentration trend recorded by the feedback features to obtain the dynamic material matching set of the target process; Using the dynamic material matching set as a scene constraint, the material feature description set is updated for delivery to obtain the delivery material set for the target process.
[0006] In a preferred embodiment, the step of aggregating the basic visual attribute parameters of the target process into a material feature description set, and combining the delivery environment indicators of the ad delivery request in the target process into real-time scene parameters of the target process, includes: Separate the foreground layer material, background layer material, and overlay layer material of the video advertisement in the target process to obtain the layer material set of the target process; The color distribution parameters, texture complexity parameters, and adaptation size ratio parameters of the layer material set are compiled into a material feature description set for the target process; The window size ratio, transmission bitrate level, and main color distribution of the content in the target process are combined into the real-time scene parameters of the target process.
[0007] In a preferred embodiment, the step of filtering the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set includes: The rectangular region of the material corresponding to the material adaptation ratio in the material feature description set is intersected with the rectangular region of the window size ratio in the real-time scene parameters to obtain the size-compliant material subset of the target process; The hue overlap matching of the main color distribution of the content in the real-time scene parameters with the color distribution parameters of the size-compliant material subset is performed to obtain the color-compliant material subset of the target process; The encoding bitrate converted from the texture complexity parameter of the hue-compliant material subset is compared with the bitrate tolerance range corresponding to the transmission bitrate level in the real-time scene parameters to determine the candidate material group for the target process.
[0008] In a preferred embodiment, the step of rendering the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate-adapted frame sequence of the target process includes: Identify the layer type corresponding to the material in the candidate material group, wherein the layer type includes background layer, foreground layer and overlay layer; A blank frame buffer is created based on the ad size of the target process, and the background layer material is written to the bottom layer of the blank frame buffer, the foreground layer material is written to the middle layer of the blank frame buffer, and the overlay layer material is written to the top layer of the blank frame buffer to obtain the composite ad frame of the target process; Based on the transmission bitrate level of the real-time scene parameters in the target process, the synthetic advertising frame is video encoded to obtain the bitrate-adapted frame sequence of the target process.
[0009] In a preferred embodiment, the determination of the current response level of the target process based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and the organization of the delivery features, visual fingerprint, and current response level of the target process into a feedback feature record of the target process, includes: The exposure duration field, completion marker field, and interaction marker field are parsed from the terminal feedback status report corresponding to the bitrate adaptation frame sequence to obtain the original delivery index group of the target process; The exposure integrity value of the target process is calculated based on the exposure duration field in the original index group and the total duration of the bitrate-adapted frame sequence. The value of the completion mark field in the original indicator group is used as the completion status value, the value of the interaction mark field is used as the interaction status value, and the exposure completeness value, the completion status value and the interaction status value are stacked in dimensions to obtain the current responsiveness of the target process. The visual feature description data corresponding to the candidate material group in the target process is taken as the visual fingerprint of the target process from the material feature description set in the target process. The real-time scene parameters of the target process, the visual fingerprint and the current response are grouped accordingly to obtain the feedback feature record of the target process.
[0010] In a preferred embodiment, the formula for calculating the current responsiveness includes: in, The current response level, The exposure integrity value is... The completion status value is... For the completion gain coefficient, This is the threshold sharpness coefficient. For complete exposure reference values, The interaction state value, For interaction gain coefficients, To amplify the index for reverse losses, The coefficient of catalytic linkage is denoted as α.
[0011] In a preferred embodiment, adjusting the applicable scenario range of the visual fingerprint based on the positive response concentration trend recorded by the feedback features to obtain the dynamic material matching set of the target process includes: Retrieve feedback feature records indexed by the visual fingerprint from the feedback feature record set to obtain a single fingerprint record set for the target process; Using the main color distribution of the content of the single fingerprint record set as the first dimension axis and the transmission bit rate level as the second dimension axis, a response distribution grid of the target process is constructed, and the response values of the single fingerprint record set are marked on the corresponding coordinate positions in the response distribution grid to obtain the fingerprint response map of the target process. Based on the spatial clustering pattern of the response values in the fingerprint response map, the suitable scenario domain for the target process is divided. The adapted scene domain is associated with the corresponding visual fingerprint to obtain the dynamic material matching set of the target process.
[0012] In a preferred embodiment, the step of dividing the target process into an adaptation scenario domain in the fingerprint response map based on the spatial clustering pattern of the response values includes: In the fingerprint response map, the coordinate points where the current response degree of the target process is greater than zero are taken as the activation points of the target process; The activation points of the free region in the fingerprint response map are taken as the starting point, and the adjacent activation points of the starting point are merged layer by layer to obtain the closed activation region of the target process. The main activation region of the target process is determined based on the number of activation points in the closed activation region, and the main color adaptation domain and transmission bit rate range of the target process are constructed based on the minimum and maximum coordinate values of the main activation region on the first and second dimension axes, respectively. The primary color adaptation domain and the transmission bit rate range are combined to form the applicable scenario range of the target process.
[0013] In a preferred embodiment, the step of updating the material feature description set using the dynamic material matching set as a scene constraint to obtain the delivery material set for the target process includes: Add a scene constraint field to the feature description data of the material feature description set, and initially set the scene constraint field to the full scene admission flag; Based on the visual fingerprint in the dynamic material matching set, the main color adaptation domain and transmission bitrate adaptation domain of the target process are overwritten into the scene constraint field to obtain the delivery material set of the target process.
[0014] To address the aforementioned problems, the present invention also provides a dynamic visual content intelligent delivery management system for video advertising, the system comprising: The data acquisition module collects the basic visual attribute parameters of the target process into a material feature description set, and combines the placement environment indicators of the advertising placement request in the target process into real-time scene parameters of the target process. The candidate material module filters the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set; The bitrate adaptation frame module renders the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate adaptation frame sequence of the target process. The feedback feature module determines the current response level of the target process based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and organizes the delivery features, visual fingerprints and the current response level of the target process into a feedback feature record of the target process. The dynamic material matching set module adjusts the applicable scenario range of the visual fingerprint based on the positive response concentration trend recorded by the feedback feature to obtain the dynamic material matching set of the target process; The material delivery module updates the material feature description set using the dynamic material matching set as a scene constraint to obtain the material delivery set for the target process.
[0015] Compared with the prior art, the present invention has the following beneficial effects:
[0016] 1. This method separates the foreground, background, and overlay layers and aggregates them into a material feature description set. It then performs a three-layer progressive screening based on real-time scene parameters such as the window size ratio, transmission bitrate level, and main color distribution of the content. Only candidate material groups with compliant size, overlapping colors, and encoding bitrates within the transmission tolerance range are retained. The ad frames are then rendered and synthesized layer by layer according to the image hierarchy and encoded according to the transmission bitrate level. This ensures that each ad frame achieves optimal performance in three dimensions: visual integrity, color coordination, and network adaptability, significantly improving the image quality and playback smoothness of a single delivery.
[0017] 2. This method further analyzes the exposure duration, completion marker, and interaction marker from the terminal's status report, calculates the exposure completeness value, completion status value, and interaction status value, and merges them into the current responsiveness. It uses visual fingerprints as index keys to aggregate historical feedback records to construct a fingerprint response map. Based on the spatial clustering pattern of activation points, it divides the adaptation scene domains and overwrites the scene constraint fields written into the material feature description set. This allows each material to automatically converge from full scene access to being deployed only within its best-performing main color adaptation domain and transmission bitrate range. This achieves continuous self-optimization of the deployment strategy based on real feedback, effectively improving the overall deployment efficiency of the material library. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating an embodiment of the intelligent delivery management method for dynamic visual materials in video advertising provided by the present invention.
[0019] Figure 2 A functional module diagram of a dynamic visual material intelligent delivery management system for video advertising provided in an embodiment of the present invention;
[0020] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0021] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0022] This application provides a method for intelligent delivery management of dynamic visual materials for video advertising. The executing entity of this method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the method for intelligent delivery management of dynamic visual materials for video advertising can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0023] Reference Figure 1 The diagram shown is a flowchart illustrating a method for intelligent delivery management of dynamic visual materials for video advertising according to an embodiment of the present invention. In this embodiment, the method includes: In this embodiment of the invention, the step of aggregating the basic visual attribute parameters of the target process into a material feature description set, and combining the placement environment indicators of the advertising placement request in the target process into the real-time scene parameters of the target process, is specifically used for: Separate the foreground layer material, background layer material, and overlay layer material of the video advertisement in the target process to obtain the layer material set of the target process; The color distribution parameters, texture complexity parameters, and adaptation size ratio parameters of the layer material set are compiled into a material feature description set for the target process; The window size ratio, transmission bitrate level, and main color distribution of the content in the target process are combined into the real-time scene parameters of the target process.
[0024] Specifically, the video advertisement file to be processed is obtained from the target process, and the video advertisement file is analyzed frame by frame. For each frame, a deep learning-based instance segmentation network is used to identify the pixel areas belonging to the foreground, background and overlay layers. The foreground layer material includes the area of the main product or person image displayed in the advertisement, the background layer material includes the environment or solid color area used as a background in the advertisement, and the overlay layer material includes the area of floating text, logo or special effects in the advertisement.
[0025] Specifically, for each layer image file in the layer material set, the color distribution parameters are first calculated by converting the image from the red-green-blue color space to the hue-saturation-brightness color space, then counting the frequency of each pixel value on the hue channel, and taking the most frequent hue values and their proportions as the main hue distribution vector. Next, the texture complexity parameters are calculated by applying the local binary mode operator after grayscale processing of the image.
[0026] Specifically, three fields are parsed from the advertising delivery request received during the target process. The first field is the window size ratio, which describes the ratio of the width to the height of the terminal device screen or player window used to display the advertisement. This ratio is directly extracted as the window size ratio. The second field is the transmission bitrate level, which describes the video data rate level that the advertising player can stably receive under the current network environment. The level is divided into three types: low, medium and high.
[0027] Furthermore, the pixel regions of each identified layer are extracted and saved as independent image files. Each image file retains its original resolution and color depth, while recording the position coordinates and layer order of each layer in the original frame. Finally, all the extracted foreground, background, and overlay layer image files are organized into a set according to the timestamp order. This set is the layer material set of the target process. This material set contains all independently operable visual elements in the video advertisement, providing basic data for subsequent feature extraction.
[0028] Furthermore, the operator iterates through each pixel and compares it with its neighboring pixels to generate a binary code. It then calculates the binary code values of all pixels to obtain a texture histogram and calculates the entropy value of the histogram as the texture complexity. The larger the entropy value, the more complex the texture. Finally, it extracts the adaptation size ratio parameters by reading the width and height pixel values of the image file, calculating the width-to-height ratio and retaining two decimal places, and arranging the calculated main color distribution vector, texture complexity value, and aspect ratio in the order of layer origin to form a structured record. The structured records of all layers are collected to obtain the material feature description set of the target process. This description set fully characterizes the visual attributes of each layer.
[0029] Furthermore, the bitrate identifier is directly extracted as the transmission bitrate bitrate. The third field is the content main color distribution, which describes the background main color of the page or application interface where the advertisement is to be placed. The parsing method is to sample the pixels in the four edge areas from the screenshot data attached to the request, calculate the mode of these pixels in the color channel as the main color value, and count the proportion of pixels occupied by the main color value. The main color value and its proportion are combined to form the content main color distribution. The above three fields are merged into a triplet data object in sequence. This object is the real-time scene parameter of the target process. This parameter reflects the terminal environment characteristics at the time of advertisement placement.
[0030] In summary, by breaking down video ads into independent layer assets, each layer can be processed, rendered in layers, and dynamically replaced in subsequent steps. This allows for personalized asset combinations for different delivery scenarios without re-encoding the entire video, significantly reducing the computational overhead of repeatedly transcoding the complete video file. At the same time, it preserves the original spatial relationships and hierarchical order between layers, providing a precise layer operation foundation for subsequent bitrate adaptation frame synthesis based on the image hierarchy.
[0031] In summary, quantifying the visual attributes of each layer of material into structured feature vectors enables computers to automatically compare, filter, and sort a large number of materials according to a unified numerical standard, avoiding inconsistencies caused by relying on subjective human judgment. At the same time, these feature parameters directly correspond to the main color distribution, transmission bitrate, and window size ratio in subsequent real-time scene parameters, providing a calculable data foundation for achieving scene-aware intelligent material matching.
[0032] In summary, by integrating key metrics of the terminal environment at the moment of ad delivery into a compact scene description object, the system can select materials that can be fully adapted without cropping based on the aspect ratio of the current window, select materials with texture complexity within the acceptable range based on the current network transmission bitrate, and select materials with overlapping hues based on the main color scheme of the current page or application. Thus, the system simultaneously meets the triple constraints of size adaptation, bitrate adaptation, and visual style adaptation during the material selection stage, avoiding problems such as image distortion, playback stuttering, or visual abruptness caused by material mismatch with the scene during playback.
[0033] In this embodiment of the invention, the step of filtering the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set is specifically used for: The rectangular region of the material corresponding to the material adaptation ratio in the material feature description set is intersected with the rectangular region of the window size ratio in the real-time scene parameters to obtain the size-compliant material subset of the target process; The hue overlap matching of the main color distribution of the content in the real-time scene parameters with the color distribution parameters of the size-compliant material subset is performed to obtain the color-compliant material subset of the target process; The encoding bitrate converted from the texture complexity parameter of the hue-compliant material subset is compared with the bitrate tolerance range corresponding to the transmission bitrate level in the real-time scene parameters to determine the candidate material group for the target process.
[0034] Specifically, the appropriate size ratio parameter corresponding to each material is extracted from the material feature description set. This parameter records the ratio of the original width to the height of the material. Based on this ratio, a rectangular area of the material with the same aspect ratio is constructed on a two-dimensional plane. At the same time, the window size ratio is extracted from the real-time scene parameters. This ratio records the ratio of the width to the height of the advertising display area. Based on this ratio, a rectangular area of the window with the same aspect ratio is constructed on the same two-dimensional plane. The rectangular area of the material and the rectangular area of the window are placed overlapping on the plane with their center points coinciding. The area of the overlapping part of the two rectangular areas is calculated.
[0035] Specifically, the content's primary color distribution is extracted from the real-time scene parameters. This distribution includes a primary color value and its pixel proportion. The color distribution parameters of each material are extracted one by one from the size-compliant material subset. These parameters record the primary color value of the material on the color channel and its proportion. The primary color value of the material is compared with the primary color value of the scene content, and the shortest arc distance between the two color values on the color wheel is calculated.
[0036] Specifically, a mapping table from texture complexity values to encoding bitrate is pre-established. This mapping table is obtained by fitting a large amount of experimental data. The higher the texture complexity, the more image details there are, and a higher encoding bitrate needs to be allocated under the same image quality requirements. Conversely, the lower the texture complexity, the lower the required encoding bitrate. The encoding bitrate value corresponding to each material is obtained by looking up the table. At the same time, the transmission bitrate level is extracted from the real-time scene parameters. This level corresponds to a specific bitrate tolerance range. The lower limit of this range is the minimum bitrate to ensure smooth playback, and the upper limit is the highest bitrate that the network can stably provide.
[0037] Furthermore, the areas of the material's rectangular region and the viewport's rectangular region are calculated separately. The ratios of the overlapping area to the area of the material's rectangular region and the overlapping area to the area of the viewport's rectangular region are taken. When both ratios are greater than zero, the material is determined to be able to fit the current viewport completely without pruning. When the ratio of the overlapping area to the area of the material's rectangular region is less than the complete fit threshold but greater than zero, the material is determined to need to be scaled proportionally before it can be fitted. When the overlapping area is zero, the material is determined to be unable to fit. The above intersection operation is performed on all materials in the material feature description set to filter out the fitable materials, i.e., all materials with an overlapping area greater than zero. The feature description data corresponding to these materials are organized into a new set, which is the size-compliant material subset of the target process.
[0038] Furthermore, the difference between the proportion of the main color tone of the material and the proportion of the main color tone of the scene is calculated. When the arc length distance on the color wheel is less than the preset proximity threshold and the proportion difference is less than the preset difference threshold, it is determined that the color tone of the material and the color tone of the scene are overlapping, that is, they belong to the same color system and have similar saturation. When the arc length distance exceeds the proximity threshold, it is determined that they are not overlapping. All materials in the size-compliant material subset are traversed and the above color overlap matching is performed. All materials with overlapping colors are selected and the feature description data corresponding to these materials are organized into a new set. This set is the color-compliant material subset of the target process.
[0039] Furthermore, the encoding bitrate of the material is compared with the bitrate tolerance range corresponding to the transmission bitrate level. When the encoding bitrate of the material falls within this range, including the lower limit or the upper limit, it is determined to be within the tolerance range, meaning that the material can be smoothly transmitted and decoded under the current network conditions. When the encoding bitrate is lower than the lower limit, it is determined that the image quality is excessively compressed and unacceptable. When the encoding bitrate is higher than the upper limit, it is determined that the network bandwidth is insufficient and will cause stuttering. All materials in the color-compliant material subset are traversed and the above tolerance range is performed to filter out all materials whose encoding bitrate is within the tolerance range. The feature description data corresponding to these materials are organized into a new set, which is the candidate material group of the target process.
[0040] In summary, by using the intersection operation of rectangular areas, materials that can be displayed completely without cropping or proportionally scaled within the current viewport size ratio are quantitatively selected. This avoids the image distortion caused by the forced stretching or cropping of materials due to mismatched aspect ratios. At the same time, using an intersection area greater than zero as the sole admission criterion ensures the visual integrity of all selected materials, eliminating the risk of incomplete or distorted display of advertising content caused by size incompatibility from the source. This provides a reliable set of materials that have been pre-verified in terms of size for subsequent layered rendering.
[0041] In summary, by calculating the shortest arc distance on the color wheel between the main color of the ad material and the main color of the scene content, as well as the difference in the proportion of the main color, ad materials with similar color scheme and saturation to the background color of the ad page or application are selected. This allows the synthesized ad frame to blend naturally with the surrounding environment without producing glaring color contrast, reducing the negative emotions of users caused by abrupt ad colors. At the same time, it improves the visual coordination between the ad and the content page, thereby indirectly increasing users' acceptance of the ad and their attention duration.
[0042] In summary, based on the positive correlation between texture complexity and encoding bitrate, materials with encoding bitrates falling within the current network's transmission bitrate tolerance range are pre-selected. This ensures that the ad frame sequence can be smoothly transmitted and decoded under given network conditions, avoiding playback stuttering or buffering delays caused by excessively fine textures leading to encoding bitrates exceeding the network limit, and also avoiding wasting available bandwidth due to excessively simple textures resulting in low bitrates. This achieves precise matching between material selection and network quality, providing a fundamental guarantee for a smooth playback experience.
[0043] In this embodiment of the invention, when rendering the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate-adapted frame sequence of the target process, the specific method is as follows: Identify the layer type corresponding to the material in the candidate material group, wherein the layer type includes background layer, foreground layer and overlay layer; A blank frame buffer is created based on the ad size of the target process, and the background layer material is written to the bottom layer of the blank frame buffer, the foreground layer material is written to the middle layer of the blank frame buffer, and the overlay layer material is written to the top layer of the blank frame buffer to obtain the composite ad frame of the target process; Based on the transmission bitrate level of the real-time scene parameters in the target process, the synthetic advertising frame is video encoded to obtain the bitrate-adapted frame sequence of the target process.
[0044] Specifically, each material is taken out one by one from the candidate material group, and the metadata recorded when the material was separated into layers is read. The metadata stores the source layer identifier of each material, which directly indicates whether the material belongs to the background layer, foreground layer or overlay layer. For materials that lack metadata tags, the type is determined by analyzing their image content.
[0045] Specifically, the window size ratio is extracted from the real-time scene parameters of the target process, and the width and height of the final ad frame are calculated by combining the target display width value carried in the ad delivery request. Based on these two pixel values, a completely transparent image buffer with all pixels initially set to zero is allocated in memory. This buffer is the blank frame buffer. Then, all materials carrying background layer tags in the candidate material group are scanned, and their pixel data is copied pixel by pixel to the corresponding position in the blank frame buffer according to the original size and original position coordinates of each background layer material to form the bottom layer image.
[0046] Specifically, the synthesized advertising frame of the target process is input into the video encoder as a single frame image. The transmission bitrate level is extracted from the real-time scene parameters. The level is divided into three levels: low, medium and high. Each level corresponds to a specific output bitrate value and a set of encoding parameters. When the transmission bitrate level is low, the encoder uses a higher compression intensity.
[0047] Furthermore, the determination method involves calculating the spatial distribution of edge density in the source image. Source images with uniform edge density distributed across the entire image and without obvious foreground object outlines are determined to be background layers. Source images with edge density concentrated in the central area of the image and having closed outlines are determined to be foreground layers. Source images with sparse edge density mainly distributed around the perimeter or corners of the image are determined to be overlay layers. After determining all source images, a layer type label is attached to each source image. The value of this label can only be one of background layer, foreground layer, or overlay layer. Thus, each source image in the candidate source image group is explicitly assigned a corresponding layer type.
[0048] Next, all materials carrying the foreground layer label are scanned, and their pixel data is written to the corresponding positions of the existing bottom layer image in the buffer. During the overwriting, the opaque pixels of the foreground layer completely replace the bottom layer pixels to form the middle layer image. Finally, all materials carrying the overlay layer label are scanned, and their pixel data is written to the top layer of the buffer. Overlay materials usually contain semi-transparent areas. When writing, a semi-transparent blending method is used, that is, each pixel of the overlay material is weighted and blended with the existing pixels in the buffer according to its transparency ratio. The blended pixels are stored as the final result. After all layers are written, the complete image in the buffer is the composite advertising frame of the target process. This composite advertising frame is a static image that contains background, foreground, and overlay elements with the correct layer relationship.
[0049] Furthermore, setting a larger quantization step size allows the transform coefficients to retain fewer high-frequency details after quantization. Simultaneously, a more efficient entropy coding mode is enabled to reduce redundant information. When the transmission bitrate is medium, the encoder uses moderate compression intensity, with the quantization step size between low and high, and entropy coding uses the standard mode. When the transmission bitrate is high, the encoder uses lower compression intensity, with the quantization step size set to the minimum to retain all image details, and some lossy compression tools are disabled. The encoder performs intra-frame compression on the synthesized advertising frame according to the set parameters, generating the corresponding encoded data packet. Since advertisements typically last several seconds, the same synthesized advertising frame needs to be repeatedly encoded into multiple consecutive frames, with each frame having identical encoded data. The encoded data of all consecutive frames are arranged chronologically into a sequence, which is the bitrate-adapted frame sequence for the target process. This sequence can be smoothly transmitted and decoded for playback at the specified transmission bitrate.
[0050] In summary, by clearly distinguishing the layer affiliation of each material, subsequent composite ad frames can be written in the correct spatial hierarchy order, avoiding logical confusion caused by incorrect layer order, such as foreground objects being obscured by the background or overlay elements being covered by the foreground. At the same time, it provides accurate layer identification for dynamically replacing a certain layer in the same ad image while keeping other layers unchanged, improving the reuse rate of materials and the efficiency of ad compositing.
[0051] In summary, the layered writing order from bottom to top ensures that the occlusion relationship between layers meets visual expectations. The background layer completely fills the bottom to form the basic image, the foreground layer covers the background to highlight the main content, and the overlay layer floats at the top to display text or special effects information. The compositing process can generate complete advertising images directly in memory without relying on any external graphics library, providing pixel-level accurate raw frame data for subsequent video encoding.
[0052] In summary, the compression intensity of video encoding is dynamically adjusted according to the current network transmission bitrate. At low bitrates, a high compression strategy is used to reduce the amount of data to adapt to narrowband environments, while at high bitrates, a low compression strategy is used to retain more image details to improve image quality. This allows the generated bitrate-adapted frame sequence to be transmitted smoothly without stuttering under various network conditions, while avoiding bandwidth waste due to excessively high bitrates or blurry images due to excessively low bitrates. This achieves the best balance between advertising image quality and transmission stability.
[0053] In this embodiment of the invention, when determining the current response level of the target process based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and organizing the delivery features, visual fingerprint, and current response level of the target process into a feedback feature record for the target process, it is specifically used for: The exposure duration field, completion marker field, and interaction marker field are parsed from the terminal feedback status report corresponding to the bitrate adaptation frame sequence to obtain the original delivery index group of the target process; The exposure integrity value of the target process is calculated based on the exposure duration field in the original index group and the total duration of the bitrate-adapted frame sequence. The value of the completion mark field in the original indicator group is used as the completion status value, the value of the interaction mark field is used as the interaction status value, and the exposure completeness value, the completion status value and the interaction status value are stacked in dimensions to obtain the current responsiveness of the target process. The visual feature description data corresponding to the candidate material group in the target process is taken as the visual fingerprint of the target process from the material feature description set in the target process. The real-time scene parameters of the target process, the visual fingerprint and the current response are grouped accordingly to obtain the feedback feature record of the target process.
[0054] Specifically, after the terminal device completes the playback of the bitrate-adapted frame sequence, the advertising player on the terminal generates a status report data packet according to the preset return protocol. This data packet is sent back to the server. After receiving the data packet, the server first verifies the integrity and source legitimacy of the data packet. After passing the verification, it uses a structured parsing method to read each field in the data packet one by one according to the field offset and length specified in the protocol.
[0055] Specifically, the value of the exposure duration field is extracted from the original target metrics group. This value represents the actual length of time the advertisement is viewed by the user. Then, the total duration of the bitrate-adapted frame sequence is calculated. The calculation method is to obtain the total number of frames contained in the bitrate-adapted frame sequence. Since the video encoding of the advertisement uses a fixed frame rate, the total number of frames is divided by the frame rate to obtain the total duration value in seconds. The exposure duration value is divided by the total duration value to obtain a ratio.
[0056] Specifically, the Boolean value of the completion mark field is directly extracted from the original index group of the delivery, and the Boolean value is converted into a numerical form. The conversion rule is that true is converted into a numerical value of one, and false is converted into a numerical value of zero. The converted numerical value is the completion status value. Similarly, the Boolean value of the interaction mark field is extracted from the original index group of the delivery, and converted into a numerical value of one or zero according to the same rule. The converted numerical value is the interaction status value. The first step of the fusion calculation is to multiply the completion gain coefficient by the completion status value to obtain the first product, subtract the full exposure reference value from the exposure integrity value to obtain the difference, multiply the difference by the threshold sharpness coefficient, square the result, add one to the result to obtain a denominator, and divide the first product by the denominator to obtain a completion contribution item.
[0057] Specifically, the implementation process of grouping the visual fingerprint and the current response rate to obtain the feedback feature record of the target process is as follows: extract the unique identifier of all materials from the candidate material group, and search for matching feature description records in the material feature description set according to these identifiers. Each feature description record contains the color distribution parameters, texture complexity parameters and adaptation size ratio parameters of the corresponding material. Concatenate these parameters into a long string according to the original order of the materials. This string is the visual fingerprint of the target process. This fingerprint uniquely identifies the visual attributes of the material combination used in this deployment.
[0058] Furthermore, the exposure duration field records the actual number of seconds from the start to the end of the ad frame sequence. During parsing, the value of this field is extracted and converted into a floating-point number. The completion flag field is a boolean flag. A true value indicates that the user watched the entire ad frame sequence until the last frame, while a false value indicates that the user exited prematurely during playback. During parsing, the original value of this flag is read directly. The interaction flag field is also a boolean flag. A true value indicates that the user had interactive behaviors such as clicking, swiping, or touching during ad playback, while a false value indicates no interaction. During parsing, the original value of this flag is also read directly. The parsed exposure duration value, completion flag boolean value, and interaction flag boolean value are stored in a temporary array in the parsing order. This array is the original indicator group for the target process.
[0059] Furthermore, when the exposure duration is greater than the total duration, the ratio is limited to an upper limit of one; when the exposure duration is less than zero, the ratio is limited to zero. In other cases, the calculated ratio is used directly. This ratio reflects the completeness of the user's viewing of the advertisement. The closer the ratio is to one, the more complete the user's viewing is. The closer the ratio is to zero, the more complete the user's viewing is. The closer the ratio is to zero, the more complete the user's viewing is. This calculated ratio is used as the output result, which is the exposure completeness value of the target process.
[0060] Further, the second step is to multiply the interaction gain coefficient by the interaction state value to obtain a second product. Subtract the exposure integrity value from one as the base, and perform a power operation with the inverse loss amplification index as the exponent to obtain the attenuation term. Then, add one to the result of multiplying the linkage catalytic coefficient by the full broadcast state value to obtain the catalytic term. Multiply the second product, the attenuation term, and the catalytic term to obtain an interaction contribution term. The third step is to add one to the full broadcast contribution term and the interaction contribution term, and then multiply by the exposure integrity value. The final calculated result is the current response of the target process.
[0061] Next, the real-time scene parameters of the target process are obtained. These parameters include three fields: window size ratio, transmission bitrate level, and content main color distribution. The current responsiveness, which was just calculated, is obtained. The responsiveness is a scalar value. The three objects, real-time scene parameters, visual fingerprint, and current responsiveness, are organized in the form of key-value pairs, where the keys are scene, fingerprint, and response, and the corresponding values are their respective complete data. These key-value pairs are encapsulated into a record object, which is the feedback feature record of the target process. This record completely saves the input conditions, material features, and user feedback results of an ad campaign.
[0062] In summary, by directly parsing three key fields from the status report returned by the terminal, real-time collection and structured storage of advertising performance data were achieved, avoiding data delays and accuracy losses caused by relying on third-party statistical platforms. At the same time, the exposure duration, completion mark, and interaction mark cover the three core dimensions of user viewing depth, completion rate, and interactive behavior, respectively, providing the original data foundation for subsequent calculation of exposure completeness value and current responsiveness, and ensuring the comprehensiveness and accuracy of feedback features.
[0063] In summary, by calculating the ratio of actual exposure time to total ad duration and implementing boundary limiting, continuous exposure time is converted into a normalized exposure completeness value. This eliminates the impact of differences in ad duration on the evaluation results, enabling the exposure completeness value to uniformly reflect the viewing progress of users on ads of any length. This provides a highly comparable input for subsequent calculations that integrate with completion status and interaction status values. At the same time, boundary limiting ensures that the calculation results are always within a reasonable range, avoiding interference from abnormal data on feedback evaluation.
[0064] In summary, by stacking three different dimensions of metrics—exposure completeness, completion status, and interaction status—into a multi-dimensional response vector, the system fully preserves multifaceted information about user responses, avoiding information loss caused by simple weighted summation. Furthermore, this vector structure allows for subsequent adjustment of the applicable scenario range of the visual fingerprint based on the central tendency of positive responses, providing fine-grained feedback for dynamically optimizing content delivery strategies. This enables the system to distinguish whether a user's high response is due to complete viewing or interactive behavior, thus achieving more accurate scenario adaptation.
[0065] In summary, by grouping real-time scene parameters, visual fingerprints, and current responsiveness into feedback feature records, a direct mapping relationship is established between the delivery environment, creative attributes, and user responses. Each record fully depicts the entire link information of an ad delivery, laying the foundation for subsequent retrieval of historical records with the same visual fingerprint from the feedback feature record set. At the same time, this record structure supports fast retrieval and aggregation analysis using visual fingerprints as index keys, enabling the system to learn the optimal applicable scene range for each creative combination based on a large amount of historical feedback data, achieving an upgrade from static delivery to dynamic adaptive delivery.
[0066] In this embodiment of the invention, the formula for calculating the current responsiveness is specifically used for: in, The current response level, The exposure integrity value is... The completion status value is... For the completion gain coefficient, This is the threshold sharpness coefficient. For complete exposure reference values, The interaction state value, For interaction gain coefficients, To amplify the index for reverse losses, The coefficient of catalytic linkage is denoted as α.
[0067] Specifically, the six parameters—completion gain coefficient, threshold sharpness coefficient, full exposure reference value, interaction gain coefficient, inverse loss amplification index, and linkage catalysis coefficient—are all obtained through regression analysis on a large amount of historical advertising data. The specific method is to use the exposure completeness value, completion status value, interaction status value, and the actual response level manually labeled or actually observed from each historical advertising record as training samples. The values of the six parameters are iteratively adjusted using a nonlinear least squares fitting method to minimize the sum of squared errors between the current response level calculated according to the formula and the actual response level. The parameter values that are fixed after convergence are stored in the system's configuration parameter library and are directly read from the configuration library each time the current response level is calculated.
[0068] Furthermore, the significance of the formula lies in integrating user feedback from three dimensions—exposure completeness, playback completion status, and interaction status—into a single current responsiveness value. The exposure completeness value acts as a base multiplier, directly amplifying the contributions of both playback completion and interaction. The playback completion status value achieves gain through a threshold function centered on a complete exposure reference value. When the exposure completeness value is exactly equal to the complete exposure reference value, the denominator of the threshold function reaches its minimum, thus peaking the playback contribution. As the exposure completeness value deviates from this reference value, the playback contribution gradually diminishes. The interaction status value achieves a reverse amplification effect through a decay factor that increases exponentially with decreasing exposure completeness. Simultaneously, the playback completion status value further amplifies the interaction contribution through a linkage catalytic coefficient. Ultimately, the current responsiveness calculated by this formula comprehensively reflects the overall response strength resulting from the synergistic effect of user viewing completeness, playback completion, and interaction.
[0069] In summary, the formula shows that when the exposure integrity value is close to the full exposure reference value and the playback completion status value is one, the playback completion contribution reaches its maximum value, making the current response significantly higher than the exposure integrity value itself. When the interaction status value is one, the current response gains additional gain due to the inverse loss amplification index when the exposure integrity value is low. This gain is further enhanced by the linkage catalytic coefficient when the playback completion status value is one. When both the playback completion status value and the interaction status value are zero, the formula degenerates into the current response equal to the exposure integrity value. When the exposure integrity value is zero, the current response is zero regardless of the playback completion and interaction status. When the exposure integrity value is one and the playback completion status value is also one, the numerator of the threshold function in the playback completion contribution term reaches its maximum value while the denominator tends to one, making the playback completion contribution term equal to the playback completion gain coefficient. At the same time, the attenuation factor in the interaction contribution term becomes zero as a result of subtracting the exposure integrity value from zero. At this point, the current response equals one plus the playback completion gain coefficient.
[0070] In this embodiment of the invention, when adjusting the applicable scenario range of the visual fingerprint based on the positive response central trend recorded by the feedback features to obtain the dynamic material matching set of the target process, it is specifically used for: Retrieve feedback feature records indexed by the visual fingerprint from the feedback feature record set to obtain a single fingerprint record set for the target process; Using the main color distribution of the content of the single fingerprint record set as the first dimension axis and the transmission bit rate level as the second dimension axis, a response distribution grid of the target process is constructed, and the response values of the single fingerprint record set are marked on the corresponding coordinate positions in the response distribution grid to obtain the fingerprint response map of the target process. Based on the spatial clustering pattern of the response values in the fingerprint response map, the suitable scenario domain for the target process is divided. The adapted scene domain is associated with the corresponding visual fingerprint to obtain the dynamic material matching set of the target process.
[0071] Specifically, the server maintains a global feedback feature record database, which stores all feedback feature records generated by all historical advertising campaigns. Each feedback feature record contains a visual fingerprint field, which is used as the primary index key of the database table. For the visual fingerprint of the current target process, the server performs an exact match retrieval operation.
[0072] Specifically, each record is retrieved from the single fingerprint record set. For each record, the real-time scene parameters it contains are parsed. The content main color distribution field and the transmission bitrate level field are extracted from the real-time scene parameters. The content main color distribution field is an integer between zero and 359, representing the angle value on the color ring, which is used as the coordinate value of the first dimension axis. The transmission bitrate level field is a discrete value of three levels: low, medium, or high, which are mapped to the values of zero, one, and two respectively as the coordinate values of the second dimension axis.
[0073] Specifically, the fingerprint response map is regarded as a two-dimensional plane, where the first dimension axis represents the distribution of the main color tone of the content, the second dimension axis represents the transmission bit rate level, and the response value at each coordinate position represents the average response intensity of the user to the visual fingerprint material under the scene combination. Connectivity analysis is performed on the fingerprint response map. The analysis method is to start from any coordinate point with a response value greater than zero and check the four adjacent coordinate points above, below, left and right of that point.
[0074] Specifically, the adaptation scene domain includes two parts: the content main color adaptation interval and the transmission bitrate adaptation level set. The content main color adaptation interval is a closed interval formed by the minimum coordinate value to the maximum coordinate value on the first dimension axis of the adaptation scene domain. The transmission bitrate adaptation level set is a set of tags for low, medium or high bitrates that are mapped back to all the bitrate values that have appeared on the second dimension axis of the adaptation scene domain. A key-value pair is constructed by using the visual fingerprint of the current target process as the key and the adaptation scene domain as the value.
[0075] Furthermore, the operation method uses the equality query statement in the structured query language to filter out all records in the feedback feature record table whose visual fingerprint field value is exactly the same as the visual fingerprint of the current target process. During the retrieval process, no sorting or filtering is performed on the records. All record rows in the query results are directly extracted. Each record row contains the real-time scene parameters, visual fingerprint and current response degree in the corresponding historical deployment. All extracted record rows are put into a new list container. All records in this container share the same visual fingerprint value. This list container is the single fingerprint record set of the target process.
[0076] Furthermore, the method for constructing the response distribution grid is to pre-create a two-dimensional array. The first dimension has a length of 360, covering all possible hue values, and the second dimension has a length of 3, covering three bitrate levels. All elements of this two-dimensional array are initially set to empty. Then, each record in the single fingerprint record set is traversed, and the exposure integrity value, playback status value, and interaction status value are extracted from the current response of the record. These three values are aggregated and calculated by multiplying the exposure integrity value by the sum of the playback status value and the interaction status value to obtain a single response value. This response value is a floating-point number between zero and one. The first-dimensional index is determined based on the main hue distribution value of the record's content, and the second-dimensional index is determined based on the transmission bitrate level mapping value. The calculated response value is filled into the corresponding position in the two-dimensional array. If there are multiple records at the same coordinate position, the arithmetic mean of all response values is taken and filled in. After traversal, each non-empty position in the two-dimensional array records a response value, and the entire two-dimensional array is the fingerprint response map of the target process.
[0077] Furthermore, if the response value of an adjacent coordinate point is also greater than zero and the difference between the response value of the current point and the current point is less than a preset difference tolerance value, then the adjacent point is assigned to the same region. Then, starting from the newly assigned point, the process continues to expand outwards until no new points can be found to be assigned. This process will form a closed connected region. The process is repeated for all points in the entire graph that have not yet been assigned and whose response values are greater than zero. Finally, several independent connected regions are obtained. For each connected region, the average response value of all coordinate points in the region is calculated. The connected region with the largest average value is taken as the main response region. The union of the content main color value range and the transmission bitrate value range corresponding to all coordinate points covered by the main response region is extracted. This union is the adaptation scene domain of the target process.
[0078] Furthermore, the key-value pair indicates that the material corresponding to this set of visual features can obtain a high user response in the delivery environment described by the adaptation scene domain. The server stores the key-value pair in a dedicated dynamic material matching library. If the same visual fingerprint already exists in the library, the original record is overwritten with the new adaptation scene domain. Finally, all key-value pairs associated with the visual fingerprint of the current target process are retrieved from the dynamic material matching library and organized into a set. Each item in the set is a pairing of a visual fingerprint with its adaptation scene domain. This set is the dynamic material matching set of the target process.
[0079] In summary, by using visual fingerprints as precise retrieval keys to filter all historical delivery records belonging to the same material combination from the global feedback feature record set, subsequent analysis is only performed on the feedback data of that specific visual fingerprint, avoiding interference between feedback data of different materials. At the same time, the single fingerprint record set gathers the multiple delivery results of the material combination in different scenarios, providing sufficient sample data for statistical analysis of the changes in its response value with the main color tone of the content and the transmission bitrate level, realizing a dimensional upgrade from single delivery feedback to multi-scenario statistical patterns.
[0080] In summary, by transforming discrete and disordered single fingerprint record sets into fingerprint response maps with a clear two-dimensional coordinate structure, each combination of content main color value and transmission bitrate level corresponds to an aggregated response value. This organizes the originally messy delivery data into a visualized response intensity distribution map. This map intuitively shows the performance differences of the same visual fingerprint material in different scene combinations, providing a structured data foundation for subsequent division of adaptation scene domains based on spatial aggregation patterns.
[0081] In summary, by analyzing the connected regions and clustering trends of response values in the fingerprint response spectrum on a two-dimensional plane, the optimal scene range for the visual fingerprint material is automatically identified, rather than relying on manually set fixed thresholds. At the same time, the analysis of spatial clustering patterns can eliminate isolated noise points and retain only statistically significant high-response areas. This ensures that the defined adaptation scene domain accurately reflects the applicable boundaries of the material in the real-world deployment environment, avoiding erroneous expansion or contraction caused by individual abnormal deployment data.
[0082] In summary, each visual fingerprint is bound to its historically verified adaptation scene domain as an entry, forming a mapping relationship from the visual features of the material to the optimal delivery scene conditions. This allows subsequent updates to the material feature description set to directly write the primary color adaptation domain and transmission bitrate adaptation domain into the scene constraint field based on this mapping. This achieves a transition from static full-scene access to dynamic scene constraint updates. At the same time, the dynamic material matching set is continuously optimized as new feedback records are added, giving the system adaptive learning capabilities.
[0083] In this embodiment of the invention, when dividing the adaptation scenario domain of the target process in the fingerprint response map based on the spatial clustering pattern of the response values, it is specifically used for: In the fingerprint response map, the coordinate points where the current response degree of the target process is greater than zero are taken as the activation points of the target process; The activation points of the free region in the fingerprint response map are taken as the starting point, and the adjacent activation points of the starting point are merged layer by layer to obtain the closed activation region of the target process. The main activation region of the target process is determined based on the number of activation points in the closed activation region, and the main color adaptation domain and transmission bit rate range of the target process are constructed based on the minimum and maximum coordinate values of the main activation region on the first and second dimension axes, respectively. The primary color adaptation domain and the transmission bit rate range are combined to form the applicable scenario range of the target process.
[0084] Specifically, each coordinate position is traversed from the two-dimensional array of the fingerprint response map. Each coordinate position corresponds to a content main color value and transmission bitrate level. The response value stored at that position is read. This response value is a single value calculated by aggregating the exposure completeness value, completion status value and interaction status value in the historical delivery. When this value is greater than zero, it indicates that the user has a positive response to the visual fingerprint material under the scene combination.
[0085] Specifically, an unvisited activation point is selected from the set of activation points as the starting point. Using this starting point as the center, its adjacent coordinates in the four directions of up, down, left, and right in the fingerprint response map are checked. If there are also activation points in the adjacent positions and the activation point has not been merged into any region, the adjacent activation point is merged into the same region as the starting point. Then, using the newly merged activation point as the new center, the adjacent activation points in the four directions are checked again. This expansion process is repeated until no unmerged adjacent activation points can be found. At this time, all merged activation points constitute a connected region, which is considered a closed activation region.
[0086] Specifically, the total number of activation points contained in each closed activation region is counted sequentially. The closed activation region with the most activation points is determined as the main activation region. When two or more regions contain the same maximum number of activation points, the region with the largest sum of response values of all activation points in these regions is selected as the main activation region. After determining the main activation region, all activation points in the main activation region are traversed, and the content color coordinate value of each activation point on the first dimension axis is recorded. The minimum value is found as the lower boundary of the main color, and the maximum value is found as the upper boundary of the main color. The lower boundary and the upper boundary together form a closed interval, which is the main color adaptation domain.
[0087] Specifically, the primary color adaptation domain is used as a value range constraint, which specifies the color range in which the visual fingerprint material can obtain a good response in the distribution of the primary color of the content. The transmission bitrate range is used as another value range constraint, which specifies the bitrate range in which the visual fingerprint material can obtain a good response in the transmission bitrate range. These two constraints are combined into a tuple.
[0088] Furthermore, when the value is equal to zero, it means that no delivery has ever been carried out or there is no positive feedback after delivery. For each coordinate position with a response value greater than zero, the coordinate position is marked as an activation point. The recording format of the activation point is a combination of the integer value of the main color tone of the content on the first dimension axis and the value of the transmission bitrate level on the second dimension axis. All the marked activation points together constitute an activation point set, and each point in the set represents a valid delivery scenario that has been verified in history.
[0089] Furthermore, if there are no adjacent active points in any of the four directions of the starting point, then the starting point itself constitutes a closed active region containing only one point. After merging a region, all merged active points are removed from the set of active points, and the next unvisited active point is selected to repeat the above process until the set of active points is empty, ultimately resulting in several closed active regions that are not connected to each other.
[0090] Furthermore, all activation points within the main activation area are traversed, and the transmission rate tier value of each activation point on the second dimension axis is recorded. Since the transmission rate tier has only three discrete values, the minimum and maximum values of all tier values that have appeared are used as the interval endpoints to construct a discrete transmission rate interval, which contains all integer tier values from the minimum tier to the maximum tier.
[0091] Furthermore, the first element of this tuple is the primary color adaptation domain, and the second element is the transmission bitrate range. This tuple defines a rectangular area that exactly covers the minimum bounding range of the main active area in the fingerprint response map. Any set of coordinates within this rectangular area represents a combination of a content primary color value and a transmission bitrate level. These combinations are the applicable deployment scenario conditions for the visual fingerprint material. Taking this tuple as the output result, the result is the applicable scenario range of the target process.
[0092] In summary, by defining coordinate points with a response greater than zero as activation points, invalid scenarios that have never been deployed or have no positive feedback after deployment are filtered out from the fingerprint response map. This allows subsequent analysis to focus only on deployment locations that have historically generated user responses, reducing the interference of noisy data on region division. At the same time, the set of activation points completely preserves the response intensity information of the visual fingerprint material in all valid scenarios, providing accurate spatial seed points for merging and closing activation regions layer by layer.
[0093] In summary, the four-neighbor connectivity merging method aggregates spatially adjacent activation points with similar response intensities into closed activation regions, realizing the transformation from discrete point sets to continuous regions. This avoids misjudgment of regions due to abnormal responses at single points. At the same time, the layer-by-layer merging ensures that each activation point belongs to only one closed region, making the region partitioning results unique and deterministic. This provides a structurally clear set of regions for subsequent identification of the main activation region from multiple closed regions.
[0094] In summary, by selecting the closed region with the most activation points as the main activation area, the extracted adaptation scene intervals are ensured to have the highest historical deployment density and statistical reliability, avoiding interval offset caused by a small number of high-response points or isolated areas. At the same time, the main color adaptation domain and transmission bitrate interval are constructed based on the minimum and maximum coordinate values of the main activation area on the first and second dimension axes, so that the generated intervals can accurately cover the minimum outer range of the main response area, without omitting effective scenes or over-extending to low-response areas.
[0095] In summary, by combining the primary color adaptation domain and the transmission bitrate range into a binary form of applicable scenario range, the applicable scenario range of the visual fingerprint material in terms of both the primary color of the content and the transmission bitrate level is fully defined. This allows the scenario constraint field written into the material feature description set to be directly replaced with this range, realizing the adaptive transformation from historical feedback data to material delivery rules. At the same time, the applicable scenario range is continuously recalculated as new feedback records are added, giving the applicable boundary of the material the ability to evolve dynamically.
[0096] In this embodiment of the invention, when updating the material feature description set using the dynamic material matching set as a scene constraint to obtain the delivery material set for the target process, it is specifically used for: Add a scene constraint field to the feature description data of the material feature description set, and initially set the scene constraint field to the full scene admission flag; Based on the visual fingerprint in the dynamic material matching set, the main color adaptation domain and transmission bitrate adaptation domain of the target process are overwritten into the scene constraint field to obtain the delivery material set of the target process.
[0097] Specifically, the material feature description set is read from the database. Each feature description data in the set corresponds to an independent visual material, which includes the material's color distribution parameters, texture complexity parameters, and adaptation size ratio parameters. Each feature description data in the set is traversed, and an attribute named "Scene Constraint Field" is added to the end of each data. This attribute is designed as a container type that can store structured data and can hold a primary color adaptation field and a transmission bitrate adaptation field. The field is initialized immediately after it is added.
[0098] Specifically, each entry is extracted one by one from the dynamic material matching set. Each entry contains a visual fingerprint and an adaptation scene domain corresponding to the visual fingerprint. The adaptation scene domain already contains two components: the main color adaptation domain and the transmission bit rate adaptation domain. Feature description data matching the current visual fingerprint is retrieved from the material feature description set.
[0099] Furthermore, the specific initialization process involves setting the primary color adaptation field in the scene constraint field to a complete closed color range from zero to 359, and setting the transmission bitrate adaptation field to a complete set including all three levels: low, medium, and high. This initialization state indicates that the material can be deployed in any scene with any primary color and any transmission bitrate when there are no historical feedback records. This initialization state is the full-scene admission flag. After completing the above operations, each feature description data in the material feature description set is accompanied by a scene constraint field that is initially set as the full-scene admission flag.
[0100] Furthermore, the matching method involves comparing whether the string formed by concatenating all visual attribute parameters in the feature description data is completely identical to the visual fingerprint. After finding matching feature description data, the scene constraint field currently attached to the feature description data is read. The original main color adaptation field in the scene constraint field is replaced with the main color adaptation field in the scene adaptation field, and the original transmission bitrate adaptation field in the scene constraint field is replaced with the transmission bitrate adaptation field in the scene adaptation field. The replacement operation is an overwrite operation, that is, the original value is discarded and the new value is directly stored in the field. When all visual fingerprints in the dynamic material matching set have completed the above overwrite operation, each feature description data in the material feature description set has a finally determined scene constraint field. This field precisely limits the scope of the applicable deployment scene for the material. The entire material feature description set is then output, and the output result is the deployment material set for the target process.
[0101] In summary, by attaching a writable scene constraint field to the feature description data of each material and initializing it as a full-scene admission flag, all materials are allowed to be deployed in any content color tone and any transmission bitrate in the absence of historical feedback data. This avoids the problem of new materials not being able to get deployment opportunities due to overly strict initial constraints. At the same time, a standardized field interface is reserved for subsequent overwriting based on dynamic material matching sets, realizing a seamless switch from open to constrained state for material deployment.
[0102] In summary, by directly overwriting the scene constraint fields of the matching material with the main color adaptation field and transmission bitrate adaptation field that have been verified by historical feedback in the dynamic material matching set, the material automatically converges from the default full-scene admission to being delivered only in the scene range where it performs best. This achieves adaptive constraint updates based on feedback learning. At the same time, the overwrite operation directly modifies the material feature description set to form the final set of delivery materials. This allows subsequent ad delivery to quickly determine the availability of materials in the current scene based on the scene constraint fields, avoiding the need to repeatedly execute the candidate material selection process for each delivery and significantly improving the efficiency of delivery decision-making.
[0103] Compared with the prior art, the present invention has the following beneficial effects:
[0104] 1. This method separates the foreground, background, and overlay layers and aggregates them into a material feature description set. It then performs a three-layer progressive screening based on real-time scene parameters such as the window size ratio, transmission bitrate level, and main color distribution of the content. Only candidate material groups with compliant size, overlapping colors, and encoding bitrates within the transmission tolerance range are retained. The ad frames are then rendered and synthesized layer by layer according to the image hierarchy and encoded according to the transmission bitrate level. This ensures that each ad frame achieves optimal performance in three dimensions: visual integrity, color coordination, and network adaptability, significantly improving the image quality and playback smoothness of a single delivery.
[0105] 2. This method further analyzes the exposure duration, completion marker, and interaction marker from the terminal's status report, calculates the exposure completeness value, completion status value, and interaction status value, and merges them into the current responsiveness. It uses visual fingerprints as index keys to aggregate historical feedback records to construct a fingerprint response map. Based on the spatial clustering pattern of activation points, it divides the adaptation scene domains and overwrites the scene constraint fields written into the material feature description set. This allows each material to automatically converge from full scene access to being deployed only within its best-performing main color adaptation domain and transmission bitrate range. This achieves continuous self-optimization of the deployment strategy based on real feedback, effectively improving the overall deployment efficiency of the material library.
[0106] like Figure 2 The diagram shown is a functional block diagram of a dynamic visual material intelligent delivery management system for video advertising provided in an embodiment of the present invention.
[0107] The intelligent dynamic visual content delivery management system 100 for video advertising described in this invention can be installed in an electronic device. Depending on the functions implemented, the intelligent dynamic visual content delivery management system 100 for video advertising may include a data acquisition module 101, a candidate content module 102, a bitrate adaptation frame module 103, a feedback feature module 104, a dynamic content matching set module 105, and a content delivery module 106. The module described in this invention can also be referred to as a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and are stored in the memory of the electronic device.
[0108] In this embodiment, the functions of each module / unit are as follows: The data acquisition module collects the basic visual attribute parameters of the target process into a material feature description set, and combines the placement environment indicators of the advertising placement request in the target process into real-time scene parameters of the target process. The candidate material module filters the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set; The bitrate adaptation frame module renders the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate adaptation frame sequence of the target process. The feedback feature module determines the current response level of the target process based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and organizes the delivery features, visual fingerprints and the current response level of the target process into a feedback feature record of the target process. The dynamic material matching set module adjusts the applicable scenario range of the visual fingerprint based on the positive response concentration trend recorded by the feedback feature to obtain the dynamic material matching set of the target process; The material delivery module updates the material feature description set using the dynamic material matching set as a scene constraint to obtain the material delivery set for the target process.
[0109] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0110] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0111] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0112] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0113] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for intelligent delivery management of dynamic visual creatives for video advertising, characterized in that, The method includes: The basic visual attribute parameters of the target process are compiled into a material feature description set, and the delivery environment indicators of the advertising delivery request in the target process are combined into the real-time scene parameters of the target process. The candidate material group for the target process is selected based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set; The materials in the candidate material group are rendered layer by layer according to the image hierarchy to obtain the bitrate-adapted frame sequence of the target process; The current response level of the target process is determined based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and the delivery characteristics, visual fingerprints and current response level of the target process are organized into the feedback feature record of the target process. The applicable scenario range of the visual fingerprint is adjusted according to the positive response concentration trend recorded by the feedback features to obtain the dynamic material matching set of the target process; Using the dynamic material matching set as a scene constraint, the material feature description set is updated for delivery to obtain the delivery material set for the target process.
2. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The process of compiling the basic visual attribute parameters of the target process into a material feature description set and combining the delivery environment indicators of the ad delivery request in the target process into real-time scene parameters of the target process includes: Separate the foreground layer material, background layer material, and overlay layer material of the video advertisement in the target process to obtain the layer material set of the target process; The color distribution parameters, texture complexity parameters, and adaptation size ratio parameters of the layer material set are compiled into a material feature description set for the target process; The window size ratio, transmission bitrate level, and main color distribution of the content in the target process are combined into the real-time scene parameters of the target process.
3. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The step of filtering the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set includes: The rectangular region of the material corresponding to the material adaptation ratio in the material feature description set is intersected with the rectangular region of the window size ratio in the real-time scene parameters to obtain the size-compliant material subset of the target process; The hue overlap matching of the main color distribution of the content in the real-time scene parameters with the color distribution parameters of the size-compliant material subset is performed to obtain the color-compliant material subset of the target process; The encoding bitrate converted from the texture complexity parameter of the hue-compliant material subset is compared with the bitrate tolerance range corresponding to the transmission bitrate level in the real-time scene parameters to determine the candidate material group for the target process.
4. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The step of rendering the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate-adapted frame sequence of the target process includes: Identify the layer type corresponding to the material in the candidate material group, wherein the layer type includes background layer, foreground layer and overlay layer; A blank frame buffer is created based on the ad size of the target process, and the background layer material is written to the bottom layer of the blank frame buffer, the foreground layer material is written to the middle layer of the blank frame buffer, and the overlay layer material is written to the top layer of the blank frame buffer to obtain the composite ad frame of the target process; Based on the transmission bitrate level of the real-time scene parameters in the target process, the synthetic advertising frame is video encoded to obtain the bitrate-adapted frame sequence of the target process.
5. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The current response level of the target process is determined by the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and the delivery features, visual fingerprint, and current response level of the target process are organized into a feedback feature record of the target process, including: The exposure duration field, completion marker field, and interaction marker field are parsed from the terminal feedback status report corresponding to the bitrate adaptation frame sequence to obtain the original delivery index group of the target process; The exposure integrity value of the target process is calculated based on the exposure duration field in the original index group and the total duration of the bitrate-adapted frame sequence. The value of the completion mark field in the original indicator group is used as the completion status value, the value of the interaction mark field is used as the interaction status value, and the exposure completeness value, the completion status value and the interaction status value are stacked in dimensions to obtain the current responsiveness of the target process. The visual feature description data corresponding to the candidate material group in the target process is taken as the visual fingerprint of the target process from the material feature description set in the target process. The real-time scene parameters of the target process, the visual fingerprint and the current response are grouped accordingly to obtain the feedback feature record of the target process.
6. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 5, characterized in that, The formula for calculating the current responsiveness includes: in, The current response level, The exposure integrity value is... The completion status value is... For the completion gain coefficient, This is the threshold sharpness coefficient. For complete exposure reference values, The interaction state value, For interaction gain coefficients, To amplify the index for reverse losses, The coefficient of catalytic linkage is denoted as α.
7. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The step of adjusting the applicable scenario range of the visual fingerprint based on the positive response central trend recorded by the feedback features to obtain the dynamic material matching set of the target process includes: Retrieve feedback feature records indexed by the visual fingerprint from the feedback feature record set to obtain a single fingerprint record set for the target process; Using the main color distribution of the content of the single fingerprint record set as the first dimension axis and the transmission bit rate level as the second dimension axis, a response distribution grid of the target process is constructed, and the response values of the single fingerprint record set are marked on the corresponding coordinate positions in the response distribution grid to obtain the fingerprint response map of the target process. Based on the spatial clustering pattern of the response values in the fingerprint response map, the suitable scenario domain for the target process is divided. The adapted scene domain is associated with the corresponding visual fingerprint to obtain the dynamic material matching set of the target process.
8. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 7, characterized in that, The step of dividing the target process into an adaptation scenario domain based on the spatial clustering pattern of the response values in the fingerprint response map includes: In the fingerprint response map, the coordinate points where the current response degree of the target process is greater than zero are taken as the activation points of the target process; The activation points of the free region in the fingerprint response map are taken as the starting point, and the adjacent activation points of the starting point are merged layer by layer to obtain the closed activation region of the target process. The main activation region of the target process is determined based on the number of activation points in the closed activation region, and the main color adaptation domain and transmission bit rate range of the target process are constructed based on the minimum and maximum coordinate values of the main activation region on the first and second dimension axes, respectively. The primary color adaptation domain and the transmission bit rate range are combined to form the applicable scenario range of the target process.
9. The intelligent delivery management method for dynamic visual materials for video advertising as described in claim 1, characterized in that, The step of updating the material feature description set using the dynamic material matching set as a scene constraint to obtain the material set for the target process includes: Add a scene constraint field to the feature description data of the material feature description set, and initially set the scene constraint field to the full scene admission flag; Based on the visual fingerprint in the dynamic material matching set, the main color adaptation domain and transmission bitrate adaptation domain of the target process are overwritten into the scene constraint field to obtain the delivery material set of the target process.
10. A dynamic visual material intelligent delivery management system for video advertising, used to implement the dynamic visual material intelligent delivery management method for video advertising as described in any one of claims 1-9, characterized in that, The system includes: The data acquisition module collects the basic visual attribute parameters of the target process into a material feature description set, and combines the placement environment indicators of the advertising placement request in the target process into real-time scene parameters of the target process. The candidate material module filters the candidate material group for the target process based on the comparison results between the real-time scene parameters and the visual feature description data in the material feature description set; The bitrate adaptation frame module renders the materials in the candidate material group layer by layer according to the image hierarchy to obtain the bitrate adaptation frame sequence of the target process. The feedback feature module determines the current response level of the target process based on the index distribution of the backpropagation state data packets corresponding to the bitrate adaptation frame sequence, and organizes the delivery features, visual fingerprints and the current response level of the target process into a feedback feature record of the target process. The dynamic material matching set module adjusts the applicable scenario range of the visual fingerprint based on the positive response concentration trend recorded by the feedback feature to obtain the dynamic material matching set of the target process; The material delivery module updates the material feature description set using the dynamic material matching set as a scene constraint to obtain the material delivery set for the target process.