An advertisement replacement method in a complex scenario

Through the two-stage neural network model STM-EOM, the problem of ads attached on curved surfaces in complex scenarios is solved, and the flexible alignment and natural bonding of advertisements on curved surfaces is achieved, and the authenticity and naturalness of the generated images are significantly improved.

CN115063511BActive Publication Date: 2025-07-04HARBIN INST OF TECH SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210752513.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-29
Publication Date
2025-07-04
Estimated Expiration
2042-06-29

AI Technical Summary

Technical Problem

The prior art is difficult to achieve flat ads in complex scenarios, especially to implant advertisements on curved surfaces with uneven arcs, and cannot effectively deal with changes in light and shadow and uneven veneers.

Method used

The two-stage neural network model STM-EOM is adopted. The spatial transformation module STM is first used to perform spatial transformation of advertising images, and then repair and optimize through the edge optimization module EOM to achieve alignment and natural joining of advertising on the curved surface.

Benefits of technology

It realizes flexible alignment and natural attachment of advertisements on curved veneers, significantly improving the authenticity and naturalness of the generated images, and can handle advertising replacement tasks in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115063511B_ABST
    Figure CN115063511B_ABST
Patent Text Reader

Abstract

The present invention provides an advertisement replacement method in a complex scenario, comprising the following steps: Step 1: Initially construct a data set for the advertisement replacement task; Step 2: Input the foreground advertisement image I 𝑓𝑜𝑟𝑒 , the mask image I 𝑚𝑎𝑠𝑘 into the spatial transformation module STM to obtain the transformed advertisement image I stn . Then, combine the background image I 𝑏𝑎𝑐𝑘 with the transformed advertisement image I stn to obtain the intermediate generated image I 𝑚𝑖d ; Step 3: Input the mask image I 𝑚𝑎𝑠𝑘 and the intermediate generated image I 𝑚𝑖d into the edge optimization module EOM, and the edge optimization module EOM processes them to generate the final result image I 𝑟𝑒𝑠 . The beneficial effects of the present invention are as follows: In the STM stage of the model, by using the STN network structure, the model has flexible spatial deformation performance, thus achieving the alignment of the foreground advertisement in terms of position, shape, and direction on the curved surface of the target arc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and in particular to an advertisement replacement method in complex scenes. Background Art

[0002] In recent years, the post-production advertising placement technology in videos has been increasingly used in commercial applications. Generally, artificial intelligence (AI) technology + image processing methods are used to place advertisements into existing movies, TV series, music videos and other video media. AI will first analyze the location in the video where the advertisement needs to be placed, then automatically model and render it, and also take into account factors such as the movement relationship of the characters in the picture, the lighting effect, and other influencing factors.

[0003] There are many companies that provide post-production advertising placement services. For example, Yingpu Technology has taken the lead in implementing AI technology for "content placement", aiming to increase the number of ad slots available for sale in video content. Yingpu Technology's intelligent video technology consists of three parts: intelligent computing technology, floating layer construction technology, and real-time placement technology. Its "Easy Placement" first searches for native advertising placement points in the video through intelligent scanning technology, and then places the brand image in scenes such as buildings, computer screens, and indoor posters in the play in various forms. For example, iQiyi's self-developed post-production advertising placement technology uses algorithms to analyze and understand the video, finds suitable placement points through algorithms or manual work, and then uses algorithms and some tools to place product information on the video screen. For example, Jilian Technology uses dynamic video training algorithms to analyze videos and dynamically analyze videos. Interactive advertising is placed through "ASMP". First, the video is structured through exclusive information processing technology (VideoAI), the scenes in the video are automatically scanned, and the placement points for interactive advertising in the video are searched. Then, with the help of the advertising creation program (VideoOS), interactive advertisements such as bubble dialogues and in-video voting are automatically placed.

[0004] Most of the existing post-production advertising implantation technologies use AI algorithms in video content understanding, scene analysis, advertising position detection and tracking. After detecting the advertising placement position, they directly replace the regular rectangular target advertisement through image processing, modeling and rendering methods. They may also consider the movement relationship of the characters in the picture, lighting effects and other influencing factors, but they do not further consider more complex situations such as the unevenness of the advertising surface.

[0005] In actual scenarios, although most advertisements have regular shapes, due to uneven attachment surfaces (such as advertisements attached to a cylinder), lens distortion, etc., in many cases, the advertisements will appear more complex in the image.

[0006] Existing advertisement insertion technologies can only operate on flat surfaces. For example, advertisements can only be attached to a flat wall surface, but for curved surfaces with arcs, such as the surface of a cylinder, automatic advertisement implantation cannot be carried out.

[0007] Although the current technologies have done well in scene analysis and detection, and relevant implementations for light and shadow changes also exist, there are no relevant methods for attaching advertisements in complex scenes such as curved surfaces. Summary of the Invention

[0008] The present invention provides an advertisement replacement method in complex scenes, including the following steps:

[0009] Step 1: Initially construct a dataset for the advertisement replacement task.

[0010] Step 2: Input the foreground advertisement image I fore , the mask image I mask into the spatial transformation module STM to obtain the transformed advertisement image I stn . Then, combine the background image I back with the transformed advertisement image I stn to obtain the intermediate generated image I mid .

[0011] Step 3: Input the mask image I mask and the intermediate generated image I mid into the edge optimization module EOM. After being processed by the edge optimization module EOM, a generated image is obtained. Combine this generated image with the intermediate generated image I mid to obtain the final generated result image I res .

[0012] As a further improvement of the present invention, in the said Step 2, it further includes:

[0013] Step 20: Conduct supervised training on the spatial transformation module STM to obtain the transformed advertisement image I stn . Through supervised training, the positioning network can calculate the transformation coordinates by inputting the mask image I mask .

[0014] Step 21: Combine the background image I back with the transformed advertisement image I stn to obtain the intermediate generated image I mid .

[0015] As a further improvement of the present invention, in the said Step 1, the dataset includes foreground advertisement pictures, scene pictures, marked advertisement area mask pictures, label pictures after deformation of the foreground advertisement pictures generated by imageology methods, and corresponding mask pictures.

[0016] As a further improvement of the present invention, in the step 20, it specifically further includes:

[0017] Step S20: Input the mask image I mask into the localization network to obtain the transformation coordinates and the reference coordinates.

[0018] Step S21: Then, generate a mapping grid from the mask image I mask by using the transformation coordinates and the reference coordinates; the grid records the pixel positions in the original advertisement image to which each pixel in the transformed image is mapped.

[0019] Step S22: Then, sample according to the grid data in the foreground advertisement image I fore to obtain the transformed advertisement image I stn .

[0020] As a further improvement of the present invention, in the step 2, the spatial transformation module STM is based on the STN network using TPS and undergoes supervised training. The formula is as follows:

[0021] I mid = STM(I fore , I mask , I back )

[0022] = sampler(Grid(LN(I mask ))), I fore ) ⊙ I mask + I back ⊙ (1 - I mask ) (2).

[0023] As a further improvement of the present invention, in the step 3, it further includes:

[0024] Step 30: Train the edge optimization module EOM to obtain the generated image;

[0025] Step 31: Combine the generated image with the intermediate generated image I mid to obtain the result image I res .

[0026] As a further improvement of the present invention, in the step 30, use the intermediate generated image I mid and the mask image I mask as the inputs of the generator, and output the generated image after passing through the generator. Use the background image I back as the positive sample of the discriminator for unsupervised adversarial training.

[0027] As a further improvement of the present invention, in the step 30, the mask image I maskThe masked image I of the black edge at the seam, i.e., the missing pixels, which needs to be obtained through image processing edge , and then connect it with the intermediate generated image I mid to input into the generator.

[0028] As a further improvement of the present invention, step 3 is expressed by the following formula:

[0029] I res = EOM(I mid , I mask )

[0030] = G(concat(I mid , I edge )) ⊙ dilate(I edge ) + I mid ⊙ (1 - dilate(I edge )) (3).

[0031] The beneficial effects of the present invention are as follows: 1. In the STM stage of the model of the present invention, by using the STN network structure, the model has flexible spatial deformation performance, thus realizing the alignment of the foreground advertisement in terms of position, shape, and direction on the curved surface of the target arc; 2. The present invention repairs and optimizes the generation result of STM through the EOM stage in the model design, making up for the errors of the STM model and making the final generation result more realistic and natural; 3. The model of the present invention provides a method for realizing advertisement implantation on the curved sticker surface as a whole, which is a task that cannot be completed by previous methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is the principle block diagram of the STM - EOM advertisement implantation model in the later stage of the advertisement replacement method of the present invention;

[0033] Figure 2 is the principle block diagram of the STM spatial transformation module of the present invention;

[0034] Figure 3 is the principle block diagram of the EOM edge optimization model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The advertisement replacement method of the present invention focuses on performing later - stage advertisement implantation in the case of complex advertisement stickers, that is, when the sticker is curved and has an arc, and the technical focus is on the advertisement replacement stage. Assuming that the advertisement placement position is known, its subsequent operations are realized.

[0036] For the late-stage advertisement implantation task in complex scenarios, we propose a two-stage neural network model STM-EOM. In the first stage, the advertisement image is spatially transformed for the implantation position, including displacement, deformation, etc., which we call the Spatial Transformer Module (STM). In the second stage, the processing and optimization of the seam between the advertisement and the background are carried out to make the synthesized image more natural and realistic, including the smoothing of the seam, the filling of the gap, the filtering of the remaining pixels, the ablation of the edge sawtooth, etc., which is called the Edge Optimization Module (EOM) in this invention. The advertisement image to be replaced in this invention is called the foreground advertisement image I fore , and the scene image containing the advertisement part is called the background image I back , and the mask image indicating the advertisement area in the corresponding scene image is I mask . The intermediate image I mid generated by STM, and the final result image I res generated by EOM. The overall flowchart of the model is shown in the figure and can be expressed by formula (1).

[0037] I res = EOM(STM(I fore , I mask , I back ), I mask ) (1)

[0038] As Figure 1 shown, an advertisement replacement method in complex scenarios disclosed by this invention includes the following steps:

[0039] Step 1: Initially construct a dataset for the complex scenario advertisement replacement task.

[0040] Advertisement-mini, which includes 660 foreground advertisement pictures, scene pictures, and mask pictures indicating the advertisement area respectively. In addition, there are 3960 label pictures and corresponding mask pictures after the deformation of the foreground advertisement pictures generated by the iconography method.

[0041] Step 2: Input the foreground advertisement image I fore , the mask image I mask into the spatial transformation module STM to obtain the transformed advertisement image I stn , and then combine the background image I back with the transformed advertisement image I stn to obtain the intermediate generated image I mid .

[0042] Step 3: Input the mask image I mask and the intermediate generated image I midInput Edge Optimization Module EOM. After being processed by the Edge Optimization Module EOM, a generated image is obtained. This generated image is combined with the intermediate generated image I mid to obtain the final generated result image I res .

[0043] The present invention is trained in two stages for STM and EOM respectively. STM uses the STN network based on TPS (Thin Plate Spline Transformation) as the basis and performs supervised training. The detailed structure is as Figure 2 shown and can be expressed by formula (2). The masked image is input into the localization network to obtain a set of transformation coordinates, and then through the generation grid and sampling stages, the spatially transformed advertisement image is obtained. Then, it is combined with the background image to obtain the intermediate generated result.

[0044] I mid = STM(I fore , I mask , I back )

[0045] = sampler(Grid(LN(I mask ))), I fore ) ⊙ I mask + I back ⊙ (1 - I mask ) (2)

[0046] Then, EOM is trained. Taking I mid and I mask as the inputs of the generator, the generator outputs the result image I res . The background image I back is used as the positive sample of the discriminator for unsupervised adversarial training. The EOM module as a whole uses the generative adversarial structure. The principle of the Generative Adversarial Network (GAN) is to train the generator to generate more realistic pictures as much as possible, and the discriminator to distinguish whether the input picture is a real picture or a generated fake picture as much as possible. The two networks perform adversarial training and make progress together. The ultimate ideal state is to reach the Nash equilibrium, and the accuracy rate of the discriminator's discrimination is 50%, that is, it cannot distinguish between real and fake images, and it is considered that the image generated by the generator is already close to the real image at this time. The detailed structure is as Figure 3 shown and can be expressed by formula (3). Among them, G is the generator, concat() is the concatenation operation, I edge is the mask image of the black edge at the seam, that is, the missing pixels, obtained through image processing, and dilate() is the dilation operation.

[0047] I res = EOM(I mid , I mask )

[0048] = G(concat(Imid , I edge )) ⊙ dilate(I edge ) + I mid ⊙ (1 - dilate(I edge )) (3)

[0049] The neural network STM - EOM model of the present invention can not only complete the late - stage advertisement implantation task in simple situations, but also complete the task in various complex scenarios, such as flat veneers, curved veneers, special deformations and other scenarios.

[0050] Straight - edged rectangle: The implantation effect of straight - edged rectangle advertisements from different observation perspectives in a simple scenario. The neural network STM - EOM model of the present invention can complete the advertisement replacement well, and the foreground and background joints are processed naturally, and the synthesized image has a high sense of reality.

[0051] Curved edge: For complex scenarios where the veneer is uneven, i.e., radian and curved surfaces, the neural network STM - EOM model of the present invention performs amazingly. The advertisement replacement method proposed by the present invention solves the problem of implanting advertisements with curved veneers in images, and also achieves good results in the processing of boundaries, and the synthesized images are more real and natural.

[0052] Special deformation: The present invention attempts to use the neural network STM - EOM model in special situations not covered by the training set. The neural network STM - EOM model still has good generation effects, indicating that the neural network STM - EOM model of the present invention has good generalization and scalability.

[0053] The beneficial effects of the present invention: 1. In the STM stage of the model of the present invention, by using the STN network structure, the model has flexible spatial deformation performance, thus realizing the alignment of the foreground advertisement in terms of position, shape and direction on the target radian curved surface; 2. The present invention repairs and optimizes the generation result of STM through the EOM stage of the model design, making up for the errors of the STM model while making the final generation result more real and natural; 3. The overall model of the present invention provides a method that can realize the advertisement implantation on the curved veneer, which is a task that previous methods could not complete.

[0054] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions or substitutions can be made, which should all be regarded as belonging to the protection scope of the present invention.

Claims

1. An advertisement replacement method in a complex scenario, characterized in that, Including the following steps: Step 1: Initially construct a dataset for the advertisement replacement task; Step 2: Input the foreground advertisement image I fore , the mask image I mask into the spatial transformation module STM to obtain the transformed advertisement image I stn . Then, combine the background image I back with the transformed advertisement image I stn to obtain the intermediate generated image I mid ; Step 3: Input the masked image I mask and the intermediate generated image I mid into the Edge Optimization Module (EOM). After being processed by the EOM, a generated image is obtained. Combine this generated image with the intermediate generated image I mid to obtain the final result image I res ; In step 3, it further includes: Step 30: Train the Edge Optimization Module (EOM) to obtain a generated image; Step 31: Combine the generated image with the intermediate generated image I mid to obtain the result image I res ; In the step 30, the intermediate generated image I mid and the mask image I mask are used as the inputs of the generator, and after passing through the generator, a generated image is output. The background image I back is used as the positive sample of the discriminator for unsupervised adversarial training; In the step 30, the mask image I mask is the mask image I of the black edge at the seam that needs to be obtained through image processing edge , and then it is connected to the intermediate generated image I mid for input to the generator; Step 3 is expressed by the following formula: I res = EOM(I mid , I mask ) = G(concat(I mid , I edge )) ⊙ dilate(I edge ) + I mid ⊙ (1 - dilate(I edge )) (3) Among them, concat() is the concatenation operation, and dilate() is the dilation operation.

2. The advertisement replacement method according to claim 1, wherein In step 2, it further includes: Step 20: Perform supervised training on the spatial transformation module STM to obtain the transformed advertisement image I stn ; Step 21: Combine the background image I back with the transformed advertisement image I stn to obtain the intermediate generated image I mid .

3. The advertisement replacement method according to claim 1, wherein In step 1, the dataset includes foreground advertisement pictures, scene pictures, labeled advertisement area mask pictures, label pictures after deformation of foreground advertisement pictures generated by the iconography method, and corresponding mask pictures.

4. The advertisement replacement method according to claim 2, wherein In step 20, it specifically further includes: Step S20: Input the mask image I mask into the localization network to obtain the transformed coordinates and the reference coordinates; Step S21: Then, for the mask image I mask generate a mapping grid by converting the coordinates and the reference coordinates; Step S22: Sample according to the grid data in the foreground advertisement image I fore to obtain the transformed advertisement image I stn .

5. The advertisement replacement method according to claim 4, wherein In step 2, the Spatial Transformation Module (STM) is based on the STN network based on TPS and undergoes supervised training, which is expressed by the following formula: I mid = STM(I fore , I mask , I back ) = sampler(Grid(LN(I mask )),I fore )⊙I mask +I back ⊙(1 - I mask ) (2).

Citation Information

Patent Citations

  • Video advertisement integration system and method based on generative adversarial network

    CN111461772A

  • Real-time target editing method based on local space conversion network

    CN114299588A