Streaming e-commerce picture material manufacturing method and system based on AIGC

Through the AIGC-based streaming e-commerce image material production method, G-Dino and SAM are used for product detection and segmentation, and background integration is combined with the BrushNet model, the problems of cumbersome and poor results are solved, and efficient and automated e-commerce image material production is achieved.

CN120259486APending Publication Date: 2025-07-04SHANGHAI WANGMAI INFORMATION TECH GRP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411765852.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The production of traditional e-commerce image materials is complicated, and the effect is poor when using the image editing software directly, and the perspective relationship between the product and the background is not matched and there is a lack of interactive relationship.

Method used

AIGC-based streaming e-commerce image material production method is adopted, G-Dino and SAM are used for product detection and segmentation, and background fusion is combined with the BrushNet model to achieve adaptive adjustment and material splicing.

Benefits of technology

It realizes efficient and automated e-commerce picture material production, reduces manpower investment, improves material quality, and solves the problems of cumbersome operations and poor results in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259486A_ABST
    Figure CN120259486A_ABST
Patent Text Reader

Abstract

The invention discloses an AIGC-based streaming e-commerce picture material manufacturing method and system. The method comprises the steps of S1, obtaining a picture containing an e-commerce product picture; s2, separating the picture containing the e-commerce product drawing by using an AIGC tool based on semantics to obtain the e-commerce product drawing; step S3, adaptive adjustment is carried out on the separated e-commerce product graph, and the adaptive adjustment comprises commodity detection: semantic-based target detection is realized based on G-Dino, and vertex data of a commodity and a bounding box of a maximum contour are extracted; calculating a scaling ratio: calculating an aspect ratio of the main body, and calculating a new width according to a target size and a preset commodity size so as to keep the original scale to scale the main body image; creating a target canvas and adjusting the position: creating a new canvas according to the target size, and moving the commodity obtained in the previous step to a preset vertex coordinate position so as to finally generate a new mask; and S4, fusing the new mask obtained in the step S3 with the background through a BrushNet model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly to a method and system for producing streaming e-commerce picture materials based on AIGC. Background Art

[0002] For the production of traditional e-commerce picture materials, a large amount of manual operations are required, such as using post-image processing software like Photoshop to edit pictures, and the operations are very cumbersome.

[0003] In addition, if the pictures of products are directly spliced with the background of Photoshop, the resulting effect is very poor. On the one hand, the perspective relationship between the product and the scene does not match. On the other hand, there is no interaction relationship between the product and the scene, and there are problems with shadow lighting, etc.

[0004] In view of this, the present invention is specifically proposed. Summary of the Invention

[0005] The purpose of the present invention is to overcome the above technical deficiencies, and provide a method and system for producing streaming e-commerce picture materials based on AIGC, so as to solve the technical problems in the related technologies that the production of e-commerce picture materials is cumbersome in operation and the effect is poor when directly using picture editing software.

[0006] To achieve the above technical purpose, the present invention adopts the following technical solutions:

[0007] According to one aspect of the present invention, there is provided a method for producing streaming e-commerce picture materials based on AIGC, including:

[0008] Step S1, obtaining a picture containing an e-commerce product picture;

[0009] Step S2, separating the picture containing the e-commerce product picture based on semantics by using an AIGC tool to separate the e-commerce product picture;

[0010] Step S3, performing adaptive adjustment on the separated e-commerce product picture, and the adaptive adjustment includes:

[0011] Commodity detection: realizing semantic-based object detection based on G-Dino, and extracting the vertex data of the commodity and the bounding box of the largest contour;

[0012] Calculating the scaling ratio: calculating the aspect ratio of the main body, and then calculating the new width according to the target size and the preset commodity size to scale the main body image while maintaining the original ratio;

[0013] Creating a target canvas and adjusting the position: creating a new canvas according to the target size, and moving the commodity obtained in the previous step to the preset vertex coordinate position, thereby finally generating a new mask;

[0014] Step S4: Fuse the new mask obtained in step S3 with the background through the BrushNet model.

[0015] Furthermore, the separation of the picture containing the e-commerce product picture by using the AIGC tool based on semantics in step S2 to separate the e-commerce product picture specifically includes:

[0016] First, perform object detection based on picture semantics through G-Dino; then segment the recognized picture contour through SAM.

[0017] Furthermore, the implementation of object detection based on picture semantics through G-Dino specifically includes:

[0018] The text features are obtained by BERT, and the feature dimension output by BERT is mapped to a dimension of 256 through an MLP layer to achieve unification with the image features;

[0019] For the image features, they are obtained by Swin Transformer, and the multi-scale feature dimensions output by Swin Transformer are unified to 256;

[0020] Both the text features and the image features have corresponding position encodings. Spatial position encoding is used for images, and sine position encoding is used for text.

[0021] Furthermore, the SAM includes an image encoder, a prompt encoder, and a mask decoder, and generates a segmentation mask according to different types of prompts.

[0022] Furthermore, the image features and the text features are output through the following formula:

[0023]

[0024] P (v) = PW (v,L) , O t2i = SoftMax(Attn)P (v) W (out,I)

[0025] Among them,

[0026] O (q) represents the output of the image query vector in the self-attention mechanism, and the superscript q represents Query;

[0027] W (q,I) is a weight matrix that converts the input I into the query vector O, where I represents the input features;

[0028] P (q) is the output of the text query vector in the self-attention mechanism, passing through the weight matrix W(q,L) Convert L into a query vector, where L represents the input feature;

[0029] Attn is the attention parameter, through O (q) and P (q) Multiply the transpose of and divide by Here, d is the scaling factor, representing the dimension of the query vector;

[0030] P (v) is the output of the value vector in the self-attention mechanism, and the superscript v represents value;

[0031] W (v,L) is the weight matrix, which converts the input L into a value vector, where L represents the input feature;

[0032] O t2i represents the mapping of the image vector from target to input, where t is the target and i is the input;

[0033] SoftMax is the activation function to obtain the attention weights;

[0034] W (out,I) is the weight matrix for converting the weighted value vector into the final output O t2i , and the superscript out represents the output feature.

[0035] Furthermore, it further includes:

[0036] Step S5, after implementing the fusion of the new mask of the commodity and the background, preset other materials on the template layer to automatically realize material splicing.

[0037] The present invention also provides a streaming e-commerce picture material production system based on AIGC, including:

[0038] A picture acquisition unit for acquiring pictures containing e-commerce product pictures;

[0039] A separation unit for separating the pictures containing e-commerce product pictures based on semantics by using AIGC tools to separate the e-commerce product pictures;

[0040] An adaptive adjustment unit for adaptively adjusting the separated e-commerce product pictures, and the adaptive adjustment includes:

[0041] Commodity detection: Based on G-Dino, semantic-based object detection is implemented to extract the vertex data of the commodity and the bounding box of the largest contour;

[0042] Calculate the scaling ratio: Calculate the aspect ratio of the main body, and then calculate the new width based on the target size and the preset product size to scale the main body image while maintaining the original ratio;

[0043] Create the target canvas and adjust the position: Create a new canvas according to the target size, and move the product obtained in the previous step to the preset vertex coordinate position, thereby finally generating a new mask;

[0044] Redraw the fusion unit for fusing the obtained new mask with the background through the BrushNet model.

[0045] Furthermore, it also includes a material automatic splicing unit for presetting other materials on the template layer after fusing the new mask of the product with the background to automatically achieve material splicing.

[0046] According to another aspect of the present invention, there is also provided a computer-readable storage medium storing one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the above method.

[0047] A method and system for producing streaming e-commerce picture materials based on AIGC provided by the present invention realizes intelligent matte extraction, background generation, content synthesis, and fine adjustment and fusion of product content through an automated process, thereby batch-producing high-quality e-commerce advertising picture materials. Compared with traditional methods, this streaming production method significantly reduces labor input, avoids the cumbersome operations of post-image processing software such as Photoshop, as well as the complex fusion process of tools such as Stable Diffusion, and realizes full automation from content extraction to finished product output, providing an efficient and low-cost picture material production solution for the e-commerce field. Description of the Drawings

[0048] Figure 1 It is a schematic flowchart of a method for producing streaming e-commerce picture materials based on AIGC provided by an embodiment of the present invention;

[0049] Figure 2 It is a schematic structural diagram of a system for producing streaming e-commerce picture materials based on AIGC provided by an embodiment of the present invention;

[0050] Figure 3 It is a structural block diagram of a terminal that can implement the method for producing streaming e-commerce picture materials based on AIGC according to an embodiment of the present invention;

[0051] Figure 4 It is an e-commerce product picture of a certain schoolbag;

[0052] Figure 5 It is a picture mask of a schoolbag generated based on G-Dino+SAM;

[0053] Figure 6 The new mask after adaptive product adjustment;

[0054] Figure 7 The merged image of the mask and the street background by BrushNet;

[0055] Figure 8 The merged image of the mask and the forest background by BrushNet;

[0056] Figure 9 For Figure 8 The rendered image after material splicing. Specific implementation manners

[0057] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0058] Embodiment 1

[0059] According to an embodiment of the present invention, a method for making streaming e-commerce picture materials based on AIGC is provided, in combination with Figure 1 , including:

[0060] Step S1, obtaining a picture containing an e-commerce product picture;

[0061] Product pictures of a certain e-commerce category can be obtained on an e-commerce website, such as the schoolbag shown in Figure 4. If directly spliced with the background of PS, the resulting effect is very poor. On the one hand, the perspective relationship between the product and the scene does not match, and on the other hand, there is no interaction relationship between the product and the scene, and there are problems with shadow lighting, etc. Therefore, it is necessary to process it by means of the following steps provided by the embodiments of the present invention.

[0062] Step S2, separating the picture containing the e-commerce product picture based on semantics by using an AIGC tool to separate the e-commerce product picture.

[0063] On the premise of ensuring that the e-commerce product remains unchanged, the AIGC tool can block the unchanged areas through image masking and redraw the non-masked areas, thus solving the problem of product image and background fusion. However, image masking generally needs to be generated manually and cannot be automated. The present invention can identify and generate image masks based on semantic understanding by GroundingDino (G-Dino) + Segment Anything Model (SAM). That is, the separation of the picture containing the e-commerce product picture by using the AIGC tool based on semantics in step S2, specifically including: first, realizing object detection based on picture semantics through G-Dino; then segmenting the identified picture contour through SAM.

[0064] The following introduces the applications of G-Dino and SAM in step S2 in this embodiment:

[0065] GroundingDino (G-Dino)

[0066] This method integrates data from two modalities, text and image, and realizes open-set object detection, that is, given a text prompt, it automatically frames the location of the object, and the object can be a category not in the training set. This method mainly realizes the above functions through a feature enhancement module, a language-guided query selection module, and a cross-modal decoding module. Grounding DINO needs to use data from both the text and image modalities simultaneously. Therefore, it is necessary to extract features from the data of the two modalities respectively and unify the feature dimensions so that they can perform cross-attention calculations with each other.

[0067] In step S2, the text features are obtained by BERT. The feature dimension output by BERT is 768, and it needs to be mapped to a dimension of 256 through an MLP layer to achieve unification with the image features. For image data, the image features are obtained by Swin Transformer. The output of Swin Transformer is multi-scale features (the number of channels is 96, 192, 384, 768 respectively), so the feature dimensions of each layer also need to be unified to 256. This can be achieved through 1*1 2D convolution, and then normalized through GroupNorm.

[0068] Both types of features need to have corresponding position encodings. The image uses Sinusoidal position embedding (spatial position encoding), and the text uses sine position embedding (sine position encoding).

[0069] Taking the image-text cross-attention as an example, the specific formula is as follows:

[0070]

[0071] P (v) = PW (v,L) , O t2i = SoftMax(Attn)P (v) W (out,I)

[0072] Among them,

[0073] O (q) represents the output of the image query vector in the self-attention mechanism, and the superscript q represents Query;

[0074] W (q,I) is a weight matrix that converts the input I into the query vector O, where I represents the input feature;

[0075] P (q) is the output of the text query vector in the self-attention mechanism. After passing through the weight matrix W (q,L) converts L into the query vector, where L represents the input feature;

[0076] Attn is the attention parameter, through O (q) and P (q) multiply the transpose of and divide by Here, d is the scaling factor, representing the dimension of the query vector;

[0077] P (v) is the output of the value vector in the self-attention mechanism, and the superscript v represents Value;

[0078] W (v,L) is the weight matrix that converts the input L into the value vector, and L represents the input feature;

[0079] O t2i represents the mapping of the image vector from Target to Input, where t is the target and i is the input;

[0080] SoftMax is the activation function to obtain the attention weights;

[0081] W (out,I) is the weight matrix used to convert the weighted value vector into the final output O t2i , and the superscript out represents the output feature.

[0082] Segment Anything Model (SAM)

[0083] The Segment Anything Model (SAM) is a model for promptable segmentation. It consists of three main components: an image encoder, a prompt encoder, and a mask decoder. These components work together to generate segmentation masks based on different types of prompts.

[0084] 1) Image Encoder: SAM utilizes a Vision Transformer (ViT) pre-trained with MAE (Masked Autoencoder) that has been adapted to handle high-resolution inputs as the image encoder. This image encoder can efficiently encode the input image. It runs once on each image and can be applied before prompting the model.

[0085] 2) Prompt Encoder: SAM considers two types of prompts: sparse prompts (such as points, boxes, and text) and dense prompts (masks). Sparse prompts (such as points and boxes) are represented using positional encoding combined with learned embeddings for each prompt type. Free-form text prompts are encoded using the text encoder in CLIP (Contrastive Language-Image Pre-training). Dense prompts (i.e., masks) are embedded using convolutions and summed element-wise with the image embedding.

[0086] 3) Mask Decoder: The mask decoder takes the image embedding, the prompt embedding, and the output tokens as inputs and efficiently maps them to segmentation masks. Inspired by previous work, SAM employs a modified Transformer decoder block that contains prompt self-attention and bidirectional cross-attention: from prompt to image embedding and vice versa. After running two decoder blocks, the image embedding is upsampled, and then the output tokens are mapped to a dynamic linear classifier through a multi-layer perceptron (MLP). This classifier calculates the mask foreground probability for each position in the image.

[0087] G-Dino + SAM Achieves Semantic Segmentation

[0088] The specific solution is as follows. First, G-Dino is used to achieve object detection based on image semantics. By inputting text, such as "bag", the outline of the schoolbag in the picture is detected. Then, SAM is used to segment the recognized picture outline, thereby achieving the goal of separating e-commerce product pictures based on semantics. The final generated picture mask is as Figure 5 shown.

[0089] After implementing product picture segmentation, it is not possible to directly redraw on the masked picture because:

[0090] 1) Size mismatch: Since the information flow materials have requirements for size, for example, the landscape pictures are generally 1280*720, and the portrait pictures are generally 1080*1920, while the Taobao product sizes are generally square, it is necessary to stretch the picture to the specified size;

[0091] 2) Product position mismatch: In addition to the products in the e-commerce materials, elements such as product logos and copywriting artistic words need to be added. The positions of these elements are generally fixed, so the position and size of the products need to be adaptively adjusted to the appropriate positions to facilitate the subsequent batch addition of other elements.

[0092] Therefore, the above problems are solved through the following step S3.

[0093] Step S3, perform adaptive adjustment on the separated e-commerce product pictures.

[0094] The adaptive adjustment includes:

[0095] Product detection: Based on G-Dino, implement semantic-based object detection to extract the vertex data of the product and the bounding box of the largest contour;

[0096] Calculate the scaling ratio: Calculate the aspect ratio of the main body, and then calculate the new width according to the target size and the preset product size to scale the main body image while maintaining the original ratio;

[0097] Create a target canvas and adjust the position: Create a new canvas according to the target size, and move the product obtained in the previous step to the preset vertex coordinate position, thus finally generating a new mask.

[0098] The finally generated new mask is as Figure 6 shown.

[0099] Step S4, fuse the obtained new mask with the background through the BrushNet model.

[0100] BrushNet is a pluggable double-branch model that can embed pixel-level mask image features into any pre-trained diffusion model to achieve semantically consistent and quality-enhanced image restoration effects, and can achieve high-quality and natural and coherent redrawing. The core innovation point of BrushNet is to process the mask image features and the noise latent vectors as two independent branches. This branch separation greatly reduces the learning burden of the model, enabling the model to finely integrate the key information of the mask image in a hierarchical manner, so it can achieve more efficient background redrawing and fusion.

[0101] Based on the background redrawing and fusion of BrushNet, fuse the mask of the e-commerce product picture produced in the above steps with backgrounds such as streets and forests. The final effect is as Figure 7 、8 as shown

[0102] Furthermore, it also includes:

[0103] Step S5, after realizing the fusion of the new mask of the commodity and the background, preset other materials on the template layer to automatically realize material splicing.

[0104] After realizing the background redrawing and fusion of the commodity, since the position of the commodity picture has been determined, by presetting the copywriting and Logo on the template layer, the commodity picture can be automatically spliced with other elements to generate the creativity for delivery. For example, based on Figure 8 , after adding other elements, the final formed effect is as Figure 9 as shown

[0105] The present invention also provides a streaming e-commerce picture material production system based on AIGC, combined with Figure 2 , including:

[0106] A picture acquisition unit for acquiring pictures containing e-commerce product pictures;

[0107] A separation unit for separating the picture containing the e-commerce product picture by using an AIGC tool based on semantics to separate the e-commerce product picture;

[0108] An adaptive adjustment unit for adaptively adjusting the separated e-commerce product picture, and the adaptive adjustment includes:

[0109] Commodity detection: Based on G-Dino, perform semantic-based object detection to extract the vertex data of the commodity and the bounding box of the maximum contour;

[0110] Calculate the scaling ratio: Calculate the aspect ratio of the main body, and then calculate the new width according to the target size and the preset commodity size to scale the main body image while maintaining the original ratio;

[0111] Create a target canvas and adjust the position: Create a new canvas according to the target size, and move the commodity obtained in the previous step to the preset vertex coordinate position, so as to finally generate a new mask;

[0112] A redrawing and fusion unit for fusing the obtained new mask with the background through the BrushNet model.

[0113] Furthermore, it also includes a material automatic splicing unit for presetting other materials on the template layer after realizing the fusion of the new mask of the commodity and the background to automatically realize material splicing.

[0114] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0115] Optionally, the specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0116] It should be noted here that the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run in the corresponding hardware environment, can be implemented by software, or can be implemented by hardware, where the hardware environment includes a network environment.

[0117] Figure 3 is a structural block diagram of a terminal according to an embodiment of the present application, as Figure 3 shown. The terminal may include: one or more (only one is shown) processors 101, a memory 103, and a transmission device 105, as Figure 3 shown. The terminal may further include an input / output device 107.

[0118] Among them, the memory 103 can be used to store software programs and modules, such as the program instructions / modules corresponding to the methods and devices in the embodiments of the present application. The processor 101 executes various functional applications and data processing by running the software programs and modules stored in the memory 103, that is, implements the above methods. The memory 103 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 103 may further include a memory remotely disposed relative to the processor 101, and these remote memories can be connected to the terminal through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0119] The above-mentioned transmission device 105 is used to receive or send data via a network, and can also be used for data transmission between the processor and the memory. Specific examples of the above-mentioned network can include a wired network and a wireless network. In one example, the transmission device 105 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers through a network cable, so as to communicate with the Internet or a local area network. In one example, the transmission device 105 is a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0120] Specifically, the memory 103 is used to store application programs.

[0121] The processor 101 can call the application programs stored in the memory 103 through the transmission device 105 to execute the various steps in the above method.

[0122] Optionally, the specific examples in this embodiment can refer to the examples described in the above embodiment, and will not be elaborated here.

[0123] Those of ordinary skill in the art can understand that the structure of the above terminal is only schematic, and the terminal can be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, and a Mobile Internet Devices (MID), a PAD and other terminal devices. Figure 3 It does not limit the structure of the above electronic device. For example, the terminal may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 3 or have a different configuration from that shown in Figure 3 shown.

[0124] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium. The storage medium can include: a flash drive, a Read Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disc, etc.

[0125] The embodiment of the present application also provides a storage medium. Optionally, in this embodiment, the above storage medium can be used to execute the program code of the above method.

[0126] Optionally, in this embodiment, the above storage medium can be located on at least one of the multiple network devices in the network shown in the above embodiment.

[0127] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the steps in the above method.

[0128] Optionally, specific examples in this embodiment may refer to the examples described in the above embodiments, and details are not described herein again.

[0129] Optionally, in this embodiment, the above storage medium may include but is not limited to: various media that can store program codes, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs.

[0130] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.

[0131] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A method for producing streaming e-commerce picture materials based on AIGC, characterized in that, Including: Step S1, obtaining a picture containing an e-commerce product picture; Step S2, based on semantics, using an AIGC tool to separate the picture containing the e-commerce product picture to separate out the e-commerce product picture; Step S3, performing adaptive adjustment on the separated e-commerce product picture, and the adaptive adjustment includes: Commodity detection: Based on G-Dino, semantic-based object detection is implemented to extract the vertex data of the commodity and the bounding box of the largest contour; Calculating the scaling ratio: Calculate the aspect ratio of the main body, and then calculate the new width according to the target size and the preset commodity size to scale the main body image while maintaining the original ratio; Creating a target canvas and adjusting the position: Create a new canvas according to the target size, and move the commodity obtained in the previous step to the preset vertex coordinate position, thereby finally generating a new mask; Step S4, fusing the new mask obtained in Step S3 with the background through the BrushNet model.

2. The method for producing streaming e-commerce picture materials based on AIGC according to claim 1, wherein In Step S2, based on semantics, using an AIGC tool to separate the picture containing the e-commerce product picture to separate out the e-commerce product picture, specifically including: First, implement semantic-based object detection of the picture through G-Dino; then segment the recognized picture contour through SAM.

3. The method for producing streaming e-commerce picture materials based on AIGC according to claim 2, wherein The implementation of semantic-based object detection of the picture through G-Dino specifically includes: The text features are obtained by BERT, and the feature dimension output by BERT is mapped to a dimension of 256 through an MLP layer to achieve unity with the image features; For the image features, they are obtained by Swin Transformer, and the multi-scale feature dimensions output by Swin Transformer are unified to 256; Both the text features and the image features have corresponding position encodings. The image uses spatial position encoding, and the text uses sine position encoding.

4. The method for producing streaming e-commerce picture materials based on AIGC according to claim 3, wherein Output the image features and text features through the following formula: P (v) = PW (v,L) , O t2i = SoftMax(Attn)P (v) W (out,I) Where, O (q) represents the output of the image query vector in the self-attention mechanism, where the superscript q represents Query; W (q,I) is a weight matrix that transforms the input I into the query vector O, where I represents the input feature; P (q) is the output of the text query vector in the self-attention mechanism, passing through the weight matrix W (q,L ) converts L into a query vector, where L represents the input feature; Attn is the attention parameter, which is multiplied by the transpose of O (q) and P (q) and then divided by Here, d is the scaling factor, representing the dimension of the query vector; P (v) is the output of the value vector in the self-attention mechanism, where the superscript v represents Value; W (v,L) is a weight matrix that transforms the input L into a value vector, where L represents the input feature; O t2i Represents the mapping from the Target to the Input of the image vector, where t is the target and i is the input; SoftMax is the activation function to obtain the attention weights; W (out,I) is the weight matrix used to convert the weighted value vector into the final output O t2i , where the superscript out represents the output feature.

5. The method for producing streaming e-commerce picture materials based on AIGC according to claim 2, wherein, The SAM includes an image encoder, a prompt encoder, and a mask decoder, and generates segmentation masks according to different types of prompts.

6. The method for producing streaming e-commerce picture materials based on AIGC according to claim 1, wherein, It also includes: Step S5, after realizing the fusion of the new mask of the commodity and the background, preset other materials on the template layer to automatically realize material splicing.

7. A streaming e-commerce picture material production system based on AIGC, characterized in that, Including: A picture acquisition unit for acquiring a picture containing an e-commerce product picture; A separation unit for separating the picture containing the e-commerce product picture based on semantics using an AIGC tool to separate out the e-commerce product picture; An adaptive adjustment unit for performing adaptive adjustment on the separated e-commerce product picture, and the adaptive adjustment includes: Commodity detection: Based on G-Dino, semantic-based object detection is implemented to extract the vertex data of the commodity and the bounding box of the largest contour; Calculating the scaling ratio: Calculate the aspect ratio of the main body, and then calculate the new width according to the target size and the preset commodity size to scale the main body image while maintaining the original ratio; Creating a target canvas and adjusting the position: Create a new canvas according to the target size, and move the commodity obtained in the previous step to the preset vertex coordinate position, thereby finally generating a new mask; A redrawing and fusing unit for fusing the obtained new mask with the background through the BrushNet model.

8. The AIGC-based streaming e-commerce picture material production system according to claim 7, wherein, It further includes a material automatic splicing unit, which is used to preset other materials on the template layer after realizing the fusion of the new mask and the background of the commodity, so as to automatically realize material splicing.

9. An electronic device, characterized in that, It includes: a processor and a memory; a computer-readable program executable by the processor is stored on the memory; when the processor executes the computer-readable program, the steps in the method according to any one of claims 1-6 are implemented.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method according to any one of claims 1-6.