Handheld article replacement method, digital human video generation method, device and equipment

By using multimodal feature fusion and diffusion generation algorithms, the problems of abrupt edge connections and unreasonable hand-object interaction logic in handheld object replacement are solved, achieving natural interaction between the target object and the hand and smooth edge transitions, thus improving the visual quality and realism of the image.

CN121708173APending Publication Date: 2026-03-20NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610195786.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, the replacement of handheld objects results in abrupt edge transitions and illogical hand-object interaction logic, leading to insufficient image realism.

Method used

A multimodal feature fusion and diffusion generation algorithm is adopted to simultaneously extract visual and semantic features, embed them into time steps, and use the diffusion generation algorithm for latent code processing and feature decoding to generate a target display image.

Benefits of technology

It achieves natural interaction between the target object and the hand, with smooth edge transitions, thus improving the visual quality and realism of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708173A_ABST
    Figure CN121708173A_ABST
Patent Text Reader

Abstract

The invention provides a handheld article replacement method, a digital human video generation method, a device and equipment. The handheld article replacement method comprises the following steps: acquiring an original display image and a target article image; identifying a handheld article in the original display image, and generating an article mask; extracting visual features of the original display image, the article mask and the target article image, and extracting semantic features of the article cue word; performing multi-modal fusion processing on the visual features and the semantic features, and embedding time steps to obtain a global feature sequence; and carrying out latent code processing and feature decoding on the global feature sequence by utilizing a diffusion generation algorithm to generate a target display image taking the target article as the handheld article. According to the method, handheld article replacement is carried out on the basis of the multi-modal features, and meanwhile, an iterative noise reduction mechanism of a diffusion generation algorithm is combined, so that the coordination of interaction between a target article and a hand after replacement is optimized, and the sense of reality and the overall visual effect of an image after replacement are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image synthesis, in particular to a handheld object replacement method, a digital human video generation method, device and equipment. BACKGROUND

[0002] In e-commerce display, virtual trial object display, commodity content creation and other scenarios, it is often necessary to replace the original handheld object in the image with a new object to achieve the needs of commodity effect simulation, new product display, creative content generation and the like.

[0003] In order to realize the replacement of the handheld object in the image, the image segmentation technology is usually used to realize the positioning of the handheld object, and then the obtained object mask is used to replace the handheld object with a new object. In related technologies, the object replacement is often only dependent on the visual features of the image, and the replacement process is mostly simple image block splicing or static mask filling. In the face of complex interaction relationships such as hand and object occlusion, pressing and the like in the handheld object scene, it is easy to have problems such as harsh edge connection and unreasonable hand and object interaction logic.

[0004] Therefore, it is urgent to provide a handheld object replacement scheme suitable for handheld object replacement and having high precision to improve the sense of reality of the image after object replacement. SUMMARY

[0005] Embodiments of the present application provide a handheld object replacement method, a digital human video generation method, device and equipment, which are used to solve the problems of harsh edge connection and unreasonable hand and object interaction logic in handheld object replacement, and the resulting lack of image reality.

[0006] In a first aspect, the embodiments of the present application provide a handheld object replacement method, which includes: obtaining an original display image and a target object image; identifying a handheld object in the original display image to generate an object mask; extracting visual features of the original display image, the object mask and the target object image, and extracting semantic features of an object prompt word; performing multi-modal fusion processing on the visual features and the semantic features and embedding time steps to obtain a global feature sequence; using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target object as the handheld object.

[0007] In a second aspect, the embodiments of the present application provide a digital human video generation method, which includes: obtaining a target display image in which a handheld object is replaced with a target object; the target display image is obtained based on the method provided in the first aspect of the present application or any implementation manner corresponding to the first aspect; inputting the target display image and associated information of the target object into a digital human driving model to obtain a digital human video for displaying the target object.

[0008] In a third aspect, the embodiments of the present application provide a handheld object replacement device, comprising: an image acquisition module configured to acquire an original display image and a target object image; a mask generation module configured to identify a handheld object in the original display image and generate an object mask; a multi-modal feature extraction module configured to extract visual features of the original display image, the object mask and the target object image, and extract semantic features of an object prompt word; a multi-modal fusion module configured to perform multi-modal fusion processing on the visual features and the semantic features and embed time steps to obtain a global feature sequence; and a target image generation module configured to perform latent code processing and feature decoding on the global feature sequence by using a diffusion generation algorithm to generate a target display image in which the target object is the handheld object.

[0009] In a fourth aspect, the embodiments of the present application provide a digital human video generation device, comprising: a display image acquisition module configured to acquire a target display image in which a handheld object is replaced by a target object; the target display image is obtained based on the method provided in the first aspect of the present application or any of the corresponding embodiments of the first aspect; and a video generation module configured to input the target display image and associated information of the target object into a digital human driving model to obtain a digital human video for displaying the target object.

[0010] In a fifth aspect, the embodiments of the present application provide an electronic device, comprising: a memory and a processor; the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory, so that the processor executes the method provided in the first aspect or the second aspect, and / or various possible embodiments corresponding to the first aspect or the second aspect.

[0011] In a sixth aspect, the embodiments of the present application provide a computer readable storage medium, which stores computer execution instructions, and the computer execution instructions are executed by a processor to implement the method provided in the first aspect or the second aspect, and / or various possible embodiments corresponding to the first aspect or the second aspect.

[0012] In a seventh aspect, the embodiments of the present application provide a computer program product, which comprises a computer program, and the computer program is executed by a processor to implement the method provided in the first aspect or the second aspect, and / or various possible embodiments corresponding to the first aspect or the second aspect.

[0013] The handheld item replacement method, digital human video generation method, apparatus, and device provided in this application address the problems of traditional handheld item replacement relying solely on visual features, having simple replacement logic, and being prone to abrupt edge transitions and unreasonable hand-object interaction logic. They employ a collaborative processing mode that integrates visual and semantic multimodal features and uses diffusion generation guidance. By simultaneously extracting visual features from the original display image, item mask, and target item image, along with semantic features from item prompts, the limitations of single visual features are overcome. Multimodal features are fused and time-step information is embedded to obtain a global feature sequence that combines scene constraints and stage guidance. Finally, a diffusion generation algorithm is used to perform latent code processing and feature decoding based on the global feature sequence, achieving accurate adaptation of the target item to the original image's hand and background. This efficiently completes the handheld item replacement, ensuring natural interaction between the target item and hand after replacement, smooth edge transitions, and improving the visual quality and realism of the replaced image. Attached Figure Description

[0014] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0015] Figure 1 A schematic diagram illustrating the application scenarios provided in this application;

[0016] Figure 2 A flowchart illustrating the handheld item replacement method provided in this application;

[0017] Figure 3 for Figure 2 A flowchart illustrating step S204 in the illustrated embodiment;

[0018] Figure 4 for Figure 2 A flowchart illustrating step S205 in the illustrated embodiment;

[0019] Figure 5 This is a schematic diagram of the structure of the handheld item replacement frame provided in the embodiments of this application;

[0020] Figure 6 A schematic diagram of the handheld article replacement device provided in this application;

[0021] Figure 7 A schematic diagram of the structure of the electronic device provided in this application.

[0022] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0023] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0024] It should be noted that the personal information (including but not limited to device information, personal attribute information, personal image, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the individual or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data comply with relevant laws, regulations and standards, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0025] First, let me explain the terms used in this application:

[0026] Mask: Binary identifier data used to accurately locate target regions in an image, clearly defining the boundary between the target region and the background region. A 1 indicates that the corresponding pixel belongs to the target region, while a 0 indicates that the corresponding pixel belongs to the background region.

[0027] Digital humans (also known as virtual humans): Anthropomorphic virtual images constructed based on technologies such as computer graphics and artificial intelligence. They have an appearance similar to real people and can express actions and transmit information through drive signals.

[0028] Digital Human Driving Model: An intelligent algorithm model used to control the movements of digital humans. It can receive multimodal inputs such as voice, text, and action commands, parse the feature information of the input data and generate corresponding driving signals to control the digital human to complete anthropomorphic actions such as lip-syncing, limb movement, and head posture adjustment.

[0029] In e-commerce live streaming, online advertising, and other similar scenarios, sellers often showcase their products through images or videos, with images of people (digital or real models) holding the product serving as the core element. To adapt to the promotional needs of different products, it is often necessary to replace the original handheld item in the image with a new target item to quickly generate diverse display materials without having to reshoot.

[0030] For example, Figure 1 A schematic diagram illustrating the application scenarios provided in this application, such as... Figure 1As shown, a user, such as an e-commerce merchant, stores the captured image material, such as the original display image img1 of a model holding product A, on the user's terminal. When the merchant needs to promote a new product, such as product B, but has not captured the corresponding handheld display image, there is no need to reorganize the model for shooting. They only need to obtain the target item image img2 of product B. They only need to upload the original display image img1 and the target item image img2 to the server through the user's terminal. The server executes the handheld item replacement method provided in this application to replace the original product A in the original display image img1 with product B, quickly generate the target display image img3 of the model holding product B, and return the target display image img3 to the user's terminal.

[0031] Users can also upload multiple target item images, and the server can perform batch replacement processing and output the target display images corresponding to each target item at once.

[0032] Users can also specify the original display images corresponding to each target item image. The server will perform the replacement according to the preset association relationship to generate personalized display images that are adapted to different original scenes and different target products, flexibly meeting the needs of creating promotional materials for multiple categories and multiple scenarios.

[0033] Traditional handheld object replacement relies on manual image editing or segmentation and replacement techniques driven solely by visual features. Manual image editing is inefficient and costly, failing to meet the needs of batch material production. Replacement methods that rely solely on visual features cannot accurately respond to personalized requirements such as the object's posture and style, nor can they handle complex interactive relationships between the hand and the object, such as occlusion and pressing. This results in replacement images generally exhibiting problems such as abrupt edge blending, contradictory hand-object interaction logic, and poor visual realism.

[0034] To address the aforementioned issues, this application provides a method for replacing handheld objects, employing a collaborative processing scheme of multimodal feature fusion and diffusion generation guidance. Specifically, visual and semantic features are extracted simultaneously and fused and embedded into a time step to obtain a global feature sequence. Then, latent code processing and feature decoding are performed based on a diffusion generation algorithm to complete the replacement of the handheld object. This effectively overcomes the limitations of relying on single visual features, improving the semantic accuracy and scene adaptability of the replacement. By guiding the generation process in stages through time steps, invalid iterations are reduced, adapting to the needs of efficient batch creation. At the same time, it ensures that the interaction between the target object and the hand is natural and the edge transitions are smooth after replacement, significantly improving the realism and visual quality of the image.

[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0036] Figure 2 A flowchart illustrating the handheld item replacement method provided in this application, as shown below. Figure 2 As shown, the method includes:

[0037] Step S201: Obtain the original display image and the target item image.

[0038] The target item image is an image containing only the target item; the item held by the person in the original display image is a different item from the target item, and is referred to as the initial item.

[0039] The original display image is stock footage showing a person holding an initial object. There is an overlap between the person and the initial object in the image to represent the hand's gripping contact with the initial object.

[0040] Users can provide the original display image directly, or a frame can be extracted from a live-action product promotion video provided by the user as the original display image.

[0041] The people in the original display images can be real people, such as relevant personnel from the seller of the initial items, models cooperating with the seller, etc., or they can be virtual characters, and the images of all related characters are authorized.

[0042] The target item image can be provided directly by the user, or the user can provide a photograph of the target item, from which the target item image can be extracted.

[0043] The user terminal uploads the original display image, the target item image, and the associated information of the target item required to drive the digital human to the cloud. The product replacement system and the digital human driving system deployed in the cloud execute the subsequent related steps to generate the target display image and drive the digital human, resulting in a digital human video used to display the target item.

[0044] Step S202: Identify the handheld item in the original display image and generate an item mask.

[0045] The item mask is used to represent the position of the handheld item, i.e., the initial item, in the original display image. Pixels with a value of 1 in the item mask belong to the handheld item, while pixels with a value of 0 do not belong to the handheld item.

[0046] By using pre-trained image segmentation models, object detection models, etc., the handheld objects in the original display image can be identified, and the original display image can be segmented at the pixel level based on the recognition results. A binary object mask can be generated based on the segmentation results.

[0047] First, key hand points in the original display image can be detected, and the hand-held region can be determined based on the detected key hand points. Then, the hand-held object can be identified within the hand-held region, and an object mask can be generated. For example, threshold segmentation or edge detection can be performed on the hand-held region to extract the object contour within the hand-held region, and an object mask can be generated based on the object contour.

[0048] Among them, the key points of the hand include the base of the palm, the base of each finger, the knuckles, and the fingertips, usually totaling 21 key points.

[0049] Step S203: Extract the visual features of the original display image, the item mask, and the target item image, as well as the semantic features of the item prompt words.

[0050] The item prompt can be either a text prompt or a visual prompt. It can be descriptive text for the initial item or a reference image of the initial item, i.e., an image of the initial item. The visual features of the item mask are the visual features of an image cropped from the original display image based on the item mask.

[0051] Visual features include, but are not limited to: contour features, edge gradient features, texture detail features, color distribution features, spatial occupancy features, and interaction type features with the hand area.

[0052] After acquiring the original display image, the complexity of the handheld object in the image can be identified. Based on the complexity, the type of object prompt can be determined, and guidance information can be generated and displayed to facilitate users providing corresponding object prompts. The complexity of the handheld object is used to characterize the complexity of its shape, pattern, etc.

[0053] For handheld items with low complexity, such as those below the first level of complexity, the item cues may consist only of text cues, such as "red coffee cup." For handheld items with high complexity, such as those greater than or equal to the first level of complexity, the item cues may include visual cues, such as a reference image of the handheld item, as well as text cues. Alternatively, the complexity can be divided into three consecutive intervals, such as a first interval, a second interval, and a third interval, using two thresholds, such as a first threshold and a second threshold. The item cues corresponding to the first interval consist only of text cues, the item cues corresponding to the second interval consist only of visual cues, and the item cues corresponding to the third interval consist of both text cues and visual cues.

[0054] Visual features of the original display image, object mask, and target object image can be extracted using any image adaptation feature extraction layer. Examples include CNN (Convolutional Neural Network), ResNet (Residual Neural Network), MobileNet (Mobile Neural Network), VAE (Variational Autoencoder), and Transformer.

[0055] The original display image, the item mask, and the target item image can be input into the VAE encoder, and the VAE encoder will output the visual features of the original display image, the item mask, and the target item image respectively.

[0056] The semantic features of item cues can be extracted using an encoder that matches the type of the cue. For text-based item cues, features can be extracted using models such as BERT (Bidirectional Encoder Representations from Transformers), CLIP-text, and RoBERTa (Robustly Optimized BERT Pretraining Approach) to obtain semantic features. For visual item cues, features can be extracted using models such as CLIP-vision, ViT (Vision Transformer), and ResNet to obtain semantic features.

[0057] An encoder for extracting semantic features of text-based item prompts can be a dual encoder consisting of a T5 (Text-To-Text Transfer Transformer) encoder and a CLIP (Contrastive Language-Image Pre-training) text encoder.

[0058] For example, the CLIP text encoder extracts the visual semantics corresponding to the item prompts, such as the color distribution corresponding to red, and outputs a 1×1×768-dimensional vector. The T5 encoder parses the long text logic of the item prompts, such as the shape description of a coffee cup, and outputs a 1×1×1024-dimensional vector. The two vectors are fused through a projection layer to obtain a 1×1×1280 (768+512)-dimensional semantic vector, which is the semantic feature.

[0059] Step S204: Perform multimodal fusion processing on visual features and semantic features and embed time steps to obtain a global feature sequence.

[0060] Visual and semantic features can be aligned across modalities and then fused to obtain fused features. Time steps are then embedded into the fused features to obtain a global feature sequence. Cross-attention can be used for fusion.

[0061] A projection layer can be used to map visual features or semantic features to the feature space where the semantic features or visual features reside, thereby achieving cross-modal alignment between the two; alternatively, two independent projection layers can be used to map visual features and semantic features to a common feature space, respectively, thereby achieving cross-modal alignment between the two.

[0062] Before fusion, the visual features can be adjusted in posture based on the grasping parameters indicated by semantic features, so that the body posture of the person and the position of the key points of the hand in the visual features are adapted to the grasping parameters.

[0063] Time step information can be converted into a time step vector through sinusoidal position encoding, and the time step vector can be concatenated with the fused features to achieve time step embedding.

[0064] Optionally, the visual and semantic features are fused in a multimodal manner and embedded with time steps to obtain a global feature sequence, including: adjusting and fusing the original feature map and mask feature map based on the target item feature map and semantic features to obtain a semantically guided feature map; the target item feature map, the original feature map, and the mask feature map are the extracted visual features of the target item image, the visual features of the original display image, and the visual features of the item mask, respectively; the time step vector converted from the time step information is concatenated with the semantically guided feature map to obtain a temporal conditional feature map; attention encoding is performed on the temporal conditional feature map to obtain a global feature sequence.

[0065] Based on the feature map of the target item, the size information of the target item can be determined. Based on the size information of the target item, the original feature map can be adjusted so that the hand, or hand, shoulder, elbow, etc. of the person in the original feature map are adapted to the target item, so that the gripping posture of the person's hand matches the posture and size of the target item.

[0066] Based on semantic features, the original feature map and the mask feature map can be offset so that the posture of the held item in the original feature map and the mask feature map matches the posture of the target item indicated in the semantic features. For example, if the original display image shows product A (held item) at a 10° angle, and the item prompt requires product B (target item) to be held at a 40° angle, product A can be rotated 30° from its original 10° angle to match the state described in the item prompt.

[0067] After the above adjustments and offset processing, the processed original feature map and mask feature map are fused with semantic features, which can be achieved using a cross-attention mechanism.

[0068] The original feature map and the mask feature map can be fused by a projection layer to map the channel dimension. After the channel dimension mapping is completed, they are concatenated and then fused with semantic features using a cross-attention mechanism to obtain a semantically guided feature map.

[0069] The time step information can be converted into a time step vector with the same channel dimension as the semantically guided feature map through sinusoidal position encoding, and then the time step vector and the semantically guided feature map can be concatenated in the channel dimension to generate a temporal conditional feature map.

[0070] The time step information includes the diffusion step number t during subsequent diffusion generation and the noise scheduling method corresponding to different diffusion step numbers. The diffusion step number t can be any positive integer between 0 and 1000, or in the case of stream-based implementation, the diffusion step number t can be between 0 and 1, such as 0.5, 0.3, etc.

[0071] During training and inference, the diffusion model follows the noise scheduling method described above, applying noise forward or backward to the image at each diffusion step. The diffusion step number t guides the model on how to inject noise at different diffusion steps. A higher diffusion step number t injects more noise into the image; when t reaches its maximum (e.g., t=1 or 1000), the image is purely noisy. For backward denoising, a lower diffusion step number t injects less noise; when t reaches its minimum (e.g., t=0), the image is clear. The diffusion step number t and its corresponding noise scheduling method can be pre-configured. During model training, the diffusion step number t is randomized, allowing the model to learn the noise adjustment method corresponding to different diffusion step numbers t. During model inference, i.e., backward denoising, the diffusion step number t iterates from high to low.

[0072] After obtaining the temporal conditional feature map, attention encoding is used to serialize the two-dimensional feature map into a global feature sequence.

[0073] Attention encoding can include spatial attention and channel attention encoding. Spatial attention encoding focuses on the areas of hand-object interaction in the temporal conditional feature map, assigning these areas higher feature weights while weakening the weights of irrelevant areas such as the background to reduce information interference. Channel attention encoding filters out the more representative channel information in the temporal conditional feature map, such as channels representing object outlines and channels representing hand-object interaction, strengthening the feature representation of these channels and suppressing the influence of other channels.

[0074] Step S205: Using the diffusion generation algorithm, perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item.

[0075] The target item is the item in the target item image.

[0076] The diffusion generation algorithm gradually adds noise to the image through a forward diffusion process until the image becomes pure noise. Through this process, it learns the degradation rules from clear images to images with different levels of noise, providing a basis for the reverse denoising process. Through the reverse denoising process, the image is restored starting from the pure noise image.

[0077] The visual features, semantic features, and temporal step information integrated in the global feature sequence are converted into a conditional guidance vector for the diffusion model. This vector is embedded in each temporal step of the reverse denoising process to clarify the optimization direction of the latent code at the current temporal step, so that the target item matches the hand posture and interaction position in the original display image, while also meeting the semantic requirements of the item prompt words.

[0078] In the latent code processing stage, the latent code is first initialized to generate a random noise latent code that matches the size of the target region, which is the area where the handheld object is located. Then, at each time step, the gradient correction vector of the noise latent code at the current time step is calculated using the conditional constraints provided by the global feature sequence, driving the latent code to gradually remove noise and obtain the target latent code. The target latent code is then input into a pre-trained decoder for decoding to obtain the target display image.

[0079] The handheld item replacement method provided in this embodiment addresses the problems of traditional handheld item replacement methods that rely solely on visual features, have simple replacement logic, and are prone to abrupt edge transitions and unreasonable hand-object interaction logic. It adopts a collaborative processing mode that integrates visual and semantic multimodal features and uses diffusion generation guidance. By simultaneously extracting visual features from the original display image, item mask, and target item image, as well as semantic features from item prompts, it overcomes the limitations of single visual features. The multimodal features are fused and time-step information is embedded to obtain a global feature sequence that combines scene constraints and stage guidance. Finally, using a diffusion generation algorithm, latent code processing and feature decoding are performed based on the global feature sequence to achieve accurate adaptation of the target item with the original image's hand and background, efficiently completing the handheld item replacement. This ensures natural interaction between the target item and the hand after replacement, smooth edge transitions, and improves the visual quality and realism of the image after item replacement.

[0080] After obtaining the target display image, the edge breakage of the target object after image fusion can be eliminated by using the Poisson fusion model to ensure a natural transition between the target object and the edge of the hand.

[0081] The target object image can be deformed based on the second object mask, such as through pose adjustment and scaling, to ensure that the pose and size of the target object are consistent with the handheld object in the original display image. Then, based on the deformed target object image, image fusion is performed between the target object image and the original display image to obtain the target display image.

[0082] Optionally, obtaining an image of the target item includes: acquiring a photograph of the target item obtained by taking a picture of the target item; removing the background from the photograph of the target item to obtain an image of the target item.

[0083] The target item photo can be a photo taken by the user of the target item in a preset background, which can be a natural background, a green screen background, or other backgrounds.

[0084] Users upload photos of the target object to the cloud via their user terminals. The cloud then removes the background from the photos, retaining only the target object portion, and stores the resulting image.

[0085] Removing the background from a photo of a target object can be done using any algorithm, such as semantic segmentation models based on deep learning, threshold segmentation algorithms, chroma key matting algorithms based on green screen backgrounds, etc. This application does not limit the specific algorithms used.

[0086] By having users provide original photos of the target item and removing the background in the cloud, the user-side operation process is greatly simplified, lowering the barrier to obtaining target item images. Users simply need to take photos of the target item against any preset background, such as a natural background or a green screen, and upload them to the cloud via their devices, eliminating the need for manual background processing. The complex background removal process is automatically handled by the cloud, ultimately outputting a clean image that retains only the target item. This saves users time and effort in manually cutting out and editing images, and avoids image defects caused by unprofessional operations.

[0087] Optionally, the original display image is obtained, including: calculating the similarity between the target item image and the handheld item in multiple display images; and determining the display image with the highest similarity as the original display image.

[0088] If there is only one display image, then there is no need to perform the aforementioned steps; the display image can be used directly as the original display image for subsequent processing.

[0089] If there are multiple display images and the display items (i.e. the items being held) in different display images are different, in order to simplify the complexity of subsequent calculation steps and improve the visual naturalness of the image after object replacement, the original display image used to synthesize the target display image can be determined from multiple display images by the similarity between the items.

[0090] The similarity can be calculated by extracting the shape features of the target item image and the items in each displayed image, and then using the shape features extracted from the two images.

[0091] Each image (including the target item image and the display image) can be tagged with items. The similarity between items in two images can be determined by the item tags. Alternatively, the display images can be filtered based on the item tags to obtain candidate display images whose items are similar to the target item type. Then, based on the extracted shape features of the items in the images, the similarity between the target item in the target item image and the display items in the candidate display images can be calculated, and the candidate display image with the highest similarity is selected as the original display image.

[0092] Figure 3 for Figure 2 The flowchart of step S204 in the illustrated embodiment is as follows: Figure 3 As shown, step S204 may specifically include the following steps:

[0093] Step S301: Determine the grasping parameters of the target object based on the target object feature map.

[0094] The gripping parameters can include information such as the parameters of the gripping region and the region mask. The gripping region is the area of ​​the target object indicated by the target object feature map that is suitable for gripping. The parameters of the gripping region can include the center coordinates, coordinate range, width, thickness, and area of ​​the gripping region.

[0095] Based on the feature map of the target object, the type and size information of the target object can be identified, and the gripping parameters of the target object can be determined based on the type and size information of the target object.

[0096] Optionally, based on the target item feature map, the grasping parameters of the target item are determined, including: performing target detection on the target item feature map to obtain the item boundary and size information of the target item; segmenting the target item feature map based on the item boundary and size information of the target item to obtain multiple segmented regions; calculating the region parameters of each segmented region, and determining the grasping region from the multiple segmented regions based on the region parameters; and determining the grasping parameters of the target item based on the determined grasping region.

[0097] The regional parameters of the segmented region include, but are not limited to, information such as area, location, and thickness.

[0098] An object detection network can be used to detect objects in the feature map of an object, obtain the boundary of the object, and output the size information of the object.

[0099] The boundaries of the target item can include the coordinates of its center point and the range of those coordinates. Dimensional information can include the length, width, and height of the target item. It can also be based on a pre-defined item category database, identifying the target item's category to determine its thickness. Thickness information characterizes the overall thickness of the target item or the overall thickness of its different components, such as the diameter of a water cup. Alternatively, when the water cup includes a body and a handle, the thickness information can include the thickness of the body and the handle, for example, a body thickness of 10cm and a handle thickness of 2cm.

[0100] Lightweight networks can be used for object detection, such as YOLOv8 (You Only Look Once version 8), YOLOX-Nano (You Only Look Once X Nano), and EfficientDet-Lite (Efficient ObjectDetector Lite).

[0101] Image segmentation networks can segment a target object's feature map based on its boundary and size information, resulting in multiple segmented regions. Specifically, the target object's feature map, boundary, and size information are input into the image segmentation network. The network then partitions the feature map into regions corresponding to the coordinates of the object's boundary, outputting multiple segmented regions. Each region corresponds to a different part of the target object, such as the upper, middle, or lower part.

[0102] Image segmentation networks can employ FCN (Fully Convolutional Networks), SAM (Segment Anything Model), Mask R-CNN (Mask Region-based Convolutional Neural Networks), and others.

[0103] After obtaining multiple segmented regions, to determine the most suitable gripping area for the target object, the first step is to define the regional parameters of each segment, such as its area, position, and thickness. The area indicator represents the proportion of the segment's area to the total area of ​​the target object. The position indicator can be represented by the distance between the center of the segment and the center of the target object. The thickness indicator can be characterized by the matching degree between the thickness information of the segment, such as its diameter, and the corresponding gripping range. A mapping relationship between object type and gripping range can be pre-established. For example, thin objects like lipstick correspond to a gripping range of 2-5cm, while a thermos cup corresponds to a gripping range of 10-25cm. For objects with complex structures, different gripping ranges can be set for different areas. The gripping ranges corresponding to various types of objects are pre-collected and stored to facilitate the determination of the target object's thickness information.

[0104] The values ​​for area, position, and thickness can all range from 0 to 1. The larger the area of ​​the target item, the larger the area index; the closer the target item is to its center, the larger the position index; and the higher the degree of matching with the corresponding grip area, the larger the thickness index. For example, when the target item is within the corresponding grip area, the thickness index can be 1.

[0105] The gripping area can be determined from multiple segmented regions based on the weighted results of area, location, and thickness metrics.

[0106] Each segmented region corresponds to a feature vector composed of area, location, and thickness metrics. The feature vector for each segmented region is input into an MLP (Multi-Layer Perceptron) network for scoring. The MLP network determines the weights of different features in the feature vector based on the target object type, and the score for the corresponding segmented region is obtained by weighting the three features. The segmented region with the highest score is identified as the grasping region, and its grasping parameters are output, such as the center coordinates, coordinate range, width, thickness, area, and region mask.

[0107] Step S302: Generate a hand feature map based on the gripping parameters and the original display image.

[0108] Among them, the hand feature map is used to characterize the hand features of the target person when grasping the target object, and the target person is the person in the original display image.

[0109] Based on the grasping parameters, the corresponding hand key points can be determined, and then the hand key points and the body key points of the target person identified from the original display image can be used to generate a hand feature map.

[0110] The key points of the hand are the core positioning points of the palm and fingers, including the base of the palm, fingertips, finger roots, and the joints of the fingers, usually 21 key points.

[0111] Optionally, a hand feature map is generated based on the grasping parameters and the original display image, including: determining the key points of the hand that are adapted to the target object based on the grasping parameters and the preset hand model; obtaining the key points of the person's body in the original display image based on the human posture constraint model, and correcting the key points of the hand based on the key points of the body; mapping the corrected key points of the hand to two-dimensional pixel coordinates, and generating a hand feature map after interpolation.

[0112] Among them, the preset hand model is a standardized digital model that integrates the physiological structural parameters of the human hand and the mechanical parameters of grasping motion. It can accurately represent the skeletal structure, joint freedom of movement, and surface morphological features of the hand. The human posture constraint model is a type of algorithmic model used to estimate, represent, and constrain the spatial position and motion pattern of human limbs. Its core function is to extract the coordinates and three-dimensional morphological information of key human joints such as shoulders, elbows, wrists, knees, and ankles from images, and to apply reasonable geometric constraints to human posture based on this information.

[0113] For example, the preset hand model can be MANO (Metric-Affine Hand Model). Using the MANO model, hand key points are calculated based on the gripping parameters obtained in the previous step, and 21 hand key points that are adapted to the gripping parameters are output.

[0114] After identifying 21 hand key points that are compatible with the gripping parameters, the hand key points are corrected based on the body features of the target person in the original display image using a human posture constraint model such as the OpenPose model.

[0115] The original display image or the target person image cropped from the original display image is input into the OpenPose model to calculate 17 body joint key points. Based on the constraints of the body joint key points on the hand key points, the hand key points are adjusted to avoid incoordination between the hand and the body during the hand pose bar process.

[0116] Specific constraint methods can include elbow flexion constraints and body distance constraints. The elbow flexion constraint involves calculating the shoulder-elbow vector based on key body joint points, and simultaneously calculating the elbow-hand vector based on key body joint points and hand key points. The angle between these two vectors (the elbow flexion angle) is then calculated. This elbow flexion angle should be between 30° and 120°. If the calculated result is outside this range, the elbow flexion is too small or too large, requiring correction of the hand key points to ensure the elbow flexion angle falls within the range of [30°, 120°]. The body distance constraint involves calculating the distance from the hand to the body's central axis based on key body joint points and hand key points. This distance should be between 15 cm and 30 cm. If the calculated result is outside this range, the hand is too close or too far from the body, requiring correction of the hand key points to ensure the distance falls within the aforementioned range. After adjusting the hand key points according to these constraints, ensuring they meet the constraints, the corrected hand key points are obtained.

[0117] Spatial Transformer Networks (STNs) can be used to map the corrected hand keypoints into two-dimensional pixel coordinates. Further interpolation processing, such as bilinear interpolation, can be performed on the mapped two-dimensional pixel coordinates to generate and output a hand feature map.

[0118] Step S303: Based on semantic features, offset the original feature map and the mask feature map.

[0119] Based on semantic features, the region where the handheld item is located in the original feature map and the mask feature map can be offset so that the state of the handheld item after offset conforms to the semantics of the item prompt, so as to facilitate item replacement.

[0120] Semantic features can be used to determine the spatial offset between the original state of the held object in the original feature map and the expected state described in the semantic features. This spatial offset can then be used to offset the original feature map and the mask feature map, such as by translation, rotation, scaling, etc.

[0121] Offset positions may cause spatial continuity breaks, feature sparsity, or overlaps with surrounding positions. Visually, this manifests as blank areas appearing at the edges of product A, or texture breaks between adjacent positions. These problems arise because the surrounding positions are not optimized simultaneously after the original feature map and mask feature map are offset. Therefore, after calculating the offset parameters, sampling prediction parameters also need to be calculated. These parameters characterize the contribution weight of surrounding feature points to the offset feature point (the feature point after offset) during subsequent feature sampling. This allows for weighted feature fusion using surrounding feature points, improving the quality of the offset processing.

[0122] Optionally, based on semantic features, the original feature map and the mask feature map are offset, including: using a cross-attention network, with semantic features as the query and the original feature map as the key, calculating attention weights and assigning them to the original feature map to obtain a semantically guided original feature map; inputting the semantically guided original feature map into an offset prediction network to obtain offset parameters and sampling weight parameters; converting the offset parameters into an offset matrix using a spatial transformation network, and performing spatial coordinate transformation on the original feature map and the mask feature map based on the offset matrix; and using the sampling weight parameters to perform weighted fusion of the offset feature points and surrounding feature points to obtain the offset original feature map and the offset mask feature map; the offset feature points are the corresponding feature points formed by the spatial coordinate transformation of the feature points in the original feature map and the mask feature map.

[0123] Cross-attention networks can extract location-related semantic information from semantic features, such as spatial descriptions like "above," "left," and "middle," or local semantics like "person's hand" and "product edge." Based on these semantics, the corresponding parts in the original feature map are weighted to give higher weights to the location-related features in the original feature map.

[0124] Migration prediction networks are used to predict migration parameters and output sampling weight parameters. Migration prediction networks can consist of pooling layers and multiple fully connected layers, or pooling layers and multiple lightweight convolutional network layers.

[0125] After inputting the semantically guided original feature map into the offset prediction network, it first undergoes global pooling through a pooling layer to compress the features into a global vector. Then, a fully connected layer or a lightweight convolutional network layer maps the global vector to a low-dimensional space and performs a transformation to obtain the offset parameters. Based on the offset parameters, sampling weight parameters are determined. These sampling weight parameters characterize the contribution weights of the feature points surrounding the original feature point (the feature point before offset) to the offset feature point (the feature point after offset) during subsequent feature sampling, i.e., the offset process.

[0126] For example, suppose a feature point M in the original feature map has surrounding feature points M1, M2, M3, and M4. After offsetting, the offset position of M is M'. To avoid the aforementioned issues of feature sparsity, subsequent feature sampling processes need to refer to M1, M2, M3, and M4 for feature fusion. The sampling weight parameter is used to characterize the weights of M1, M2, M3, and M4 respectively during feature fusion.

[0127] The sampling weight parameters can be calculated using bilinear interpolation. The core principle is to assign weights based on the distance between the offset feature point (the feature point being offset) and its surrounding feature points; the closer the distance, the higher the weight. Using the example above, assuming the distances between M´ and M1, M2, M3, and M4 are 4, 3, 2, and 1 respectively, then the weights assigned to M1, M2, M3, and M4 are 0.4, 0.3, 0.2, and 0.1 respectively.

[0128] In the bilinear interpolation calculation process described above, the sampling range (i.e., the number of surrounding feature points selected) can be fixed. A fixed sampling range usually only achieves good results when the shape is simple and the degree of offset is small; for complex shapes and large degrees of offset, a dynamic sampling method can be used, that is, the sampling range is dynamically adjusted based on the offset parameter.

[0129] The sampling orientation of dynamic sampling can be determined based on preset rules, as follows: The larger the translation distance, the larger the sampling range. For example, when the translation distance is < 2 tokens (pixel blocks), the sampling range can take 4 surrounding feature points; when the translation distance is ≥ 2 tokens, the sampling range can take 8 surrounding feature points. The larger the rotation angle, the larger the sampling range. For example, when the rotation angle is < 30°, the sampling range takes 4 surrounding feature points; when the rotation angle is ≥ 30°, the sampling range takes 8 surrounding feature points. When the scaling ratio is greater than or less than a preset value (i.e., the greater the scaling degree), the larger the sampling range. For example, when the scaling ratio is > 1.2 (enlargement) or < 0.8 (reduction), the sampling range takes 8 surrounding feature points; when 0.8 ≤ scaling ratio ≤ 1.2, the sampling range takes 4 surrounding feature points.

[0130] Furthermore, dynamic sampling can be performed based on the location of feature points in the image. For example, if M is located at the edge of product A, the sampling range is increased by 2 points to prioritize covering the original features in the edge direction and avoid edge breakage. If M is located in the solid color area of ​​product A, the sampling range can be reduced by 2 points to reduce redundant calculations. If M is located at the boundary between product A and the background (such as the bottom of a cup contacting the table), the sampling range is fixed to the minimum, such as 4 points, to avoid introducing background features.

[0131] In addition to dynamic sampling based on preset rules, the sampling range can also be calculated using a pre-trained prediction network.

[0132] The offset parameters are converted into an offset matrix using a spatial transformation network, and the original feature map and the mask feature map are then transformed in spatial coordinates based on the offset matrix. Using sampling weight parameters, the offset feature points are weighted and fused with surrounding feature points within the sampling range to obtain the offset original feature map and the offset mask feature map.

[0133] Step S304: Based on the hand feature map, the pose of the offset original feature map is adjusted, and the pose-adjusted original feature map is fused with the offset mask feature map to obtain a joint feature map.

[0134] After obtaining the hand feature map and completing the offset, the offset original feature map is further adjusted using the hand feature map, so that the person's hand and body posture in the original feature map matches the hand feature map, thus completing the posture adjustment.

[0135] The original feature map after pose adjustment and the offset mask feature map can be mapped and concatenated along the channel dimension to obtain a joint feature map.

[0136] Optionally, based on the hand feature map, pose adjustment is performed on the offset original feature map, including: extracting style information and pose information from the offset original feature map; correcting the hand features in the pose information based on the hand feature map; concatenating the corrected hand features with the style information to obtain a concatenated original feature map; generating a product grasping mask and a hand feature mask based on the grasping region; extracting product grasping feature sequences and hand feature sequences from the concatenated original feature map using the product grasping feature sequence as the query and the hand feature sequence as the key; calculating attention weights through a cross-attention network and adjusting the hand feature sequences based on the attention weights, mapping them to the original feature map to obtain the pose-adjusted original feature map.

[0137] The offset original feature map can be input into two parallel convolutional layers to extract style and pose information respectively. Then, the pose information and hand feature map are input into an MLP network for weight allocation. The MLP network, based on dynamic weight allocation rules learned during training, calculates weights for the features representing the hand portion in the pose information and the features represented by the hand feature map, and performs a weighted sum based on these calculated weights to correct the pose information, resulting in corrected pose information. For example, for small or complex objects, which are difficult to hold, a higher weight is assigned to the hand feature map based on the grasping posture to enhance the accuracy of hand grasping; conversely, for regular large items, a lower weight can be assigned to the hand feature map to preserve the consistency between the item and the hand as a whole. Finally, the corrected pose information is concatenated with the aforementioned style information to obtain a concatenated original feature map.

[0138] The aforementioned steps only adjusted the hand posture, but there may still be some details that need optimization when the target object is actually grasped, especially the contact position between the fingers and the target object. Therefore, it is necessary to generate a product grasping mask and a hand feature mask based on the grasping area, so as to extract and stitch the feature map corresponding to the original feature map through the mask, and use a cross-attention network to perform weighted fusion on the extracted feature map to obtain the original feature map in which the target object and the hand maintain a better grasping posture, that is, the original feature map after posture adjustment.

[0139] The product gripping mask is the mask corresponding to the gripping area, and the hand feature mask is the mask corresponding to the hand feature map.

[0140] Based on the product grasp mask and hand feature mask, two feature maps are obtained by stitching together the product grasp and hand parts from the original feature map. Based on these two feature maps, product grasp feature sequences and hand feature sequences are obtained. Using a cross-attention network, the product grasp feature sequence is used as the query object Q, and the hand feature sequence is used as the query object, i.e., key K and value V. Attention weights for the hand feature sequence are calculated based on the product grasp feature sequence, and the hand feature sequence is weighted according to the calculation results. This yields a hand feature sequence corrected based on the product contact position. Mapping this sequence yields the original feature map showing the optimal grip posture between the target item and the hand—the posture-adjusted original feature map.

[0141] First, convolutional layers can be used to map the pose-adjusted original feature map to the target object feature map in terms of channel dimensions, maintaining consistency (e.g., mapping to 512 or 1024 dimensions). The channels of the offset mask feature map can also be mapped to fixed values, such as 64 or 128 dimensions. Alternatively, fully connected layers can be used to ensure the channel dimensions of the semantic features are consistent with those of the target object feature map. The channel-mapped original feature map and the mask feature map are then concatenated to obtain a joint feature map. The channel-mapped target object feature map and semantic features are then output separately to obtain a semantically guided feature map in subsequent steps.

[0142] Step S305: Based on the cross-attention mechanism, the semantic features, target item feature map and joint feature map are interacted to generate a semantically guided feature map.

[0143] First, the joint feature map and the target feature map are concatenated along the channel dimension to obtain a concatenated feature map, which is then mapped to a concatenated sequence. Cross-attention calculation is performed on the concatenated sequence and semantic features. After the calculation is completed, the concatenated sequence is weighted based on the calculated attention weight matrix to obtain a feature map with semantic guidance.

[0144] Optionally, based on a cross-attention mechanism, semantic features and joint feature maps are interacted to generate semantically guided feature maps, including: concatenating the joint feature map and the target item feature map along the channel dimension to obtain a concatenated feature map, and mapping the concatenated feature map to a concatenated sequence; using semantic features as queries and the concatenated sequence as keys, attention weights are calculated through a cross-attention network to obtain an attention weight matrix; based on the attention weight matrix, the concatenated sequence is adjusted to obtain a preliminary fused feature map; the similarity between the original item and the target item in multiple spatial locations in the preliminary fused feature map is calculated; based on the similarity in each spatial location, the weight of the target item in each spatial location is determined, and the weight of the main region corresponding to the item mask is adjusted to obtain a weight vector; based on the weight vector, the features of the original item and the features of the target item in the preliminary fused feature map are weighted and fused to obtain a semantically guided feature map.

[0145] After obtaining the concatenated sequence, an attention weight matrix is ​​calculated using semantic features as the query Q and the concatenated sequence as the key K and value V. Higher weights indicate that the corresponding image region needs to adhere more strictly to the semantic feature constraints. Weights are assigned to the concatenated sequence based on the attention weight matrix, and the features at corresponding positions in the concatenated sequence are weighted according to the assigned weights. The weighted concatenated sequence is then converted into a feature map, resulting in a preliminary fused feature map.

[0146] Since the initial fusion feature map contains features of both the original item and the target item, and the original item is the handheld item in the original display image, the aforementioned fusion process only injected textual conditions, i.e. semantic features, without distinguishing the different emphases in the item replacement process. Therefore, it is still necessary to further adjust the main area of ​​the item based on the similarity between the original item and the target item.

[0147] Let's take product A as the source item and product B as the target item as an example. The weights corresponding to the features of product A are denoted as WA, and the weights corresponding to the features of product B are denoted as WB. For the same position, WA + WB must satisfy 1.

[0148] In the initial fusion feature map, the feature corresponding to product A is denoted as FA, and the feature corresponding to product B is denoted as FB. For FA and FB, the similarity between them at each spatial location is calculated, and a similarity map is generated based on the similarity calculation results. This similarity map is used to represent the similarity between FA and FB at each spatial location. In the above similarity map, the closer the similarity value is to 1, the more similar the features of product A and product B are at that location, and the easier it is to replace product A with product B in that part; conversely, the closer the similarity value is to 0, the greater the feature difference between product A and product B at that location, and the more difficult it is to replace product A with product B in that part.

[0149] Based on the similarity map above, WA and WB can be calculated using the following formula:

[0150] WB(h,w)=α⋅S(h,w)+β; WA(h,w)=1-WB(h,w);

[0151] The above WA(h,w) and WB(h,w) represent the WA and WB of the spatial location (h,w) respectively, S(h,w) characterizes the similarity of a spatial location (h,w), and α and β are learnable parameters obtained through training.

[0152] The above calculations result in the following: regions with larger S values ​​have higher WB values, and regions with smaller S values ​​have lower WB values. Specifically, the more similar product A is to product B, i.e., the larger the S value, the easier it is to perform the replacement task. Replacing product A with product B in this region results in less visual difference, so a higher WB value can be set for this region. Conversely, the greater the difference between product A and product B, i.e., the smaller the S value, the more difficult it is to perform the replacement task. Directly performing the replacement task will result in more obvious visual breaks, so a lower WB value can be set for this region to retain more of product A's features for the transition.

[0153] For example, product A includes a gray metallic cup body and a dark wooden handle, while product B includes a gray plastic cup body and a white plastic handle. After similarity calculation, the feature similarity between product A and product B in the cup body region is 0.8, resulting in a WB of 0.9 and a WA of 0.1. Similarly, the feature similarity between product A and product B in the handle region is 0.2, resulting in a WB of 0.9 and a WA of 0.1. For other regions, the similarity may be moderate, resulting in a WB and WA of 0.5 for each region.

[0154] The similarity between two features can be calculated using cosine distance, Euclidean distance, Manhattan distance, Jaccard similarity coefficient, Pearson correlation coefficient, etc.

[0155] Because the target object itself may have significant differences between adjacent parts, such as the cup body and handle in the previous example, if the WB values ​​of adjacent areas differ too much, it will cause abrupt boundary issues at the junction of adjacent areas. Therefore, further spatial continuity processing is required.

[0156] For the WB of each spatial location or region calculated above, the WB of multiple adjacent locations is smoothed, for example, the average of the WB of these multiple adjacent locations is used as the WB of the center location.

[0157] The main area is the region where the handheld item is located, as indicated by the item mask. For the weight (WB) of the target item in the main area, the WB of each spatial location within the main area can be increased, for example, by 0.2, to ensure the accuracy of the main area replacement. No further adjustments are needed to the WB of spatial locations outside the main area.

[0158] For example, the WB of each spatial location within the main body area can be uniformly set to 0.8, and the larger of the WB of each spatial location within the main body area calculated based on similarity can be taken as the final WB of each spatial location.

[0159] After calculating WA and WB through the aforementioned steps, the weight vectors (WA and WB for each spatial location) are obtained. Based on WA and WB, the features of the original item and the target item in the preliminary fused feature map are weighted and fused to obtain a semantically guided feature map. The channel dimension of the semantically guided feature map is consistent with the channel dimension of the original feature map or the target item feature map.

[0160] Step S306: The time step vector converted from the time step information is concatenated with the semantically guided feature map to obtain the time condition feature map.

[0161] Step S307: Based on the attention layer composed of a multi-head self-attention structure and a normalization layer, extract the visual label sequence and text label sequence from the temporal conditional feature map.

[0162] Temporal conditional feature maps can be input in parallel into the image preprocessing subunit and the text preprocessing subunit. Through differential feature extraction from the two subunits, visual label sequences and text label sequences can be obtained respectively.

[0163] The image preprocessing subunit and the text preprocessing subunit adopt a parallel structure and are used to extract the image stream and text stream corresponding to the aforementioned temporal condition feature map, respectively.

[0164] The image preprocessing subunit consists of an image projection layer and an image self-attention layer connected in series, and is used to capture the dynamic relationship between spatial morphology and time step in the temporal condition feature map.

[0165] The image projection layer consists of a 3×3 convolutional layer, a batch normalization layer, and residual connections. The 3×3 convolutional layer extracts local spatial features and maps dimensions from the temporal conditional feature map. The batch normalization layer stabilizes the feature distribution and accelerates training convergence. The residual connections effectively avoid the vanishing gradient problem in deep networks. After processing by this layer, the two-dimensional temporal conditional feature map is transformed into a one-dimensional sequence of visual tokens. Each visual token corresponds to a spatial location in the feature map and carries the visual features and temporal step information at that location.

[0166] The image self-attention layer consists of a multi-head self-attention structure and a normalization layer. The image self-attention layer allows each token in the visual token sequence to autonomously attend to other spatially adjacent and temporally consecutive visual tokens. Through the weight allocation of multi-head self-attention, it can accurately capture the dynamic relationships between visual features, such as the spatial morphological change of a hand from "open" to "held," and the temporal correspondence between the target object and the contact area of ​​the hand. The normalization layer then standardizes the features after attention calculation to obtain a visual tag sequence.

[0167] The text preprocessing subunit consists of a text projection layer and a text self-attention layer connected in series, and is used to mine the semantic information embedded in the temporal condition feature map and the corresponding relationship between the time steps.

[0168] The text projection layer consists of a 1×1 convolutional layer, a global average pooling layer, and residual connections. The 1×1 convolutional layer is responsible for cross-channel feature fusion and dimensionality compression of the temporal conditional feature map. The global average pooling layer aggregates global feature information and weakens the interference of local spatial noise. The residual connections ensure the integrity of feature transmission. After processing by this layer, the semantic information carried in the temporal conditional feature map is extracted and transformed into a one-dimensional text token sequence. Each text token corresponds to a set of semantic features and is associated with the corresponding time step information.

[0169] The text self-attention layer is similar to the image self-attention layer, both consisting of a multi-head self-attention structure and a normalization layer. The core function of the text self-attention layer is to allow each token in the text token sequence to autonomously attend to other semantically related and temporally connected text tokens. Through attention weight allocation, temporal associations of semantic features can be established, such as the matching relationship between the attribute semantics of the text prompt "red ceramic cup" and the hand-holding action at different time steps, and the temporal correspondence between the target item style semantics and scene features. The normalization layer then standardizes the semantic features after attention calculation, resulting in a text token sequence.

[0170] It is important to emphasize that the above preprocessing sub-unit approach does not simply separate the fused visual, semantic, and temporal features in the temporal conditional feature map. Instead, it achieves targeted feature extraction through differentiated weight parameters designed for the two sub-units. For the image preprocessing sub-unit, its weight parameters are optimized for visual information and spatiotemporal correlation, focusing on preserving the spatial location, morphological details, and temporal step correlation of features. For the text preprocessing sub-unit, its weight parameters are optimized for semantic information and temporal correlation, strengthening the matching relationship between feature category attributes, style descriptions, and temporal steps.

[0171] Step S308: The visual marker sequence and the text marker sequence are aligned and fused across modalities through dual-path cross-attention calculation.

[0172] The dual-path cross-attention calculation includes attention weight calculation with visual tag sequences as queries and text tag sequences as keys, as well as attention weight calculation with text tag sequences as queries and visual tag sequences as keys.

[0173] Visual tag sequences and text tag sequences can be input into the dual-stream cross-attention subunit for dual-path cross-attention calculation.

[0174] The dual-stream cross-attention subunit consists of multiple layers of dual-stream attention blocks. Each layer of dual-stream attention block includes a dual-path cross-attention subunit and two sets of post-processing subunits (corresponding to the visual tag sequence and the text tag sequence, respectively).

[0175] Each dual-stream attention block includes two independent processing paths. Path 1 (text path) uses the visual marker sequence as the query object Q and the text marker sequence as the query object (key K and value V), calculating a text weight matrix. This text weight matrix represents the weight allocation of the text marker sequence based on the visual marker sequence, and outputs a weight-adjusted text marker sequence. Path 2 (image path) uses the text marker sequence as the query object and the visual marker sequence as the query object, calculating an image weight matrix. This image weight matrix represents the weight allocation of the image marker sequence based on the text marker sequence, and outputs a weight-adjusted visual marker sequence.

[0176] After the above calculations are completed, the original visual label sequence is weighted to obtain the updated visual label sequence, and the original text label sequence is weighted to obtain the updated text label sequence. The updated visual label sequence and the updated text label sequence are then input into their respective post-processing subunits, namely the image post-processing subunit and the text post-processing subunit.

[0177] The image post-processing subunit and the text post-processing subunit have the same structure but different and independent parameters. Both include a normalization layer, a feedforward network layer, and a gating processing layer.

[0178] In the image post-processing subunit, the normalization layer can use AdaLN (Adaptive Layer Normalization) to dynamically normalize the image features in the updated visual label sequence based on their distribution characteristics (such as strong spatial correlation and stable numerical range), and incorporate the aforementioned time steps to adapt to the needs of the diffusion generation stage, such as strengthening the global structure in the early stage and preserving details in the later stage.

[0179] Feedforward network layers are used to enhance the spatial correlation of image features (e.g., the positional dependence of the digital human arm and the goods, the continuity of background texture, etc.) and capture details (such as light and shadow gradation, material differences) through nonlinear transformations.

[0180] Structurally, the feedforward network layer employs a structure of fully connected layers, non-linear activation functions (such as GELU and ReLU), and fully connected layers. The first fully connected layer maps the normalized labeled sequence to a higher dimension, expanding the feature representation space. Then, the activation function introduces non-linearity, enabling the model to learn complex feature relationships. Finally, the second fully connected layer maps the high-dimensional features back to the original dimension.

[0181] The gating processing layer learns the gating parameters and dynamically adjusts the relative weights of the initial label and the label after the aforementioned series of processing for each label. Based on the relative weights, it outputs the final label of the dual-stream attention block, thereby avoiding excessive influence between text features and image features during the dual-stream cross-attention calculation process.

[0182] The text post-processing subunit is the same as above. The normalization layer uses AdaLN to optimize the distribution consistency of text features in the text feature sequence, avoiding interference with interaction due to excessively large feature amplitudes of individual high-frequency words. The feedforward network layer focuses on strengthening the semantic logic of text features, capturing abstract semantics in language through nonlinear transformations. The structure of the feedforward network layer and subsequent gating processing layers is the same as above and will not be described again.

[0183] After the visual and text tag sequences are processed by their respective post-processing subunits, the calculation and output of the current layer's dual-stream attention block are completed. Simultaneously, this process is input into the next layer's dual-stream attention block, repeating the above steps to progressively deepen the cross-modal alignment of the image and text. This process is repeated iteratively through multiple layers of dual-stream attention blocks, ultimately outputting a deeply interactive visual feature sequence and a processed text feature sequence.

[0184] Step S309: Perform three-dimensional rotation position encoding on the processed visual feature sequence to obtain an image feature sequence with position information.

[0185] The image feature sequence (i.e., the processed visual feature sequence) output by the dual-stream cross-attention subunit is input into the 3D rotation coding subunit. The 3D rotation coding subunit consists of one or more rotation position coding layers, which are used to embed each feature block (token) in the above image feature sequence with position encoding based on rotation position coding (RoPE).

[0186] The output of the 3D rotation coding subunit is an image feature sequence containing positional information, i.e., an image feature sequence with positional information. In this sequence, the feature vector of each position already contains the positional information in 3D space, providing a feature representation with positional information for subsequent single-stream self-attention calculation.

[0187] Step S310: Generate semantic constraints based on the processed text feature sequence.

[0188] The processed text feature sequence can be converted into semantic constraints, such as a global semantic condition vector, through pooling, which can then be used as constraints for subsequent single-stream self-attention computation.

[0189] Step S311: Input the image feature sequence with location information into the attention layer composed of a multi-layer multi-head self-attention structure and a normalization layer. Under the constraints of semantic constraints, the attention weight is calculated, and after residual connection, feedforward network and layer normalization processing, the global feature sequence is output.

[0190] Image feature sequences with location information can be input into a single-stream self-attention subunit, which then outputs a global feature sequence under semantic constraints.

[0191] The single-stream self-attention subunit consists of multiple layers of single-stream self-attention blocks. Each single-stream self-attention block includes a self-attention layer and a feedforward network layer. The self-attention layer adopts a multi-head self-attention structure. After the image feature sequence with location information is input, it first passes through multiple linear projection layers to convert the image feature sequence with location information into corresponding query (Q) matrices, key (K) matrices, and value (V) matrices. Then, the query (Q) matrix, key (K) matrix, and value (V) matrix are split into multiple attention heads according to the feature dimension. Each attention head corresponds to a Q sub-matrix, K sub-matrix, and value V sub-matrix, respectively. The attention weights of the above multiple attention heads are calculated independently to obtain the corresponding weight matrix, thereby completing the weight allocation for that attention head. After the calculation is completed, the outputs of each attention head are concatenated to obtain the first output sequence of the self-attention layer.

[0192] For the first output sequence, a residual connection is made between it and the input image feature sequence with location information to avoid the loss of feature information in the deep network. The resulting sequence is then normalized through layer normalization to achieve a stable distribution of features, thus obtaining the second output sequence.

[0193] The second output sequence is fed into a feedforward network layer for processing. This feedforward network layer consists of two fully connected layers (FC) and a non-linear activation function (such as GELU or ReLU). The first fully connected layer maps the input second output sequence to a higher dimension, expanding the feature representation space. Then, the activation function introduces non-linearity, allowing the model to learn complex feature relationships. Finally, the second fully connected layer maps the high-dimensional features back to the original dimension, obtaining the third output sequence, which is then output.

[0194] For the third output sequence, a residual concatenation is performed with the second output sequence, and the resulting sequence is subjected to layer normalization to obtain the fourth output sequence, which is the output of the single-stream self-attention block in this layer.

[0195] The fourth output sequence is input into the next layer of single-stream self-attention block, and the above operation is repeated. After iterative calculation through multiple layers of single-stream self-attention blocks, the final output feature sequence is the global feature sequence finally output by the attention encoding unit.

[0196] In the aforementioned attention calculation process, the processed text feature sequence output by the aforementioned dual-stream cross-attention subunit is no longer used as input, but is only retained as a constraint condition to perform semantic constraints during the attention calculation process of the aforementioned image feature sequence with location information.

[0197] Optionally, the semantic constraints are modulated by parameters generated by the first fully connected network to adjust the parameters of the linear projection layers corresponding to the queries, keys, and values ​​in the attention layer; or, the semantic constraints are supervised by a supervision matrix generated by the second fully connected network to be multiplied with the weight matrix output by at least one attention layer.

[0198] First, the processed text feature sequence is converted into a global semantic conditional vector through pooling. This global semantic conditional vector then supervises the single-stream self-attention computation of the image feature sequence with location information in the following two ways:

[0199] Method 1: The global semantic conditional vector is used to generate a parameter modulation factor through a first fully connected network. This parameter modulation factor is then input into the projection parameters of the linear projection layers corresponding to Q, K, and V, respectively, for interaction. For example, the parameter modulation factor is multiplied or added to the projection parameters to complete the conditional injection into the linear projection layers. Subsequently, when the image feature sequence with location information is input into the linear projection layers, textual semantic preferences will be assigned to the generated results of Q, K, and V. For example, if the text describes a product tilted 30° clockwise, after the above processing, the Q generated by the image feature sequence with location information will pay more attention to the relative angle of the product. In subsequent attention calculations, the feature blocks (tokens) that match the angle will have higher weights.

[0200] Method 2: Generate a supervision matrix from the global semantic condition vector through a second fully connected network. Multiply the supervision matrix with the weight matrix output by the self-attention layer in at least one single-stream self-attention block. Use the result as the basis for weight allocation of image feature sequences with location information by that layer.

[0201] The aforementioned supervision matrix is ​​used to calibrate the weight vector of a single-stream self-attention block. For example, if the text description includes a red apple, the generated supervision matrix indicates that the red and apple parts should be given priority. In the matrix calculation, higher values ​​(such as 0.9) are assigned to the apple region and the color region in the original weight vector, while lower values ​​(such as 0.1) are assigned to other regions not described in the text.

[0202] In the above operation, the weight matrix of each single-stream self-attention block can be multiplied with the supervision matrix (i.e., each layer is supervised), or specific layers can be selected for supervision. Typically, two to three layers of single-stream self-attention blocks in the shallow layer (usually the first 1 / 3 to 1 / 2 of the attention blocks) and two to three layers in the deep layer (usually the last 1 / 3 to 1 / 2 of the attention blocks) can be selected for multiplication with the supervision matrix, while the remaining single-stream self-attention blocks can retain the original weight matrix.

[0203] The two methods mentioned above can be used individually or simultaneously. The effect of both is to ensure that the region corresponding to the text description is given priority during the single-stream self-attention calculation process.

[0204] The attention encoding module uses a combination of a two-stream attention module and a single-stream self-attention module. The number of layers in the two-stream attention block and the single-stream self-attention block is typically 12 to 32. The aforementioned global feature vectors serve as conditions for generation in the subsequent diffusion generation process, thus imposing mandatory task constraints on the generation process.

[0205] In this implementation, during the semantically guided feature map generation process, the grasping parameters are first determined based on the target object feature map, and a hand feature map is generated to ensure the interaction adaptability between the hand and the target object. Then, the original feature map and the mask feature map are offset and adjusted through semantic features, and the pose calibration and feature fusion are completed by combining the hand feature map, so that the joint feature map not only meets the semantic requirements but also conforms to the geometric logic of hand-object interaction. Finally, the cross-attention mechanism is used to realize the deep interaction between semantic features, target object features, and joint features, so that the generated semantically guided feature map simultaneously meets the multiple requirements of visual form adaptation, semantic attribute matching, and natural hand-object interaction, which greatly improves the generation quality of subsequent target display images and avoids defects such as object floating and hand-object incoordination. For the attention encoding process of temporal conditional feature maps, a multi-head self-attention structure is used to separate and extract visual and textual label sequences. Then, cross-modal accurate alignment is achieved through dual-path cross-attention. This ensures the preservation of both visual spatiotemporal correlation and semantic temporal correlation. Furthermore, the positional information of the features is enhanced through three-dimensional rotational position encoding. Combined with multi-layer attention optimization under semantic constraints and the collaborative processing of residual connections and feedforward networks, the final output global feature sequence has spatial accuracy, semantic consistency, and temporal coherence. This provides a comprehensive and accurate constraint basis for subsequent diffusion generation and effectively avoids problems such as spatial misalignment and semantic deviation in the generated results.

[0206] Figure 4 for Figure 2 The flowchart of step S205 in the illustrated embodiment is shown. Step S205 can be executed by the diffusion generation unit, such as... Figure 4As shown, step S205 may specifically include the following steps:

[0207] Step S401: Perform dimensional compression processing on the global feature sequence according to the preset latent code channel dimension to obtain the latent code feature vector.

[0208] For example, the preset latent code channel dimension can be 4 channels.

[0209] The dimension of the input global feature sequence can be compressed using a linear projection layer, and the resulting latent code feature vector is obtained after dimensionality reduction.

[0210] Step S402: Perform sequence recombination on the latent code feature vector to obtain the latent code feature map.

[0211] The latent code feature vectors can be reassembled using a sequence recombination layer to obtain a latent code feature map. Since the latent code feature vectors obtained after dimensionality reduction of the global feature vectors are still one-dimensional feature sequences, subsequent processing needs to be performed according to a two-dimensional spatial feature map. Therefore, the latent code feature vectors need to be reassembled according to spatial arrangement to obtain a two-dimensional latent code feature map.

[0212] Step S403: Perform local channel weighted fusion on the latent code feature map to generate a conditional latent code map that matches the time step.

[0213] A local fusion layer can be used to perform local channel-weighted fusion of the feature values ​​of multiple channels corresponding to each spatial location in the latent code feature map, thereby avoiding information conflicts within channels and making the overall image transition smoother. The latent code feature map after local fusion is called the conditional latent code map.

[0214] Since the aforementioned global feature vector embeds time step information, the conditional latent code map obtained from the global feature vector can represent the conditional latent code map corresponding to different time step information, that is, the ideal state of denoising the noisy latent code at a certain time step. Based on this, the gradient correction vector can be predicted by calculating the difference between the conditional latent code map and the noisy latent code at the current time step. The gradient correction vector mathematically represents the rate of change of the latent code per unit time. In the application scenario of this embodiment, it represents how to correct the noisy latent code at the current time step between two adjacent time steps so that it conforms to the optimal path of the conditional latent code map corresponding to the current time step.

[0215] Step S404: Initialize and generate a random noise latent code map with the same dimension as the conditional latent code map, as the noisy latent code map for the first time step.

[0216] Step S405: Predict the gradient correction vector for the next time step by differentiating the conditional latent code map with the noisy latent code at the corresponding time step.

[0217] Step S406: Based on the predicted gradient correction vector, perform iterative updates of the noisy latent code map to obtain the target latent code map.

[0218] The latent code update subunit is used to update the noisy latent code at the current time step according to the aforementioned gradient correction vector, and generate the noisy latent code for the next time step. The noisy latent code is updated as follows: Z(t−1) = Zt + ΔZ×Δt; where Z(t−1) represents the latent code for the next time step; Zt represents the noisy latent code at the current time step; ΔZ represents the gradient correction vector predicted at the current time step; and Δt represents the interval between two adjacent time steps.

[0219] Based on the above calculations, through iterative processing of the noisy latent code, the noisy latent code can be corrected along the gradient correction vector direction at each step of the diffusion generation process, gradually transforming the image from pure noise into a clear image. After the iteration is completed, the latent code update subunit outputs the target latent code map, which is the latent code representation of the target display image where the handheld object is replaced with the target object.

[0220] Step S407: Decode the target latent code image to obtain the target display image.

[0221] By inputting the target latent code image into a decoder such as a VAE decoder, the target display image can be obtained.

[0222] In this embodiment, the global feature sequence is condensed into a conditional latent code map carrying visual, semantic, and spatiotemporal constraints through linear projection dimensionality reduction, sequence recombination, and local channel weighted fusion. This clarifies the ideal target for denoising at each time step, avoiding blind exploration of the latent code in high-dimensional space. Secondly, the gradient correction vector is calculated by differentiating the conditional latent code map from the current noisy latent code, providing a precise correction path for iteration. This ensures that each update efficiently approaches the target state, reducing ineffective iterations. Furthermore, the global feature sequence has pre-integrated key information such as hand posture, target object shape, and semantic description, providing a clear framework for the diffusion process. Iterations only require fine-tuning within this framework, significantly reducing computational complexity. Strong constraints are applied throughout the process, replacing unguided exploration with precise priors, ensuring a clear direction for each step of the latent code iteration, ultimately reducing time steps and resource consumption.

[0223] Optionally, the target latent code is decoded to obtain the target display image, including: inputting the target latent code into the decoder to obtain a preliminary display image; and performing edge blending, illumination blending and shadow blending processing on the preliminary display image in sequence to obtain the target display image.

[0224] Edge blending can be achieved using the Poisson blending model. The Poisson blending model is a pixel-level fusion algorithm based on image gradient information to preserve the texture features of the target object while ensuring the target object blends seamlessly with the background, avoiding the abrupt edge distortion caused by traditional fusion methods.

[0225] Edge transition regions are determined based on the object mask. These regions are defined as areas formed by extending a certain number of pixels inwards and outwards from the object edge in the object mask. Gradient weights for the object and background sub-regions within the edge transition region are set, for example, 0.7 and 0.3 respectively, along with the number of iterations, such as 50 or 100. During each iteration of the smoothing process, the pixel gradients of the edge transition region are read. The pixel gradients of the object and background sub-regions are then weighted and summed according to the preset weights to obtain the fused gradient. Based on this fused gradient, the brightness and color values ​​of the current pixel are adjusted using Poisson variance calculation to ensure continuous gradient changes with adjacent pixels.

[0226] Pixel gradients can be calculated using any gradient operator, such as the Sobel operator or the Prewitt operator.

[0227] By introducing the Poisson blending model and combining it with the area to be processed precisely defined by the object mask, pixel-level smoothing of the target object and background edges is achieved, making the gradient of the target object and the hand edge natural and continuous, avoiding problems such as hard edge fragmentation, color banding and texture blurring that are common in traditional blending methods.

[0228] After using the Poisson fusion model for edge smoothing, visual attributes such as lighting and shadows can also be optimized.

[0229] When generating the item mask, different markers can be used to distinguish different interactions between the hand and the held item. For example, occlusion interactions can be marked in red or marked with a first marker, pressing interactions can be marked in yellow or marked with a second marker, the area where the held item is located can be marked in blue or marked with a fourth marker, and other areas can be left unmarked or marked with a second marker, thus eliminating the need for edge blending.

[0230] Since the aforementioned item masks have already been differentiated by different markers based on the different interactions between the finger and the item during the generation process, the initial displayed image can be subjected to gradient processing based on these different markers during the fusion process, with the following corresponding rules:

[0231] For the fourth marker (corresponding to the handheld object itself), the internal gradient information must be completely consistent with the initial displayed image. A complete replacement process can be performed, that is, replacing the pixels of the fourth marker with the pixels at the corresponding positions of the target object. For areas without markers or with the third marker, the original pixels are retained.

[0232] For the first marker (corresponding to the edge of the occluded finger), it is necessary to strictly prevent the target object from intruding into the finger area. Therefore, in actual processing, the gradient information of the finger edge at the position corresponding to the first marker is used as a constraint, and the gradient information from the target object area to the finger edge is smoothly transitioned through the calculation of the Poisson fusion model. Specifically, the transparency processing performed on the aforementioned pressed edge area (the area corresponding to the first mark) includes: 1) spreading a certain distance (usually 1 to 5 pixels) from the edge corresponding to the first mark to both sides to form a first pressed buffer area; 2) taking the fingertip key point inside this area as the origin, calculating the distance from each pixel to the neighboring fingertip key point to obtain the pressed distance field information; 3) calculating the transparency weight corresponding to different pixels based on the pressed distance field and the preset transparency calculation model; the transparency calculation model can be referred to as: α = A- B×exp(-d² / (2σ²)), where α is the transparency weight, and this value is inversely proportional to the transparency visual effect; σ is the diffusion coefficient; d is the distance from the pixel to the neighboring fingertip key point; A and B are constants, for example, A is 0.9 and B is 0.6.

[0233] For the second marker (corresponding to the edge of the pressed finger), this part needs to comprehensively consider the allocation of the target object and the finger. Therefore, we introduce gradient weight allocation calculations for this region, as follows:

[0234] The second press buffer area is formed by spreading a certain distance (usually 1 to 5 pixels) to both sides of the edge of the target item corresponding to the area. This second press buffer area overlaps at least partially with the first press buffer area in the item mask. Thus, the first press buffer area and the second press buffer area together constitute three areas: the item area completely located in the target item, the finger area completely located in the finger, and the overlapping area that overlaps with each other.

[0235] The gradient weights for finger transparency and / or target object are adjusted within the aforementioned range, specifically including:

[0236] Within the overlapping region, the gradient weight of the target item and the transparency of the finger gradually decrease as the finger extends towards the target item. The closer the finger is to the target item, the higher its transparency, indicating that the finger becomes more transparent closer to the edge, resulting in lower finger realism. Simultaneously, the gradient weight of the target item increases, enhancing its visual proportion and better showcasing product details. Conversely, the closer the finger is to the inside, the lower its transparency, indicating that the finger becomes more realistic closer to the inside, resulting in higher finger realism. At the same time, the gradient weight of the target item gradually decreases, reducing its visual proportion and ensuring the tactile sensation of pressing the finger is preserved.

[0237] In the item area, the gradient weight of the target item and the transparency of the finger in the finger area still follow the above change principle, that is, it increases from the inside (finger) to the outside (item) and decreases from the outside to the inside.

[0238] In actual processing, the gradient of the finger can be calculated from the original image and the gradient of the target object from the target object image using preset gradient operators (such as the Sobel operator and the Prewitt operator). Different markers can be obtained from the object mask to complete the above calculations. Through these gradient changes, the harsh edges between the target object, the hand, and the background can be eliminated, making the pixel gradient of the target object and the gradient of the hand edge natural and continuous, thus solving the problem of the seamlessness in the contact area between the finger and the target object.

[0239] Lighting blending can be performed by a lighting blending subunit. The lighting blending subunit uses lighting models such as the Retinex model and CycleGAN to improve the consistency of lighting effects between the target object and the hand based on the image obtained by edge blending.

[0240] The input to the lighting blending unit includes the image obtained by edge blending and the background image of the person obtained by removing the handheld object from the original display image.

[0241] Shadow blending can be performed by a shadow blending subunit. This subunit uses a shadow blending model such as CycleGAN to improve the consistency of shadow effects between the target object and the background, building upon the image obtained after lighting blending.

[0242] The input to the shadow blending model includes the image obtained after lighting blending processing, and the pure background image after removing the handheld items and people from the original display image, and the output is the target display image.

[0243] Based on the object mask, the held object in the original display image is removed to obtain the background image of the person. Based on the background image of the person, the lighting parameters of the target object in the edge-smoothed image are adjusted to obtain the target display image. The edge-smoothed image is the image obtained after edge blending or edge smoothing processing.

[0244] Among them, lighting parameters include brightness, light source direction, color temperature, hue, reflectivity, and high light intensity.

[0245] Based on the previously acquired object mask, the area where the held object is located in the original display image is accurately located. Pixels in this area are then removed using image segmentation technology, resulting in a background image that retains only the person and the scene background. Subsequently, using the background image as a lighting reference, the lighting parameters of the target object in the edge-smoothed image are adjusted accordingly.

[0246] It can extract the global average brightness, color temperature parameters, and light source direction of a person's background image, enabling adjustments to the lighting properties of target objects in an image with smoothed edges. For example, if the overall brightness of the person's background image is too high, the brightness value of the target object is simultaneously increased to avoid brightness discontinuities between the object and the background; if the person's background image has a cool lighting tone, the color temperature curve of the target object is adjusted to ensure its color tendency is consistent with the person's background image; at the same time, combined with the light source direction, the highlight and reflection positions on the target object's surface are corrected to ensure that the highlight area matches the background light source direction and the reflection intensity matches the diffuse / spectral reflection characteristics of the background.

[0247] Based on lighting models such as the Retinex model and the CycleGAN model, the lighting parameters of target objects in the image with smoothed edges can be adjusted in a targeted manner, using the background image of the person as the lighting reference.

[0248] By using background images of people to adjust the lighting parameters of target objects, the lighting and color of the target objects in the final target display image are highly adapted to the background environment, improving the visual integration and realism of the objects and the scene.

[0249] Shadow blending processing specifically includes: using the background image as a shadow reference, extracting shadow parameters of scene elements in the background image, such as shadow direction, length ratio, transparency, blur radius, etc., and adjusting the shadow of the target object according to these shadow parameters so that the shadow direction of the target object is consistent with that of the background image. It can also scale the shadow length of the target object according to the size of the target object and its distance from the light source, and adjust the shadow transparency and blur radius.

[0250] A shadow model such as CycleGAN can be used to learn the generation rules of native shadows in a scene based on the shadow parameters of scene elements in the background image; then, the inherent shadows of the target object can be corrected based on these rules, so that the shadow brightness in the target display image is uniform and the transition is natural.

[0251] Edge blending of the initial display image can be performed by an edge blending subunit. This subunit can perform preliminary edge blending processing on the initial display image based on the item mask. For example, a Poisson blending model can be used to process different regions. The purpose of Poisson blending is to adjust the gradient of the image (the gradient refers to the rate of change of pixel values ​​at a point in the image along the x and y axes, specifically reflecting the degree of abrupt changes in brightness, color, and contour of the corresponding area in the image) under preset constraints, so that the gradient changes between different parts of the image are continuous and smooth, resulting in a more ideal edge connection between the product and the finger.

[0252] Figure 5This is a schematic diagram of the structure of the handheld item replacement frame provided in the embodiments of this application. The handheld item replacement frame is used to perform the handheld item replacement method provided in the foregoing embodiments, such as... Figure 5 As shown, the handheld item replacement framework includes: a feature extraction unit, a multimodal feature fusion unit, an attention encoding unit, a diffusion generation unit, a feature decoding unit, and a pixel fusion unit.

[0253] The feature extraction unit takes the original display image, the item mask, the target item image, and the item prompt as input, and extracts features from each of the inputs, outputting the original feature map, the mask feature map, the target item feature map, and semantic features.

[0254] The feature extraction unit may include an image feature extraction unit and a text semantic extraction unit. The image feature extraction unit, sampling a VAE encoder, is used to extract visual features from the input original display image, the item mask, and the target item image, outputting an original feature map, a mask feature map, and a target item feature map. The text semantic extraction unit employs a dual encoder consisting of a CLIP text encoder and a T5 encoder to extract semantic features from item prompts.

[0255] The multimodal fusion unit is used to perform feature fusion processing between text and image based on the features extracted above. The multimodal fusion unit includes a gesture pose adjustment unit, a dynamic offset unit, a text conditional fusion unit, and a time step embedding unit.

[0256] The gesture pose adjustment unit includes an object detection subunit and a hand pose matching subunit. The object detection subunit detects the main body of the target object and determines the most suitable area for grasping based on the main body of the target object, outputting grasping parameters. The hand pose matching subunit generates a hand feature map based on the grasping parameters.

[0257] The object detection subunit can include an object detection network, an image segmentation network, and a grasping position detection network. The object detection network, such as YOLOv8, detects the boundaries of the object and outputs its size information. It can also determine the object's thickness information based on a pre-defined database. The image segmentation network, such as the SAM network, segments the object based on its size and thickness information, resulting in multiple segmented regions. The grasping position detection network calculates scores for multiple segmented regions to determine the most suitable grasping region and outputs its grasping parameters. The grasping position detection network uses an MLP network, which consists of multiple fully connected layers. It first calculates the feature vector for each segmented region (this calculation process is independent of the MLP and is implemented solely through numerical computation) to determine the feature vector of each segmented region, including area, position, and thickness information. The feature vector corresponding to each segmented region is then input into the MLP network for scoring. The segmented region with the highest score is determined as the grasping region (i.e., the most suitable grasping region), and its grasping parameters are output.

[0258] The hand pose matching subunit comprises a hand pose matching network, a body pose matching network, and a spatial transformation network. The hand pose matching network samples the MANO model to calculate hand keypoints based on grasping parameters, outputting 21 hand keypoints adapted to the grasping parameters. The body pose matching network includes the OpenPose model, used to correct the aforementioned hand keypoints based on the body features of the target person image (or the original display image). Hand keypoint correction can be performed using elbow flexion constraints and body distance constraints.

[0259] The Spatial Transformation Network (STN) is used to map the corrected hand keypoints into two-dimensional pixel coordinates, and then perform bilinear interpolation based on the pixel coordinates to generate and output a hand feature map.

[0260] The dynamic offset unit is used to offset the region of the held object in the original feature map and the mask feature map based on the semantic features of the object cue words. The dynamic offset unit includes an offset prediction layer, a feature sampling layer, a pose adjustment layer, and a projection layer. The offset prediction layer is used to calculate the spatial offset between the original state of the held object in the original feature map and the expected state corresponding to the semantic features, based on the semantic features, to obtain the offset parameters and sampling weight parameters.

[0261] The offset prediction layer consists of a cross-attention network and an offset prediction network. The cross-attention network takes the original feature map and semantic features as input, uses the semantic features as the query object Q, and the original feature map as the query object (key K and value V). It calculates the attention weights of the original feature map based on the semantic features, and assigns weights to the original feature map based on the calculation results, thus obtaining an original feature map with textual semantic guidance.

[0262] The offset prediction network consists of pooling layers and multiple fully connected layers (fully connected layers can also be replaced by lightweight convolutional networks). After the original feature map guided by text semantics is input into the offset prediction network, it first undergoes global pooling to compress the features into a global vector. Then, the fully connected layers map this global vector to low-dimensional parameters, and further convert these low-dimensional parameters into specific parameters, namely the offset parameters. After calculating the offset parameters, the offset prediction layer also needs to calculate sampling weight parameters. These sampling weight parameters characterize the contribution weight of the feature points surrounding the original feature point (the feature point before offset) to the offset feature point (the feature point after offset) during subsequent feature sampling.

[0263] The feature sampling layer includes a spatial transformation network, which first converts the bias parameters into a bias matrix. Based on this bias matrix, it transforms the spatial coordinates of the features in the original feature map and the mask feature map, thus obtaining the original feature map and the mask feature map after offset adjustment. On this basis, the feature sampling layer further adjusts the sampling weight parameters according to the aforementioned parameters, performing a weighted sum based on the sampling range selected by the offset prediction layer and the corresponding sampling weight parameters. This completes the feature fusion of the offset feature points with surrounding feature points. The fused offset feature points significantly improve spatial continuity and avoid the aforementioned problems of feature sparsity and overlap, ultimately yielding the original feature map and the mask feature map after spatial position adjustment.

[0264] The pose adjustment layer is used to further refine the original feature map after spatial positioning adjustment, utilizing the hand feature map. The pose adjustment layer can include a pose fusion network and an object contact correction network. The pose fusion network fuses the hand feature map and the original feature map after spatial positioning adjustment. First, it extracts style and pose information from the original feature map using two parallel 1×1 convolutional layers; then, it inputs the hand feature map and pose information into an MLP network to correct the pose information; finally, it concatenates the corrected pose information with the style information to obtain the original feature map with complete pose adjustment.

[0265] The object contact correction network is used to extract the object grasping part (corresponding to the grasping area) and the hand part (corresponding to the hand feature map area) from the original feature map after posture adjustment, to obtain the object grasping feature sequence and the hand feature sequence. Through the cross-attention network, the object grasping feature sequence is used as the query object Q and the hand feature sequence is used as the query object (key K and value V). The attention weight of the hand feature sequence is calculated based on the object grasping feature sequence, and the hand feature sequence is weighted according to the calculation result to obtain the hand feature sequence after correction based on the object contact position. After mapping, the original feature map of the target object and the hand maintaining a better grip posture can be obtained, that is, the original feature map after posture adjustment.

[0266] The projection layer is used to map the pose-adjusted original feature map, the previously processed mask feature map, the target object feature map, and the semantic features along the channel dimension. After completing the channel dimension mapping, the original feature map and the mask feature map are concatenated into a joint feature map, and the target object feature map and semantic features are output separately.

[0267] The text conditional fusion unit is used to generate semantically guided feature maps by interacting semantic features with joint feature maps through a cross-attention structure.

[0268] The temporal step embedding unit is used to concatenate the temporal step vector with the semantically guided feature map in the channel dimension to generate a temporal conditional feature map.

[0269] Attention encoding units are used to perform attention encoding on temporal conditional feature maps to obtain global feature sequences.

[0270] The diffusion generation unit generates the target latent code map of the target display image based on the global feature sequence. The feature decoding unit decodes the target latent code map to obtain the preliminary display image.

[0271] The pixel fusion unit is used to perform edge fusion, lighting fusion, and shadow fusion on the above preliminary display image, so that the edges, lighting and shadows of the target object after replacement are more in line with the real holding state, and the final target display image is obtained.

[0272] This application also provides a digital human video generation method, including: acquiring a target display image that replaces a handheld item with a target item; obtaining the target display image based on the handheld item replacement method provided in any of the foregoing embodiments of this application; and inputting the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0273] The associated information of the target item can be in the form of at least one of text, semantics, and instructions.

[0274] Taking related information as text as an example, we can first perform semantic understanding on the related information based on a natural language processing model, extract features such as core selling points, functional attributes, and emotional tendencies, and generate keyword tags and language scripts; based on the keyword tags, we can determine the interactive actions of the digital human from a preset action library; we can then perform layer fusion between the target display image and the selected or default digital human image, synchronously play the voice corresponding to the language script, and control the digital human to synchronously perform the corresponding interactive actions to generate a coherent video, which is the digital human video used to display the target item.

[0275] If the associated information is speech, it can be converted into text first, and then the aforementioned method can be used to obtain the digital human video. Alternatively, features can be extracted from the speech to obtain rhythm features, intonation features, emotional features, and semantic keywords; lip-syncing signals can be generated based on rhythm and intonation features, and semantic keywords can be used to match the display parts of the target item to automatically generate targeted action instructions; at the same time, the amplitude and posture of the digital human's movements can be adjusted according to the emotional features of the speech to control the digital human to complete coordinated lip, limb, and head movements, and then the digital human's product-selling video can be output after being fused with the target display image.

[0276] If the associated information is an instruction, then the digital human's action instructions, expression instructions, and visual presentation instructions can be directly parsed from the instruction. Then, through the instruction matching engine, the action instructions are accurately associated with the preset action library, the expression instructions are converted into the digital human's tone parameters and lip-sync rhythm parameters, and the visual presentation instructions are converted into eye control parameters. Subsequently, the target display image and the digital human image are composited into layers, and the digital human is driven to perform corresponding actions and expressions according to the instruction sequence to generate a coherent digital human video that meets the instruction requirements.

[0277] By replacing items, new display materials for new items are automatically synthesized, and digital human-driven models are linked to generate promotional videos for the target items. There is no need to shoot images of each item individually. By replacing items in the original images, the materials required for new items are automatically synthesized, which greatly reduces the cost and time of material preparation, improves the efficiency of material generation, and thus improves the efficiency of digital human video generation, making it suitable for mass product promotion scenarios.

[0278] Figure 6 A schematic diagram of the handheld article replacement device provided in this application is shown below. Figure 6As shown, the handheld item replacement device provided in this embodiment includes: an image acquisition module for acquiring an original display image and a target item image; a mask generation module for identifying the handheld item in the original display image and generating an item mask; a multimodal feature extraction module for extracting visual features from the original display image, the item mask, and the target item image, as well as semantic features from the item prompts; a multimodal fusion module for performing multimodal fusion processing on the visual and semantic features and embedding time steps to obtain a global feature sequence; and a target image generation module for using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item.

[0279] In one possible implementation, the multimodal fusion module includes: a semantic fusion unit, used to adjust and fuse the original feature map and the mask feature map based on the target item feature map and semantic features to obtain a semantically guided feature map; the target item feature map, the original feature map, and the mask feature map are the extracted visual features of the target item image, the visual features of the original display image, and the visual features of the item mask, respectively; a temporal step embedding unit, used to concatenate the temporal step vector converted from temporal step information with the semantically guided feature map to obtain a temporal conditional feature map; and an encoding unit, used to perform attention encoding on the temporal conditional feature map to obtain a global feature sequence.

[0280] In one possible implementation, the semantic fusion unit includes: a gripping parameter determination subunit, used to determine the gripping parameters of the target object based on the target object feature map; a hand feature map generation subunit, used to generate a hand feature map based on the gripping parameters and the original display image; an offset subunit, used to offset the original feature map and the mask feature map based on semantic features; a pose adjustment subunit, used to adjust the pose of the offset original feature map based on the hand feature map, and fuse the pose-adjusted original feature map with the offset mask feature map to obtain a joint feature map; and an attention interaction subunit, used to interact with the semantic features, the target object feature map, and the joint feature map based on a cross-attention mechanism to generate a semantically guided feature map.

[0281] In one possible implementation, the gripping parameter determination subunit is specifically used for: performing target detection on the target item feature map to obtain the item boundary and size information of the target item; segmenting the target item feature map based on the item boundary and size information of the target item to obtain multiple segmented regions; calculating the region parameters of each segmented region, and determining the gripping region from the multiple segmented regions based on the region parameters; and determining the gripping parameters of the target item based on the determined gripping region.

[0282] In one possible implementation, the hand feature map generation subunit is specifically used for: determining key hand points that are compatible with the target object based on grasping parameters and a preset hand model; obtaining key body points of the person in the original display image based on a human posture constraint model, and correcting the key hand points based on the key body points; mapping the corrected key hand points to two-dimensional pixel coordinates, and generating a hand feature map after interpolation.

[0283] In one possible implementation, the offset subunit is specifically used for: calculating attention weights and allocating them to the original feature map using a cross-attention network, with semantic features as the query and the original feature map as the key, to obtain a semantically guided original feature map; inputting the semantically guided original feature map into an offset prediction network to obtain offset parameters and sampling weight parameters; converting the offset parameters into an offset matrix using a spatial transformation network, and performing spatial coordinate transformation on the original feature map and the mask feature map based on the offset matrix; and using the sampling weight parameters to perform weighted fusion of the offset feature points and surrounding feature points to obtain the offset original feature map and the offset mask feature map; the offset feature points are the corresponding feature points formed by the spatial coordinate transformation of the feature points in the original feature map and the mask feature map.

[0284] In one possible implementation, the posture adjustment subunit is specifically used for: extracting style information and posture information from the offset original feature map; correcting the hand features in the posture information based on the hand feature map; concatenating the corrected hand features with the style information to obtain a concatenated original feature map; generating a product grasping mask and a hand feature mask based on the grasping region; extracting product grasping feature sequences and hand feature sequences from the concatenated original feature map using the product grasping mask and hand feature mask respectively; using the product grasping feature sequence as the query and the hand feature sequence as the key, calculating attention weights through a cross-attention network and adjusting the hand feature sequence based on the attention weights, mapping it to the original feature map to obtain the posture-adjusted original feature map.

[0285] In one possible implementation, the attention interaction subunit is specifically used for: concatenating the joint feature map and the target item feature map along the channel dimension to obtain a concatenated feature map, and mapping the concatenated feature map to a concatenated sequence; calculating attention weights through a cross-attention network using semantic features as queries and the concatenated sequence as keys to obtain an attention weight matrix; adjusting the concatenated sequence based on the attention weight matrix to obtain a preliminary fused feature map; calculating the similarity between the original item and the target item in multiple spatial locations in the preliminary fused feature map; the original item being the handheld item in the original display image; determining the weight of the target item in each spatial location based on the similarity at each spatial location, and adjusting the weight of the main body region corresponding to the item mask to obtain a weight vector; and performing weighted fusion of the features of the original item and the target item in the preliminary fused feature map based on the weight vector to obtain a semantically guided feature map.

[0286] In one possible implementation, the encoding unit is specifically used for: extracting visual and text marker sequences from the temporal conditional feature map based on an attention layer composed of a multi-head self-attention structure and a normalization layer; performing cross-modal alignment and feature fusion on the visual and text marker sequences through dual-path cross-attention calculation; the dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key; performing three-dimensional rotational position encoding on the processed visual feature sequence to obtain an image feature sequence with positional information; generating semantic constraints based on the processed text feature sequence; inputting the image feature sequence with positional information into the attention layer composed of a multi-layer multi-head self-attention structure and a normalization layer, performing attention weight calculation under the semantic constraints, and outputting a global feature sequence after residual connection, feedforward network, and layer normalization processing.

[0287] In one possible implementation, the semantic constraints are modulated by parameters generated by a first fully connected network to adjust the parameters of the linear projection layer corresponding to the query, key, and value in the attention layer; or, the semantic constraints are multiplied by a supervision matrix generated by a second fully connected network with the weight matrix output by at least one attention layer.

[0288] In one possible implementation, the target image generation module includes: a dimensionality compression unit, used to perform dimensionality compression processing on the global feature sequence according to a preset latent code channel dimension to obtain a latent code feature vector; a sequence recombination unit, used to recombine the latent code feature vector to obtain a latent code feature map; a local channel fusion unit, used to perform local channel weighted fusion on the latent code feature map to generate a conditional latent code map matching the time step; an initialization unit, used to initialize and generate a random noise latent code map with the same dimension as the conditional latent code map, as the noisy latent code map for the first time step; a gradient prediction unit, used to predict the gradient correction vector for the next time step by calculating the difference between the conditional latent code map and the noisy latent code map of the corresponding time step; a latent code update unit, used to iteratively update the noisy latent code map based on the predicted gradient correction vector to obtain a target latent code map; and a target image generation unit, used to decode the target latent code map to obtain a target display image.

[0289] In one possible implementation, the target image generation unit is specifically used to: input the target latent code into the decoder to obtain a preliminary display image; and sequentially perform edge blending, illumination blending, and shadow blending processing on the preliminary display image to obtain the target display image.

[0290] The handheld item replacement device provided in this embodiment can perform the handheld item replacement method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0291] This application also provides a digital human video generation device, including: a display image acquisition module, used to acquire a target display image in which a handheld item is replaced with a target item; the target display image is obtained based on the handheld item replacement method provided in any embodiment of this application; and a video generation module, used to input the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0292] The digital human video generation device provided in this embodiment can execute the digital human video generation method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0293] Figure 7 A schematic diagram of the structure of the electronic device provided in this application. Figure 7 As shown, the electronic device 70 provided in this embodiment includes at least one processor 701 and a memory 702. Optionally, the device 70 further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus 704.

[0294] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to perform the above-described method.

[0295] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0296] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0297] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0298] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0299] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0300] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0301] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0302] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0303] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0304] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0305] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0306] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0307] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0308] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for replacing a handheld item, characterized in that, include: Obtain the original display image and the target item image; Identify the handheld item in the original display image and generate an item mask; Visual features of the original display image, the item mask, and the target item image are extracted, as well as semantic features of the item prompt words; The visual features and semantic features are fused using a multimodal process and embedded with time steps to obtain a global feature sequence; Using a diffusion generation algorithm, the global feature sequence is processed by latent code and feature decoding to generate a target display image with the target item as the handheld item.

2. The method according to claim 1, characterized in that, The process of multimodal fusion of the visual features and semantic features and embedding time steps to obtain a global feature sequence includes: Based on the target item feature map and semantic features, the original feature map and the mask feature map are adjusted and fused to obtain a semantically guided feature map; the target item feature map, the original feature map, and the mask feature map are respectively the extracted visual features of the target item image, the visual features of the original display image, and the visual features of the item mask. The time step vector, converted from the time step information, is concatenated with the semantically guided feature map to obtain the time conditional feature map. Attention encoding is performed on the temporal conditional feature map to obtain the global feature sequence.

3. The method according to claim 2, characterized in that, The process involves adjusting and fusing the original feature map and mask feature map based on the target item feature map and semantic features to obtain a semantically guided feature map, including: Based on the feature map of the target item, the grasping parameters of the target item are determined; Based on the gripping parameters and the original display image, a hand feature map is generated; Based on the semantic features, the original feature map and the mask feature map are offset; Based on the hand feature map, the pose of the offset original feature map is adjusted, and the pose-adjusted original feature map is fused with the offset mask feature map to obtain a joint feature map; Based on the cross-attention mechanism, the semantic features, the target item feature map, and the joint feature map are interacted to generate the semantically guided feature map.

4. The method according to claim 3, characterized in that, The determination of the grasping parameters of the target item based on its feature map includes: Target detection is performed on the feature map of the target item to obtain the item boundary and size information of the target item; Based on the object boundary and size information of the target object, the feature map of the target object is segmented to obtain multiple segmentation regions; Calculate the region parameters of each of the segmented regions, and determine the gripping region from the plurality of segmented regions based on the region parameters; Based on the determined gripping area, the gripping parameters of the target item are determined.

5. The method according to claim 3, characterized in that, The step of generating a hand feature map based on the gripping parameters and the original displayed image includes: Based on the gripping parameters and the preset hand model, determine the key hand points that are compatible with the target item; Based on the human posture constraint model, the key points of the person's body in the original display image are obtained, and the key points of the hands are corrected based on the key points of the body. The corrected hand key points are mapped to two-dimensional pixel coordinates, and the hand feature map is generated after interpolation.

6. The method according to claim 3, characterized in that, The offsetting of the original feature map and the mask feature map based on the semantic features includes: Using a cross-attention network, with the semantic features as the query and the original feature map as the key, attention weights are calculated and assigned to the original feature map to obtain the semantically guided original feature map. The semantically guided original feature map is input into the offset prediction network to obtain offset parameters and sampling weight parameters. The offset parameters are converted into an offset matrix through a spatial transformation network, and the original feature map and the mask feature map are transformed into spatial coordinates based on the offset matrix. The offset feature points and surrounding feature points are weighted and fused using the sampling weight parameters to obtain the original feature map and the mask feature map after offset. The offset feature points are the corresponding feature points formed by the spatial coordinate transformation of the feature points in the original feature map and the mask feature map.

7. The method according to claim 4, characterized in that, The step of adjusting the pose of the offset original feature map based on the hand feature map includes: Extract style and pose information from the offset original feature map; Based on the hand feature map, the hand features in the posture information are corrected; The corrected hand features are then combined with the style information to obtain the original composite feature map. Based on the gripping area, generate a product gripping mask and a hand feature mask; Using the product gripping mask and the hand feature mask, product gripping feature sequence and hand feature sequence are extracted from the stitched original feature map, respectively; Using the product grasping feature sequence as the query and the hand feature sequence as the key, an attention weight is calculated through a cross-attention network, and the hand feature sequence is adjusted based on the attention weight and mapped to the original feature map to obtain the original feature map after posture adjustment.

8. The method according to claim 3, characterized in that, The method based on cross-attention interacts with the semantic features, the target item feature map, and the joint feature map to generate the semantically guided feature map, including: The joint feature map and the target item feature map are concatenated along the channel dimension to obtain a concatenated feature map, and the concatenated feature map is mapped to a concatenated sequence; Using the semantic features as the query and the concatenated sequence as the key, attention weights are calculated through a cross-attention network to obtain an attention weight matrix; Based on the attention weight matrix, the spliced ​​sequence is adjusted to obtain a preliminary fused feature map; Calculate the similarity between the original item and the target item in multiple spatial locations in the preliminary fused feature map; the original item is the handheld item in the original display image; Based on the similarity at each of the spatial locations, the weight of the target item at each of the spatial locations is determined, and the weight of the main area located at the object mask is adjusted to obtain a weight vector; Based on the weight vector, the features of the original item and the features of the target item in the preliminary fused feature map are weighted and fused to obtain the semantically guided feature map.

9. The method according to claim 2, characterized in that, The step of performing attention encoding on the temporal conditional feature map to obtain the global feature sequence includes: Based on an attention layer composed of a multi-head self-attention structure and a normalization layer, visual marker sequences and text marker sequences are extracted from the temporal conditional feature map. The visual marker sequence and the text marker sequence are aligned and fused across modalities through a dual-path cross-attention calculation. The dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key. The processed visual feature sequence is subjected to three-dimensional rotational position encoding to obtain an image feature sequence with positional information; Based on the processed text feature sequence, semantic constraints are generated. The image feature sequence with location information is input into an attention layer consisting of a multi-layer multi-head self-attention structure and a normalization layer. Attention weights are calculated under the semantic constraints. After residual connection, feedforward network and layer normalization processing, the global feature sequence is output.

10. The method according to claim 9, characterized in that, The semantic constraints are modulated by parameters generated by the first fully connected network, which are used to adjust the parameters of the linear projection layers corresponding to the queries, keys, and values ​​in the attention layer; or, the semantic constraints are supervised by a supervision matrix generated by the second fully connected network, which is used to multiply the weight matrix output by at least one attention layer.

11. The method according to any one of claims 1-10, characterized in that, The process of using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item includes: According to the preset latent code channel dimension, the global feature sequence is subjected to dimensional compression processing to obtain the latent code feature vector; The latent code feature vector is reassembled to obtain a latent code feature map; The latent code feature map is subjected to local channel weighted fusion to generate a conditional latent code map that matches the time step; Initialize and generate a random noise latent code map with the same dimension as the conditional latent code map, and use it as the noisy latent code map for the first time step; The gradient correction vector for the next time step is predicted by differentiating the conditional latent code map with the noisy latent code map at the corresponding time step. The target latent code map is obtained by iteratively updating the noisy latent code map based on the predicted gradient correction vector. The target latent code image is decoded to obtain the target display image.

12. The method according to claim 11, characterized in that, Decoding the target latent code to obtain the target display image includes: The target latent code is input into the decoder to obtain a preliminary display image; The initial display image is sequentially processed with edge blending, lighting blending, and shadow blending to obtain the target display image.

13. A method for generating digital human videos, characterized in that, include: A target display image is obtained to replace the held item with the target item; the target display image is obtained based on the method provided by any one of claims 1-12; The target display image and the associated information of the target item are input into the digital human driving model to obtain a digital human video for displaying the target item.

14. A handheld item replacement device, characterized in that, include: The image acquisition module is used to acquire the original display image and the target item image; A mask generation module is used to identify handheld items in the original display image and generate an item mask. The multimodal feature extraction module is used to extract visual features of the original display image, the item mask, and the target item image, as well as semantic features of the item prompt words; A multimodal fusion module is used to perform multimodal fusion processing on the visual features and the semantic features and embed time steps to obtain a global feature sequence; The target image generation module is used to perform latent code processing and feature decoding on the global feature sequence using a diffusion generation algorithm to generate a target display image with the target item as the handheld item.

15. A digital human video generation device, characterized in that, include: The image acquisition module is used to acquire the target display image to replace the held item with the target item; The target display image is obtained based on the method provided by any one of claims 1-12; The video generation module is used to input the target display image and the associated information of the target item into the digital human driving model to obtain a digital human video for displaying the target item.

16. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-13.

17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-13.

18. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-13.