Hand-held article image generation method, digital human video generation method and device

By using two image segmentation models working together, occluded elements are accurately located and removed, and missing parts of the item are completed, solving the problem of incomplete item outlines in handheld item replacement and achieving high-quality image replacement results.

CN121708172APending Publication Date: 2026-03-20NANJING SILICON INTELLIGENCE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610195784.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, when replacing handheld objects, the object outline is incomplete and the boundaries are confused, which affects the realism of the image. Furthermore, single image segmentation models are easily affected by occlusion interference, resulting in insufficient positioning accuracy.

Method used

A dual image segmentation model works in tandem. The first model accurately segments the handheld object to obtain a basic mask, while the second model combines object clues to remove occluding elements and fill in missing parts to generate a high-fidelity mask, ultimately achieving accurate fusion of the target object and the original image.

Benefits of technology

It improves the visual quality and realism of images after item replacement, with natural edge blending, solves the limitations of single-model segmentation, and enables efficient batch material production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708172A_ABST
    Figure CN121708172A_ABST
Patent Text Reader

Abstract

The invention provides a handheld article image generation method and device and a digital human video generation method and device. The handheld article image generation method comprises the following steps: acquiring an original display image and a target article image; segmenting the handheld article in the original display image through a first image segmentation model to obtain a first article mask; performing shielding element removal and handheld article complementation on the handheld image corresponding to the first article mask through a second image segmentation model under the guidance of the article cue word to obtain a second article mask; based on the first article mask and the second article mask, fusing the target article image to the original display image to obtain a target display image; the handheld article in the target display image is the target article. Through cooperative work of double image segmentation models, natural connection of a target object and an original image is guaranteed, and the problem of replacement distortion caused by a single segmentation model is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image synthesis technology, and in particular to a method for generating images of handheld objects, a method and apparatus for generating digital human videos. Background Technology

[0002] In e-commerce displays, virtual product demonstrations, and product content creation, it is often necessary to replace the original handheld item in an image with a new item to achieve product effect simulation, new product display, and creative content generation.

[0003] To replace a handheld object in an image, image segmentation techniques are typically used to locate the object, and then the resulting object mask is used to replace it with a new object. However, many related techniques employ a single image segmentation model to segment the handheld object, leading to problems such as incomplete object outlines and blurred boundaries between the hand and the object after replacement, severely impacting the realism of the resulting image.

[0004] Therefore, there is an urgent need to provide a method for accurately segmenting handheld objects in images to improve the realism of images after object replacement. Summary of the Invention

[0005] This application provides a method for generating images of handheld items, a method and apparatus for generating digital human videos, to solve the problem of low segmentation accuracy of handheld items in images, resulting in poor image quality after item replacement.

[0006] In a first aspect, embodiments of this application provide a method for generating a handheld object image, comprising: acquiring an original display image and a target object image; segmenting the handheld object in the original display image using a first image segmentation model to obtain a first object mask; performing occlusion element removal and handheld object completion on the handheld image corresponding to the first object mask using a second image segmentation model under the guidance of an object prompt word to obtain a second object mask; and fusing the target object image into the original display image based on the first object mask and the second object mask to obtain a target display image; wherein the handheld object in the target display image is the target object.

[0007] In one possible implementation, guided by an item prompt, a second image segmentation model removes occlusion elements and completes the handheld image corresponding to the first item mask to obtain a second item mask. This includes: identifying hand key points in the original display image based on a gesture recognition model; inputting the item prompt, hand key points, the first item mask, and the original display image into the second image segmentation model; cropping the original display image based on the first item mask using the second image segmentation model to obtain a handheld image; removing occlusion elements from the handheld image based on the item prompt and hand key points, and completing the occluded parts of the handheld item to obtain a second item mask representing the area where the handheld item is located.

[0008] In one possible implementation, the target item image is fused to the original display image based on a first item mask and a second item mask to obtain the target display image. This includes: determining the hand edge based on hand key points and the first item mask; determining a replacement region based on the second item mask and the hand edge; mapping the pixels of the target item image to the replacement region to replace the pixels in the replacement region of the original display image; and performing adaptation processing on the pixels in the region corresponding to the hand edge of the original display image to obtain the target display image. The adaptation processing includes at least one of the following: edge protection processing and transparency processing.

[0009] In one possible implementation, the hand edge includes an occlusion edge and a pressing edge; determining the hand edge based on hand key points and a first object mask includes: identifying the interaction area between the hand and the held object based on the hand key points and the first object mask; determining a first pixel in the interaction area of ​​the original display image that is in an occlusion state and a second pixel in a pressing state; and determining the occlusion edge and the pressing edge based on the first pixel and the second pixel, respectively.

[0010] In one possible implementation, determining a first pixel in an occluded state and a second pixel in a pressed state in the interactive area of ​​the original display image includes: for each pixel in the interactive area of ​​the original display image, determining whether a pixel is a first pixel or a second pixel based on at least one of the following: the distance between the pixel and the edge of the held object, the feature value of the pixel, and the pressure confidence of the pixel; the edge of the held object is obtained based on a second object mask, and the pressure confidence is used to characterize the confidence of the corresponding pixel in the pressure applied to the held object.

[0011] In one possible implementation, the pixels of the original display image located at the edge of the hand are subjected to adaptation processing to obtain the target display image, including: performing edge protection processing on the pixels of the original display image located in the area corresponding to the occlusion edge, and performing transparency processing on the pixels of the original display image located in the area corresponding to the pressing edge, to obtain the target display image.

[0012] In one possible implementation, the target item image is fused to the original display image based on a first item mask and a second item mask to obtain a target display image. This includes: determining occlusion edges and pressing edges based on hand key points and the first item mask; marking the first item mask and the second item mask based on the occlusion edges and pressing edges to generate a marked item mask; extracting visual features from the original display image, the marked item mask, and the target item image, as well as semantic features from the item prompts; performing multimodal fusion processing on the visual features and semantic features and embedding them into a time step to obtain a global feature sequence; and using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item.

[0013] In one possible implementation, visual and semantic features are fused using a multimodal process and embedded with time steps to obtain a global feature sequence, including: determining the grasping parameters of the target object based on the target object feature map; the target object feature map being the extracted visual features of the target object image; generating a hand feature map based on the grasping parameters and the original display image; shifting the original feature map and the mask feature map based on semantic features; the original feature map and the mask feature map being the extracted visual features of the original display image and the visual features of the marked object mask, respectively; adjusting the pose of the shifted original feature map based on the hand feature map, and fusing the pose-adjusted original feature map with the shifted mask feature map to obtain a joint feature map; interacting with the semantic features, the target object feature map, and the joint feature map based on a cross-attention mechanism to generate a semantically guided feature map; concatenating the time step vector transformed from time step information with the semantically guided feature map to obtain a temporal conditional feature map; and performing attention encoding on the temporal conditional feature map to obtain a global feature sequence.

[0014] In one possible implementation, attention encoding is performed on the temporal conditional feature map to obtain a global feature sequence, including: extracting visual and text marker sequences from the temporal conditional feature map using an attention layer composed of a multi-head self-attention structure and a normalization layer; performing cross-modal alignment and feature fusion on the visual and text marker sequences through dual-path cross-attention calculation; the dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key; performing three-dimensional rotational position encoding on the processed visual feature sequence to obtain an image feature sequence with positional information; generating semantic constraints based on the processed text feature sequence; inputting the image feature sequence with positional information into the attention layer composed of a multi-layer multi-head self-attention structure and a normalization layer, calculating attention weights under the semantic constraints, and outputting the global feature sequence after residual connection, feedforward network, and layer normalization processing.

[0015] In one possible implementation, a diffusion generation algorithm is used to process the global feature sequence using latent code and decode the features to generate a target display image with the target item as the handheld item. This includes: processing the global feature sequence using the diffusion generation algorithm to obtain a target latent code; inputting the target latent code into a decoder to obtain a preliminary display image; smoothing the pixels in the preliminary display image located in the area corresponding to the occlusion edge; and making the pixels in the preliminary display image located in the area corresponding to the pressing edge transparent to obtain the target display image.

[0016] In one possible implementation, edge protection processing is performed on pixels located in the region corresponding to the occlusion edge, including: spreading the occlusion edge to both sides by a first preset distance to form a first edge buffer region; and blurring each pixel in the edge buffer region based on the distance between each pixel in the first edge buffer region and the occlusion edge.

[0017] In one possible implementation, the pixels located in the area corresponding to the pressing edge are made transparent, including: spreading the pressing edge to both sides by a second preset distance to form a second edge buffer area; determining the transparency weight of each pixel in the edge buffer area based on the distance between each pixel in the second edge buffer area and the nearest hand key point; and making each pixel in the edge buffer area transparent based on the transparency weight of each pixel in the edge buffer area.

[0018] In one possible implementation, the first image segmentation model is pre-trained and then fine-tuned using a training set of images of handheld objects; the second image segmentation model is not pre-trained and is trained based on the training set of images of handheld objects.

[0019] Secondly, embodiments of this application provide a digital human video generation method, including: acquiring a target display image that replaces a handheld item with a target item; obtaining the target display image based on the method provided in the first aspect of this application or any implementation corresponding to the first aspect; inputting the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0020] Thirdly, embodiments of this application provide a handheld item image generation device, comprising: an image acquisition module for acquiring an original display image and a target item image; a first segmentation module for segmenting the handheld item in the original display image using a first image segmentation model to obtain a first item mask; a second segmentation module for performing occlusion element removal and handheld item completion on the handheld image corresponding to the first item mask using a second image segmentation model under the guidance of item prompts to obtain a second item mask; and a handheld item replacement module for fusing the target item image into the original display image based on the first item mask and the second item mask to obtain a target display image; wherein the handheld item in the target display image is the target item.

[0021] Fourthly, embodiments of this application provide a digital human video generation apparatus, comprising: a display image acquisition module, configured to acquire a target display image in which a handheld item is replaced with a target item; the target display image is obtained based on the method provided in the first aspect of this application or any implementation thereof; and a video generation module, configured to input the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0022] Fifthly, embodiments of this application provide an electronic device, including: a memory, a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the method provided in the first or second aspect above, and / or, various possible implementations corresponding to the first or second aspect.

[0023] Sixthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the first or second aspect above, and / or various possible implementations corresponding to the first or second aspect.

[0024] In a seventh aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the first or second aspect above, and / or various possible implementations corresponding to the first or second aspect.

[0025] The handheld object image generation method, digital human video generation method, and apparatus provided in this application address the problems of occlusion interference and incomplete object contour extraction in traditional handheld object replacement using a single segmentation model. They employ a dual-image segmentation model collaborative working mode. The first image segmentation model accurately locates the original handheld object to obtain a basic mask, i.e., the first object mask. Then, using the second image segmentation model combined with object prompts, occluding elements are intelligently removed and missing parts of the object are filled in, generating a high-fidelity second object mask. Finally, based on the dual masks, the target object and the original image are accurately fused, efficiently completing the handheld object replacement, ensuring a natural connection between the target object and the original image, and improving the visual quality and realism of the image after object replacement. Attached Figure Description

[0026] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0027] Figure 1 A schematic diagram illustrating the application scenarios provided in this application;

[0028] Figure 2 Flowchart of the method for generating handheld object images provided in this application Figure 1 ;

[0029] Figure 3 Flowchart of the method for generating handheld object images provided in this application Figure 2 ;

[0030] Figure 4 A schematic diagram showing the distribution of pixels for each marker in the original display image provided in this application;

[0031] Figure 5 A schematic diagram of the handheld item image generation device provided in this application;

[0032] Figure 6 A schematic diagram of the structure of the electronic device provided in this application.

[0033] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0035] It should be noted that the personal information (including but not limited to device information, personal attribute information, personal image, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the individual or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data comply with relevant laws, regulations and standards, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0036] First, let me explain the terms used in this application:

[0037] Mask: Binary identifier data used to accurately locate target regions in an image, clearly defining the boundary between the target region and the background region. A 1 indicates that the corresponding pixel belongs to the target region, while a 0 indicates that the corresponding pixel belongs to the background region.

[0038] Digital humans (also known as virtual humans): Anthropomorphic virtual images constructed based on technologies such as computer graphics and artificial intelligence. They have an appearance similar to real people and can express actions and transmit information through drive signals.

[0039] Digital Human Driving Model: An intelligent algorithm model used to control the movements of digital humans. It can receive multimodal inputs such as voice, text, and action commands, parse the feature information of the input data and generate corresponding driving signals to control the digital human to complete anthropomorphic actions such as lip-syncing, limb movement, and head posture adjustment.

[0040] In e-commerce live streaming, online advertising, and other similar scenarios, sellers often showcase their products through images or videos, with images of people (digital or real models) holding the product serving as the core element. To adapt to the promotional needs of different products, it is often necessary to replace the original handheld item in the image with a new target item to quickly generate diverse display materials without having to reshoot.

[0041] For example, Figure 1 A schematic diagram illustrating the application scenarios provided in this application, such as... Figure 1As shown, a user, such as an e-commerce merchant, stores the captured image material, such as the original display image img1 of a model holding product A, on the user's terminal. When the merchant needs to promote a new product, such as product B, but has not captured the corresponding handheld display image, there is no need to reorganize the model for shooting. They only need to obtain the target item image img2 of product B. They only need to upload the original display image img1 and the target item image img2 to the server through the user's terminal. The server executes the handheld item image generation method provided in this application, which can replace the original product A in the original display image img1 with product B, quickly generate the target display image img3 of holding product B, and return the target display image img3 to the user's terminal.

[0042] Users can also upload multiple target item images, and the server can perform batch replacement processing and output the target display images corresponding to each target item at once.

[0043] Users can also specify the original display images corresponding to each target item image. The server will perform the replacement according to the preset association relationship to generate personalized display images that are adapted to different original scenes and different target products, flexibly meeting the needs of creating promotional materials for multiple categories and multiple scenarios.

[0044] Traditional methods for replacing handheld objects rely on manual image editing or single image segmentation techniques. Manual image editing is inefficient and costly, and cannot meet the needs of batch material production. Single segmentation models are easily affected by hand occlusion, resulting in insufficient object positioning accuracy and problems such as edge distortion and abrupt transitions in the replaced image.

[0045] To address the aforementioned issues, this application provides a method for generating images of handheld objects. It employs a dual image segmentation model working collaboratively. The first model accurately segments the original handheld object to obtain a first mask, while the second model, combined with prompts, removes occlusions and completes the object to generate a second mask. Based on the dual masks, the target object is naturally integrated with the original image, thus completing the handheld object replacement. This method effectively overcomes the limitations of single-model segmentation, improving object positioning accuracy and mask integrity; it automates the replacement process, efficiently adapting to batch material production; and the replaced image exhibits natural edge blending and excellent visual effects.

[0046] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0047] Figure 2 Flowchart of the method for generating handheld object images provided in this application Figure 1 ,like Figure 2 As shown, the method includes:

[0048] Step S201: Obtain the original display image and the target item image.

[0049] The target item image is an image containing only the target item; the item held by the person in the original display image is a different item from the target item, and is referred to as the initial item.

[0050] The original display image is stock footage showing a person holding an initial object. There is an overlap between the person and the initial object in the image to represent the hand's gripping contact with the initial object.

[0051] Users can provide the original display image directly, or a frame can be extracted from a live-action product promotion video provided by the user as the original display image.

[0052] The people in the original display images can be real people, such as relevant personnel from the seller of the initial items, models cooperating with the seller, etc., or they can be virtual characters, and the images of all related characters are authorized.

[0053] The target item image can be provided directly by the user, or the user can provide a photograph of the target item, from which the target item image can be extracted.

[0054] The user terminal uploads the original display image, the target item image, and the associated information of the target item required to drive the digital human to the cloud. The product replacement system and the digital human driving system deployed in the cloud execute the subsequent related steps to generate the target display image and drive the digital human, resulting in a digital human video used to display the target item.

[0055] Step S202: Using the first image segmentation model, the handheld item in the original display image is segmented to obtain the first item mask.

[0056] Step S203: Guided by the item prompt words, the second image segmentation model removes occluded elements and completes the handheld image corresponding to the first item mask by using the second image segmentation model to obtain the second item mask.

[0057] In this model, both the first item mask (Mask1) and the second item mask (Mask2) are used to represent the position of the handheld item in the original display image. The first image segmentation model and the second image segmentation model are different image segmentation models. Occlusion elements are elements that occlude the handheld item, including the hand itself, and may also include other elements.

[0058] The first item mask, Mask1, is a mask for the initial item and the overlapping area between the person and the initial item in the original display image, to more completely represent the area where the initial item is located and the interaction area between the initial item and the hand.

[0059] Item prompts can be text-based or visual-based. They can be descriptive text about the initial item or a reference image of the initial item.

[0060] The handheld image corresponding to the first item mask Mask1 is a local image of the pixel area covered by the first item mask Mask1 in the original display image, that is, a local image composed of all pixels where the value of the first item mask Mask1 is 1, cropped from the original display image.

[0061] After acquiring the original display image, the complexity of the handheld object in the image can be identified. Based on the complexity, the type of object prompt can be determined, and guidance information can be generated and displayed to facilitate users providing corresponding object prompts. The complexity of the handheld object is used to characterize the complexity of its shape, pattern, etc.

[0062] For handheld items with low complexity, such as those below the first level of complexity, the item cues may consist only of text cues, such as "red coffee cup." For handheld items with high complexity, such as those greater than or equal to the first level of complexity, the item cues may include visual cues, such as a reference image of the handheld item, as well as text cues. Alternatively, the complexity can be divided into three consecutive intervals, such as a first interval, a second interval, and a third interval, using two thresholds, such as a first threshold and a second threshold. The item cues corresponding to the first interval consist only of text cues, the item cues corresponding to the second interval consist only of visual cues, and the item cues corresponding to the third interval consist of both text cues and visual cues.

[0063] The handheld object in the original display image can be identified by a pre-trained first image segmentation model, and the original display image can be segmented at the pixel level based on the recognition result. A binarized first object mask Mask1 can be generated based on the segmentation result.

[0064] The first item mask (Mask1) includes the area of ​​the hand interacting with the held item. That is, pixels with a value of 1 in the first item mask (Mask1) not only cover the area corresponding to the held item but also include the part of the hand that touches the held item. In the second item mask (Mask2), by culling and completing the occluded parts of the held item, pixels with a value of 1 can accurately represent the area where the held item is located.

[0065] After obtaining the first item mask Mask1, the first item mask Mask1, the original display image and the item prompt are input into the second image segmentation model. The second image segmentation model uses the area indicated by the first item mask Mask1 as the spatial constraint and the encoding of the item prompt as the condition vector to guide the image segmentation process, and obtains a more accurate second item mask Mask2 that indicates the area where the held item is located.

[0066] Optionally, the first image segmentation model is pre-trained and then fine-tuned using a training set of images of handheld objects; the second image segmentation model is not pre-trained and is trained based on a training set of images of handheld objects.

[0067] For example, the first image segmentation model can be U 2 The model could be a U-Squared Network (U-Squared Network), a Mask R-CNN (Mask Region-based Convolutional Neural Network), or another image segmentation model.

[0068] For example, the second image segmentation model can be a CLIP-based segmentation model (CLIPseg, CLIP-based Segmentation Model), such as the CLIP-Dissect model. Leveraging CLIP's cross-modal semantic understanding capabilities and combining it with object prompts, it can accurately identify and remove hand-occluded areas, thus completing the handheld object. Alternatively, the second image segmentation model can be a derivative model combining SAM (Segment Anything Model) and CLIP. By utilizing SAM's interactive segmentation capabilities and CLIP's text-guided features, it can complete the removal of occluded elements and the generation of a complete object mask.

[0069] First, key hand points in the original display image can be detected, and the hand-held area can be determined based on the detected key hand points. Then, the first image segmentation model segments the hand-held item within the hand-held area to generate a first item mask. The second image segmentation model, guided by item guide words, performs occlusion removal, hand-held item completion and segmentation within the hand-held area to generate a second item mask.

[0070] Among them, the key points of the hand include the base of the palm, the base of each finger, the knuckles, and the fingertips, usually totaling 21 key points.

[0071] Step S204: Based on the first item mask and the second item mask, the target item image is fused into the original display image to obtain the target display image; the handheld item in the target display image is the target item.

[0072] The target item is the item in the target item image.

[0073] By using pixel-level image fusion, the position of the handheld item in the original display image indicated by the second item mask Mask2 can be used as a spatial constraint to fuse the target display image with the original display image. The interaction edge between the hand and the handheld item, jointly indicated by the first item mask Mask1 and the second item mask Mask2, can be used to achieve smooth processing of the image fusion boundary. This allows the handheld item in the original display image to be replaced with the target item, resulting in a target display image for displaying the target item.

[0074] After obtaining the target display image, the edge breakage of the target object after image fusion can be eliminated by using the Poisson fusion model to ensure a natural transition between the target object and the edge of the hand.

[0075] The target object image can be deformed based on the second object mask, such as through pose adjustment and scaling, to ensure that the pose and size of the target object are consistent with the handheld object in the original display image. Then, based on the deformed target object image, image fusion is performed between the target object image and the original display image to obtain the target display image.

[0076] The handheld object image generation method provided in this embodiment addresses the problems of occlusion interference and incomplete object contour extraction in traditional handheld object replacement using a single segmentation model. It employs a collaborative working mode of dual image segmentation models. The first image segmentation model accurately locates the original handheld object to obtain a basic mask, i.e., the first object mask. Then, using the second image segmentation model combined with object prompts, occluding elements are intelligently removed and missing parts of the object are filled in, generating a high-fidelity second object mask. Finally, based on the dual masks, the target object and the original image are accurately fused, efficiently completing the handheld object replacement, ensuring a natural connection between the target object and the original image, and improving the visual quality and realism of the replaced image.

[0077] Figure 3 Flowchart of the method for generating handheld object images provided in this application Figure 2 In this embodiment Figure 2 Based on the embodiments, the method for generating images of handheld items is described in detail, such as... Figure 3 As shown, the method includes:

[0078] Step S301: Obtain the original display image and the target item image.

[0079] Optionally, obtaining an image of the target item includes: acquiring a photograph of the target item obtained by taking a picture of the target item; removing the background from the photograph of the target item to obtain an image of the target item.

[0080] The target item photo can be a photo taken by the user of the target item in a preset background, which can be a natural background, a green screen background, or other backgrounds.

[0081] Users upload photos of the target object to the cloud via their user terminals. The cloud then removes the background from the photos, retaining only the target object portion, and stores the resulting image.

[0082] Removing the background from a photo of a target object can be done using any algorithm, such as semantic segmentation models based on deep learning, threshold segmentation algorithms, chroma key matting algorithms based on green screen backgrounds, etc. This application does not limit the specific algorithms used.

[0083] By having users provide original photos of the target item and removing the background in the cloud, the user-side operation process is greatly simplified, lowering the barrier to obtaining target item images. Users simply need to take photos of the target item against any preset background, such as a natural background or a green screen, and upload them to the cloud via their devices, eliminating the need for manual background processing. The complex background removal process is automatically handled by the cloud, ultimately outputting a clean image that retains only the target item. This saves users time and effort in manually cutting out and editing images, and avoids image defects caused by unprofessional operations.

[0084] Optionally, the original display image is obtained, including: calculating the similarity between the target item image and the handheld item in multiple display images; and determining the display image with the highest similarity as the original display image.

[0085] If there is only one display image, then there is no need to perform the aforementioned steps; the display image can be used directly as the original display image for subsequent processing.

[0086] If there are multiple display images and the display items (i.e. the items being held) in different display images are different, in order to simplify the complexity of subsequent calculation steps and improve the visual naturalness of the image after object replacement, the original display image used to synthesize the target display image can be determined from multiple display images by the similarity between the items.

[0087] The similarity can be calculated by extracting the shape features of the target item image and the items in each displayed image, and then using the shape features extracted from the two images.

[0088] Each image (including the target item image and the display image) can be tagged with items. The similarity between items in two images can be determined by the item tags. Alternatively, the display images can be filtered based on the item tags to obtain candidate display images whose items are similar to the target item type. Then, based on the extracted shape features of the items in the images, the similarity between the target item in the target item image and the display items in the candidate display images can be calculated, and the candidate display image with the highest similarity is selected as the original display image.

[0089] Step S302: Using the first image segmentation model, the handheld item in the original display image is segmented to obtain the first item mask.

[0090] Step S303: Identify key hand points in the original display image based on the gesture recognition model.

[0091] Gesture recognition models can include MediaPipe Hands, OpenPose, YOLO-Pose, etc.

[0092] The key points of the hand are the core positioning points of the palm and fingers, including the base of the palm, fingertips, finger roots, and the joints of the fingers, usually 21 key points.

[0093] Step S304: Input the item prompt, hand key points, first item mask, and original display image into the second image segmentation model.

[0094] Step S305: The original display image is cropped based on the first item mask by the second image segmentation model to obtain the handheld image. Based on the item prompt and key points of the hand, the occluding elements of the handheld image are removed, and the occluding parts of the handheld item are filled in to obtain the second item mask used to represent the area where the handheld item is located.

[0095] The item prompt can be in natural language or structured instructions, used to inform the second image segmentation model that the object to be segmented is a handheld item. It can also include descriptive information about the handheld item, such as its type and color, for example, a red coffee cup. For handheld items with complex shapes and patterns, the item prompt can also include a reference image of the handheld item.

[0096] When performing image segmentation, the second image segmentation model quickly locates the area provided by the first object mask Mask1. Then, based on the object prompt words, it removes the hand and other parts that do not belong to the initial object, such as objects that accidentally entered the frame when the original display image was taken. Based on the object prompt words, it fills in the parts of the initial object that were originally covered. The filled-in initial object is then separated from the original display image to obtain the second object mask Mask2.

[0097] The second image segmentation model takes item cues, a first item mask (Mask1), hand key points, and the original display image as input, and a second item mask (Mask2) as output. During image segmentation, firstly, based on the first item mask (Mask1), the handheld region is determined. Then, combining the item cues, items that interact with the hand (e.g., contact, wrapping) are identified from the handheld region. Pixel-level classification removes parts that do not belong to the handheld items (i.e., the initial items), generating a preliminary segmentation result. Finally, based on the item cues, the occluded parts of the initial items in the preliminary segmentation result are completed, and the edges of the completed initial items are smoothed to obtain the second item mask (Mask2).

[0098] Step S306: Determine the edge of the hand based on the key points of the hand and the first object mask.

[0099] Among them, the hand edge is the edge where the hand interacts with the held object.

[0100] Based on the key points of the hand, the outline of the hand can be determined. Then, based on the distribution of pixels with a value of 1 in the first object mask Mask1, the boundary that contacts the hand object can be filtered out from the hand outline to obtain the edge of the hand.

[0101] When a hand comes into contact with an object, there are two states: obscuring and pressing. The obscuring state occurs when the hand merely covers the object without any actual contact or with minimal force, preventing deformation of the contact area. Conversely, the pressing state indicates actual contact between the hand and the object, causing deformation of the contact area. Based on this, the hand edge can be further divided into obscuring edges (the portion of the hand edge in an obscuring state) and pressing edges (the portion of the hand edge in a pressing state) according to the interaction state with the object.

[0102] Optionally, the hand edge includes an occlusion edge and a pressing edge; determining the hand edge based on hand key points and a first object mask includes: identifying the interaction area between the hand and the held object based on hand key points and the first object mask; determining a first pixel in the interaction area of ​​the original display image that is in an occlusion state and a second pixel in a pressing state; and determining the occlusion edge and the pressing edge based on the first pixel and the second pixel, respectively.

[0103] The interactive area is the area where the hand overlaps or touches the object being held.

[0104] Using the key points of the hand as spatial anchors, target pixels that spatially overlap with the key points of the hand within the first item mask Mask1 are selected. After integrating the target pixels, the interaction area between the hand and the held item is obtained.

[0105] In the interactive area, the state of pixels belonging to the hand (or pixels) is identified, namely, the pressed state or the occluded state. Based on the identification results, the hand pixels (pixels belonging to the hand) in the interactive area are divided into the first pixel (pixels in the occluded state) and the second pixel (pixels in the pressed state).

[0106] The occlusion edge is obtained based on the boundary line between the first pixel and the surrounding object pixels; the pressing edge is obtained based on the boundary line between the second pixel and the surrounding object pixels. Both the occlusion edge and the pressing edge belong to the outer contour of the hand.

[0107] By recognizing the pressed and occluded edges through the interaction between the hand and the object, it is easy to perform adaptive processing on the corresponding areas of different edges. This ensures that the edges are connected naturally when the object is replaced, restoring the integrity of the object and retaining the realism of the interaction between the hand and the object.

[0108] Optionally, determining the first pixel in the original display image that is occluded and the second pixel in the pressed state in the interactive area includes: for each pixel in the original display image located in the interactive area, determining whether the pixel is the first pixel or the second pixel based on at least one of the following: the distance between the pixel and the edge of the handheld object, the feature value of the pixel, and the pressure confidence of the pixel; the edge of the handheld object is obtained based on the second object mask, and the pressure confidence is used to characterize the confidence of the corresponding pixel in the pressure applied to the handheld object.

[0109] The edge of the handheld item is the edge of the handheld item, which is composed of pixels with a value of 1 adjacent to pixels with a value of 0 in the second item mask Mask2.

[0110] The feature values ​​of pixels within the interactive area can include features such as color distribution, texture, brightness gradient, and edge density, which are used to determine the interaction state between the hand corresponding to the pixel and the held object. Based on these feature values, it can identify whether features such as epidermal diffusion features and white squeezing features exist. If they exist, the object is in a pressing state; if they do not exist, the object is in an occluded state.

[0111] The epidermal diffusion feature is used to characterize the diffuse texture changes of hand skin texture caused by pressing and adhering to the surface of an object. The corresponding pixel feature values ​​will show the characteristics of skin texture and object surface texture mixing and blurred diffusion at the edges.

[0112] The white squeeze feature is used to characterize the phenomenon of increased local brightness on the skin of the hand due to squeezing an object. The corresponding pixel feature value will be characterized by a significant increase in the proportion of white component in the color channel and a brightness value higher than the surrounding pixels.

[0113] The pressure confidence score of a pixel is used to quantify the probability that the area where the pixel is located belongs to the hand pressure area. The value is usually 0~100% or 0~1. The higher the value, the higher the confidence that the pixel belongs to the hand pressure area.

[0114] For each pixel within the interactive area, if the pixel meets any of the following conditions, the pixel is determined to be in an occluded state and belongs to the first pixel: the distance from the edge of the held object is greater than or equal to a preset value such as 2 pixels, there is no skin diffusion feature or no white squeezing feature, and the pressure confidence is less than or equal to the first confidence, such as 0.3.

[0115] For each pixel within the interactive area, if the pixel meets any of the following conditions, the pixel is determined to be in a pressed state and belongs to the second pixel: the distance to the edge of the held object is less than a preset value, there is a skin diffusion feature or a white squeezing feature, and the pressure confidence is greater than or equal to the second confidence, such as 0.7.

[0116] By integrating multiple parameters such as distance, feature value, and pressure confidence, the system can identify the state of pixels within the interactive area, improving the accuracy of identification and providing a foundation for subsequent differentiated adaptive processing of different types of edges.

[0117] Step S307: Determine the replacement area based on the second item mask and the edge of the hand.

[0118] Step S308: Map the pixels of the target item image to the replacement area, replace the pixels in the replacement area of ​​the original display image, and perform adaptation processing on the pixels in the area corresponding to the hand edge of the original display image to obtain the target display image.

[0119] The adaptation processing includes at least one of the following: edge protection processing and transparency processing.

[0120] The replacement area can be obtained by subtracting the hand region corresponding to the hand edge from the area consisting of pixels with a value of 1 in the second item mask Mask2. Within the replacement area, the pixels of the original displayed image are directly replaced with the pixels of the corresponding positions in the target item image.

[0121] A third marker, such as a blue marker, can be assigned to the pixels in the replacement area, which facilitates subsequent image fusion by performing corresponding processing based on the different assigned markers.

[0122] During image fusion, for pixels marked in blue in the original display image, the pixel value of the pixel is directly replaced with the pixel value of the corresponding position in the target item image, thereby initially replacing the handheld item in the original display image with the target item.

[0123] In order to achieve a natural connection between the target object and the hand, in addition to the replacement process mentioned above, it is also necessary to implement a differentiated adaptation strategy for the pixels located near the edge of the hand in different scenes.

[0124] The edge of the hand can be extended outward by 2 to 5 pixels to obtain a buffer area. Within the buffer area, the pixel adaptation processing method is determined according to parameters such as the distance between the pixel and the edge of the held object, the feature value of the pixel, and the distance between the pixel and the key points of the hand. For example, edge protection processing or transparency processing can be used, and the specific processing intensity can be determined, such as transparency weight, edge enhancement or blur amplitude.

[0125] In this implementation, a gesture recognition model was introduced to identify key hand points. The identified key hand points can assist the second image segmentation model in more accurately separating the hand region and the object region, more completely restoring the spatial position and outline of the held object, and improving the generation accuracy and boundary fit of the second object mask. In the image fusion stage, the detected hand edges were used to accurately divide the region and implement targeted differentiated processing strategies to achieve intelligent fusion. By edge protection and transparency processing of the corresponding areas of the hand edges, the connection between the target object and the hand not only conforms to the physical interaction logic but also has a natural visual transition effect. The final fused image output shows a significant improvement in detail realism and overall coordination.

[0126] Optionally, the pixels of the original display image located at the edge of the hand are subjected to adaptation processing to obtain the target display image, including: performing edge protection processing on the pixels of the original display image located in the area corresponding to the occlusion edge, and performing transparency processing on the pixels of the original display image located in the area corresponding to the pressing edge, to obtain the target display image.

[0127] Edge protection processing is used to strengthen the boundary between the hand and the target object, avoiding jagged edges or color intrusion, and weakening splicing marks; transparency processing is used to make the pixels in the corresponding area show a gradual degree of transparency by grading the transparency weight.

[0128] Edge protection processing can include smoothing processes, such as Gaussian blurring.

[0129] Optionally, based on the first and second item masks, the target item image is fused into the original display image to obtain the target display image, including: determining occlusion edges and pressing edges based on hand key points and the first item mask; marking the first and second item masks based on the occlusion edges and pressing edges to generate a marked item mask; extracting visual features from the original display image, the marked item mask, and the target item image, as well as semantic features from the item prompts; performing multimodal fusion processing on the visual and semantic features and embedding them into a time step to obtain a global feature sequence; and using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item.

[0130] The item mask can include four channels to indicate different areas, such as the pressed area, the occluded area, the body area, and other areas, where the pressed area is the area corresponding to the pressed edge, such as... Figure 4 The second edge buffer region, Buffer2, is defined as the occlusion region, which corresponds to the occlusion edge. Figure 4 The first edge buffer area Buffer1; the main body area is the area corresponding to the second item mask minus the pressed area and the occluded area, as shown in the example. Figure 4 The replacement area is defined as follows: other areas are the remaining areas of the first item mask excluding the pressed area, the occluded area, and the main body area. Different areas can be marked with different tags, while other areas can be left unmarked.

[0131] Visual features include, but are not limited to: contour features, edge gradient features, texture detail features, color distribution features, spatial occupancy features, and interaction type features with the hand area.

[0132] It can utilize any image-adaptive feature extraction layer to extract visual features from the original display image, the labeled object mask, and the target object image. Examples include CNN (Convolutional Neural Network), ResNet (Residual Neural Network), MobileNet (Mobile Neural Network), VAE (Variational Autoencoder), and Transformer.

[0133] The original display image, the marked item mask, and the target item image can be input into the VAE encoder, and the VAE encoder will output the visual features of the original display image, the marked item mask, and the target item image respectively.

[0134] The semantic features of item cues can be extracted using an encoder that matches the type of the cue. For text-based item cues, features can be extracted using models such as BERT (Bidirectional Encoder Representations from Transformers), CLIP-text, and RoBERTa (Robustly Optimized BERT Pretraining Approach) to obtain semantic features. For visual item cues, features can be extracted using models such as CLIP-vision, ViT (Vision Transformer), and ResNet to obtain semantic features.

[0135] An encoder for extracting semantic features of text-based item prompts can be a dual encoder consisting of a T5 (Text-To-Text Transfer Transformer) encoder and a CLIP (Contrastive Language-Image Pre-training) text encoder.

[0136] For example, the CLIP text encoder extracts the visual semantics corresponding to the item prompts, such as the color distribution corresponding to red, and outputs a 1×1×768-dimensional vector. The T5 encoder parses the long text logic of the item prompts, such as the shape description of a coffee cup, and outputs a 1×1×1024-dimensional vector. The two vectors are fused through a projection layer to obtain a 1×1×1280 (768+512)-dimensional semantic vector, which is the semantic feature.

[0137] Visual and semantic features can be aligned across modalities and then fused to obtain fused features. Time steps are then embedded into the fused features to obtain a global feature sequence. Cross-attention can be used for fusion.

[0138] A projection layer can be used to map visual features or semantic features to the feature space where the semantic features or visual features reside, thereby achieving cross-modal alignment between the two; alternatively, two independent projection layers can be used to map visual features and semantic features to a common feature space, respectively, thereby achieving cross-modal alignment between the two.

[0139] Before fusion, the visual features can be adjusted in posture based on the grasping parameters indicated by semantic features, so that the body posture of the person and the position of the key points of the hand in the visual features are adapted to the grasping parameters.

[0140] Time step information can be converted into a time step vector through sinusoidal position encoding, and the time step vector can be concatenated with the fused features to achieve time step embedding.

[0141] The visual features, semantic features, and temporal step information integrated in the global feature sequence are converted into a conditional guidance vector for the diffusion model. This vector is embedded in each temporal step of the reverse denoising process to clarify the optimization direction of the latent code at the current temporal step, so that the target item matches the hand posture and interaction position in the original display image, while also meeting the semantic requirements of the item prompt words.

[0142] In the latent code processing stage, the latent code is first initialized to generate a random noise latent code that matches the size of the target region, which is the area where the handheld object is located. Then, at each time step, the gradient correction vector of the noise latent code at the current time step is calculated using the conditional constraints provided by the global feature sequence, driving the latent code to gradually remove noise and obtain the target latent code. The target latent code is then input into a pre-trained decoder for decoding to obtain the target display image.

[0143] Optionally, multimodal fusion processing is performed on visual and semantic features and embedded with time steps to obtain a global feature sequence, including: determining the grasping parameters of the target object based on the target object feature map; the target object feature map is the extracted visual features of the target object image; generating a hand feature map based on the grasping parameters and the original display image; shifting the original feature map and the mask feature map based on semantic features; the original feature map and the mask feature map are respectively the extracted visual features of the original display image and the feature map of the marked object mask; adjusting the pose of the shifted original feature map based on the hand feature map, and fusing the pose-adjusted original feature map with the shifted mask feature map to obtain a joint feature map; interacting with the semantic features, the target object feature map, and the joint feature map based on a cross-attention mechanism to generate a semantically guided feature map; concatenating the time step vector transformed from time step information with the semantically guided feature map to obtain a temporal conditional feature map; and performing attention encoding on the temporal conditional feature map to obtain a global feature sequence.

[0144] The grasping parameters may include information such as the parameters of the grasping region and the region mask. The grasping region is the area of ​​the target object indicated by the target object feature map that is suitable for grasping. The parameters of the grasping region may include the center coordinates, coordinate range, width, thickness, and area of ​​the grasping region. The hand feature map is used to characterize the hand features of the target person when grasping the target object; the target person is the person in the original display image.

[0145] Based on the feature map of the target object, the type and size information of the target object can be identified, and the gripping parameters of the target object can be determined based on the type and size information of the target object.

[0146] Optionally, based on the target item feature map, the grasping parameters of the target item are determined, including: performing target detection on the target item feature map to obtain the item boundary and size information of the target item; segmenting the target item feature map based on the item boundary and size information of the target item to obtain multiple segmented regions; calculating the region parameters of each segmented region, and determining the grasping region from the multiple segmented regions based on the region parameters; and determining the grasping parameters of the target item based on the determined grasping region.

[0147] The regional parameters of the segmented region include, but are not limited to, information such as area, location, and thickness.

[0148] An object detection network can be used to detect objects in the feature map of an object, obtain the boundary of the object, and output the size information of the object.

[0149] The boundaries of the target item can include the coordinates of its center point and the range of those coordinates. Dimensional information can include the length, width, and height of the target item. It can also be based on a pre-defined item category database, identifying the target item's category to determine its thickness. Thickness information characterizes the overall thickness of the target item or the overall thickness of its different components, such as the diameter of a water cup. Alternatively, when the water cup includes a body and a handle, the thickness information can include the thickness of the body and the handle, for example, a body thickness of 10cm and a handle thickness of 2cm.

[0150] Lightweight networks can be used for object detection, such as YOLOv8 (You Only Look Once version 8), YOLOX-Nano (You Only Look Once X Nano), and EfficientDet-Lite (Efficient ObjectDetector Lite).

[0151] Image segmentation networks can segment a target object's feature map based on its boundary and size information, resulting in multiple segmented regions. Specifically, the target object's feature map, boundary, and size information are input into the image segmentation network. The network then partitions the feature map into regions corresponding to the coordinates of the object's boundary, outputting multiple segmented regions. Each region corresponds to a different part of the target object, such as the upper, middle, or lower part.

[0152] Image segmentation networks can employ FCN (Fully Convolutional Networks), SAM (Segment Anything Model), Mask R-CNN (Mask Region-based Convolutional Neural Networks), and others.

[0153] After obtaining multiple segmented regions, to determine the most suitable gripping area for the target object, the first step is to define the regional parameters of each segment, such as its area, position, and thickness. The area indicator represents the proportion of the segment's area to the total area of ​​the target object. The position indicator can be represented by the distance between the center of the segment and the center of the target object. The thickness indicator can be characterized by the matching degree between the thickness information of the segment, such as its diameter, and the corresponding gripping range. A mapping relationship between object type and gripping range can be pre-established. For example, thin objects like lipstick correspond to a gripping range of 2-5cm, while a thermos cup corresponds to a gripping range of 10-25cm. For objects with complex structures, different gripping ranges can be set for different areas. The gripping ranges corresponding to various types of objects are pre-collected and stored to facilitate the determination of the target object's thickness information.

[0154] The values ​​for area, position, and thickness can all range from 0 to 1. The larger the area of ​​the target item, the larger the area index; the closer the target item is to its center, the larger the position index; and the higher the degree of matching with the corresponding grip area, the larger the thickness index. For example, when the target item is within the corresponding grip area, the thickness index can be 1.

[0155] The gripping area can be determined from multiple segmented regions based on the weighted results of area, location, and thickness metrics.

[0156] Each segmented region corresponds to a feature vector composed of area, location, and thickness metrics. The feature vector for each segmented region is input into an MLP (Multi-Layer Perceptron) network for scoring. The MLP network determines the weights of different features in the feature vector based on the target object type, and the score for the corresponding segmented region is obtained by weighting the three features. The segmented region with the highest score is identified as the grasping region, and its grasping parameters are output, such as the center coordinates, coordinate range, width, thickness, area, and region mask.

[0157] Based on the grasping parameters, the corresponding hand key points can be determined, and then the hand key points and the body key points of the target person identified from the original display image can be used to generate a hand feature map.

[0158] The key points of the hand are the core positioning points of the palm and fingers, including the base of the palm, fingertips, finger roots, and the joints of the fingers, usually 21 key points.

[0159] Optionally, a hand feature map is generated based on the grasping parameters and the original display image, including: determining the key points of the hand that are adapted to the target object based on the grasping parameters and the preset hand model; obtaining the key points of the person's body in the original display image based on the human posture constraint model, and correcting the key points of the hand based on the key points of the body; mapping the corrected key points of the hand to two-dimensional pixel coordinates, and generating a hand feature map after interpolation.

[0160] For example, the preset hand model can be MANO (Metric-Affine Hand Model).

[0161] The original display image or the target person image cropped from the original display image is input into the OpenPose model to calculate 17 body joint key points. Based on the constraints of the body joint key points on the hand key points, the hand key points are adjusted to avoid incoordination between the hand and the body during the hand pose bar process.

[0162] Specific constraint methods can include elbow flexion constraints and body distance constraints. The elbow flexion constraint involves calculating the shoulder-elbow vector based on key body joint points, and simultaneously calculating the elbow-hand vector based on key body joint points and hand key points. The angle between these two vectors (the elbow flexion angle) is then calculated. This elbow flexion angle should be between 30° and 120°. If the calculated result is outside this range, the elbow flexion is too small or too large, requiring correction of the hand key points to ensure the elbow flexion angle falls within the range of [30°, 120°]. The body distance constraint involves calculating the distance from the hand to the body's central axis based on key body joint points and hand key points. This distance should be between 15 cm and 30 cm. If the calculated result is outside this range, the hand is too close or too far from the body, requiring correction of the hand key points to ensure the distance falls within the aforementioned range. After adjusting the hand key points according to these constraints, ensuring they meet the constraints, the corrected hand key points are obtained.

[0163] Spatial Transformer Networks (STNs) can be used to map the corrected hand keypoints into two-dimensional pixel coordinates. Further interpolation processing, such as bilinear interpolation, can be performed on the mapped two-dimensional pixel coordinates to generate and output a hand feature map.

[0164] Based on semantic features, the region where the handheld item is located in the original feature map and the mask feature map can be offset so that the state of the handheld item after offset conforms to the semantics of the item prompt, so as to facilitate item replacement.

[0165] Semantic features can be used to determine the spatial offset between the original state of the held object in the original feature map and the expected state described in the semantic features. This spatial offset can then be used to offset the original feature map and the mask feature map, such as by translation, rotation, scaling, etc.

[0166] Optionally, based on semantic features, the original feature map and the mask feature map are offset, including: using a cross-attention network, with semantic features as the query and the original feature map as the key, calculating attention weights and assigning them to the original feature map to obtain a semantically guided original feature map; inputting the semantically guided original feature map into an offset prediction network to obtain offset parameters and sampling weight parameters; converting the offset parameters into an offset matrix using a spatial transformation network, and performing spatial coordinate transformation on the original feature map and the mask feature map based on the offset matrix; and using the sampling weight parameters to perform weighted fusion of the offset feature points and surrounding feature points to obtain the offset original feature map and the offset mask feature map; the offset feature points are the corresponding feature points formed by the spatial coordinate transformation of the feature points in the original feature map and the mask feature map.

[0167] Cross-attention networks can extract location-related semantic information from semantic features, such as spatial descriptions like "above," "left," and "middle," or local semantics like "person's hand" and "product edge." Based on these semantics, the corresponding parts in the original feature map are weighted to give higher weights to the location-related features in the original feature map.

[0168] Migration prediction networks are used to predict migration parameters and output sampling weight parameters. Migration prediction networks can consist of pooling layers and multiple fully connected layers, or pooling layers and multiple lightweight convolutional network layers.

[0169] After inputting the semantically guided original feature map into the offset prediction network, it first undergoes global pooling through a pooling layer to compress the features into a global vector. Then, a fully connected layer or a lightweight convolutional network layer maps the global vector to a low-dimensional space and performs a transformation to obtain the offset parameters. Based on the offset parameters, sampling weight parameters are determined. These sampling weight parameters characterize the contribution weights of the feature points surrounding the original feature point (the feature point before offset) to the offset feature point (the feature point after offset) during subsequent feature sampling, i.e., the offset process.

[0170] For example, suppose a feature point M in the original feature map has surrounding feature points M1, M2, M3, and M4. After offsetting, the offset position of M is M'. To avoid the aforementioned issues of feature sparsity, subsequent feature sampling processes need to refer to M1, M2, M3, and M4 for feature fusion. The sampling weight parameter is used to characterize the weights of M1, M2, M3, and M4 respectively during feature fusion.

[0171] The sampling weight parameters can be calculated using bilinear interpolation. The core principle is to assign weights based on the distance between the offset feature point (the feature point being offset) and its surrounding feature points; the closer the distance, the higher the weight. Using the example above, assuming the distances between M´ and M1, M2, M3, and M4 are 4, 3, 2, and 1 respectively, then the weights assigned to M1, M2, M3, and M4 are 0.4, 0.3, 0.2, and 0.1 respectively.

[0172] In the bilinear interpolation calculation process described above, the sampling range (i.e., the number of surrounding feature points selected) can be fixed. A fixed sampling range usually only achieves good results when the shape is simple and the degree of offset is small; for complex shapes and large degrees of offset, a dynamic sampling method can be used, that is, the sampling range is dynamically adjusted based on the offset parameter.

[0173] The sampling orientation of dynamic sampling can be determined based on preset rules, as follows: The larger the translation distance, the larger the sampling range. For example, when the translation distance is < 2 tokens (pixel blocks), the sampling range can take 4 surrounding feature points; when the translation distance is ≥ 2 tokens, the sampling range can take 8 surrounding feature points. The larger the rotation angle, the larger the sampling range. For example, when the rotation angle is < 30°, the sampling range takes 4 surrounding feature points; when the rotation angle is ≥ 30°, the sampling range takes 8 surrounding feature points. When the scaling ratio is greater than or less than a preset value (i.e., the greater the scaling degree), the larger the sampling range. For example, when the scaling ratio is > 1.2 (enlargement) or < 0.8 (reduction), the sampling range takes 8 surrounding feature points; when 0.8 ≤ scaling ratio ≤ 1.2, the sampling range takes 4 surrounding feature points.

[0174] Furthermore, dynamic sampling can be performed based on the location of feature points in the image. For example, if M is located at the edge of product A, the sampling range is increased by 2 points to prioritize covering the original features in the edge direction and avoid edge breakage. If M is located in the solid color area of ​​product A, the sampling range can be reduced by 2 points to reduce redundant calculations. If M is located at the boundary between product A and the background (such as the bottom of a cup contacting the table), the sampling range is fixed to the minimum, such as 4 points, to avoid introducing background features.

[0175] In addition to dynamic sampling based on preset rules, the sampling range can also be calculated using a pre-trained prediction network.

[0176] The offset parameters are converted into an offset matrix using a spatial transformation network, and the original feature map and the mask feature map are then transformed in spatial coordinates based on the offset matrix. Using sampling weight parameters, the offset feature points are weighted and fused with surrounding feature points within the sampling range to obtain the offset original feature map and the offset mask feature map.

[0177] After obtaining the hand feature map and completing the offset, the offset original feature map is further adjusted using the hand feature map, so that the person's hand and body posture in the original feature map matches the hand feature map, thus completing the posture adjustment.

[0178] The original feature map after pose adjustment and the offset mask feature map can be mapped and concatenated along the channel dimension to obtain a joint feature map.

[0179] Optionally, based on the hand feature map, pose adjustment is performed on the offset original feature map, including: extracting style information and pose information from the offset original feature map; correcting the hand features in the pose information based on the hand feature map; concatenating the corrected hand features with the style information to obtain a concatenated original feature map; generating a product grasping mask and a hand feature mask based on the grasping region; extracting product grasping feature sequences and hand feature sequences from the concatenated original feature map using the product grasping feature sequence as the query and the hand feature sequence as the key; calculating attention weights through a cross-attention network and adjusting the hand feature sequences based on the attention weights, mapping them to the original feature map to obtain the pose-adjusted original feature map.

[0180] The offset original feature map can be input into two parallel convolutional layers to extract style and pose information respectively. Then, the pose information and hand feature map are input into an MLP network for weight allocation. The MLP network, based on dynamic weight allocation rules learned during training, calculates weights for the features representing the hand portion in the pose information and the features represented by the hand feature map, and performs a weighted sum based on these calculated weights to correct the pose information, resulting in corrected pose information. For example, for small or complex objects, which are difficult to hold, a higher weight is assigned to the hand feature map based on the grasping posture to enhance the accuracy of hand grasping; conversely, for regular large items, a lower weight can be assigned to the hand feature map to preserve the consistency between the item and the hand as a whole. Finally, the corrected pose information is concatenated with the aforementioned style information to obtain a concatenated original feature map.

[0181] The aforementioned steps only adjusted the hand posture, but there may still be some details that need optimization when the target object is actually grasped, especially the contact position between the fingers and the target object. Therefore, it is necessary to generate a product grasping mask and a hand feature mask based on the grasping area, so as to extract and stitch the feature map corresponding to the original feature map through the mask, and use a cross-attention network to perform weighted fusion on the extracted feature map to obtain the original feature map in which the target object and the hand maintain a better grasping posture, that is, the original feature map after posture adjustment.

[0182] The product gripping mask is the mask corresponding to the gripping area, and the hand feature mask is the mask corresponding to the hand feature map.

[0183] Based on the product grasp mask and hand feature mask, two feature maps are obtained by stitching together the product grasp and hand parts from the original feature map. Based on these two feature maps, product grasp feature sequences and hand feature sequences are obtained. Using a cross-attention network, the product grasp feature sequence is used as the query object Q, and the hand feature sequence is used as the query object, i.e., key K and value V. Attention weights for the hand feature sequence are calculated based on the product grasp feature sequence, and the hand feature sequence is weighted according to the calculation results. This yields a hand feature sequence corrected based on the product contact position. Mapping this sequence yields the original feature map showing the optimal grip posture between the target item and the hand—the posture-adjusted original feature map.

[0184] First, convolutional layers can be used to map the pose-adjusted original feature map to the target object feature map in terms of channel dimensions, maintaining consistency (e.g., mapping to 512 or 1024 dimensions). The channels of the offset mask feature map can also be mapped to fixed values, such as 64 or 128 dimensions. Alternatively, fully connected layers can be used to ensure the channel dimensions of the semantic features are consistent with those of the target object feature map. The channel-mapped original feature map and the mask feature map are then concatenated to obtain a joint feature map. The channel-mapped target object feature map and semantic features are then output separately to obtain a semantically guided feature map in subsequent steps.

[0185] First, the joint feature map and the target feature map are concatenated along the channel dimension to obtain a concatenated feature map, which is then mapped to a concatenated sequence. Cross-attention calculation is performed on the concatenated sequence and semantic features. After the calculation is completed, the concatenated sequence is weighted based on the calculated attention weight matrix to obtain a feature map with semantic guidance.

[0186] Optionally, based on a cross-attention mechanism, semantic features and joint feature maps are interacted to generate semantically guided feature maps, including: concatenating the joint feature map and the target item feature map along the channel dimension to obtain a concatenated feature map, and mapping the concatenated feature map to a concatenated sequence; using semantic features as queries and the concatenated sequence as keys, attention weights are calculated through a cross-attention network to obtain an attention weight matrix; based on the attention weight matrix, the concatenated sequence is adjusted to obtain a preliminary fused feature map; the similarity between the original item and the target item in multiple spatial locations in the preliminary fused feature map is calculated; based on the similarity in each spatial location, the weight of the target item in each spatial location is determined, and the weight of the main body region (i.e., the aforementioned replacement region or ontology region) corresponding to the marked item mask is adjusted to obtain a weight vector; based on the weight vector, the features of the original item and the features of the target item in the preliminary fused feature map are weighted and fused to obtain a semantically guided feature map.

[0187] After obtaining the concatenated sequence, an attention weight matrix is ​​calculated using semantic features as the query Q and the concatenated sequence as the key K and value V. Higher weights indicate that the corresponding image region needs to adhere more strictly to the semantic feature constraints. Weights are assigned to the concatenated sequence based on the attention weight matrix, and the features at corresponding positions in the concatenated sequence are weighted according to the assigned weights. The weighted concatenated sequence is then converted into a feature map, resulting in a preliminary fused feature map.

[0188] Since the initial fusion feature map contains features of both the original item and the target item, and the original item is the handheld item in the original display image, the aforementioned fusion process only injected textual conditions, i.e. semantic features, without distinguishing the different emphases in the item replacement process. Therefore, it is still necessary to further adjust the main area of ​​the item based on the similarity between the original item and the target item.

[0189] Let's take product A as the source item and product B as the target item as an example. The weights corresponding to the features of product A are denoted as WA, and the weights corresponding to the features of product B are denoted as WB. For the same position, WA + WB must satisfy 1.

[0190] In the initial fusion feature map, the feature corresponding to product A is denoted as FA, and the feature corresponding to product B is denoted as FB. For FA and FB, the similarity between them at each spatial location is calculated, and a similarity map is generated based on the similarity calculation results. This similarity map is used to represent the similarity between FA and FB at each spatial location. In the above similarity map, the closer the similarity value is to 1, the more similar the features of product A and product B are at that location, and the easier it is to replace product A with product B in that part; conversely, the closer the similarity value is to 0, the greater the feature difference between product A and product B at that location, and the more difficult it is to replace product A with product B in that part.

[0191] Based on the similarity map above, WA and WB can be calculated using the following formula:

[0192] WB(h,w)=α⋅S(h,w)+β; WA(h,w)=1-WB(h,w);

[0193] The above WA(h,w) and WB(h,w) represent the WA and WB of the spatial location (h,w) respectively, S(h,w) characterizes the similarity of a spatial location (h,w), and α and β are learnable parameters obtained through training.

[0194] The above calculations result in the following: regions with larger S values ​​have higher WB values, and regions with smaller S values ​​have lower WB values. Specifically, the more similar product A is to product B, i.e., the larger the S value, the easier it is to perform the replacement task. Replacing product A with product B in this region results in less visual difference, so a higher WB value can be set for this region. Conversely, the greater the difference between product A and product B, i.e., the smaller the S value, the more difficult it is to perform the replacement task. Directly performing the replacement task will result in more obvious visual breaks, so a lower WB value can be set for this region to retain more of product A's features for the transition.

[0195] For example, product A includes a gray metallic cup body and a dark wooden handle, while product B includes a gray plastic cup body and a white plastic handle. After similarity calculation, the feature similarity between product A and product B in the cup body region is 0.8, resulting in a WB of 0.9 and a WA of 0.1. Similarly, the feature similarity between product A and product B in the handle region is 0.2, resulting in a WB of 0.9 and a WA of 0.1. For other regions, the similarity may be moderate, resulting in a WB and WA of 0.5 for each region.

[0196] The similarity between two features can be calculated using cosine distance, Euclidean distance, Manhattan distance, Jaccard similarity coefficient, Pearson correlation coefficient, etc.

[0197] Because the target object itself may have significant differences between adjacent parts, such as the cup body and handle in the previous example, if the WB values ​​of adjacent areas differ too much, it will cause abrupt boundary issues at the junction of adjacent areas. Therefore, further spatial continuity processing is required.

[0198] For the WB of each spatial location or region calculated above, the WB of multiple adjacent locations is smoothed, for example, the average of the WB of these multiple adjacent locations is used as the WB of the center location.

[0199] The main area is the region where the handheld item, indicated by the item mask, is located. For the weight (WB) of the target item within the main area, the WB of each spatial location within the main area can be increased, for example, by 0.2, to ensure the accuracy of the main area replacement. For spatial locations outside the main area, no further adjustment is needed.

[0200] For example, the WB of each spatial location within the main body area can be uniformly set to 0.8, and the larger of the WB of each spatial location within the main body area calculated based on similarity can be taken as the final WB of each spatial location.

[0201] After calculating WA and WB through the aforementioned steps, the weight vectors (WA and WB for each spatial location) are obtained. Based on WA and WB, the features of the original item and the target item in the preliminary fused feature map are weighted and fused to obtain a semantically guided feature map. The channel dimension of the semantically guided feature map is consistent with the channel dimension of the original feature map or the target item feature map.

[0202] Optionally, attention encoding is performed on the temporal conditional feature map to obtain a global feature sequence, including: extracting visual and text marker sequences from the temporal conditional feature map using an attention layer composed of a multi-head self-attention structure and a normalization layer; performing cross-modal alignment and feature fusion on the visual and text marker sequences through dual-path cross-attention calculation; the dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key; performing three-dimensional rotation position encoding on the processed visual feature sequence to obtain an image feature sequence with positional information; generating semantic constraints based on the processed text feature sequence; inputting the image feature sequence with positional information into the attention layer composed of a multi-layer multi-head self-attention structure and a normalization layer, calculating attention weights under the semantic constraints, and outputting the global feature sequence after residual connection, feedforward network, and layer normalization processing.

[0203] Temporal conditional feature maps can be input in parallel into the image preprocessing subunit and the text preprocessing subunit. Through differential feature extraction from the two subunits, visual label sequences and text label sequences can be obtained respectively.

[0204] The image preprocessing subunit and the text preprocessing subunit adopt a parallel structure and are used to extract the image stream and text stream corresponding to the aforementioned temporal condition feature map, respectively.

[0205] The image preprocessing subunit consists of an image projection layer and an image self-attention layer connected in series, and is used to capture the dynamic relationship between spatial morphology and time step in the temporal condition feature map.

[0206] The image projection layer consists of a 3×3 convolutional layer, a batch normalization layer, and residual connections. The 3×3 convolutional layer extracts local spatial features and maps dimensions from the temporal conditional feature map. The batch normalization layer stabilizes the feature distribution and accelerates training convergence. The residual connections effectively avoid the vanishing gradient problem in deep networks. After processing by this layer, the two-dimensional temporal conditional feature map is transformed into a one-dimensional sequence of visual tokens. Each visual token corresponds to a spatial location in the feature map and carries the visual features and temporal step information at that location.

[0207] The image self-attention layer consists of a multi-head self-attention structure and a normalization layer. The image self-attention layer allows each token in the visual token sequence to autonomously attend to other spatially adjacent and temporally consecutive visual tokens. Through the weight allocation of multi-head self-attention, it can accurately capture the dynamic relationships between visual features, such as the spatial morphological change of a hand from "open" to "held," and the temporal correspondence between the target object and the contact area of ​​the hand. The normalization layer then standardizes the features after attention calculation to obtain a visual tag sequence.

[0208] The text preprocessing subunit consists of a text projection layer and a text self-attention layer connected in series, and is used to mine the semantic information embedded in the temporal condition feature map and the corresponding relationship between the time steps.

[0209] The text projection layer consists of a 1×1 convolutional layer, a global average pooling layer, and residual connections. The 1×1 convolutional layer is responsible for cross-channel feature fusion and dimensionality compression of the temporal conditional feature map. The global average pooling layer aggregates global feature information and weakens the interference of local spatial noise. The residual connections ensure the integrity of feature transmission. After processing by this layer, the semantic information carried in the temporal conditional feature map is extracted and transformed into a one-dimensional text token sequence. Each text token corresponds to a set of semantic features and is associated with the corresponding time step information.

[0210] The text self-attention layer is similar to the image self-attention layer, both consisting of a multi-head self-attention structure and a normalization layer. The core function of the text self-attention layer is to allow each token in the text token sequence to autonomously attend to other semantically related and temporally connected text tokens. Through attention weight allocation, temporal associations of semantic features can be established, such as the matching relationship between the attribute semantics of the text prompt "red ceramic cup" and the hand-holding action at different time steps, and the temporal correspondence between the target item style semantics and scene features. The normalization layer then standardizes the semantic features after attention calculation, resulting in a text token sequence.

[0211] It is important to emphasize that the above preprocessing sub-unit approach does not simply separate the fused visual, semantic, and temporal features in the temporal conditional feature map. Instead, it achieves targeted feature extraction through differentiated weight parameters designed for the two sub-units. For the image preprocessing sub-unit, its weight parameters are optimized for visual information and spatiotemporal correlation, focusing on preserving the spatial location, morphological details, and temporal step correlation of features. For the text preprocessing sub-unit, its weight parameters are optimized for semantic information and temporal correlation, strengthening the matching relationship between feature category attributes, style descriptions, and temporal steps.

[0212] Visual tag sequences and text tag sequences can be input into the dual-stream cross-attention subunit for dual-path cross-attention calculation.

[0213] The dual-stream cross-attention subunit consists of multiple layers of dual-stream attention blocks. Each layer of dual-stream attention block includes a dual-path cross-attention subunit and two sets of post-processing subunits (corresponding to the visual tag sequence and the text tag sequence, respectively).

[0214] Each dual-stream attention block includes two independent processing paths. Path 1 (text path) uses the visual marker sequence as the query object Q and the text marker sequence as the query object (key K and value V), calculating a text weight matrix. This text weight matrix represents the weight allocation of the text marker sequence based on the visual marker sequence, and outputs a weight-adjusted text marker sequence. Path 2 (image path) uses the text marker sequence as the query object and the visual marker sequence as the query object, calculating an image weight matrix. This image weight matrix represents the weight allocation of the image marker sequence based on the text marker sequence, and outputs a weight-adjusted visual marker sequence.

[0215] After the above calculations are completed, the original visual label sequence is weighted to obtain the updated visual label sequence, and the original text label sequence is weighted to obtain the updated text label sequence. The updated visual label sequence and the updated text label sequence are then input into their respective post-processing subunits, namely the image post-processing subunit and the text post-processing subunit.

[0216] The image post-processing subunit and the text post-processing subunit have the same structure but different and independent parameters. Both include a normalization layer, a feedforward network layer, and a gating processing layer.

[0217] In the image post-processing subunit, the normalization layer can use AdaLN (Adaptive Layer Normalization) to dynamically normalize the image features in the updated visual label sequence based on their distribution characteristics (such as strong spatial correlation and stable numerical range), and incorporate the aforementioned time steps to adapt to the needs of the diffusion generation stage, such as strengthening the global structure in the early stage and preserving details in the later stage.

[0218] Feedforward network layers are used to enhance the spatial correlation of image features (e.g., the positional dependence of the digital human arm and the goods, the continuity of background texture, etc.) and capture details (such as light and shadow gradation, material differences) through nonlinear transformations.

[0219] Structurally, the feedforward network layer employs a structure of fully connected layers, non-linear activation functions (such as GELU and ReLU), and fully connected layers. The first fully connected layer maps the normalized labeled sequence to a higher dimension, expanding the feature representation space. Then, the activation function introduces non-linearity, enabling the model to learn complex feature relationships. Finally, the second fully connected layer maps the high-dimensional features back to the original dimension.

[0220] The gating processing layer learns the gating parameters and dynamically adjusts the relative weights of the initial label and the label after the aforementioned series of processing for each label. Based on the relative weights, it outputs the final label of the dual-stream attention block, thereby avoiding excessive influence between text features and image features during the dual-stream cross-attention calculation process.

[0221] The text post-processing subunit is the same as above. The normalization layer uses AdaLN to optimize the distribution consistency of text features in the text feature sequence, avoiding interference with interaction due to excessively large feature amplitudes of individual high-frequency words. The feedforward network layer focuses on strengthening the semantic logic of text features, capturing abstract semantics in language through nonlinear transformations. The structure of the feedforward network layer and subsequent gating processing layers is the same as above and will not be described again.

[0222] After the visual and text tag sequences are processed by their respective post-processing subunits, the calculation and output of the current layer's dual-stream attention block are completed. Simultaneously, this process is input into the next layer's dual-stream attention block, repeating the above steps to progressively deepen the cross-modal alignment of the image and text. This process is repeated iteratively through multiple layers of dual-stream attention blocks, ultimately outputting a deeply interactive visual feature sequence and a processed text feature sequence.

[0223] The image feature sequence (i.e., the processed visual feature sequence) output by the dual-stream cross-attention subunit is input into the 3D rotation coding subunit. The 3D rotation coding subunit consists of one or more rotation position coding layers, which are used to embed each feature block (token) in the above image feature sequence with position encoding based on rotation position coding (RoPE).

[0224] The output of the 3D rotation coding subunit is an image feature sequence containing positional information, i.e., an image feature sequence with positional information. In this sequence, the feature vector of each position already contains the positional information in 3D space, providing a feature representation with positional information for subsequent single-stream self-attention calculation.

[0225] The processed text feature sequence can be converted into semantic constraints, such as a global semantic condition vector, through pooling, which can then be used as constraints for subsequent single-stream self-attention computation.

[0226] Image feature sequences with location information can be input into a single-stream self-attention subunit, which then outputs a global feature sequence under semantic constraints.

[0227] The single-stream self-attention subunit consists of multiple layers of single-stream self-attention blocks. Each single-stream self-attention block includes a self-attention layer and a feedforward network layer. The self-attention layer adopts a multi-head self-attention structure. After the image feature sequence with location information is input, it first passes through multiple linear projection layers to convert the image feature sequence with location information into corresponding query (Q) matrices, key (K) matrices, and value (V) matrices. Then, the query (Q) matrix, key (K) matrix, and value (V) matrix are split into multiple attention heads according to the feature dimension. Each attention head corresponds to a Q sub-matrix, K sub-matrix, and value V sub-matrix, respectively. The attention weights of the above multiple attention heads are calculated independently to obtain the corresponding weight matrix, thereby completing the weight allocation for that attention head. After the calculation is completed, the outputs of each attention head are concatenated to obtain the first output sequence of the self-attention layer.

[0228] For the first output sequence, a residual connection is made between it and the input image feature sequence with location information to avoid the loss of feature information in the deep network. The resulting sequence is then normalized through layer normalization to achieve a stable distribution of features, thus obtaining the second output sequence.

[0229] The second output sequence is fed into a feedforward network layer for processing. This feedforward network layer consists of two fully connected layers (FC) and a non-linear activation function (such as GELU or ReLU). The first fully connected layer maps the input second output sequence to a higher dimension, expanding the feature representation space. Then, the activation function introduces non-linearity, allowing the model to learn complex feature relationships. Finally, the second fully connected layer maps the high-dimensional features back to the original dimension, obtaining the third output sequence, which is then output.

[0230] For the third output sequence, a residual concatenation is performed with the second output sequence, and the resulting sequence is subjected to layer normalization to obtain the fourth output sequence, which is the output of the single-stream self-attention block in this layer.

[0231] The fourth output sequence is input into the next layer of single-stream self-attention block, and the above operation is repeated. After iterative calculation through multiple layers of single-stream self-attention blocks, the final output feature sequence is the global feature sequence finally output by the attention encoding unit.

[0232] In the aforementioned attention calculation process, the processed text feature sequence output by the aforementioned dual-stream cross-attention subunit is no longer used as input, but is only retained as a constraint condition to perform semantic constraints during the attention calculation process of the aforementioned image feature sequence with location information.

[0233] Optionally, the semantic constraints are modulated by parameters generated by the first fully connected network to adjust the parameters of the linear projection layers corresponding to the queries, keys, and values ​​in the attention layer; or, the semantic constraints are supervised by a supervision matrix generated by the second fully connected network to be multiplied with the weight matrix output by at least one attention layer.

[0234] First, the processed text feature sequence is converted into a global semantic conditional vector through pooling. This global semantic conditional vector then supervises the single-stream self-attention computation of the image feature sequence with location information in the following two ways:

[0235] Method 1: The global semantic conditional vector is used to generate a parameter modulation factor through a first fully connected network. This parameter modulation factor is then input into the projection parameters of the linear projection layers corresponding to Q, K, and V, respectively, for interaction. For example, the parameter modulation factor is multiplied or added to the projection parameters to complete the conditional injection into the linear projection layers. Subsequently, when the image feature sequence with location information is input into the linear projection layers, textual semantic preferences will be assigned to the generated results of Q, K, and V. For example, if the text describes a product tilted 30° clockwise, after the above processing, the Q generated by the image feature sequence with location information will pay more attention to the relative angle of the product. In subsequent attention calculations, the feature blocks (tokens) that match the angle will have higher weights.

[0236] Method 2: Generate a supervision matrix from the global semantic condition vector through a second fully connected network. Multiply the supervision matrix with the weight matrix output by the self-attention layer in at least one single-stream self-attention block. Use the result as the basis for weight allocation of image feature sequences with location information by that layer.

[0237] The aforementioned supervision matrix is ​​used to calibrate the weight vector of a single-stream self-attention block. For example, if the text description includes a red apple, the generated supervision matrix indicates that the red and apple parts should be given priority. In the matrix calculation, higher values ​​(such as 0.9) are assigned to the apple region and the color region in the original weight vector, while lower values ​​(such as 0.1) are assigned to other regions not described in the text.

[0238] In the above operation, the weight matrix of each single-stream self-attention block can be multiplied with the supervision matrix (i.e., each layer is supervised), or specific layers can be selected for supervision. Typically, two to three layers of single-stream self-attention blocks in the shallow layer (usually the first 1 / 3 to 1 / 2 of the attention blocks) and two to three layers in the deep layer (usually the last 1 / 3 to 1 / 2 of the attention blocks) can be selected for multiplication with the supervision matrix, while the remaining single-stream self-attention blocks can retain the original weight matrix.

[0239] The two methods mentioned above can be used individually or simultaneously. The effect of both is to ensure that the region corresponding to the text description is given priority during the single-stream self-attention calculation process.

[0240] The attention encoding module uses a combination of a two-stream attention module and a single-stream self-attention module. The number of layers in the two-stream attention block and the single-stream self-attention block is typically 12 to 32. The aforementioned global feature vectors serve as conditions for generation in the subsequent diffusion generation process, thus imposing mandatory task constraints on the generation process.

[0241] In this implementation, during the semantically guided feature map generation process, the grasping parameters are first determined based on the target object feature map, and a hand feature map is generated to ensure the interaction adaptability between the hand and the target object. Then, the original feature map and the mask feature map are offset and adjusted through semantic features, and the pose calibration and feature fusion are completed by combining the hand feature map, so that the joint feature map not only meets the semantic requirements but also conforms to the geometric logic of hand-object interaction. Finally, the cross-attention mechanism is used to realize the deep interaction between semantic features, target object features, and joint features, so that the generated semantically guided feature map simultaneously meets the multiple requirements of visual form adaptation, semantic attribute matching, and natural hand-object interaction, which greatly improves the generation quality of subsequent target display images and avoids defects such as object floating and hand-object incoordination. For the attention encoding process of temporal conditional feature maps, a multi-head self-attention structure is used to separate and extract visual and textual label sequences. Then, cross-modal accurate alignment is achieved through dual-path cross-attention. This ensures the preservation of both visual spatiotemporal correlation and semantic temporal correlation. Furthermore, the positional information of the features is enhanced through three-dimensional rotational position encoding. Combined with multi-layer attention optimization under semantic constraints and the collaborative processing of residual connections and feedforward networks, the final output global feature sequence has spatial accuracy, semantic consistency, and temporal coherence. This provides a comprehensive and accurate constraint basis for subsequent diffusion generation and effectively avoids problems such as spatial misalignment and semantic deviation in the generated results.

[0242] Optionally, a diffusion generation algorithm is used to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item. This includes: using the diffusion generation algorithm to perform latent code processing on the global feature sequence to obtain the target latent code; inputting the target latent code into the decoder to obtain a preliminary display image; smoothing the pixels in the region corresponding to the occlusion edge of the preliminary display image, and making the pixels in the region corresponding to the pressing edge of the preliminary display image transparent to obtain the target display image.

[0243] Optionally, a diffusion generation algorithm is used to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item. This includes: compressing the global feature sequence according to a preset latent code channel dimension to obtain a latent code feature vector; recombining the latent code feature vector to obtain a latent code feature map; performing local channel weighted fusion on the latent code feature map to generate a conditional latent code map matching the time step; initializing and generating a random noise latent code map with the same dimension as the conditional latent code map as the noisy latent code map for the first time step; predicting the gradient correction vector for the next time step by calculating the difference between the conditional latent code map and the noisy latent code map of the corresponding time step; and iteratively updating the noisy latent code map based on the predicted gradient correction vector to obtain the target latent code map.

[0244] By iteratively processing the noisy latent code, the diffusion generation process can correct the noisy latent code along the gradient correction vector at each step, gradually transforming the image from pure noise into a clear image. After iteration, the latent code update subunit outputs a target latent code map, which is the latent code representation of the target display image where the handheld object is replaced with the target object.

[0245] By inputting the target latent code image into a decoder such as a VAE decoder, a preliminary display image can be obtained.

[0246] Optionally, edge protection processing is performed on pixels located in the area corresponding to the occlusion edge, including: spreading the occlusion edge to both sides by a first preset distance to form a first edge buffer area; and blurring each pixel in the edge buffer area based on the distance between each pixel in the first edge buffer area and the occlusion edge.

[0247] The first preset distance ranges from 1 to 5 pixels.

[0248] Blur processing can be performed using bilateral filtering, mean blurring, Gaussian blurring, or other blurring techniques.

[0249] The blur intensity can be adjusted based on the distance between the pixel and the occlusion edge, eliminating jagged edges and rough edges while maintaining the sense of boundary between the hand and the target object. The closer to the occlusion edge, the weaker the blur intensity, for example, a smaller Gaussian blur radius; the farther from the occlusion edge, the stronger the blur intensity, for example, a larger Gaussian blur radius.

[0250] Taking Gaussian blur processing as an example, for each pixel within the first edge buffer region, a weighted average of the surrounding pixels is calculated with the pixel as the center and the corresponding blur radius as the range (pixels closer to the center have higher weights). This average value is then used to replace the original pixel value, thereby achieving smooth blurring and eliminating abrupt changes between pixels. The blur radius is determined based on the distance between the pixel and the occlusion edge.

[0251] Taking a first preset distance of 3 pixels as an example, the blur radius of pixels on the edge of the occlusion is 0.5. When the distance from the occlusion edge is 1 pixel, the blur radius is 1; when the distance from the occlusion edge is greater than 1 pixel, the blur radius is 2.

[0252] Optionally, the pixels located in the area corresponding to the pressing edge are made transparent, including: spreading the pressing edge to both sides by a second preset distance to form a second edge buffer area; determining the transparency weight of each pixel in the edge buffer area based on the distance between each pixel in the second edge buffer area and the nearest hand key point; and making each pixel in the edge buffer area transparent based on the transparency weight of each pixel in the edge buffer area.

[0253] The second preset distance ranges from 1 to 5 pixels.

[0254] Calculate the distance from each pixel within the second edge buffer area to the nearest fingertip key point to obtain the pressure distance field information of the second edge buffer area (including the distance from each pixel to the nearest fingertip key point); calculate the transparency weight corresponding to each pixel based on the pressure distance field information and the preset transparency calculation model.

[0255] The higher the transparency weight, the lower the transparency of the pixel and the weaker the transparent visual effect. The transparency weight can be inversely correlated with the distance from the pixel to the nearest fingertip key point. That is, the closer the pixel is to the nearest fingertip key point, the greater the transparency weight and the more realistic the visual effect of the pixel. This creates an effect where the transparency gradually increases from the center area of ​​the fingertip pressing outwards. This avoids affecting the presentation of the core features of the target object during the subsequent fusion process. At the same time, it can make the overall effect more closely match the finger deformation during the actual pressing process, achieving a more realistic grip and finger wrapping effect.

[0256] The transparency calculation model can be: α = A - B × exp(-d 2 / (2σ 2 )), where α represents the transparency weight, σ is the second preset distance; A and B are both constants; exp() is the exponential function, which is an exponential operation with the natural constant e as the base; d represents the distance from the pixel to the nearest fingertip key point.

[0257] A first marker, such as a red marker, can be assigned to pixels located at the occlusion edge in the original display image, and a second marker, such as a yellow marker, can be assigned to the same pixels. During the image fusion stage, pixels with the third marker are subjected to replacement processing, pixels with the first marker are subjected to edge protection processing, and pixels with the second marker are subjected to transparency processing. Specifically, the replacement processing involves replacing the pixel value of a pixel with the pixel value of the corresponding pixel in the target display image.

[0258] For example, Figure 4 A schematic diagram illustrating the distribution of pixels assigned to each marker in the original display image provided in this application, as shown below. Figure 4As shown, the handheld item in the original display image is a coffee cup. In the area corresponding to the first item mask (Mask1), the edges of the thumb, index finger, and palm are occlusion edges, forming a first edge buffer area (Buffer1) through expansion, which is marked in red. In the first item mask (Mask1), the edges of the middle, ring, and little fingers are pressing edges, forming a second edge buffer area (Buffer2) through expansion, corresponding to the aforementioned pressing edges, which is marked in yellow. In the second item mask (Mask2), the remaining portion of the coffee cup's edge after completion, including its interior, excluding the first and second edge buffer areas (Buffer1 and Buffer2), is the replacement area, uniformly marked in blue. The remaining parts of the original display image, such as the interiors of the fingers and palm, and areas outside the area corresponding to the first item mask (Mask1), are not marked; that is, these remaining parts do not require processing and their original pixel values ​​are retained.

[0259] By assigning values ​​to the aforementioned different markers, a four-channel item mask is obtained, with each of the four channels corresponding to the pixel values ​​under the various markers. Therefore, when fusing the target item image into the original display image to replace the held item with the target item, the four-channel item mask (with the same size as the original display image) can be used to perform corresponding processing on the pixels of different markers to obtain the target display image.

[0260] This application also provides a digital human video generation method, including: acquiring a target display image that replaces a handheld item with a target item; obtaining the target display image based on the handheld item image generation method provided in any of the foregoing embodiments of this application; and inputting the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0261] The associated information of the target item can be in the form of at least one of text, semantics, and instructions.

[0262] Taking related information as text as an example, we can first perform semantic understanding on the related information based on a natural language processing model, extract features such as core selling points, functional attributes, and emotional tendencies, and generate keyword tags and language scripts; based on the keyword tags, we can determine the interactive actions of the digital human from a preset action library; we can then perform layer fusion between the target display image and the selected or default digital human image, synchronously play the voice corresponding to the language script, and control the digital human to synchronously perform the corresponding interactive actions to generate a coherent video, which is the digital human video used to display the target item.

[0263] If the associated information is speech, it can be converted into text first, and then the aforementioned method can be used to obtain the digital human video. Alternatively, features can be extracted from the speech to obtain rhythm features, intonation features, emotional features, and semantic keywords; lip-syncing signals can be generated based on rhythm and intonation features, and semantic keywords can be used to match the display parts of the target item to automatically generate targeted action instructions; at the same time, the amplitude and posture of the digital human's movements can be adjusted according to the emotional features of the speech to control the digital human to complete coordinated lip, limb, and head movements, and then the digital human's product-selling video can be output after being fused with the target display image.

[0264] If the associated information is an instruction, then the digital human's action instructions, expression instructions, and visual presentation instructions can be directly parsed from the instruction. Then, through the instruction matching engine, the action instructions are accurately associated with the preset action library, the expression instructions are converted into the digital human's tone parameters and lip-sync rhythm parameters, and the visual presentation instructions are converted into eye control parameters. Subsequently, the target display image and the digital human image are composited into layers, and the digital human is driven to perform corresponding actions and expressions according to the instruction sequence to generate a coherent digital human video that meets the instruction requirements.

[0265] By replacing items, new display materials for new items are automatically synthesized, and digital human-driven models are linked to generate promotional videos for the target items. There is no need to shoot images of each item individually. By replacing items in the original images, the materials required for new items are automatically synthesized, which greatly reduces the cost and time of material preparation, improves the efficiency of material generation, and thus improves the efficiency of digital human video generation, making it suitable for mass product promotion scenarios.

[0266] Figure 5 This is a schematic diagram of the structure of the handheld item image generation device provided in this application, such as... Figure 5 As shown, the handheld item image generation device provided in this embodiment includes: an image acquisition module 510, used to acquire an original display image and a target item image; a first segmentation module 520, used to segment the handheld item in the original display image using a first image segmentation model to obtain a first item mask; a second segmentation module 530, used to remove occluding elements and complete the handheld image corresponding to the first item mask using a second image segmentation model under the guidance of item prompts to obtain a second item mask; and a handheld item replacement module 540, used to fuse the target item image into the original display image based on the first item mask and the second item mask to obtain a target display image; the handheld item in the target display image is the target item.

[0267] In one possible implementation, the second segmentation module 530 is specifically used for: recognizing hand key points in the original display image based on a gesture recognition model; inputting the item prompt, hand key points, first item mask, and the original display image into a second image segmentation model; cropping the original display image based on the first item mask via the second image segmentation model to obtain a handheld image; removing occluding elements from the handheld image based on the item prompt and hand key points, and completing the occluded parts of the handheld item to obtain a second item mask representing the area where the handheld item is located.

[0268] In one possible implementation, the handheld item replacement module 540 includes: a hand edge determination unit, used to determine the hand edge based on hand key points and a first item mask; a replacement region determination unit, used to determine a replacement region based on a second item mask and the hand edge; and an image fusion unit, used to map pixels of the target item image to the replacement region, replace pixels in the replacement region of the original display image, and perform adaptation processing on pixels in the original display image located in the region corresponding to the hand edge to obtain the target display image; the adaptation processing includes at least one of the following: edge protection processing and transparency processing.

[0269] In one possible implementation, the hand edge includes an occlusion edge and a pressing edge; the hand edge determination unit includes: an interaction area recognition subunit, used to identify the interaction area between the hand and the held object based on hand key points and a first object mask; a state detection subunit, used to determine a first pixel in the interaction area of ​​the original display image that is in an occlusion state and a second pixel in a pressing state; and an edge segmentation subunit, used to determine the occlusion edge and the pressing edge based on the first pixel and the second pixel, respectively.

[0270] In one possible implementation, the state detection subunit is specifically used to: for each pixel in the interactive area of ​​the original display image, determine whether the pixel is a first pixel or a second pixel based on at least one of the following: the distance between the pixel and the edge of the handheld object, the feature value of the pixel, and the pressure confidence of the pixel; the edge of the handheld object is obtained based on a second object mask, and the pressure confidence is used to characterize the confidence of the corresponding pixel in the pressure applied to the handheld object.

[0271] In one possible implementation, the image fusion unit includes: a replacement subunit for mapping pixels of the target item image to a replacement area to replace pixels in the replacement area of ​​the original display image; an edge protection subunit for performing edge protection processing on pixels in the original display image located in the area corresponding to the occlusion edge; and a transparency processing subunit for performing transparency processing on pixels in the original display image located in the area corresponding to the pressing edge to obtain the target display image.

[0272] In one possible implementation, the handheld item replacement module 540 includes: an occlusion and pressing edge determination unit, used to determine occlusion edges and pressing edges based on hand key points and a first item mask; a marker mask generation unit, used to perform marker processing on the first item mask and the second item mask based on the occlusion edges and pressing edges to generate a marker item mask; a feature extraction unit, used to extract visual features of the original display image, the marker item mask, and the target item image, as well as semantic features of the item prompt words; a multimodal fusion unit, used to perform multimodal fusion processing on the visual features and semantic features and embed time steps to obtain a global feature sequence; and a diffusion generation unit, used to perform latent code processing and feature decoding on the global feature sequence using a diffusion generation algorithm to generate a target display image with the target item as the handheld item.

[0273] In one possible implementation, the multimodal fusion unit is specifically configured to: determine the grasping parameters of the target item based on the target item feature map; the target item feature map is the extracted visual features of the target item image; generate a hand feature map based on the grasping parameters and the original display image; offset the original feature map and the mask feature map based on semantic features; the original feature map and the mask feature map are respectively the extracted visual features of the original display image and the visual features of the marked item mask; adjust the pose of the offset original feature map based on the hand feature map, and fuse the pose-adjusted original feature map with the offset mask feature map to obtain a joint feature map; interact with the semantic features, the target item feature map, and the joint feature map based on a cross-attention mechanism to generate a semantically guided feature map; concatenate the time step vector transformed from time step information with the semantically guided feature map to obtain a temporal conditional feature map; and perform attention encoding on the temporal conditional feature map to obtain a global feature sequence.

[0274] In one possible implementation, the multimodal fusion unit, when performing attention encoding on the temporal conditional feature map to obtain a global feature sequence, specifically performs the following: Extracting visual and text marker sequences from the temporal conditional feature map using an attention layer composed of a multi-head self-attention structure and a normalization layer; performing cross-modal alignment and feature fusion on the visual and text marker sequences through dual-path cross-attention calculation; the dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key; performing three-dimensional rotational position encoding on the processed visual feature sequence to obtain an image feature sequence with positional information; generating semantic constraints based on the processed text feature sequence; inputting the image feature sequence with positional information into the attention layer composed of a multi-layer multi-head self-attention structure and a normalization layer, calculating attention weights under the semantic constraints, and outputting the global feature sequence after residual connections, a feedforward network, and layer normalization processing.

[0275] In one possible implementation, the diffusion generation unit includes: a diffusion generation subunit, used to perform latent code processing on the global feature sequence using a diffusion generation algorithm to obtain a target latent code; a decoding subunit, used to input the target latent code into a decoder to obtain a preliminary display image; an edge protection subunit, used to smooth the pixels in the preliminary display image located in the area corresponding to the occlusion edge; and a transparency processing unit, used to perform transparency processing on the pixels in the preliminary display image located in the area corresponding to the pressing edge to obtain the target display image. In one possible implementation, the edge protection subunit is specifically used to: diffuse the occlusion edge to both sides by a first preset distance to form a first edge buffer area; and blur each pixel in the edge buffer area based on the distance between each pixel in the first edge buffer area and the occlusion edge.

[0276] In one possible implementation, the transparency processing subunit is specifically used to: spread the pressing edge to both sides by a second preset distance to form a second edge buffer area; determine the transparency weight of each pixel in the edge buffer area based on the distance between each pixel in the second edge buffer area and the nearest hand key point; and perform transparency processing on each pixel in the edge buffer area based on the transparency weight of each pixel in the edge buffer area.

[0277] In one possible implementation, the first image segmentation model is pre-trained and then fine-tuned using a training set of images of handheld objects; the second image segmentation model is not pre-trained and is trained based on the training set of images of handheld objects.

[0278] The handheld item image generation device provided in this embodiment can execute the handheld item image generation method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0279] This application also provides a digital human video generation device, including: a display image acquisition module, used to acquire a target display image of a handheld item being replaced with a target item; the target display image is obtained based on the handheld item image generation method provided in any embodiment of this application; and a video generation module, used to input the target display image and the associated information of the target item into a digital human driving model to obtain a digital human video for displaying the target item.

[0280] The digital human video generation device provided in this embodiment can execute the digital human video generation method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0281] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus 604.

[0282] In a specific implementation, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to perform the above-described method.

[0283] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0284] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0285] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0286] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0287] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0288] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0289] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0290] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0291] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0292] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0293] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0294] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0295] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.

[0296] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.

Claims

1. A method for generating an image of a handheld object, characterized in that, include: Obtain the original display image and the target item image; The handheld item in the original display image is segmented using a first image segmentation model to obtain a first item mask; Guided by the item prompt, the second image segmentation model removes occluded elements and completes the handheld image corresponding to the first item mask by using the second image segmentation model to obtain the second item mask. Based on the first item mask and the second item mask, the target item image is fused into the original display image to obtain the target display image; the handheld item in the target display image is the target item.

2. The method according to claim 1, characterized in that, The process of removing occluded elements and completing the handheld image corresponding to the first item mask using a second image segmentation model guided by item prompts, to obtain the second item mask, includes: The hand key points in the original display image are identified based on the gesture recognition model; The item prompt, the hand key points, the first item mask, and the original display image are input into the second image segmentation model; The original display image is cropped based on the first item mask using the second image segmentation model to obtain the handheld image. Based on the item prompt and the key points of the hand, the occluding elements of the handheld image are removed, and the occluding parts of the handheld item are filled in to obtain the second item mask used to represent the area where the handheld item is located.

3. The method according to claim 2, characterized in that, The step of fusing the target item image into the original display image based on the first item mask and the second item mask to obtain the target display image includes: Based on the key points of the hand and the first item mask, the edge of the hand is determined; Based on the second item mask and the edge of the hand, the replacement area is determined; The pixels of the target item image are mapped to the replacement area, replacing the pixels of the replacement area in the original display image, and the pixels of the original display image located in the area corresponding to the edge of the hand are adapted to obtain the target display image; the adaptation process includes at least one of the following: edge protection processing and transparency processing.

4. The method according to claim 3, characterized in that, The hand edge includes an obscuring edge and a pressing edge; determining the hand edge based on the hand key points and the first item mask includes: Based on the key points of the hand and the first item mask, the interaction area between the hand and the held item is identified; Determine the first pixel of the original displayed image that is in an occluded state and the second pixel that is in a pressed state within the interactive area; The occlusion edge and the pressing edge are determined based on the first pixel and the second pixel, respectively.

5. The method according to claim 4, characterized in that, Determining the first pixel in the interactive area of ​​the original displayed image that is in an occluded state and the second pixel that is in a pressed state includes: For each pixel in the interactive area of ​​the original display image, it is determined whether the pixel is the first pixel or the second pixel based on at least one of the following: the distance between the pixel and the edge of the handheld object, the feature value of the pixel, and the pressure confidence of the pixel. The edge of the handheld object is obtained based on the second object mask, and the pressure confidence is used to characterize the confidence of the corresponding pixel in the pressure applied to the handheld object.

6. The method according to claim 4, characterized in that, The step of performing adaptation processing on the pixels of the original display image located at the edge of the hand to obtain the target display image includes: Edge protection processing is performed on the pixels of the original display image located in the area corresponding to the occlusion edge, and transparency processing is performed on the pixels of the original display image located in the area corresponding to the pressing edge, to obtain the target display image.

7. The method according to claim 2, characterized in that, The step of fusing the target item image into the original display image based on the first item mask and the second item mask to obtain the target display image includes: determining the occlusion edge and the pressing edge based on the hand key points and the first item mask; Based on the occlusion edge and the pressing edge, the first item mask and the second item mask are marked to generate a marked item mask; the visual features of the original display image, the marked item mask, and the target item image are extracted, as well as the semantic features of the item prompt words are extracted; The visual features and semantic features are fused using a multimodal process and embedded with time steps to obtain a global feature sequence; Using a diffusion generation algorithm, the global feature sequence is processed by latent code and feature decoding to generate a target display image with the target item as the handheld item.

8. The method according to claim 7, characterized in that, The process of multimodal fusion of the visual features and semantic features and embedding time steps to obtain a global feature sequence includes: Based on the target item feature map, the grasping parameters of the target item are determined; the target item feature map is the extracted visual features of the target item image. Based on the gripping parameters and the original display image, a hand feature map is generated; Based on the semantic features, the original feature map and the mask feature map are offset; the original feature map and the mask feature map are the extracted visual features of the original display image and the visual features of the marked item mask, respectively. Based on the hand feature map, the pose of the offset original feature map is adjusted, and the pose-adjusted original feature map is fused with the offset mask feature map to obtain a joint feature map; Based on the cross-attention mechanism, the semantic features, the target item feature map, and the joint feature map interact to generate a semantically guided feature map; The time step vector, converted from the time step information, is concatenated with the semantically guided feature map to obtain the time conditional feature map. Attention encoding is performed on the temporal conditional feature map to obtain the global feature sequence.

9. The method according to claim 8, characterized in that, The step of performing attention encoding on the temporal conditional feature map to obtain the global feature sequence includes: Based on an attention layer composed of a multi-head self-attention structure and a normalization layer, visual marker sequences and text marker sequences are extracted from the temporal conditional feature map. The visual marker sequence and the text marker sequence are aligned and fused across modalities through a dual-path cross-attention calculation. The dual-path cross-attention calculation includes attention weight calculation with the visual marker sequence as the query and the text marker sequence as the key, and attention weight calculation with the text marker sequence as the query and the visual marker sequence as the key. The processed visual feature sequence is subjected to three-dimensional rotational position encoding to obtain an image feature sequence with positional information; Based on the processed text feature sequence, semantic constraints are generated. The image feature sequence with location information is input into an attention layer consisting of a multi-layer multi-head self-attention structure and a normalization layer. Attention weights are calculated under the semantic constraints. After residual connection, feedforward network and layer normalization processing, the global feature sequence is output.

10. The method according to claim 7, characterized in that, The process of using a diffusion generation algorithm to perform latent code processing and feature decoding on the global feature sequence to generate a target display image with the target item as the handheld item includes: The target latent code is obtained by processing the global feature sequence using a diffusion generation algorithm. The target latent code is input into the decoder to obtain a preliminary display image; The pixels in the preliminary display image located in the area corresponding to the occlusion edge are smoothed, and the pixels in the preliminary display image located in the area corresponding to the pressing edge are made transparent to obtain the target display image.

11. The method according to claim 6, characterized in that, Perform edge protection processing on pixels located in the region corresponding to the occlusion edge, including: The obstructing edge is spread out to both sides by a first preset distance to form a first edge buffer area; Based on the distance between each pixel in the first edge buffer area and the occlusion edge, each pixel in the edge buffer area is blurred.

12. The method according to claim 6, characterized in that, Perform transparency processing on pixels located in the area corresponding to the pressed edge, including: The pressing edge is spread outwards by a second preset distance to form a second edge buffer area; The transparency weight of each pixel in the second edge buffer region is determined based on the distance between each pixel in the second edge buffer region and the nearest hand key point. Based on the transparency weight of each pixel within the edge buffer area, each pixel within the edge buffer area is made transparent.

13. The method according to any one of claims 1-12, characterized in that, The first image segmentation model is obtained by pre-training and then fine-tuning it using a training set of images of handheld objects; The second image segmentation model was not pre-trained and was trained based on a training set of images of handheld objects.

14. A method for generating digital human videos, characterized in that, include: Get the target display image that replaces the held item with the target item; The target display image is obtained based on the method provided in any one of claims 1-13; The target display image and the associated information of the target item are input into the digital human driving model to obtain a digital human video for displaying the target item.

15. A device for generating images of handheld objects, characterized in that, include: The image acquisition module is used to acquire the original display image and the target item image; The first segmentation module is used to segment the handheld item in the original display image using a first image segmentation model to obtain a first item mask; The second segmentation module is used to remove occluded elements and complete the handheld item in the handheld image corresponding to the first item mask by means of the second image segmentation model under the guidance of the item prompt words, so as to obtain the second item mask; The handheld item replacement module is used to fuse the target item image into the original display image based on the first item mask and the second item mask to obtain a target display image; the handheld item in the target display image is the target item.

16. A digital human video generation device, characterized in that, include: The image acquisition module is used to acquire the target display image to replace the held item with the target item; The target display image is obtained based on the method provided in any one of claims 1-13; The video generation module is used to input the target display image and the associated information of the target item into the digital human driving model to obtain a digital human video for displaying the target item.

17. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-14.

18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-14.

19. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-14.