Virtual backpack on method, device and storage medium

Through a two-stage processing flow, the reasonable position and posture of the backpack in the human body image are first determined, and then the key visual attributes are refined and transferred. This solves the problem of high fidelity in spatial layout and detailed features in virtual try-on, and achieves a natural and realistic body synthesis effect.

CN121685073BActive Publication Date: 2026-05-19DONSON TIMES INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
DONSON TIMES INFORMATION TECH CO LTD
Filing Date
2026-02-11
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing virtual try-on technology struggles to achieve a balance between reasonable spatial positioning, orientation, occlusion relationships, and high-fidelity reproduction of details such as color, brand logo, and material texture in bag synthesis, resulting in a lack of realism in the try-on results.

Method used

Through a two-stage processing flow, the image understanding model first achieves accurate spatial positioning and pose synthesis of the backpack in the human body image, and then the detail transfer model performs local redrawing of the target area to ensure accurate reproduction of high-fidelity details such as color, brand logo and material texture.

Benefits of technology

It achieves a balance between high realism and high detail fidelity when trying on a virtual backpack, generating a natural and realistic composite effect when wearing the backpack.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121685073B_ABST
    Figure CN121685073B_ABST
Patent Text Reader

Abstract

The application discloses a virtual backpack on-body method, device and storage medium, comprising: receiving a user image and a backpack image, and obtaining a try-on instruction based on the user image and the backpack image; inputting the user image and the backpack image into an image understanding and editing model, and identifying the try-on instruction to synthesize the user image and the backpack image into an initial synthesized image through the image understanding and editing model; extracting key detail features in the backpack image, and determining a target region where the key detail features are located in the initial synthesized image; taking the key detail features and the target region as inputs of a detail migration and redrawing model, and outputting an on-body synthesized image based on area redrawing processing of the detail migration and redrawing model. The application realizes the technical effect of unifying high realism and high detail fidelity in virtual backpack try-on through two-stage collaborative processing, that is, first determining a natural and reasonable position and posture of the backpack, and then migrating key visual attributes in detail.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual try-on technology, and in particular to a method, device and storage medium for virtual backpack try-on. Background Technology

[0002] With the rapid development of e-commerce and the fashion industry, consumers' demand for online previews of how products actually look on them is increasing. In the category of accessories such as bags, users not only want to see the style of the product itself, but also expect to intuitively experience how it matches their clothing, body shape, and posture. However, while current virtual try-on technology has made breakthroughs in categories such as clothing and eyewear, it still faces significant shortcomings when dealing with accessories like bags, which have complex spatial interactions and detailed features. Current common display methods mostly rely on static images or simple augmented reality (AR) overlay technology, which struggles to realistically simulate the drape, occlusion, and light and shadow effects created by the contact between the bag strap and the body, resulting in a lack of realism in the try-on results.

[0003] Looking further, solutions involving virtual wearables or image synthesis in related technical fields often struggle to balance spatial positioning accuracy with high fidelity in detail. For example, some hardware-adaptive technologies focus on improving the physical structure of wearable devices to enhance wearing comfort, but do not address the natural synthesis of virtual objects and human images; some environment-aware wearable devices focus on navigation and interaction functions, without involving image-level editing and rendering of virtual objects. Furthermore, general image processing algorithms, such as style transfer or super-resolution reconstruction techniques, while capable of overall image optimization, are not specifically designed for the task of "synthesizing a specified object onto a human body while preserving its key visual attributes," leading to distortions in details such as color, logos, and textures.

[0004] Overall, there is currently no technical solution that can simultaneously solve the problem of how to reasonably arrange the bag in a human body image and reproduce its core features such as color, brand logo, and material with high fidelity. This has become a key technical bottleneck that restricts the further improvement of the online bag try-on experience.

[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0006] The main purpose of this application is to provide a method, device and storage medium for trying on a virtual backpack, which aims to solve the technical problem that existing virtual backpacks cannot simultaneously guarantee a reasonable spatial layout and high fidelity of key details during trial use.

[0007] To achieve the above objectives, this application proposes a method for virtual backpack wearing, the method comprising:

[0008] Receive user image and backpack image;

[0009] In response to the trial carrying instruction of the user image and the backpack image, the user image and the backpack image are input into the image understanding and editing model, and the user image and the backpack image are synthesized into an initial composite image by the image understanding and editing model. The initial composite image includes the spatial layout and occlusion relationship between the user and the backpack.

[0010] Extract key detail features from the backpack image, and determine the target region where the key detail features are located in the initial synthesized image;

[0011] Using the key detail features and the target region as input to the detail transfer and redrawing model, the region redrawing process is performed based on the detail transfer and redrawing model. The result of the region redrawing process replaces the target region in the initial composite image to obtain the upper body composite image.

[0012] In one embodiment, in response to a trial carrying instruction of the user image and the backpack image, the step of inputting the user image and the backpack image into an image understanding and editing model, and compositing the user image and the backpack image into an initial composite image through the image understanding and editing model, includes:

[0013] In response to the trial strap instruction from the user image and the backpack image, the trial strap instruction is semantically parsed to obtain trial strap keywords;

[0014] The trial band keywords are converted into spatial constraints, and the spatial constraints are injected into the image understanding and editing model as prior knowledge. The spatial constraints are executed through the attention mechanism of the image understanding and editing model to locate the body parts that match the trial band instructions in the user image and output the location results.

[0015] Based on the location results and the trial pack keywords, the image understanding and editing model synthesizes the backpack image and the user image into the initial composite image.

[0016] In one embodiment, the step of executing the spatial constraints through the attention mechanism of the image understanding and editing model, locating the body part matching the trial band instruction in the user image, and outputting the location result includes:

[0017] The image understanding and editing model queries the spatial constraints corresponding to the trial band keywords in the pre-built trial band knowledge graph and extracts visual features from the user image.

[0018] The spatial constraints and visual features are fused and interacted in a multimodal manner, and spatial perception editing instruction features are obtained based on the interaction results.

[0019] The positioning result is adjusted and output based on the spatial perception editing instruction features.

[0020] In one embodiment, the step of executing the spatial constraints through the attention mechanism of the image understanding and editing model, locating the body part matching the trial band instruction in the user image, and outputting the location result includes:

[0021] The image understanding and editing model encodes the spatial constraints into attention query vectors and maps the visual features of the user image into key vectors and value vectors.

[0022] Calculate the similarity between the attention query vector and the key vector, and generate an attention weight distribution based on the calculation results;

[0023] The attention weight distribution and the value vector are iteratively processed through a multi-layer attention mechanism to obtain the weighted aggregated feature information of the value vector;

[0024] Based on the weighted aggregation features, the spatial coordinates of the body parts in the user image related to the trial band instruction are determined, and the spatial coordinates are used as the positioning result.

[0025] In one embodiment, the step of synthesizing the backpack image and the user image into the initial synthesized image based on the positioning result and the trial carry keywords using the image understanding and editing model includes:

[0026] Based on the spatial coordinates of the body parts determined by the positioning results, spatial transformation parameters are generated to adapt to the backpack image;

[0027] The backpack image is subjected to three-dimensional pose simulation and two-dimensional projection processing based on the spatial transformation parameters. The spatial pose of the backpack image is adapted to the orientation and curvature of the body parts according to the processing results.

[0028] Based on the depth information of the human body parts in the user image, calculate the expected occlusion relationship between the backpack image and the human body parts;

[0029] The pixel coverage is processed by the expected occlusion relationship, and the contact area is feathered and rendered with light and shadow to generate the initial composite image.

[0030] In one embodiment, the step of extracting key detail features from the backpack image and determining the target region where the key detail features are located in the initial synthesized image includes:

[0031] Visual features of the backpack image at different scales are extracted by a pre-trained multi-branch convolutional network. The first branch convolutional network is used to extract color and material features, and the second branch convolutional network is used to extract logo and contour structure features.

[0032] The extracted visual features are fused and compressed to obtain the key detail features;

[0033] Locate the backpack region in the initial synthesized image, and generate an initial backpack mask based on the location result;

[0034] The high-response regions of the key detail features are spatially correlated with the initial knapsack mask, and the feature regions are selected from the spatial correlation results as the target regions.

[0035] In one embodiment, the steps of using the key detail features and the target region as input to a detail transfer and redrawing model, performing region redrawing processing based on the detail transfer and redrawing model, and replacing the target region in the initial composite image with the result of the region redrawing processing to obtain the upper body composite image include:

[0036] The initial synthesized image is spatially aligned with the key detail features;

[0037] Within the mask range corresponding to the target region, the key detail features are fused to the corresponding positions through feature injection to obtain the redrawn region image;

[0038] The redrawn region image is used to reconstruct a local region image within the target region, and the synthesized upper body image is obtained based on the reconstruction result.

[0039] In one embodiment, the step of fusing the key detail features to the corresponding position of the initial synthesized image through feature injection within the mask range corresponding to the target region includes:

[0040] The key detail features are subjected to feature transformation processing, and source features that are adapted to the visual characteristics of the target region are generated based on the processing results;

[0041] The background features of the target region are weighted and fused with the source features to obtain a comprehensive feature. The comprehensive feature is then decoded by a feature reconstruction network to generate a redrawn region image.

[0042] In addition, to achieve the above objectives, this application also proposes a virtual backpack wearing device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the virtual backpack wearing method described above.

[0043] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the virtual backpack wearing method described above.

[0044] One or more technical solutions proposed in this application have at least the following technical effects:

[0045] The technical solution of this application receives a user image and a backpack image; in response to a trial carrying instruction from the user image and the backpack image, the user image and the backpack image are input into an image understanding and editing model, and the user image and the backpack image are synthesized into an initial composite image through the image understanding and editing model. The initial composite image includes the spatial layout and occlusion relationship between the user and the backpack; key detail features are extracted from the backpack image, and the target region where the key detail features are located is determined in the initial composite image; the key detail features and the target region are used as input to a detail transfer and redrawing model, and a region redrawing process is performed based on the detail transfer and redrawing model. The result of the region redrawing process replaces the target region in the initial composite image to obtain an upper body composite image.

[0046] This application achieves a unified technical effect of high realism and high detail fidelity when trying on a virtual backpack by first determining the natural and reasonable position and posture of the backpack through two-stage collaborative processing, and then refining the transfer of key visual attributes. Attached Figure Description

[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a flowchart illustrating the first embodiment of the virtual backpack wearing method of this application;

[0050] Figure 2 This is a detailed step diagram of step S20 in the first embodiment described above;

[0051] Figure 3 This is a detailed step diagram of step S30 in the first embodiment above;

[0052] Figure 4 This is a detailed step diagram of step S40 in the first embodiment described above;

[0053] Figure 5 An initial composite image of the user image and the backpack image;

[0054] Figure 6 The upper body composite image is a redrawn version of the initial composite image;

[0055] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the virtual backpack wearing method in the embodiments of this application.

[0056] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0057] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0058] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0059] The main solution of this application embodiment is as follows: receiving a user image and a backpack image; responding to a trial carrying instruction from the user image and the backpack image, inputting the user image and the backpack image into an image understanding and editing model, and using the image understanding and editing model to synthesize the user image and the backpack image into an initial synthesized image, the initial synthesized image including the spatial layout and occlusion relationship between the user and the backpack; extracting key detail features from the backpack image, and determining the target region where the key detail features are located in the initial synthesized image; using the key detail features and the target region as input to a detail transfer and redrawing model, performing region redrawing processing based on the detail transfer and redrawing model, and replacing the target region in the initial synthesized image with the region redrawing processing result to obtain an upper body synthesized image.

[0060] Existing virtual try-on methods cannot achieve a balance between reasonable spatial positioning, orientation, occlusion relationships, and high-fidelity reproduction of details such as color, brand logo, and material texture in bag synthesis.

[0061] This application provides a solution that uses a two-stage processing flow. First, an image understanding model achieves accurate spatial positioning and pose synthesis of the backpack in a human body image. Then, a detail transfer model performs local redrawing of the target area. While maintaining reasonable spatial relationships, it ensures the accurate reproduction of high-fidelity details such as color, brand logo, and material texture, ultimately achieving a natural and realistic body synthesis effect.

[0062] Based on this, the embodiments of this application provide a method for virtual backpack wearing, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the virtual backpack wearing method of this application. In this embodiment, the virtual backpack wearing method includes steps S10 to S40:

[0063] Step S10: Receive the user image and the backpack image;

[0064] In this embodiment, in the initial stage of the virtual backpack try-on method, the user image and backpack image are received. This process constitutes the basic data entry point for the entire virtual try-on system. Specifically, it completes the basic processing of data reception and instruction parsing, which focuses on the standardized reception process of the user image and backpack image, as well as the try-on instruction generation mechanism based on multimodal information. Input data is received through a highly scalable distributed image acquisition interface. This interface uses an asynchronous communication protocol (such as HTTP / RESTful API or WebSocket) and can handle concurrent requests from multiple users simultaneously, ensuring the efficiency and stability of data reception. The input data consists of the user image and backpack image. The user image is usually obtained from the camera of a smart device or uploaded files. It is an RGB three-channel digital image containing complete human posture information, in formats including JPEG or PNG, and its resolution must meet a minimum of 1080p to ensure the accuracy of subsequent processing. The backpack image is a PNG format product image with a transparency channel (Alpha channel). This format helps to preserve the outline details of the backpack and avoid background interference during virtual overlay. During the receiving process, the system performs preliminary data verification, such as checking the integrity of image files, format compatibility, and consistency of metadata (such as EXIF ​​information). If anomalies are found (such as corrupted files or insufficient resolution), an automatic retransmission or user prompt mechanism is triggered to maintain the reliability of the data pipeline.

[0065] The image preprocessing module, as an extension of the receiving step, immediately standardizes the input data after successful reception. This includes scaling the image resolution to a uniform 1024×768 pixels, a size that balances processing efficiency and visual quality while adapting to most display devices. For user images, preprocessing begins with human detection bounding box localization, using deep learning-based YOLO or SSD algorithms to quickly identify human regions in the image and extract key pose points (such as shoulders and waist) to aid in subsequent alignment. For backpack images, foreground segmentation is performed, using a convolutional neural network (CNN) segmentation model to remove the background, retaining only the backpack itself to ensure seamless integration during virtual try-on. Furthermore, the receiving process integrates metadata management, such as attaching timestamps, user identifiers, and session IDs to each image for easy data flow tracking and troubleshooting. The entire design of step S10 not only enhances the robustness of data reception but also lays a solid foundation for subsequent virtual rendering and user experience optimization. Through this modular processing, the system can flexibly adapt to various application scenarios, such as e-commerce try-on or AR interaction. In summary, this embodiment highlights the key role of step S10 in the method flow by refining the technical details of the receiving process, ensuring high-quality input of the user image and backpack image, thereby driving the entire virtual backpack loading process to operate efficiently.

[0066] Step S20: In response to the trial carrying instruction of the user image and the backpack image, the user image and the backpack image are input into the image understanding and editing model. The image understanding and editing model is used to synthesize the user image and the backpack image into an initial synthesized image. The initial synthesized image includes the spatial layout and occlusion relationship between the user and the backpack.

[0067] In this embodiment, in response to the trial-wearing instruction of the user image and the backpack image, the user image and the backpack image are input into an image understanding and editing model, and the image understanding and editing model synthesizes the user image and the backpack image into an initial composite image. (Please see below for more details.) Figure 5 , Figure 5 This is the initial composite image of the user image and the backpack image.

[0068] In response to the try-on instruction based on the user's image and the backpack image, the image understanding and editing model is activated. The image understanding and editing model is a Transformer-based multimodal architecture specifically designed for virtual try-on tasks, containing three core components: a visual encoder, a text encoder, and an image generator. It is responsible for receiving and coordinating the processing of input information from both image and text modalities.

[0069] At the beginning of the data processing pipeline, the visual encoder performs depth feature extraction on both the user image and the backpack image. For the user image, a pre-trained convolutional neural network (such as ResNet-50) is used to extract its multi-scale visual feature map to capture information from low-level texture to high-level semantics. In parallel, a dedicated human body parsing network (such as SCHP) is used to obtain pixel-level segmentation masks for body parts (such as the torso, arms, and back regions). This provides crucial structural priors for subsequent judgments of spatial layout and occlusion relationships. Simultaneously, the backpack image undergoes a similar visual encoding process. However, to accommodate the potential diversity and non-rigid deformation of backpack products, its encoder additionally integrates deformable convolutional layers, enabling more flexible capture and understanding of the geometric features of different backpack shapes.

[0070] The text encoder is specifically designed to parse and understand the semantic meaning of the trial strap instructions. It encodes the instruction text using a pre-trained language model (such as BERT or its variants), identifying key semantic information in the instructions (such as action descriptions and positional descriptions like "wear on both shoulders" and "place on the back"), and converting the natural language description into a machine-understandable high-dimensional constraint vector. This vector is essentially a digital abstraction of the synthetic target.

[0071] Next comes the crucial step of aligning and fusing multimodal features. The spatial constraint vector output by the text encoder is finely aligned with the user image feature map extracted by the visual encoder in the feature space. Specifically, the model employs a cross-attention mechanism to calculate the correlation weight between the spatial constraint vector and each spatial location in the user image feature map. Based on the calculated attention weight distribution, the model can determine the potential placement area of ​​the backpack in the human image, initially establishing the spatial layout. This determination process is not instantaneous but employs an iterative optimization strategy, gradually fine-tuning the attention weights through multiple iterations, enabling the image understanding and editing model to progressively focus on the body part that best matches the description of the trial pack instruction.

[0072] After initial spatial localization, the model proceeds to the feature fusion and geometric correction stage. The image understanding and editing model uses a spatial transformation network to perform geometric correction on the extracted backpack features. This includes adaptive deformation based on the surface curvature of the target body parts (such as the curvature of the back), adjusting the backpack size based on human depth information estimated from the user image to conform to perspective principles, and initially adjusting the backpack's color saturation and contrast based on the overall lighting conditions of the user image. The corrected backpack features and user image features are then fused at multiple levels in the decoder. The fusion process uses a gated recurrent unit (GRU) to control the information flow, ensuring a natural transition at the synthesis boundary (especially the area where the backpack contacts the human body).

[0073] Furthermore, the image generator is based on a state-of-the-art diffusion model architecture, performing detail optimization and image reconstruction on the basis of initial fusion. This generation process uses a denoising diffusion probability model as its core, gradually removing noise from a random noise image through a preset number of iterations (e.g., 50 steps), continuously injecting the fusion features and spatial constraints obtained in the preceding steps. In each denoising step, the image understanding and editing model references spatial constraints to ensure the accuracy of the backpack's position, while utilizing prior knowledge learned through adversarial training strategies to guarantee the realism of the synthesized effect. For example, it simulates the interaction between the backpack straps and clothing folds, and generates reasonable shadows based on human posture and light source direction. Finally, the initial synthesized image generated by the image understanding and editing model recognizing and executing the trial-wearing instruction not only places the backpack in the correct spatial position but also accurately simulates the complex spatial layout between the user and the backpack (e.g., the backpack's direction and distance relative to the body) and physical occlusion relationships (e.g., an arm in front of the backpack, or the backpack partially obscured by the shoulder), and incorporates basic lighting effects.

[0074] The entire synthesis process typically runs on a distributed computing cluster equipped with high-performance GPUs, with each image taking approximately 3-5 seconds to generate, achieving a good balance between efficiency and quality. The image understanding and editing model is trained using a large-scale virtual try-on dataset. Through end-to-end optimization, it minimizes pixel-level reconstruction loss, perceptual loss, and a geometric consistency loss specifically designed for accuracy in spatial layout and occlusion relationships, thereby ensuring the stability of the final generated initial synthesized image in terms of visual realism and spatial plausibility.

[0075] Step S30: Extract key detail features from the backpack image and determine the target region where the key detail features are located in the initial synthesized image;

[0076] In this embodiment, the extraction of detailed features from the backpack and the localization of the target region are completed, providing accurate input data for subsequent detail transfer. Specifically, the feature extraction module is based on a deep convolutional neural network architecture, while the target region localization combines instance segmentation and feature matching techniques.

[0077] In the detailed feature extraction stage, a multi-branch feature pyramid network is used to process the backpack image. Specifically, the first branch focuses on low-level visual features, extracting color distribution, material texture, and edge contour features through stacked convolutional layers; the second branch is responsible for high-level semantic features, capturing the overall shape and structure of the backpack and the semantic information of the brand logo through global pooling layers; the third branch is optimized for specific tasks, using separable convolutions to extract fine features of the logo area. The output features of the three branches are concatenated along a specific dimension to form a complete representation of the key detailed features.

[0078] Furthermore, based on the aforementioned feature extraction network and the ImageNet pre-trained model, fine-tuning was performed using a professional dataset containing 500,000 backpack images. The training process employed a multi-task learning strategy, simultaneously optimizing classification loss, reconstruction loss, and contrastive loss to ensure that the extracted features were both discriminative and retained reconstruction capability. The feature extraction network ultimately outputs a 1024-dimensional feature vector and a corresponding spatial feature map, which preserves the relative positional information of the features.

[0079] In the target region localization stage, the initial synthesized image is segmented into instances, and the pixel regions of the backpack in the synthesized image are accurately identified using the Mask R-CNN architecture. The segmentation network uses ResNet-101 as the backbone network, combined with a feature pyramid network to enhance multi-scale detection capabilities. The segmentation results undergo morphological post-processing, including hole filling and edge smoothing, to obtain an accurate binary mask.

[0080] Next, a feature matching algorithm is used to establish the correspondence between the backpack image and the backpack region in the initial composite image. The algorithm first calculates the cosine similarity between the key detail features and the features of the composite backpack region, and then uses a random sampling consensus algorithm to remove incorrect matching points. Based on correct matching point pairs, detailed areas that need to be retained can be accurately identified, such as the brand logo location, special texture areas, and color gradient areas.

[0081] Finally, a detailed target region description file is generated, including the coordinate data of the binary mask, the feature vectors of the key detail features, and the relative position mapping of the detail regions within the mask. All data is encapsulated in JSON format to ensure accurate parsing and use by subsequent processing modules. The entire processing process fully considers the balance between computational efficiency and accuracy, keeping processing time within an acceptable range while ensuring feature quality.

[0082] Step S40: Using the key detail features and the target region as input to the detail transfer and redrawing model, perform region redrawing processing based on the detail transfer and redrawing model, and replace the target region in the initial composite image with the region redrawing processing result to obtain the upper body composite image.

[0083] In this embodiment, the key detail features and the target region are used as input to the detail transfer and redrawing model. Based on this model, region redrawing processing is performed, and the result replaces the target region in the initial synthesized image, ultimately yielding a high-fidelity composite upper-body image. The aim is to accurately transfer the detailed features of the original backpack to the target region of the synthesized image, thereby generating a commercially viable upper-body image. The detail transfer and redrawing model is based on a conditional generative adversarial network architecture, specifically optimized for local region details to achieve high-quality synthesis. Please see [link / reference]. Figure 6 , Figure 6 The upper body composite image is a redrawn version of the initial composite image.

[0084] At the beginning of the data processing flow, the input data is aligned and prepared. The key detail features and the target region mask, along with the initial synthesized image, are input to the spatial alignment module. The spatial alignment module calculates the pixel-level correspondence between the backpack image and the synthesized backpack region using an optical flow estimation algorithm. The alignment process employs a dense matching strategy, finding the optimal matching point for each target pixel in the original backpack image to establish a complete coordinate mapping matrix.

[0085] Furthermore, the detail transfer and redraw model injects the key detail features into the intermediate layer of the generator network through a spatial adaptive normalization layer. Specifically, the key detail features undergo dimensionality transformation, converting them to a scale and number of channels that match the feature map of the generator network. Then, guided by the target region mask, an attention gating mechanism is used to control the intensity of feature injection, ensuring that the detail features only function within the target region and do not affect the surrounding background region.

[0086] The detail transfer and redrawing model's region redrawing process adopts a phased generation strategy: the first phase is to reconstruct the basic structure, generating the basic shape and main color distribution of the backpack in the target area through the encoder-decoder network of the U-Net architecture; the second phase is to enhance details, using a conditional adversarial generative network to specifically optimize high-frequency details such as logos and textures; the third phase is to perform multi-scale fusion, seamlessly stitching the redrawn area with the original background, and eliminating boundary artifacts through the Laplacian pyramid fusion algorithm.

[0087] For model training, a progressive training strategy was adopted, first pre-training on a large clothing dataset and then fine-tuning on a specialized backpack try-on dataset. The loss function incorporates multiple constraints, including pixel-level L1 loss, perceptual feature loss, adversarial loss, and detail preservation loss. Among them, the detail preservation loss ensures high consistency of key visual attributes by calculating the similarity between the redrawn region and the original backpack image in the feature space.

[0088] Finally, the region redrawing result obtained from the detail transfer and redrawing model is used to replace the target region in the initial composite image. Following this, post-processing optimization is performed, including edge smoothing based on bilateral filtering, color correction according to ambient lighting, and multi-frame stabilization (in video sequence applications). Ultimately, the composite upper body image is output based on the region redrawing process using the detail transfer and redrawing model. The output composite upper body image not only retains all the key visual features of the backpack but also seamlessly integrates with the lighting, shadows, and perspective relationships of the user image, achieving commercial-grade visual effects. The entire redrawing process has an average processing time of 2.3 seconds on a single V100 GPU, meeting the performance requirements of practical applications.

[0089] In summary, by using a two-stage collaborative processing approach, first determining the natural and reasonable position and posture of the backpack, and then refining the transfer of key visual attributes, a technical effect of achieving a high degree of realism and high detail fidelity during virtual backpack try-on was achieved.

[0090] Furthermore, you can also view Figure 2 , Figure 2 This is a detailed step diagram of step S20 in the first embodiment described above, based on the shown... Figure 2 In response to the trial carrying instruction of the user image and the backpack image, the user image and the backpack image are input into the image understanding and editing model, and the user image and the backpack image are synthesized into an initial composite image by the image understanding and editing model, including steps S21-23:

[0091] Step S21: In response to the trial carrying instruction from the user image and the backpack image, perform semantic parsing on the trial carrying instruction to obtain trial carrying keywords;

[0092] Step S22: Convert the trial bandage keywords into spatial constraints, and inject the spatial constraints as prior knowledge into the image understanding and editing model; wherein, the spatial constraints are executed through the attention mechanism of the image understanding and editing model to locate the body parts in the user image that match the trial bandage instruction, and output the location results;

[0093] Step S23: Based on the positioning results and the trial pack keywords, the image understanding and editing model synthesizes the backpack image and the user image into the initial synthesized image.

[0094] In this embodiment, the image understanding and editing model is executed according to a structured process, from instruction parsing to image synthesis. This process includes four key technical stages: semantic parsing, spatial constraint injection, attention localization, and geometric synthesis.

[0095] In the semantic parsing stage, the trial-wearing instructions are semantically parsed to obtain trial-wearing keywords. This process employs a pre-trained language model based on the BERT architecture to perform deep semantic analysis on the trial-wearing instructions. The language model first converts the input text into a high-dimensional vector representation through a word embedding layer, and then captures the contextual semantic relationships through a 12-layer Transformer encoder. The parsing module is specifically optimized for professional terminology in the wearable field, and domain-adaptive training enhances the recognition accuracy for wearable verbs such as "single-shoulder carrying," "crossbody," and "hand-carrying." Finally, the output layer uses a conditional random field model for sequence labeling to accurately extract the trial-wearing keywords based on three dimensions: trial-wearing method, target body part, and spatial relationship.

[0096] Subsequently, the trial-wearing keywords are converted into spatial constraints, and these spatial constraints are injected as prior instructions into the image understanding and editing model. A built-in wearable knowledge graph stores compatibility rules for different body parts and backpack types, as well as spatial parameter ranges corresponding to various trial-wearing methods. Upon receiving the trial-wearing keywords, the graph query engine retrieves relevant spatial constraint templates using a graph traversal algorithm, including allowed placement coordinate intervals, reasonable rotation angle ranges, and size scaling ratios. This abstract constraint is converted into a specific numerical representation through a spatial parameterization module, forming the structured spatial constraints. During the injection process, the multimodal fusion layer of the image understanding and editing model aligns the spatial constraints from the text source with visual features, achieving semantic-to-spatial mapping through a cross-modal attention mechanism. Then, the attention mechanism of the image understanding and editing model executes the spatial constraints, locates the body part matching the trial-wearing instruction in the user image, and outputs the location result.

[0097] Specifically, the visual Transformer module in the image understanding and editing model encodes the spatial constraints into query vectors and maps the visual features of the user image into key-value pairs. In the multi-head attention layer, the similarity score between the query vector and the key vector is calculated to generate an attention weight heatmap. This heatmap clearly indicates the body regions in the user image most relevant to the fitting requirements, such as the right shoulder region for a shoulder bag fitting. The localization process employs a coarse-to-fine strategy, first determining the approximate body region, then progressively narrowing the localization range through a cascaded refinement network, ultimately outputting precise body part bounding box coordinates as the localization result.

[0098] Finally, based on the positioning results and the trial keywords, the backpack image and the user image are synthesized into the initial composite image. The geometric transformation engine first calculates the spatial transformation parameters required for the backpack image, including a non-rigid deformation matrix based on the curvature of the body part and a scale factor based on the human body proportions. The rendering pipeline then performs a 3D pose simulation of the backpack, generating 2D projections from different viewpoints using differentiable rendering technology. The occlusion processing module accurately calculates the occlusion relationship between the backpack and the body based on the depth information map of the user image, determining which areas should be preserved or hidden. Finally, the image synthesizer uses a physically based rendering method to blend the processed backpack image with the user background, while applying ambient occlusion and soft shadow techniques to enhance realism, generating the initial composite image that meets all spatial constraints.

[0099] Based on the above Figure 2 The content described in step S22 is further refined. This step involves executing the spatial constraints through the attention mechanism of the image understanding and editing model, locating the body part matching the trial band instruction in the user image, and outputting the location result. This includes steps S22-1 to S22-3:

[0100] Step S22-1: The image understanding and editing model queries the spatial constraints corresponding to the trial band keywords in the pre-constructed trial band knowledge graph, and extracts visual features from the user image.

[0101] Step S22-2: Perform multimodal fusion interaction between the spatial constraints and the visual features, and obtain spatial perception editing instruction features based on the interaction results;

[0102] Step S22-3: Adjust the positioning result according to the spatial perception editing instruction features and then output it.

[0103] In this embodiment, during the critical stage of generating and injecting spatial constraints, a combination of knowledge-driven and data-driven approaches is used to achieve a precise conversion from semantics to space. This conversion process constructs a complete mapping link from abstract keywords to specific image editing parameters.

[0104] In the steps of querying the structured spatial constraints of the trial strap keywords in the pre-constructed trial strap knowledge graph and extracting visual features from the user image, the trial strap knowledge graph is constructed using a resource description framework, containing 386 entity nodes and 892 relation edges, covering professional knowledge in three dimensions: human anatomy, backpack typology, and physical contact mechanics. The graph query engine uses the SPARQL query language to perform multi-hop inference based on the input trial strap keywords. For example, from "single-shoulder bag," structured spatial constraints such as the applicable shoulder area, the expected suspension angle range, and the contact area with the torso can be derived. Simultaneously, visual features are extracted from the user image. The visual feature extraction module processes the user image based on a deep residual network, enhancing the feature response to body edges and posture key points through a bottleneck attention mechanism, and outputting a 1024-dimensional feature vector containing spatial information.

[0105] Next, a step is performed to perform multimodal fusion interaction between the spatial constraints and the visual features, and to obtain spatially perceived editing instruction features based on the interaction results. Spatially perceived editing instructions are generated through multimodal fusion, wherein a dual-stream encoder architecture is used to configure the fusion network. The text stream encoder converts the structured spatial constraints into 128-dimensional semantic embedding vectors, and the visual stream encoder projects the visual features of the user image onto the same embedding space. In the multimodal interaction layer, a dot-product-based attention mechanism is used to calculate the correlation matrix between text semantics and visual features, and a soft alignment strategy is used to establish the correspondence between semantic concepts and image regions. The interaction process adopts a hierarchical design: the bottom layer handles the matching of geometric constraints and pose features, the middle layer handles the consistency of material properties and lighting conditions, and the top layer handles the aesthetic evaluation of the overall composition. Finally, based on the interaction results, the spatially perceived editing instruction features that fuse semantic understanding and visual perception are output.

[0106] Finally, spatially aware features are injected into the image generation path. The generator of the image understanding and editing model adopts a U-Net architecture, inserting a spatially adaptive normalization layer at the skip connection between the encoder and decoder. This normalization layer uses the spatially aware editing instruction features as conditional input, dynamically calculates affine transformation parameters, and adjusts the mean and variance distribution of the feature map. Specifically, the spatial parameter prediction network maps the 128-dimensional editing instruction features into localization offsets and confidence weights, used to dynamically adjust the initial localization boxes generated from the initial visual features. This adjustment process is accomplished through a lightweight regression network, which learns how to fine-tune the initial localization results based on the editing instruction features, making them more accurately aligned to the body parts indicated by the knowledge graph. The adjusted localization results (such as precise body part bounding box coordinates) are finally output, providing accurate spatial guidance for subsequent image synthesis steps.

[0107] Based on the above Figure 2 The content described in step S23 is further refined. This step involves executing the spatial constraints through the attention mechanism of the image understanding and editing model, locating the body part matching the trial band instruction in the user image, and outputting the location result. This includes steps S23-1 to S23-3:

[0108] Step S23-1: The spatial constraints are encoded into attention query vectors through the image understanding and editing model, and the visual features of the user image are mapped into key vectors and value vectors.

[0109] Step S23-2: Calculate the similarity between the attention query vector and the key vector, and generate an attention weight distribution based on the calculation results;

[0110] Step S23-3: Iteratively calculate the attention weight distribution and the value vector using a multi-layer attention mechanism to obtain the weighted aggregated feature information of the value vector;

[0111] Step S23-4: Determine the spatial coordinates of the body parts in the user image related to the trial band instruction based on the weighted aggregation features, and use the spatial coordinates as the positioning result.

[0112] In this embodiment, the conversion process from spatial constraints to precise coordinates of body parts is realized based on the attention mechanism. This conversion process combines modern Transformer architecture and computer vision technology to ensure the accuracy and robustness of the positioning results.

[0113] The spatial constraint encoder uses a fully connected neural network to transform the spatial constraints into a 512-dimensional attention query vector. The encoding process includes a three-layer perceptron and layer normalization operations to ensure the query vector has sufficient expressive power. Simultaneously, the visual feature mapping module processes the user image based on the Vision Transformer architecture, segmenting the input image into a 16×16 tile sequence and converting it into a 256-dimensional embedding representation through linear projection. In the Transformer encoder, these embedding representations are further processed into key vectors and value vectors, where the key vectors are used to calculate the attention distribution, and the value vectors are used for feature aggregation. The position encoder adds absolute positional information to each tile, enabling the model to understand the spatial structure of the image.

[0114] A scaled dot product attention mechanism is employed to calculate the dot product similarity between the attention query vector and the key vector of each tile. The similarity score is normalized using a softmax function to generate the probabilistic attention weight distribution. To improve the expressive power of the attention weight heatmap, a multi-head attention mechanism is used, splitting the 512-dimensional query vector into eight 64-dimensional heads. Each head independently calculates its attention distribution, and the results are then concatenated and merged. Visualization of the attention weights shows that the model can accurately focus on body parts related to the "try on" instruction; for example, for the "right shoulder bag" instruction, the weight peak is concentrated in the right deltoid muscle region.

[0115] Subsequently, when determining the final spatial coordinates through iterative optimization, a six-layer Transformer decoder stack is used to implement the multi-layer attention mechanism. Each layer receives the output of the previous layer and uses the updated attention weight distribution to iteratively calculate and aggregate the value vector. During the iteration process, the lower-level layers process coarse localization information and identify the approximate body region; the middle-level layers refine the localization range and eliminate interference areas; the higher-level layers perform pixel-level refinement, and finally output the weighted aggregated feature information of the value vector containing high-level semantics and fine spatial information.

[0116] The coordinate regression head uses a fully connected layer to map the final weighted aggregated features into four bounding box parameters (top-left x and y coordinates, width, and height) and a confidence score. The post-processing module applies non-maximum suppression to eliminate overlap detections and verifies the reasonableness of the localization results based on human skeleton priors. Finally, the spatial coordinates of the body region are output as the localization result, achieving a localization accuracy of 92.3% mAP on the PASCAL VOC dataset.

[0117] Based on the above Figure 2The content of step S24 is further refined, namely, the step of synthesizing the backpack image and the user image into the initial synthesized image based on the positioning result and the trial carry keywords using the image understanding and editing model, including steps S24-1 to S24-4:

[0118] Step S24-1: Based on the spatial coordinates of the body parts determined by the positioning results, generate spatial transformation parameters for adapting the backpack image;

[0119] Step S24-2: Perform three-dimensional pose simulation and two-dimensional projection processing on the backpack image according to the spatial transformation parameters, and make the spatial pose of the backpack image match the orientation and curvature of the body parts according to the processing results.

[0120] Step S24-3: Calculate the expected occlusion relationship between the backpack image and the human body part based on the depth information of the human body part in the user image;

[0121] Step S24-4: Process pixel coverage through the expected occlusion relationship, and perform edge feathering and lighting rendering on the contact area to generate the initial composite image.

[0122] In this embodiment, a natural and realistic compositing effect of the backpack image is achieved through geometric transformation, 3D simulation and physical rendering techniques. This compositing process takes into account multiple dimensions such as spatial geometry, occlusion relationship and lighting consistency.

[0123] The parameter calculation engine constructs a local coordinate system based on the spatial coordinates of the body parts and the 21 joint point information output by the human pose estimation network. The spatial adaptation algorithm calculates the correspondence between the backpack anchor points and the body contact surface, and then solves for the optimal rigid body transformation parameters, including the rotation matrix, translation vector, and scaling factor, using the least squares method. For non-rigid deformation, a thin-plate spline interpolation algorithm is used to calculate the deformation field, enabling the backpack to adapt to changes in the curvature of the body surface. The final output spatial transformation parameters include a 4×4 transformation matrix, a 9-dimensional rotation vector, a 3-dimensional translation vector, and a 3-dimensional scaling factor. These parameters provide the mathematical basis for subsequent geometric transformations.

[0124] The built-in backpack 3D database contains CAD models of 127 common backpack types, each model containing approximately 5000 vertices and 10000 triangles. The 3D rendering engine first retrieves the corresponding 3D model based on the backpack type, and then applies the aforementioned spatial transformation parameters to transform the model. The posture simulation module considers gravity and fabric physics, simulating the natural drooping of the backpack straps and their deformation upon contact with the body through a mass spring system. The projection stage uses a perspective camera model, adjusting projection parameters according to the orientation and curvature of the body parts, ultimately rendering the 3D model into a 2D projected image consistent with the user's viewpoint. The entire processing is accelerated on the GPU, with a single-frame rendering time controlled within 33 milliseconds.

[0125] The depth estimation module uses the MiDaS deep learning model to recover dense depth maps from monocular user images, achieving an accuracy of REL 0.078 on the NYU Depth V2 dataset. The occlusion analysis algorithm compares the depth values ​​of the backpack image projection region with the corresponding background region, determining occlusion relationships through depth testing: when the backpack depth value is less than the background depth value, the backpack is fully visible; when there are regions where the backpack depth value is greater than the background depth value, these regions are marked as occluded areas. The system specifically handles semi-transparent occlusion cases, such as the visibility of the backpack behind sheer clothing, calculating the final pixel value through transparency blending.

[0126] Finally, the final image compositing and rendering were completed. The pixel compositor employed depth-based layer fusion technology, compositing layers sequentially from far to near according to depth order. For occluded boundary areas, a bilateral filter performed edge-aware smoothing to eliminate jagged edges. The lighting and shadow rendering module used ambient occlusion technology to calculate the shadow intensity of the area where the backpack contacted the body and simulated ambient lighting effects using a convolutional neural network. The color consistency correction module matched the color distribution of the backpack image with the lighting conditions of the user image, ensuring that the composite backpack matched the scene's lighting direction and that the shadows were soft and natural. The final output of the initial composite image achieved a realism score of 4.5 / 5.0 in a visual perception survey, proving that the compositing effect met commercial quality standards.

[0127] Furthermore, you can also view Figure 3 , Figure 3 This is a detailed step diagram of step S30 in the first embodiment described above, based on... Figure 3 The step of extracting key detail features from the backpack image and determining the target region where the key detail features are located in the initial synthesized image includes steps S31-34:

[0128] Step S31: Visual features of the backpack image at different scales are extracted by a pre-trained multi-branch convolutional network. The first branch convolutional network is used to extract color and material features, and the second branch convolutional network is used to extract logo and contour structure features.

[0129] Step S32: The extracted visual features are fused and compressed to obtain the key detail features;

[0130] Step S33: Locate the backpack region in the initial synthesized image and generate an initial backpack mask based on the location result;

[0131] Step S34: Spatially correlate the high-response regions of the key detail features with the initial knapsack mask, and select the feature regions as the target regions from the spatial correlation results.

[0132] In this example, a hierarchical processing strategy is used to achieve a complete transformation from the original image to the precisely located region. This transformation process ensures the accurate capture and localization of key visual attributes through the coordinated work of four sub-steps: feature extraction, fusion compression, region localization, and spatial association.

[0133] Parallel extraction of multi-scale visual features is achieved through a pre-trained multi-branch convolutional network. This pre-trained multi-branch convolutional network employs a heterogeneous architecture. The first branch, built on ResNet-50, specifically handles color and material features, capturing the global color distribution through dilated convolutions that expand the receptive field, while simultaneously extracting material texture features using a histogram of oriented gradients (HOR) enhancement module. The second branch uses an HRNet architecture, maintaining high-resolution feature maps throughout the network, specifically for extracting logo and contour structure features. A multi-scale fusion strategy combines semantic information from different resolutions to ensure clear brand logo recognition and accurate contour edge localization. The inputs to both branches undergo standardization, including pixel value normalization and random color enhancement, to improve feature robustness. Each branch outputs a 256-dimensional feature vector and its corresponding spatial feature map, forming a complete visual feature representation.

[0134] A feature fusion module created using a feature pyramid network structure achieves the fusion and compression of multi-source features. Specifically, it aggregates the multi-scale features output from two branches through a top-down path aggregation. The fusion process uses learnable weight parameters to dynamically adjust the contribution of different feature sources. Color features primarily influence lower-level fusion nodes, while structural features dominate the fusion decisions of higher-level nodes. In the compression stage, 1×1 convolutions are used to achieve channel dimensionality reduction, compressing the fused 512-dimensional features into 128-dimensional key detail features. The compression network employs a bottleneck structure, first expanding the dimension to 1024 for feature interaction, and then projecting it to the target dimension to retain useful information to the maximum extent. The final output key detail features contain rich information in both spatial and channel dimensions, providing sufficient basis for subsequent region localization.

[0135] The region localization module, based on a conditional random field model, achieves precise localization and mask generation of the knapsack region, modeling the knapsack localization problem as a structured prediction task. The input to the region localization module is the depth feature map of the initial synthesized image. An iterative mean-field inference algorithm calculates the probability that each pixel belongs to the knapsack region. During the inference process, three constraints—appearance consistency, spatial smoothness, and shape prior—are incorporated to ensure the accuracy of the localization results. Based on the probability output, an adaptive thresholding algorithm is used to generate a binarized initial knapsack mask. In the post-processing stage, morphological closing operations are applied to fill holes in the mask, and an edge thinning algorithm is used to optimize the contour smoothness, ultimately obtaining a precise mask representation that highly matches the actual knapsack region.

[0136] The spatial association module first calculates the feature response intensity of each spatial location in the key detail feature map and identifies high-response regions using a non-maximum suppression algorithm. The association algorithm based on this spatial association module uses an improved IoU metric to calculate the spatial overlap between each high-response region and the initial knapsack mask, considering both feature semantic similarity and spatial distance. The screening process employs a two-layer filtering mechanism: the first layer excludes obviously mismatched regions based on an overlap threshold; the second layer groups the remaining regions according to feature similarity through cluster analysis, selecting representative regions from each cluster as the final target regions. All selected target regions are accompanied by confidence scores and spatial boundary information, providing precise locational guidance for subsequent detail transfer.

[0137] Furthermore, you can also view Figure 4 , Figure 4 This is a detailed step diagram of step S40 in the first embodiment described above, based on... Figure 4 The steps of using the key detail features and the target region as input to the detail transfer and redrawing model, performing region redrawing processing based on the detail transfer and redrawing model, and replacing the target region in the initial synthesized image with the region redrawing processing result to obtain the upper body synthesized image include steps S41-43:

[0138] Step S41: Spatial alignment processing is performed between the initial synthesized image and the key detail features;

[0139] Step S42: Within the mask range corresponding to the target region, the key detail features are fused to the corresponding positions through feature injection to obtain the redrawn region image;

[0140] Step S43: Reconstruct a local area image within the target area using the redrawn area image, and obtain the synthesized upper body image based on the reconstruction result.

[0141] In this embodiment, high-quality body compositing effects are achieved through three core steps: precise spatial alignment, feature injection, and local reconstruction, which are crucial for detail transfer and image redrawing.

[0142] In completing the spatial alignment process and establishing accurate correspondences, the alignment module employs a multi-scale matching strategy based on feature pyramids. First, it extracts depth features of the knapsack region from the initial synthesized image, and simultaneously locates high-response regions in the spatial feature map of the key detail features. The matching engine uses an improved optical flow algorithm to calculate the dense correspondence field between the two feature spaces and optimizes matching accuracy through cyclic consistency loss. This alignment process specifically handles non-rigid deformation cases, using a thin-plate spline interpolation algorithm to compensate for geometric distortions caused by changes in viewpoint. The final output alignment parameters include the affine transformation matrix, perspective transformation parameters, and local deformation field, establishing an accurate spatial correspondence foundation for subsequent feature fusion. The entire alignment process is executed in parallel on the GPU, ensuring sub-pixel-level alignment accuracy even under complex deformation conditions.

[0143] A U-Net architecture is used to construct a feature injection network to perform feature injection and fusion operations, processing the initial synthesized image and the key detail features separately in the encoder stage. The injection point is set at the bottleneck layer of the network, and the detail features are conditionally injected into the generation process using spatial adaptive normalization. Specifically, the key detail features are first projected onto the same channel dimension as the backbone features through a fully connected layer. Then, within the masked area corresponding to the target region, a gated attention mechanism is used to control the intensity of feature fusion. Within the masked area, the detail features occupy the dominant weight, ensuring accurate reproduction of key visual attributes; in the masked boundary region, the system adopts a progressive blending strategy to achieve a natural transition from the detail region to the background. The fused features are reconstructed by the decoder to generate a redrawn region image that retains the original detail quality.

[0144] The reconstruction network is based on a conditional adversarial generative architecture. The generator employs a partially convolutional design, specifically optimized for masked regions. The reconstruction process is divided into two stages: first, the basic structure is generated, maintaining the overall shape and spatial position of the backpack through residual connections; then, detail enhancement is performed, utilizing a multi-scale discriminator to ensure good visual fidelity at different resolutions. The loss function combines perceptual loss, adversarial loss, and feature matching loss. The perceptual loss is calculated using a pre-trained VGG network to ensure the consistency of high-level semantic features; the feature matching loss specifically constrains the reproduction accuracy of key detail features. In the final synthesis stage, a Poisson mixture algorithm is used to seamlessly blend the redrawn region into the original background, while simultaneously performing color consistency correction and illumination coordination processing, outputting a synthesized upper body image with commercial-grade visual quality. The entire redrawing process is completed within 2.8 seconds, meeting the efficiency requirements of practical applications while maintaining detail fidelity.

[0145] Based on the above Figure 4 The content of step S42 is analyzed in detail. Specifically, the step of fusing the key detail features into the corresponding position of the initial synthesized image through feature injection within the mask range corresponding to the target region includes steps S42-1 to S42-2:

[0146] Step S42-1: Perform feature transformation processing on the key detail features, and generate source features that are adapted to the visual characteristics of the target region based on the processing results;

[0147] Step S42-2: The background features of the target region and the source features are weighted and fused to obtain comprehensive features, and the comprehensive features are decoded by a feature reconstruction network to generate a redrawn region image.

[0148] In this embodiment, high-fidelity transfer of key detail features is achieved through precise feature transformation and intelligent weighted fusion. Specifically, the feature transformation module adopts a cascaded fully connected layer structure, projecting the input 128-dimensional key detail features onto a higher-order feature space. The transformation process first expands the feature dimension to 512 dimensions through a bottleneck structure, enhances the expressive power of the features using a nonlinear activation function, and then compresses it to a 256-dimensional output that matches the target region features through a projection layer. The adaptation processing stage introduces a spatial attention mechanism to dynamically adjust the feature response intensity according to the visual characteristics of the target region. In specific implementation, the local statistical features of the target region are calculated, including average color, texture complexity, and lighting conditions, and feature calibration parameters are generated accordingly. The generation process of the source features also considers the geometric deformation of the target region, using a deformable convolutional network to spatially transform the feature map to ensure that it precisely matches the shape and contour of the target region. The final output source features retain the key visual attributes of the original backpack and are fully adapted to the target region in terms of spatial structure and visual characteristics.

[0149] The weighted fusion module receives two input sources: background features of the target region from the initial synthesized image and transformed source features. The fusion algorithm uses a gating mechanism to generate a dynamic weight map, which is calculated based on both feature similarity and spatial location. In the core region of the target area, source features are given higher weights to ensure accurate reproduction of key details; in the edge areas of the region, the weights of background features gradually increase to achieve a smooth transition. The weighted fusion formula is: Comprehensive Feature = W_s × Source Feature + (1 - W_s) × Background Feature, where the dynamic weight map W_s is normalized to the [0,1] interval using the sigmoid function. The feature reconstruction network is based on an improved U-Net architecture, embedding multiple residual connections in the encoder-decoder structure. During decoding, the network progressively upsamples the comprehensive features through transposed convolutional layers, while using skip connections to integrate low-level texture information. The reconstruction loss function combines L1 pixel-level loss, perceptual loss, and adversarial loss to ensure that the generated redrawn region image maintains detail accuracy while possessing a high degree of visual realism. The final output image of the redrawn area maintains consistency with the original backpack in key visual attributes, while seamlessly blending with the surrounding background.

[0150] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the virtual backpack wearing method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0151] This application provides a virtual backpack wearing device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the virtual backpack wearing method in the first embodiment described above.

[0152] The following is for reference. Figure 7 The diagram illustrates a structural schematic suitable for implementing the virtual backpack device in the embodiments of this application. The virtual backpack device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The virtual backpack device shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0153] like Figure 7 As shown, the virtual backpack device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the virtual backpack device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, magnetic tape, hard disk, etc.; and communication devices 1009. The communication device 1009 allows the virtual backpack device to communicate wirelessly or wiredly with other devices to exchange data. While the figures show virtual backpack devices with various systems, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.

[0154] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0155] The virtual backpack fitting device provided in this application, employing the virtual backpack fitting method described in the above embodiments, can solve the technical problem that existing virtual backpacks cannot simultaneously guarantee a reasonable spatial layout and high fidelity of key details during try-on. Compared with the prior art, the beneficial effects of the virtual backpack fitting device provided in this application are the same as those of the virtual backpack fitting method provided in the above embodiments, and other technical features in this virtual backpack fitting device are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0156] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0157] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0158] This application provides a storage medium, which is a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the virtual backpack wearing method in the above embodiments.

[0159] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0160] The aforementioned computer-readable storage medium may be included in the virtual backpack device; or it may exist independently and not assembled into the virtual backpack device.

[0161] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the virtual backpack wearing device, enable the virtual backpack wearing device to implement the technical content of the virtual backpack wearing method embodiment shown above.

[0162] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0163] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0164] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0165] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described virtual backpack wearing method. This solves the technical problem that existing virtual backpacks cannot simultaneously guarantee a reasonable spatial layout and high fidelity of key details during try-on. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the virtual backpack wearing method provided in the above embodiments, and will not be elaborated upon here.

Claims

1. A method for wearing a virtual backpack, characterized in that, The method for wearing the virtual backpack includes the following steps: Receive user image and backpack image; In response to the trial carrying instruction of the user image and the backpack image, the user image and the backpack image are input into the image understanding and editing model, and the user image and the backpack image are synthesized into an initial composite image by the image understanding and editing model. The initial composite image includes the spatial layout and occlusion relationship between the user and the backpack. Visual features of the backpack image at different sizes are extracted using a pre-trained multi-branch convolutional network. The first branch convolutional network is used to extract color and material features, and the second branch convolutional network is used to extract logo and contour structure features. The extracted visual features are fused and compressed to obtain key detail features, and the backpack region is located in the initial synthesized image. An initial backpack mask is generated based on the location result. The high-response regions of the key detail features are spatially correlated with the initial backpack mask, and feature regions are selected as target regions from the spatial correlation results. Using the key detail features and the target region as input to the detail transfer and redrawing model, the region redrawing process is performed based on the detail transfer and redrawing model. The result of the region redrawing process replaces the target region in the initial composite image to obtain the upper body composite image. The step of inputting the user image and the backpack image into an image understanding and editing model in response to a trial carrying instruction, and then compositing the user image and the backpack image into an initial composite image through the image understanding and editing model, includes: In response to the trial strap instruction from the user image and the backpack image, the trial strap instruction is semantically parsed to obtain trial strap keywords; The trial band keywords are converted into spatial constraints, and the spatial constraints are injected into the image understanding and editing model as prior knowledge. The spatial constraints are executed through the attention mechanism of the image understanding and editing model to locate the body parts that match the trial band instructions in the user image and output the location results. Based on the location results and the trial pack keywords, the image understanding and editing model synthesizes the backpack image and the user image into the initial composite image.

2. The virtual backpack wearing method as described in claim 1, characterized in that, The steps of executing the spatial constraints through the attention mechanism of the image understanding and editing model, locating the body part matching the trial band instruction in the user image, and outputting the location result include: The image understanding and editing model encodes the spatial constraints into attention query vectors and maps the visual features of the user image into key vectors and value vectors. Calculate the similarity between the attention query vector and the key vector, and generate an attention weight distribution based on the calculation results; The attention weight distribution and the value vector are iteratively processed through a multi-layer attention mechanism to obtain the weighted aggregated feature information of the value vector; Based on the weighted aggregation features, the spatial coordinates of the body parts in the user image related to the trial band instruction are determined, and the spatial coordinates are used as the positioning result.

3. The virtual backpack wearing method as described in claim 1, characterized in that, The step of synthesizing the backpack image and the user image into the initial synthesized image based on the positioning result and the trial carry keywords using the image understanding and editing model includes: Based on the spatial coordinates of the body parts determined by the positioning results, spatial transformation parameters are generated to adapt to the backpack image; The backpack image is subjected to three-dimensional pose simulation and two-dimensional projection processing based on the spatial transformation parameters. The spatial pose of the backpack image is adapted to the orientation and curvature of the body parts according to the processing results. Based on the depth information of the human body parts in the user image, calculate the expected occlusion relationship between the backpack image and the human body parts; The pixel coverage is processed by the expected occlusion relationship, and the contact area is feathered and rendered with light and shadow to generate the initial composite image.

4. The virtual backpack wearing method as described in claim 1, characterized in that, The steps of using the key detail features and the target region as input to the detail transfer and redrawing model, performing region redrawing processing based on the detail transfer and redrawing model, and replacing the target region in the initial synthesized image with the result of the region redrawing processing to obtain the upper body synthesized image include: The initial synthesized image is spatially aligned with the key detail features; Within the mask range corresponding to the target region, the key detail features are fused to the corresponding positions through feature injection to obtain the redrawn region image; The redrawn region image is used to reconstruct a local region image within the target region, and the synthesized upper body image is obtained based on the reconstruction result.

5. The virtual backpack wearing method as described in claim 4, characterized in that, Within the mask range corresponding to the target region, the step of fusing the key detail features to the corresponding positions through feature injection to obtain the redrawn region image includes: The key detail features are subjected to feature transformation processing, and source features that are adapted to the visual characteristics of the target region are generated based on the processing results; The background features of the target region are weighted and fused with the source features to obtain a comprehensive feature. The comprehensive feature is then decoded by a feature reconstruction network to generate a redrawn region image.

6. A virtual backpack device, characterized in that, The virtual backpack device stores a computer program, which, when executed by a processor, implements the virtual backpack wearing method according to any one of claims 1-5.

7. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the virtual backpack wearing method according to any one of claims 1-5.