An e-commerce personalization fitting method and system for a multi-modal query

CN122798516APending Publication Date: 2026-09-22ZHIFU TECH (FUJIAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611233478.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-14
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

试衣过程缺少布料物理形变仿真运算环节,不能依据人体体态完成服饰结构形态的调整,也无法提取标准化纹理映射关系用于图像生成

Benefits of technology

[0091]This approach integrates human body images and natural language text describing clothing to form multimodal query data. Multimodal feature encoding is performed on heterogeneous image and text data to generate a joint feature vector with a unified dimension. This joint feature vector is then used as the retrieval basis for matching and filtering within an e-commerce clothing database. By combining visual and semantic features in the selection of candidate clothing, the limitations of a single retrieval method are overcome, expanding the feature dimensions available for the retrieval process and allowing it to consider both the physical attributes of human body image information and textual descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122798516A_ABST
    Figure CN122798516A_ABST
Patent Text Reader

Abstract

This invention discloses a multimodal query-based personalized virtual try-on method and system for e-commerce, relating to the field of virtual try-on simulation technology. The method includes collecting multimodal query data composed of user human body images and natural language text describing clothing; generating joint feature vectors by multimodal feature encoding of the multimodal query data; retrieving a list of candidate clothing from an e-commerce clothing database based on the joint feature vectors; constructing 3D virtual reconstruction models of the candidate clothing; spatially aligning the 3D clothing models with the human body images and their poses; running a fabric physical deformation simulation process; extracting texture mapping relationships under the clothing's wearing state; generating virtual try-on result images based on the texture mapping relationships; and outputting try-on feedback to achieve multimodal query-driven virtual try-on in e-commerce scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual try-on simulation technology, specifically a multimodal query e-commerce personalized try-on method and system. Background Technology

[0002] Existing virtual fitting applications in the e-commerce sector mostly employ a single modality for search interaction, supporting only text keyword searches or single human image searches, and are incompatible with heterogeneous input formats such as images and natural language text. A single information modality cannot fully represent a user's body shape characteristics and clothing preferences, easily leading to matching biases during the search process and failing to adapt to diverse user input expressions. Conventional search architectures lack a unified feature encoding mechanism for image and text modalities; the two types of modal data are independent of each other, making it difficult to form a unified feature vector for database clothing matching.

[0003] Current virtual try-on processes generally use 2D Lycra mapping and overlay methods to present clothing, without performing 3D solid reconstruction of the garments or possessing the spatial alignment logic between the 3D model and the human body posture. The try-on process lacks a fabric physical deformation simulation step, cannot adjust the garment's structure based on the human body posture, and cannot extract standardized texture mapping relationships for image generation. The 2D presentation mode lacks three-dimensional structure and physical rule constraints, making it difficult for the presented try-on state to accurately reflect the actual wearing condition.

[0004] E-commerce virtual try-on services require standardized capabilities in multimodal heterogeneous data fusion encoding and retrieval, clothing 3D reconstruction, human pose spatial alignment, and fabric physical simulation and texture mapping generation. Current technologies lack a multimodal joint feature encoding retrieval architecture and a complete processing flow for 3D reconstruction pose alignment and physical deformation simulation, making it difficult to meet the operational standards of high-precision, structured e-commerce virtual try-on services. Summary of the Invention

[0005] This invention aims to solve at least one of the technical problems existing in the prior art;

[0006] Therefore, this invention proposes a multimodal query method for personalized online clothing try-on in e-commerce, comprising:

[0007] Obtain multimodal query data input by the user;

[0008] The multimodal query data includes at least one human image uploaded by the user and at least one piece of natural language text describing the desired clothing;

[0009] Multimodal feature encoding is performed on the multimodal query data to generate a joint feature vector;

[0010] Retrieve a list of candidate clothing items from the e-commerce clothing database based on the joint feature vector;

[0011] Perform 3D virtual reconstruction on each garment in the candidate clothing list to generate a 3D clothing model;

[0012] Spatially align the 3D clothing model with the human body posture in the human body image;

[0013] Perform physical deformation simulation of fabric between the spatially aligned 3D clothing model and the human body image;

[0014] Extract the texture mapping relationship of clothing after wearing from the results of fabric physical deformation simulation;

[0015] A virtual fitting result image is generated based on the texture mapping relationship;

[0016] The virtual try-on result image is output as the try-on feedback corresponding to the multimodal query data.

[0017] Furthermore, the steps for obtaining multimodal query data input by the user specifically include:

[0018] Receive at least one human image uploaded by the user on the e-commerce front-end interface;

[0019] The at least one human body image includes a frontal human body image, a back human body image, and at least one side human body image;

[0020] Simultaneously receive at least one natural language text describing the desired clothing input by the user on the same e-commerce front-end interface;

[0021] The natural language text includes at least one of the following: clothing category description, clothing color description, clothing material description, and clothing pattern description;

[0022] The received at least one human image and at least one natural language text describing the desired clothing are packaged into a multimodal query data packet according to the chronological order of user input.

[0023] Assign a unique session identifier to the multimodal query data packet;

[0024] The unique session identifier is associated with and stored in the multimodal query data packet.

[0025] Further, the step of performing multimodal feature encoding on the multimodal query data to generate a joint feature vector specifically includes:

[0026] Separate the at least one human image and the at least one natural language text describing the desired clothing from the multimodal query data;

[0027] Perform forward computation of a convolutional neural network on the at least one human image to extract multi-layer convolutional feature maps for each human image;

[0028] Perform word embedding encoding on the at least one natural language text describing the desired clothing, converting each word into a dense word vector;

[0029] Perform forward computation using a bidirectional long short-term memory network on the dense word vectors to extract text context feature vectors;

[0030] The multi-layer convolutional feature map is converted into a global feature vector of human image through an adaptive pooling layer;

[0031] The global feature vector of the human face and the text context feature vector are concatenated along the feature dimension.

[0032] Perform a multilayer perceptron transformation on the concatenated feature vectors to output the joint feature vector.

[0033] Furthermore, the step of retrieving a candidate clothing list from the e-commerce clothing database based on the joint feature vector specifically includes:

[0034] Obtain the apparel feature vector of each garment from the pre-stored e-commerce apparel database;

[0035] The clothing feature vector and the joint feature vector are in the same feature space dimension;

[0036] Calculate the cosine similarity between the joint feature vector and the garment feature vector of each garment;

[0037] Sort all clothing items in descending order of cosine similarity;

[0038] Select the top preset number of clothing items from the sorting results as the initial candidate clothing set;

[0039] The initial candidate clothing set is filtered by clothing category to remove clothing whose categories do not match the clothing category description in the natural language text.

[0040] The filtered clothing items are reordered according to cosine similarity to generate the candidate clothing list.

[0041] Furthermore, the step of performing 3D virtual reconstruction on each garment in the candidate garment list to generate a 3D garment model specifically includes:

[0042] Retrieve the currently pending clothing identifier from the candidate clothing list;

[0043] Based on the clothing identifier, the multi-view display image set and size specification parameters of the clothing are read from the e-commerce clothing database;

[0044] A multi-view stereo matching algorithm is executed on the multi-view display image set to calculate the three-dimensional point cloud of the clothing surface;

[0045] Perform Poisson surface reconstruction on the three-dimensional point cloud to generate an initial triangular mesh model;

[0046] The initial triangular mesh model is non-uniformly scaled according to the size specification parameters so that the girth of the scaled triangular mesh model is consistent with the size specification parameters.

[0047] The scaled triangular mesh model is used as the 3D clothing model.

[0048] Furthermore, the step of spatially aligning the 3D clothing model with the human pose in the human body image specifically includes:

[0049] Perform human key point detection on the human image and output a set of coordinates of key points of the two-dimensional human skeleton;

[0050] Estimate the three-dimensional posture parameters of the human body based on the set of coordinates of key points of the two-dimensional human skeleton.

[0051] The human body three-dimensional posture parameters include the rotation angle and three-dimensional spatial position of each joint point;

[0052] A parameterized human body mesh model is constructed based on the aforementioned three-dimensional human body posture parameters;

[0053] Calculate the scaling ratio between the bounding box of the 3D clothing model and the bounding box of the parametric human body mesh model;

[0054] The vertex coordinates of the 3D clothing model are uniformly scaled according to the scaling ratio.

[0055] The spatial alignment is achieved by projecting each vertex of the scaled 3D clothing model to the nearest neighbor point of the parameterized human body mesh model.

[0056] Furthermore, the step of simulating the physical deformation of the fabric between the spatially aligned 3D clothing model and the human body image specifically includes:

[0057] Initialize all mesh vertices of the spatially aligned 3D clothing model as point masses;

[0058] By setting spring constraints between adjacent mass points, a mass-spring system is formed.

[0059] The spring constraints include three categories: structural springs, shear springs, and bending springs.

[0060] Collision detection boundary conditions are set based on the surface normals of the parameterized human body mesh model;

[0061] A gravitational force and a frictional force between the clothing and the human body are applied to the mass spring system.

[0062] The equations of motion of the point-spring system are solved iteratively using the implicit Euler integral method.

[0063] After each iteration of the solution is completed, the positions of all particles are projected outside the collision detection boundary conditions.

[0064] The iteration stops when the velocity change of all particles is lower than the preset kinetic energy convergence threshold, and the deformed 3D clothing mesh is output.

[0065] Furthermore, the steps for extracting the texture mapping relationship of clothing after wearing from the results of fabric physical deformation simulation specifically include:

[0066] Obtain the three-dimensional spatial coordinates of the three vertices of each triangular facet in the deformed three-dimensional clothing mesh;

[0067] Obtain the original texture coordinates of the three vertices of the same triangular facet in the 3D clothing model before deformation;

[0068] Calculate the affine transformation matrix from the original texture coordinate space to the three-dimensional space coordinates for each triangular facet;

[0069] The affine transformation matrix is ​​stored as the texture mapping parameters of the triangular facet.

[0070] All texture mapping parameters of the triangles are organized into a texture mapping parameter list according to the triangle patch index order;

[0071] The texture mapping parameter list is used as the texture mapping relationship after the clothing is worn.

[0072] Furthermore, the step of generating the virtual fitting result image based on the texture mapping relationship specifically includes:

[0073] Create a blank canvas for the target fitting image;

[0074] The size of the blank target fitting image canvas is the same as the size of the human body image;

[0075] Perform the following operations on each triangular facet in the deformed 3D clothing mesh:

[0076] Extract the pixel block of the human body region covered by the triangular facet from the human body image;

[0077] The human body region pixel blocks are inversely transformed according to the texture mapping parameters of the triangular facet to generate clothing texture pixel blocks;

[0078] The clothing texture pixel block is drawn onto the screen projection area corresponding to the triangular facet on the blank target fitting image canvas;

[0079] After traversing all the triangular facets, a weighted average smoothing filter is performed on the pixels at the intersection of different triangular facets on the target fitting image canvas.

[0080] The smoothed and filtered target fitting image canvas is used as the virtual fitting result image.

[0081] The step of outputting the virtual try-on result image as the try-on feedback corresponding to the multimodal query data specifically includes:

[0082] Encode the virtual try-on result image into a web-compatible image format;

[0083] Generate a fitting feedback data packet containing the virtual fitting result image;

[0084] The fitting feedback data package also includes the product link address of each garment in the candidate clothing list;

[0085] The user's current e-commerce front-end interface instance is located based on the unique session identifier.

[0086] Send the fitting feedback data packet to the found e-commerce front-end interface instance;

[0087] The virtual try-on result image is rendered and displayed on the e-commerce front-end interface instance;

[0088] A clickable button displaying the product link address is rendered in the area below the virtual try-on result image.

[0089] Furthermore, the present invention also includes a multimodal query e-commerce personalized try-on system, the system including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements the steps of the multimodal query e-commerce personalized try-on method described above.

[0090] Compared with the prior art, the beneficial effects of the present invention are:

[0091] This approach integrates human body images and natural language text describing clothing to form multimodal query data. Multimodal feature encoding is performed on heterogeneous image and text data to generate a joint feature vector with a unified dimension. This joint feature vector is then used as the retrieval basis for matching and filtering within an e-commerce clothing database. By combining visual and semantic features in the selection of candidate clothing, the limitations of a single retrieval method are overcome, expanding the feature dimensions available for the retrieval process and allowing it to consider both the physical attributes of human body image information and textual descriptions.

[0092] The retrieved candidate garments are reconstructed in 3D virtual form to create a complete 3D garment model with a full spatial structure. The coordinates and structure are aligned spatially according to the human posture in the human body image, establishing a positional correspondence between the 3D garment model and the human body posture. Simulation calculations of fabric physical deformation are performed, and the garment structure is adjusted according to physical rules to generate a garment distribution state that conforms to the human body posture.

[0093] The texture mapping relationship of clothing under wearing conditions is extracted from the physical deformation simulation results of the fabric. Using fixed mapping parameters as the basis for image generation, the virtual try-on result image is generated and output. The entire processing flow establishes a complete processing link from multimodal feature encoding, clothing retrieval, 3D reconstruction, pose alignment to physical simulation and texture mapping generation. It standardizes the data processing standards of each stage of virtual try-on, maintains the data coherence between multimodal query input and try-on image output, and adapts to the standardized virtual try-on business operation mode in e-commerce scenarios. Attached Figure Description

[0094] Figure 1 This is a state diagram of a multimodal query method for personalized virtual try-on in e-commerce, as described in this invention.

[0095] Figure 2 A flowchart for multimodal query data acquisition and associated storage;

[0096] Figure 3 A flowchart for generating joint feature vectors for multimodal feature encoding. Detailed Implementation

[0097] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0098] See Figure 1The overall implementation scheme of this multimodal query-based personalized clothing try-on method for e-commerce is as follows: First, acquire multimodal query data input by the user. This multimodal query data includes at least one human image uploaded by the user and at least one natural language text describing the desired clothing. Second, perform multimodal feature encoding processing on the acquired multimodal query data to generate a joint feature vector. Third, retrieve a candidate clothing list from the e-commerce clothing database based on the generated joint feature vector. Fourth, perform 3D virtual reconstruction on each garment in the candidate clothing list to generate a corresponding 3D clothing model. Fifth, spatially align the generated 3D clothing model with the human pose in the user-uploaded human image. Sixth, perform fabric physical deformation simulation between the spatially aligned 3D clothing model and the human image. Seventh, extract the texture mapping relationship after the clothing is worn from the results of the fabric physical deformation simulation. Eighth, generate a virtual try-on result image based on the extracted texture mapping relationship. Finally, output the generated virtual try-on result image to the user as the try-on feedback corresponding to the multimodal query data.

[0099] In one embodiment of the present invention, when obtaining multimodal query data input by the user, refer to... Figure 2 The system receives at least one human image uploaded by a user on the e-commerce front-end interface. This at least one human image includes a frontal human image, a back human image, and at least one side human image. Simultaneously, it receives at least one piece of natural language text input by the user on the same e-commerce front-end interface, describing the desired clothing. This natural language text includes at least one of the following: clothing category description, clothing color description, clothing material description, and clothing style description. The received at least one human image and at least one piece of natural language text describing the desired clothing are packaged into a multimodal query data packet according to the chronological order of user input. A unique session identifier is assigned to this multimodal query data packet, and the unique session identifier is associated with and stored in conjunction with the multimodal query data packet.

[0100] In practice, users upload at least one human image through the e-commerce front-end interface. The at least one human image includes a frontal human image, a back human image, and at least one side human image. For example, a user uploads a full-body frontal photo, a full-body back photo, and a full-body left-side photo. Simultaneously, users input at least one piece of natural language text describing the desired clothing on the same e-commerce front-end interface. The natural language text includes at least one of the following: clothing category description, clothing color description, clothing material description, and clothing style description. For example, the text input by the user is "long-sleeved dress, red, silk material, A-line style".

[0101] In some embodiments, the received at least one human body image and at least one natural language text describing the desired clothing are packaged into a multimodal query data packet according to the chronological order of user input. In this case, the user first uploads a frontal human body image, then a back human body image, then a side human body image, and finally inputs the natural language text. The frontal human body image file, the back human body image file, the side human body image file, and the string of the natural language text are then merged into a binary data packet according to the chronological order. The header of the binary data packet records the length and type identifier of each data segment. Optionally, a unique session identifier is assigned to the multimodal query data packet. The unique session identifier is generated in the following manner:

[0102]

[0103] in: This represents the unique session identifier. Indicates a secure hash algorithm. Represents the user account identifier. This represents the Unix timestamp when the multimodal query data packet was received. This represents a randomly generated 128-bit random number. This indicates a byte string concatenation operation.

[0104] It is understood that the unique session identifier is associated with the multimodal query data packet and stored. A key-value pair storage entry is created in the e-commerce backend server. The key of the key-value pair storage entry is the unique session identifier, and the value of the key-value pair storage entry is the complete binary content of the multimodal query data packet, while recording the storage timestamp.

[0105] In some embodiments, when the natural language text simultaneously contains clothing category description, clothing color description, clothing material description, and clothing pattern description, the above four descriptions are extracted into independent text fragments according to the order of appearance in the natural language input by the user, and then all text fragments are concatenated into a continuous text string according to the order of appearance and stored in the multimodal query data package.

[0106] Optionally, before packaging the multimodal query data packet, each received human image undergoes format standardization processing, adjusting the pixel dimensions of the human image to a preset standard width and height, and converting the color space to RGB three-channel format. The preset standard width and height are 512 pixels. The natural language text is preprocessed by removing leading and trailing whitespace characters and compressing consecutive whitespace characters. It is understood that the unique session identifier is simultaneously written into a session management queue, maintained by the e-commerce backend server. Each session management queue element contains the unique session identifier and the corresponding reception time of the multimodal query data packet, used in subsequent steps to quickly locate the e-commerce frontend interface instance currently used by the user based on the unique session identifier.

[0107] In one embodiment of the present invention, when performing multimodal feature encoding to generate a joint feature vector on multimodal query data, refer to... Figure 3 The process involves separating at least one human image and at least one natural language text describing the desired clothing from multimodal query data. A convolutional neural network is used to perform forward computation on the separated human image to extract multi-layer convolutional feature maps for each image. Word embedding encoding is performed on the separated natural language text describing the desired clothing, converting each word in the text into a dense word vector. A bidirectional long short-term memory network is then used to perform forward computation on the converted dense word vectors to extract text context feature vectors. The extracted multi-layer convolutional feature maps are then converted into a global human image feature vector through an adaptive pooling layer. This global human image feature vector is then concatenated with the text context feature vector along the feature dimension. A multilayer perceptron transformation is performed on the concatenated feature vector to finally output a joint feature vector.

[0108] When retrieving a candidate clothing list from an e-commerce clothing database based on the joint feature vector, the process first involves obtaining the clothing feature vector of each garment pre-stored in the database. This garment feature vector shares the same feature space dimension as the joint feature vector. The cosine similarity between the joint feature vector and the garment feature vector of each individual garment is calculated, and all garments are sorted in descending order of their calculated cosine similarity. A predetermined number of garments at the top of the sorted list are selected as the initial candidate clothing set. This initial candidate clothing set is then filtered by clothing category, removing garments whose categories do not match the descriptions in the natural language text. The remaining garments are then re-sorted according to their cosine similarity to generate the final candidate clothing list.

[0109] In specific implementation, when performing multimodal feature encoding to generate joint feature vectors on the multimodal query data, at least one human body image and at least one piece of natural language text describing the desired clothing are separated from the multimodal query data. For example, a frontal human body image, a back human body image, a left-side human body image, and the text "white shirt, cotton, slim fit" are parsed from the multimodal query data package. A convolutional neural network is used to perform forward computation on the separated at least one human body image to extract multi-layer convolutional feature maps for each human body image. The convolutional neural network uses a ResNet-50 network pre-trained on the ImageNet dataset. The frontal human body image, the back human body image, and the left-side human body image are respectively input into the ResNet-50 network, and feature maps with sizes of 28×28×512 and 14×14×1024 are extracted at the output of the third and fourth convolutional blocks of the network model, respectively.

[0110] In some embodiments, word embedding encoding is performed on the separated natural language text describing the desired clothing, converting each word in the text into a dense word vector. Using a pre-trained 300-dimensional GloVe word embedding matrix, each word in the natural language text "white shirt, cotton, slim fit" ("white", "shirt", "cotton", "slim fit", "fit") is mapped to a 300-dimensional real vector. The resulting dense word vectors are then processed by a bidirectional long short-term memory network (LSTM) to extract the text context feature vector. The bidirectional LSTM network contains two independent LSTM layers: one processes the dense word vector sequence from front to back, and the other processes it from back to front. The last hidden state in both directions is concatenated to obtain a 600-dimensional text context feature vector.

[0111] Optionally, the extracted multi-layer convolutional feature maps are converted into a global human image feature vector through an adaptive pooling layer. For each human image, the extracted multi-layer convolutional feature map is passed through an adaptive average pooling layer. The adaptive average pooling layer converts the feature map of arbitrary size into a fixed-size 1×1×C output, where C is the number of channels of the feature map. The pooling outputs of the frontal human image, the back human image, and the side human image are concatenated along the channel dimension to form a total human image feature vector. Then, the total human image feature vector is concatenated with the text context feature vector along the feature dimension. The dimension of the total human image feature vector is (512+1024)×3=4608, and the dimension of the text context feature vector is 600. After concatenation, an initial fusion vector of 5208 dimensions is obtained.

[0112] It is understood that a multilayer perceptron transformation is performed on the concatenated feature vector to output the joint feature vector. The multilayer perceptron contains two fully connected layers and one ReLU activation layer. The first fully connected layer maps the 5208-dimensional input to a 1024-dimensional output. After passing through the ReLU activation function, the input is fed into the second fully connected layer, which maps the 1024-dimensional input to a 512-dimensional output. The 512-dimensional output is the joint feature vector.

[0113] In specific implementation, when retrieving a candidate clothing list from the e-commerce clothing database based on the joint feature vector, the clothing feature vector of each garment pre-stored in the e-commerce clothing database is obtained. The clothing feature vector and the joint feature vector are in the same feature space dimension, i.e., both are 512-dimensional vectors. The clothing feature vector is pre-calculated and stored by inputting multiple display images and corresponding text descriptions of each garment from the e-commerce clothing database into the same feature encoding network. The cosine similarity between the joint feature vector and the clothing feature vector of each garment is calculated. The cosine similarity is calculated as follows:

[0114]

[0115] in: Represents the joint feature vector, Indicates the first item in the e-commerce apparel database The feature vector of a garment. This represents the dot product operation of two vectors. Denotes the L2 norm of the joint eigenvector. Indicates the first The L2 norm of the feature vector of a garment.

[0116] In some embodiments, all clothing items are sorted in descending order of their calculated cosine similarity. For example, if there are 10,000 clothing items in an e-commerce clothing database, after calculating the cosine similarity of each item, a quicksort algorithm is used to arrange the 10,000 clothing items in descending order of their cosine similarity values. A preset number of clothing items at the top of the sorting results are selected as the initial candidate clothing set. The preset number is 200, that is, the top 200 clothing items with the highest cosine similarity are selected to form the initial candidate clothing set.

[0117] Optionally, an apparel category filtering operation is performed on the initial candidate apparel set to remove apparel whose category does not match the apparel category description in the natural language text. The apparel category description extracted from the natural language text is "shirt". The category tag of each apparel in the initial candidate apparel set stored in the e-commerce apparel database is checked. If the apparel category tag is "T-shirt", "dress", "pants" or "outerwear", the apparel is removed from the set. If the apparel category tag is "shirt" or "shirt", the apparel is retained. The remaining apparel after filtering is re-sorted according to cosine similarity to generate the final candidate apparel list. For example, if 45 apparel remain after filtering, these 45 apparel are arranged from high to low according to cosine similarity to form the candidate apparel list.

[0118] It is understood that when the natural language text simultaneously contains descriptions of clothing category, clothing color, clothing material, and clothing pattern, clothing color filtering, clothing material filtering, and clothing pattern filtering are further performed after the clothing category filtering step. The clothing color filtering uses a mapping table from color name to RGB color space for comparison, the clothing material filtering uses a predefined set of material keywords for matching, and the clothing pattern filtering uses a set of pattern keywords for matching. After filtering layer by layer, a list of candidate clothing that matches all the user's description conditions is obtained.

[0119] In one embodiment of the present invention, when performing 3D virtual reconstruction to generate a 3D clothing model for each garment in the candidate clothing list, the identifier of the garment to be processed is first retrieved from the candidate clothing list. Based on the clothing identifier, a set of multi-view display images and size specifications of the garment are read from an e-commerce clothing database. A multi-view stereo matching algorithm is performed on the read multi-view display image set to calculate the 3D point cloud of the garment surface. Poisson surface reconstruction is performed on the calculated 3D point cloud to generate an initial triangular mesh model. Based on the read size specifications, the initial triangular mesh model is non-uniformly scaled to ensure that the girth of the scaled triangular mesh model is consistent with the size specifications. The scaled triangular mesh model is then used as the 3D clothing model.

[0120] When spatially aligning a 3D clothing model with a human pose in a human image, the process begins by detecting key points on the human image and outputting a set of coordinates for key points on a 2D human skeleton. Based on this set, 3D pose parameters are estimated, including the rotation angles and 3D spatial positions of each joint. A parametric human mesh model is then constructed based on these estimated parameters. The scaling ratio between the bounding box of the 3D clothing model and the bounding box of the parametric human mesh model is calculated, and the vertex coordinates of the 3D clothing model are uniformly scaled according to this calculated scaling ratio. Finally, each vertex of the scaled 3D clothing model is projected to its nearest neighbor on the parametric human mesh model, completing the spatial alignment.

[0121] In specific implementation, a 3D virtual reconstruction is performed on each garment in the candidate garment list to generate a 3D garment model. The current garment identifier to be processed is retrieved from the candidate garment list. For example, the garment identifier ranked first in the candidate garment list is "ITEM_87392". Based on the garment identifier "ITEM_87392", a set of multi-view display images and size specifications of the garment are read from the e-commerce garment database. The set of multi-view display images contains garment images taken from 12 different angles, and the size specifications include chest circumference, waist circumference, garment length, and shoulder width.

[0122] In some embodiments, a multi-view stereo matching algorithm is executed on the read multi-view display image set to calculate the three-dimensional point cloud of the clothing surface. The multi-view stereo matching algorithm adopts a deep learning-based MVSNet architecture. Twelve multi-view display images are input into the MVSNet network, and a point cloud data containing approximately 150,000 three-dimensional points is output. Each three-dimensional point contains X, Y, and Z coordinate values. Poisson surface reconstruction is performed on the calculated three-dimensional point cloud to generate an initial triangular mesh model. The depth parameter of the Poisson surface reconstruction is set to 8. After solving the Poisson equation, an initial triangular mesh model composed of approximately 50,000 triangular facets is obtained.

[0123] Optionally, the initial triangular mesh model is non-uniformly scaled according to the read size specifications to ensure that the circumference of the scaled triangular mesh model is consistent with the size specifications. For example, if the size specifications include a chest circumference of 88 cm, a waist circumference of 70 cm, a garment length of 92 cm, and a shoulder width of 38 cm, the original chest circumference, waist circumference, garment length, and shoulder width of the initial triangular mesh model are first calculated. Then, the chest circumference scaling factor, waist circumference scaling factor, garment length scaling factor, and shoulder width scaling factor are calculated respectively. All vertex coordinates of the initial triangular mesh model are multiplied by their respective dimensions by the corresponding scaling factor to obtain the scaled triangular mesh model. The scaled triangular mesh model is then used as a 3D clothing model. It can be understood that when the multi-view image set includes images of the internal structure of clothing, the multi-view stereo matching algorithm additionally reconstructs the inner surface to generate a closed triangular mesh model.

[0124] In specific implementation, the three-dimensional clothing model is spatially aligned with the human posture in the human body image, human key point detection is performed on the human body image, and a set of coordinates of two-dimensional human skeleton key points is output. A human key point detector based on a convolutional pose machine is used. After inputting the human body image, the coordinates of 18 two-dimensional human skeleton key points are output. The 18 two-dimensional human skeleton key points include the left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, left and right ankles, as well as the center points of the head, neck and torso.

[0125] Optionally, the three-dimensional human posture parameters are estimated based on the detected set of coordinates of key points of the two-dimensional human skeleton. The three-dimensional human posture parameters include the rotation angle and three-dimensional spatial position of each joint. The inverse kinematics optimization method is adopted, using a standard human skeleton template as the initial estimate. The rotation angle and three-dimensional spatial position of each joint are iteratively adjusted by minimizing the reprojection error, so that the error between the reprojected two-dimensional key points and the detected coordinates of the two-dimensional human skeleton key points is minimized. The three-dimensional human posture parameters include the quaternion rotation representation of 16 joints and the three-dimensional spatial coordinate values ​​of 16 joints.

[0126] In some embodiments, a parametric human body mesh model is constructed based on the estimated three-dimensional human body pose parameters. Using the SMPL parametric human body model, the estimated rotation angles of the 16 joints are input into the pose parameter vector of the SMPL model. The estimated three-dimensional spatial coordinates of the 16 joints are used to determine the global translation of the SMPL model. The SMPL model outputs a parametric human body mesh model containing 6890 vertices and 13776 triangular faces. The bounding box of the three-dimensional clothing model and the parametric human body mesh model are calculated. The scaling ratios between the bounding boxes are calculated. The minimum corner coordinates of the bounding box of the 3D clothing model are (-25.3, -12.8, -3.6), and the maximum corner coordinates are (25.3, 82.4, 18.2). The minimum corner coordinates of the bounding box of the parametric human body mesh model are (-22.1, -11.5, -5.2), and the maximum corner coordinates are (22.1, 170.3, 28.7). The scaling ratios in the X-axis direction, Y-axis direction, and Z-axis direction are calculated, and the scaling ratio is taken as the maximum value among the scaling ratios in the three directions.

[0127] It is understood that the vertex coordinates of the 3D clothing model are uniformly scaled according to the calculated scaling ratio, that is, each vertex coordinate in the 3D clothing model is multiplied by the scaling ratio, while maintaining the relative shape inside the 3D clothing model. Then, each vertex of the scaled 3D clothing model is projected to its nearest neighbor onto the nearest surface point of the parametric human body mesh model to complete spatial alignment. For a vertex coordinate in the scaled 3D clothing model... Traverse all triangular faces of the parameterized human body mesh model and calculate The projection point corresponding to the shortest distance to each triangular facet ,Will The coordinates are updated to After traversing all vertices of the 3D clothing model, each vertex of the 3D clothing model is attached to the surface of the parametric human body mesh model, as shown in Table 1.

[0128] Table 1: Comparison Table of Scaling Parameters for 3D Clothing Models

[0129] Chest direction 84.6 cm 88 cm 1.0402 Waist direction 66.2 cm 70 cm 1.0574 Length direction 88.5 cm 92 cm 1.0395 Shoulder width direction 36.8 cm 38 cm 1.0326

[0130] The formula is an expression for calculating the scaling ratio:

[0131]

[0132] in: Indicates the uniform scaling factor. This refers to the half-width value of the chest in the size specification parameters. This represents the original half-width value of the chest area in the 3D clothing model. This refers to the garment length value in the size specification parameters. This represents the original garment length value of the three-dimensional clothing model. This refers to the shoulder thickness value in the dimensional specification parameters. This represents the original shoulder thickness value of the three-dimensional clothing model. This indicates the operation of taking the maximum value of the three values ​​in parentheses.

[0133] In one embodiment of the present invention, when performing fabric physical deformation simulation between a spatially aligned 3D clothing model and a human body image, all mesh vertices of the spatially aligned 3D clothing model are initialized as mass points. Spring constraints are set between adjacent mass points to form a mass-spring system, which includes three types of spring constraints: structural springs, shear springs, and bending springs. Collision detection boundary conditions are set according to the surface normal of the parameterized human body mesh model, and gravitational force and frictional force between the clothing and the human body are applied to the mass-spring system. The equations of motion of the mass-spring system are iteratively solved using the implicit Euler integral method. After each iteration, the positions of all mass points are projected outside the collision detection boundary conditions. The iteration stops when the velocity change of all mass points is lower than a preset kinetic energy convergence threshold, and the deformed 3D clothing mesh is output.

[0134] When extracting the texture mapping relationship of clothing after wearing from the results of fabric physical deformation simulation, the three-dimensional spatial coordinates of the three vertices of each triangle in the deformed 3D clothing mesh are obtained, along with the original texture coordinates of the three vertices of the same triangle in the undeformed 3D clothing model. For each triangle, an affine transformation matrix from the original texture coordinate space to the 3D spatial coordinates is calculated, and this affine transformation matrix is ​​stored as the texture mapping parameter for that triangle. The texture mapping parameters of all triangles are organized into a texture mapping parameter list according to their index order, and this list is used as the texture mapping relationship after clothing wear.

[0135] In the specific implementation, a physical deformation simulation of the fabric is performed between the spatially aligned 3D clothing model and the human body image. All mesh vertices of the spatially aligned 3D clothing model are initialized as mass points. The spatially aligned 3D clothing model contains 18,762 mesh vertices and 37,456 triangular facets. Spring constraints are set between adjacent mass points to form a mass point spring system. The spring constraints include three types: structural springs, shear springs, and bending springs. Structural springs connect two adjacent vertices within the same triangular facet, shear springs connect two vertices that share an edge but do not belong to the same triangular facet, and bending springs connect two non-adjacent vertices that span two triangular facets.

[0136] In some embodiments, collision detection boundary conditions are set according to the surface normal of the parametric human body mesh model. For each triangular facet of the surface of the parametric human body mesh model, its unit normal vector is calculated, and the positive direction of the unit normal vector is designated as the external direction of the human body. The collision detection boundary conditions stipulate that any mass in the mass-spring system cannot penetrate any triangular facet of the parametric human body mesh model during movement, that is, the projection of the mass in the collision detection boundary conditions must be located on the outer side of the triangular facet. A gravitational field force and a frictional force between the clothing and the human body are applied to the mass-spring system. The magnitude of the gravitational field force is the mass of the mass multiplied by the gravitational acceleration constant 9.8 m / s², and the direction is vertically downward. The magnitude of the frictional force is proportional to the relative velocity of the contact point between the mass and the surface of the parametric human body mesh model, with a proportionality coefficient of 0.3.

[0137] Optionally, the implicit Euler integral method is used to iteratively solve the equations of motion of the point mass spring system. The equations of motion of the point mass spring system are: the mass matrix multiplied by the acceleration vector equals the external force vector minus the spring constraint force vector. The iterative update rule of the implicit Euler integral method is as follows:

[0138]

[0139] in: This represents the mass diagonal matrix of a point mass spring system, where each diagonal element corresponds to the mass value of a point mass. This indicates the time step, which has a value of 0.005 seconds. The Jacobian matrix represents the elastic force vector with respect to the particle position vector. This represents the vector of changes in the velocity of a particle within a time step. Indicates at time particle position vector and velocity vector The resultant force vector under the conditions; after each iteration, the positions of all particles are projected outside the collision detection boundary conditions. Specifically, each particle is traversed. If the position of a particle is inside the triangular facet of the parameterized human body mesh model, the particle position is moved outside the triangular facet along the normal direction of the triangular facet, with a moving distance of 0.001 meters.

[0140] It is understood that the iteration stops when the velocity change of all particles is lower than the preset kinetic energy convergence threshold, which is 0.001 joules. The total kinetic energy of all particles is calculated. If the absolute value of the difference between the total kinetic energy of the current iteration step and the total kinetic energy of the previous iteration step is less than 0.001 joules, then convergence is determined, the iteration process stops, and the deformed 3D clothing mesh is output, as shown in Table 2.

[0141] Table 2: Setting Table of Constraint Parameters for Three Types of Springs in a Point Mass Spring System

[0142] Structural springs 1500 Euclidean distance between two points Adjacent vertices of the same triangular face shear spring 800 Euclidean distance between two points Adjacent triangular faces share edges corresponding to vertices. Bending spring 400 Euclidean distance between two points Non-adjacent vertices spanning two triangular faces

[0143] In practice, the texture mapping relationship after the garment is worn is extracted from the results of the fabric physical deformation simulation. The three-dimensional spatial coordinates of the three vertices of each triangular facet in the deformed three-dimensional garment mesh are obtained. For example, the coordinates of the three vertices of triangle number 237 after deformation are... , , Simultaneously, the original texture coordinates of the three vertices of the same triangular facet in the 3D clothing model before deformation are obtained. The 3D clothing model before deformation stores the original texture coordinates of each vertex. The original texture coordinates of the three vertices of the 237th triangular facet are as follows: , , .

[0144] In some embodiments, an affine transformation matrix from the original texture coordinate space to the three-dimensional space coordinates is calculated for each triangular facet. For the 237th triangular facet, two 3×3 matrices are constructed. The first matrix consists of the homogeneous form of the original texture coordinates of the three vertices, and the second matrix consists of the three-dimensional space coordinates of the three vertices. A 2×3 affine transformation matrix is ​​solved by the least squares method, such that the original texture coordinates are multiplied by the affine transformation matrix to obtain the three-dimensional space coordinates. The affine transformation matrix is ​​stored as the texture mapping parameters of the triangular facet. The texture mapping parameters contain 6 floating-point numbers, which correspond to the linear part and translation part of the affine transformation, respectively.

[0145] Optionally, the texture mapping parameters of all triangles are organized into a texture mapping parameter list according to the triangle patch index order. The texture mapping parameter list is an array of length 37456, where the array contains the first... The element stores the first... The six floating-point numbers corresponding to each triangular facet are used; the texture mapping parameter list is used as the texture mapping relationship after the clothing is worn. It can be understood that when a triangular facet in the deformed 3D clothing mesh undergoes a large deformation, the nonlinear part of the affine transformation matrix is ​​implicitly expressed by the relative position changes of the three vertices of the triangular facet before and after the deformation, without the need for additional calculation of the strain tensor.

[0146] Alternatively, the texture mapping parameter list is stored in a compressed format. For three consecutive triangular faces that share the same texture mapping parameter, only one copy of the parameter and the repetition count identifier are stored.

[0147] In one embodiment of the present invention, when generating a virtual fitting result image based on texture mapping, a blank target fitting image canvas is first created, the size of which is the same as the size of the human body image. For each triangular facet in the deformed 3D clothing mesh, the following operations are performed: a pixel block representing the human body region covered by the triangular facet is extracted from the human body image; this pixel block is then inversely transformed according to the texture mapping parameters of the triangular facet to generate a clothing texture pixel block; this clothing texture pixel block is then drawn onto the screen projection area corresponding to the triangular facet on the blank target fitting image canvas. After traversing all triangular facets, a weighted average smoothing filter is applied to the pixels at the intersections of different triangular facets on the target fitting image canvas; the smoothed target fitting image canvas is then used as the virtual fitting result image.

[0148] When outputting the virtual try-on result image as try-on feedback corresponding to multimodal query data, the virtual try-on result image is first encoded into a web-compatible image format to generate a try-on feedback data packet containing the virtual try-on result image. This try-on feedback data packet also contains the product link address for each garment in the candidate clothing list. The currently used e-commerce front-end interface instance is located based on a unique session identifier, and the generated try-on feedback data packet is sent to the found e-commerce front-end interface instance. The virtual try-on result image is rendered and displayed on the e-commerce front-end interface instance, and a clickable button for the product link address is rendered and displayed in the area below the virtual try-on result image.

[0149] In a specific implementation, a virtual fitting result image is generated based on the texture mapping relationship, and a blank target fitting image canvas is created. The size of the blank target fitting image canvas is the same as the size of the human body image. For example, if the human body image is 1024 pixels wide and 1536 pixels high, then an RGB three-channel blank image canvas with a width of 1024 pixels and a height of 1536 pixels is created, and the initial RGB value of each pixel is (0,0,0).

[0150] In some embodiments, the following operations are performed on each triangular facet in the deformed 3D clothing mesh: The pixel block of the human body region covered by the triangular facet is extracted from the human body image. The projection area of ​​the 1250th triangular facet in the deformed 3D clothing mesh in screen space is a triangle. The screen coordinates of the three vertices of this triangle are (320, 450), (380, 420), and (350, 490). Using these three vertices as boundaries, the bounding box covered by the triangular facet is calculated from the upper left corner (320, 420) to the lower right corner (380, 490). All pixels within this rectangular region are cropped from the human body image to obtain a human body region pixel block with a width of 60 pixels and a height of 70 pixels. Then, the human body region pixel block is mapped according to the texture parameters of the triangular facet. The data undergoes an inverse transformation to generate clothing texture pixel blocks. The texture mapping parameter is a 2×3 affine transformation matrix. The inverse matrix of this affine transformation matrix is ​​calculated, and the coordinates of each pixel in the human body region pixel block are multiplied by the inverse matrix to obtain the corresponding clothing texture coordinates. The color value at this coordinate is sampled from the original clothing texture image corresponding to the 3D clothing model before deformation, and these color values ​​are filled into a new pixel block. The size of the new pixel block is the same as that of the human body region pixel block, thus generating the clothing texture pixel block. The clothing texture pixel block is drawn onto the screen projection area corresponding to the triangular facet on the blank target fitting image canvas. Specifically, each pixel in the clothing texture pixel block is written to the corresponding position on the blank target fitting image canvas according to the original screen projection coordinates, overwriting the original pixel value at that position. Optionally, when multiple triangular facets are written to the same screen pixel position, a depth testing mechanism is used to compare the average depth value of the triangular facets corresponding to the written pixel, and the pixel written by the triangular facet with the smaller depth value is retained. The depth value is obtained by interpolation from the Z coordinate of the vertices of the deformed 3D clothing mesh.

[0151] In the specific implementation, after traversing all triangular facets, a weighted average smoothing filter is performed on the pixels at the intersections of different triangular facets on the target fitting image canvas. First, pixels belonging to the projection boundaries of two or more triangular facets are detected among all pixels on the target fitting image canvas. For each boundary pixel, the color values ​​of all non-zero pixels in its 3×3 neighborhood are collected, and a weighted average of these color values ​​is calculated. The weighting coefficient is the reciprocal of the Euclidean distance from the neighboring pixels to the boundary pixel. The original color value of the boundary pixel is replaced with the weighted average color value. The smoothed target fitting image canvas is then used as the virtual fitting result image. It can be understood that the kernel function of the weighted average smoothing filter is a Gaussian kernel, and the standard deviation parameter of the Gaussian kernel is set to 1.2 pixels.

[0152] In some embodiments, the virtual try-on result image is output as the try-on feedback corresponding to the multimodal query data. The virtual try-on result image is encoded into a web-compatible image format, namely JPEG format, with the encoding quality parameter set to 85. The pixel matrix of the virtual try-on result image is converted into a JPEG binary data stream. A try-on feedback data packet containing the virtual try-on result image is generated. The try-on feedback data packet is a JSON structure. The first field of the JSON structure is named "tryon_image", and its value is the string of the JPEG binary data stream after Base64 encoding. The second field of the JSON structure is named "product_list", and its value is an array. Each element in the array is a product object, and each product object contains three attributes: "product_id", "product_name", and "product_url".

[0153] Optionally, the fitting feedback data packet also includes the product link address of each garment in the candidate clothing list. The candidate clothing list contains a total of 8 garments. For each garment, the corresponding Uniform Resource Locator (URL) of the product details page is read from the e-commerce clothing database, and the URL is filled into the "product_url" attribute of the product object.

[0154] In specific implementation, the user's currently used e-commerce front-end interface instance is located based on the unique session identifier, which is "a3f5c2e1b8d4f6a7". The active session record corresponding to this identifier is retrieved from the session management queue maintained by the e-commerce back-end server. The active session record stores the WebSocket connection identifier of the user's currently used e-commerce front-end interface instance. The virtual try-on feedback data packet is sent to the found e-commerce front-end interface instance. Through the long connection established by the WebSocket connection identifier, the JSON structure is sent to the e-commerce front-end interface instance running in the user's browser.

[0155] It is understood that, when the virtual try-on result image is rendered and displayed on the e-commerce front-end interface instance, after receiving the try-on feedback data packet, the e-commerce front-end interface instance decodes the Base64 encoded string in the JSON structure into binary data, creates an Image object and sets its src attribute to the BlobURL of the decoded JPEG data, and adds the Image object to the specified location in the DOM tree. Inside the container, the specified The size of the container is the same as the size of the virtual try-on result image.

[0156] Alternatively, a clickable button displaying the product link address is rendered in the area below the virtual try-on result image. For each product object in the product_list array of the try-on feedback data package, an HTML anchor element is created. Set the href attribute of the anchor element to the product_url value of the product object, set the target attribute of the anchor element to "_blank", add a button element inside the anchor element, set the text content of the button element to the product_name of the product object, add all the generated anchor elements in sequence to the parent container of the virtual try-on result image, arrange them in the lower area of ​​the virtual try-on result image, and separate each anchor element with a line break or horizontal spacing.

[0157] The formula is the calculation expression for weighted average smoothing filtering:

[0158]

[0159] in: Indicates the coordinates after smoothing filtering The pixel color value at that location, Represents the column coordinates of a pixel. Represents the row coordinates of a pixel. This represents the half-width of the smoothing filter kernel, where the half-width is 1. This represents the lateral offset of a neighboring pixel relative to the center pixel. This represents the vertical offset of a neighboring pixel relative to the center pixel. Indicates pixel offset Gaussian weight values ​​at the location, Indicates the coordinates before smoothing filtering The pixel color value at that location.

[0160] The above embodiments are only used to illustrate the technical methods of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of the present invention without departing from the spirit and scope of the technical methods of the present invention.

Claims

1. A multimodal query method for personalized online clothing try-on in e-commerce, characterized in that, include: Obtain multimodal query data input by the user; The multimodal query data includes at least one human image uploaded by the user and at least one piece of natural language text describing the desired clothing; Multimodal feature encoding is performed on the multimodal query data to generate a joint feature vector; Retrieve a list of candidate clothing items from the e-commerce clothing database based on the joint feature vector; Perform 3D virtual reconstruction on each garment in the candidate clothing list to generate a 3D clothing model; Spatially align the 3D clothing model with the human body posture in the human body image; Perform physical deformation simulation of fabric between the spatially aligned 3D clothing model and the human body image; Extract the texture mapping relationship of clothing after wearing from the results of fabric physical deformation simulation; A virtual fitting result image is generated based on the texture mapping relationship; The virtual try-on result image is output as the try-on feedback corresponding to the multimodal query data.

2. The e-commerce personalized try-on method based on multimodal query according to claim 1, characterized in that, The steps to obtain multimodal query data input by the user include: Receive at least one human image uploaded by the user on the e-commerce front-end interface; The at least one human body image includes a frontal human body image, a back human body image, and at least one side human body image; Simultaneously receive at least one natural language text describing the desired clothing input by the user on the same e-commerce front-end interface; The natural language text includes at least one of the following: clothing category description, clothing color description, clothing material description, and clothing pattern description; The received at least one human image and at least one natural language text describing the desired clothing are packaged into a multimodal query data packet according to the chronological order of user input. Assign a unique session identifier to the multimodal query data packet; The unique session identifier is associated with and stored in the multimodal query data packet.

3. The e-commerce personalized try-on method based on multimodal query according to claim 2, characterized in that, The steps of performing multimodal feature encoding on the multimodal query data to generate a joint feature vector specifically include: Separate the at least one human image and the at least one natural language text describing the desired clothing from the multimodal query data; Perform forward computation of a convolutional neural network on the at least one human image to extract multi-layer convolutional feature maps for each human image; Perform word embedding encoding on the at least one natural language text describing the desired clothing, converting each word into a dense word vector; Perform forward computation using a bidirectional long short-term memory network on the dense word vectors to extract text context feature vectors; The multi-layer convolutional feature map is converted into a global feature vector of human image through an adaptive pooling layer; The global feature vector of the human face and the text context feature vector are concatenated along the feature dimension. Perform a multilayer perceptron transformation on the concatenated feature vectors to output the joint feature vector.

4. The e-commerce personalized try-on method based on multimodal query according to claim 3, characterized in that, The steps of retrieving a candidate clothing list from an e-commerce clothing database based on the joint feature vector specifically include: Obtain the apparel feature vector of each garment from the pre-stored e-commerce apparel database; The clothing feature vector and the joint feature vector are in the same feature space dimension; Calculate the cosine similarity between the joint feature vector and the garment feature vector of each garment; Sort all clothing items in descending order of cosine similarity; Select the top preset number of clothing items from the sorting results as the initial candidate clothing set; The initial candidate clothing set is filtered by clothing category to remove clothing whose categories do not match the clothing category description in the natural language text. The filtered clothing items are reordered according to cosine similarity to generate the candidate clothing list.

5. The e-commerce personalized try-on method based on multimodal query according to claim 4, characterized in that, The steps of performing 3D virtual reconstruction on each garment in the candidate garment list to generate a 3D garment model specifically include: Retrieve the currently pending clothing identifier from the candidate clothing list; Based on the clothing identifier, the multi-view display image set and size specification parameters of the clothing are read from the e-commerce clothing database; A multi-view stereo matching algorithm is executed on the multi-view display image set to calculate the three-dimensional point cloud of the clothing surface; Perform Poisson surface reconstruction on the three-dimensional point cloud to generate an initial triangular mesh model; The initial triangular mesh model is non-uniformly scaled according to the size specification parameters so that the girth of the scaled triangular mesh model is consistent with the size specification parameters. The scaled triangular mesh model is used as the 3D clothing model.

6. The e-commerce personalized try-on method based on multimodal query according to claim 5, characterized in that, The step of spatially aligning the 3D clothing model with the human pose in the human body image specifically includes: Perform human key point detection on the human image and output a set of coordinates of key points of the two-dimensional human skeleton; Estimate the three-dimensional posture parameters of the human body based on the set of coordinates of key points of the two-dimensional human skeleton. The human body three-dimensional posture parameters include the rotation angle and three-dimensional spatial position of each joint point; A parameterized human body mesh model is constructed based on the aforementioned three-dimensional human body posture parameters; Calculate the scaling ratio between the bounding box of the 3D clothing model and the bounding box of the parametric human body mesh model; The vertex coordinates of the 3D clothing model are uniformly scaled according to the scaling ratio. The spatial alignment is achieved by projecting each vertex of the scaled 3D clothing model to the nearest neighbor point of the parameterized human body mesh model.

7. The e-commerce personalized try-on method based on multimodal query according to claim 6, characterized in that, The steps for simulating the physical deformation of fabric between the spatially aligned 3D clothing model and the human body image specifically include: Initialize all mesh vertices of the spatially aligned 3D clothing model as point masses; By setting spring constraints between adjacent mass points, a mass-spring system is formed. The spring constraints include three categories: structural springs, shear springs, and bending springs. Collision detection boundary conditions are set based on the surface normals of the parameterized human body mesh model; A gravitational force and a frictional force between the clothing and the human body are applied to the mass spring system. The equations of motion of the point-spring system are solved iteratively using the implicit Euler integral method. After each iteration of the solution is completed, the positions of all particles are projected outside the collision detection boundary conditions. The iteration stops when the velocity change of all particles is lower than the preset kinetic energy convergence threshold, and the deformed 3D clothing mesh is output.

8. The e-commerce personalized try-on method based on multimodal query according to claim 7, characterized in that, The steps for extracting the texture mapping relationship of clothing after wearing from the results of fabric physical deformation simulation specifically include: Obtain the three-dimensional spatial coordinates of the three vertices of each triangular facet in the deformed three-dimensional clothing mesh; Obtain the original texture coordinates of the three vertices of the same triangular facet in the 3D clothing model before deformation; Calculate the affine transformation matrix from the original texture coordinate space to the three-dimensional space coordinates for each triangular facet; The affine transformation matrix is ​​stored as the texture mapping parameters of the triangular facet. All texture mapping parameters of the triangles are organized into a texture mapping parameter list according to the triangle patch index order; The texture mapping parameter list is used as the texture mapping relationship after the clothing is worn.

9. The e-commerce personalized try-on method based on multimodal query according to claim 8, characterized in that, The step of generating a virtual fitting result image based on the texture mapping relationship specifically includes: Create a blank canvas for the target fitting image; The size of the blank target fitting image canvas is the same as the size of the human body image; Perform the following operations on each triangular facet in the deformed 3D clothing mesh: Extract the pixel block of the human body region covered by the triangular facet from the human body image; The human body region pixel blocks are inversely transformed according to the texture mapping parameters of the triangular facet to generate clothing texture pixel blocks; The clothing texture pixel block is drawn onto the screen projection area corresponding to the triangular facet on the blank target fitting image canvas; After traversing all the triangular facets, a weighted average smoothing filter is performed on the pixels at the intersection of different triangular facets on the target fitting image canvas. The smoothed and filtered target fitting image canvas is used as the virtual fitting result image; The step of outputting the virtual try-on result image as the try-on feedback corresponding to the multimodal query data specifically includes: Encode the virtual try-on result image into a web-compatible image format; Generate a fitting feedback data packet containing the virtual fitting result image; The fitting feedback data package also includes the product link address of each garment in the candidate clothing list; The user's current e-commerce front-end interface instance is located based on the unique session identifier. Send the fitting feedback data packet to the found e-commerce front-end interface instance; The virtual try-on result image is rendered and displayed on the e-commerce front-end interface instance; A clickable button displaying the product link address is rendered in the area below the virtual try-on result image.

10. A multimodal query-based personalized virtual try-on system for e-commerce, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the e-commerce personalized try-on method for multimodal query as described in any one of claims 1 to 9.