Model training data construction method and apparatus
By acquiring images from a virtual item display platform and using a multimodal question response model for image matching and filtering, the problem of slow training data construction in existing technologies is solved, enabling rapid and high-quality training dataset construction and virtual wearable item image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2025-01-14
- Publication Date
- 2026-07-14
AI Technical Summary
Existing virtual try-on technology requires a lot of manpower and financial resources to build training data, resulting in slow training sample generation and an inability to quickly build high-quality training datasets.
By acquiring original item images and images of objects wearing them from a virtual item display platform, key point information of the objects and item categories are extracted. A multimodal question response model is then used for image matching and filtering to construct a high-quality training dataset.
It enables the rapid and accurate construction of training datasets for virtual wearable items, improves the quality of training data, and generates high-quality images of virtual wearable items.
Smart Images

Figure CN122391773A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method and apparatus for constructing model training data. Background Technology
[0002] The main goal of virtual try-on technology is to generate an image of a model wearing the target clothing, ensuring that the fine details of the clothing are preserved and that it blends seamlessly with the surrounding environment. In existing technological solutions, the training data required for virtual try-on models is typically constructed by manually taking flat lay images and model photos of the same garment to build paired data, or by acquiring a large number of product and model images from the internet and then manually selecting flat lay images and model photos belonging to the same garment to form a matching pair.
[0003] In existing technical solutions, whether it is to take photos of clothing and models wearing them or to crawl a large number of images from the Internet and then manually select them to build matching data, a lot of manpower and financial resources are required. It is impossible to quickly build a large number of training samples, which restricts the training of virtual fitting models. Summary of the Invention
[0004] This application provides a method and apparatus for constructing model training data, which can quickly and accurately construct a batch training dataset of virtual wearable items, thereby training an item wearing model based on the training dataset; and automatically generating images of a target object trying on virtual wearable items based on the item wearing model.
[0005] On the one hand, this application provides a method for constructing model training data, the method comprising:
[0006] Obtain the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item.
[0007] Extract the object key point information corresponding to each object wearing image in the object wearing image set, as well as the category of the wearing item corresponding to each object wearing image;
[0008] Determine the matching result between the object key point information corresponding to each object's wearing image and the category of the worn item;
[0009] The object wearing images in the object wearing image set whose matching results represent successful matching results are determined as the initial screening wearing images;
[0010] Images that meet preset conditions are selected from the initial screened images of the wearable objects to obtain target wearable images;
[0011] Based on the original set of object images and the target wearing image, a training dataset for the virtual wearable object is constructed; the training dataset is used to train the object wearing model.
[0012] On the other hand, a model training data construction apparatus is provided, the apparatus comprising:
[0013] The image set acquisition module is used to acquire the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item.
[0014] The wearable item category determination module is used to extract the object key point information corresponding to each object wearable image in the object wearable image set, as well as the wearable item category corresponding to each object wearable image;
[0015] The matching result determination module is used to determine the matching result between the object key point information corresponding to each object's wearing image and the wearing item category;
[0016] The initial screening wearable image determination module is used to determine the wearable images of objects in the object wearable image set whose matching results represent successful matching results as initial screening wearable images;
[0017] The target wearable image determination module is used to select images that meet preset conditions from the initial screened object wearable images to obtain the target wearable image;
[0018] The training data construction module is used to construct a training dataset for the virtual wearable item based on the original item image set and the target wearing image; the training dataset is used to train the item wearing model.
[0019] In one exemplary embodiment, the image set acquisition module includes:
[0020] A comprehensive image set acquisition unit is used to acquire a comprehensive image set corresponding to the virtual wearable item in the virtual item display platform; the comprehensive image set includes the original item image set and the object wearing image set.
[0021] The image recognition result prediction unit is used to input the comprehensive image set and the first question text into the multimodal question response model, perform response result prediction processing, and obtain the image recognition result of each image in the comprehensive image set; the first question text is used to describe whether the virtual wearable item in each image of the comprehensive image set is worn on the preset object.
[0022] The image set segmentation unit is used to segment the comprehensive image set into the original item image set and the object wearing image set based on the image recognition results of each image in the comprehensive image set.
[0023] In one exemplary embodiment, the apparatus further includes:
[0024] The sample image acquisition module is used to acquire a comprehensive image set of samples corresponding to multiple sample virtual wearable items; the comprehensive image set of samples includes the original sample item image of each sample virtual wearable item and the sample object wearing image obtained by the sample object wearing each sample virtual wearable item;
[0025] The sample result determination module is used to input each sample image and the corresponding sample question text in the comprehensive sample image set into a multimodal large model for response result prediction processing, so as to obtain the sample response result corresponding to each sample image; the sample question text includes the first question text.
[0026] The correction module is used to correct the erroneous results in the sample response results based on the manual inspection results of each sample response result, and to fine-tune the multimodal large model based on the correction results to obtain the multimodal problem response model.
[0027] In one exemplary embodiment, the correction module includes:
[0028] An accuracy determination unit is used to obtain the manual inspection results of each sample response and determine the accuracy of the multimodal large model based on the manual inspection results.
[0029] The sample result correction unit is used to, when the accuracy is less than a preset threshold, take the sample response results that are characterized as correct results by the manual inspection as the first sample result, and correct the sample response results that are characterized as incorrect results by the manual inspection to obtain the second sample result.
[0030] The sample image acquisition unit is used to acquire a first sample image corresponding to the first sample result and a second sample image corresponding to the second sample result in the sample composite image set;
[0031] A training set construction unit is used to construct a training set based on the first sample image and the second sample image;
[0032] The model fine-tuning unit is used to input the training set into the multimodal large model, and to fine-tune the multimodal large model based on the first sample result label annotated by the first sample image and the first sample result label annotated by the first sample image to obtain a multimodal question response model.
[0033] In one exemplary embodiment, the model fine-tuning unit includes:
[0034] The sub-unit is used to fine-tune the multimodal large model and use the model at the end of training as the current fine-tuning model;
[0035] The verification image set acquisition subunit is used to acquire the current verification image set, input the current verification image set and the sample question text into the current fine-tuning model for response result prediction processing, and obtain the current sample prediction result.
[0036] The current accuracy determination subunit is used to determine the current accuracy of the current fine-tuning model based on the results of manual inspection of the current sample prediction results.
[0037] The fine-tuning training subunit is used to correct erroneous results in the current sample prediction results when the current accuracy is less than the preset threshold, and to fine-tune the current fine-tuning model based on the current correction results.
[0038] The model determination subunit is used to determine the current fine-tuned model as the multimodal problem response model when the current accuracy is greater than or equal to the preset threshold.
[0039] In an exemplary embodiment, the sample question text includes a second question text, which describes whether the virtual wearable item in each original item image in the original item image set is complete; the training data construction module includes:
[0040] The response prediction unit is used to input the original item image set and the second question text into the multimodal question response model, perform response result prediction processing, and obtain the response result corresponding to each original item image in the original item image set;
[0041] An image filtering unit is used to filter out complete original images of the virtual wearable item from the original item image set, thereby obtaining a filtered item image set;
[0042] The target item image determination unit is used to determine the target item image based on the filtered item image set;
[0043] The training dataset construction unit is used to construct a training dataset for the virtual wearable item based on the target item image and the target wearing image.
[0044] In an exemplary embodiment, the sample question text includes a third question text, which is used to confirm the wearable item category corresponding to each filtered item image in the filtered item image set. The target item image determination unit includes:
[0045] The response prediction subunit is used to input the set of filtered item images and the third question text into the multimodal question response model, perform response result prediction processing, and obtain the wearable item category of each filtered item image in the set of filtered item images;
[0046] The labeling subunit is used to label each filtered item image in the filtered item image set with the wearable item category label of each filtered item image, so as to obtain the target item image.
[0047] In one exemplary embodiment, the wearable item category determination module includes:
[0048] The item category acquisition unit is used to acquire the item category corresponding to each object's item image.
[0049] The object wearable image feature extraction unit is used to input the object wearable image into the comprehensive image extraction model for each object wearable image in the object wearable image set, and extract the image features in the object wearable image based on the image feature extraction network of the comprehensive image extraction model to obtain the object wearable image features;
[0050] The key point heatmap output unit is used to detect object key points in the object wearable image features based on the key point detection network of the comprehensive image extraction model and output an object key point heatmap, wherein the object key point heatmap represents the object key point information.
[0051] The key point correlation graph output unit is used to extract the relative positional relationship between key points of the object in the wearable image features based on the correlation graph generation network of the comprehensive image extraction model, and output the key point correlation graph according to the relative positional relationship.
[0052] In one exemplary embodiment, the matching result determination module includes:
[0053] The location information determination unit is used to extract the location information of each object key point based on the object key point heatmap of each object key point for each object wearing image in the object wearing image set.
[0054] The pose determination unit is used to determine the pose of the preset object in the object wearing image based on the key point correlation diagram and the position information of each object key point;
[0055] The matching result determination unit is used to determine the matching result between the pose of the preset object in the object wearing image and the category of the worn item.
[0056] In an exemplary embodiment, the sample question text further includes a fourth question text, which describes whether the preset object in the initial screening object's wearing image is a frontal image. The target wearing image determination module includes:
[0057] A first-result determination unit is used to input the wear image of the initially screened object and the text of the fourth question into the multimodal question response model, perform response result prediction processing, and obtain the first-response result corresponding to the wear image of the initially screened object;
[0058] The initial screening unit is used to identify the image of the initial screening object, which is characterized by a frontal image of the preset object, as a screened image based on the first response result, and determine it as an initial screening image that meets the preset conditions.
[0059] A target image determination unit is used to determine the target wearable image based on the initial screened image;
[0060] In an exemplary embodiment, the sample question text further includes a fifth question text, which describes whether the preset object in the initial screening image is in a standing posture. The target image determination unit includes:
[0061] The secondary result determination subunit is used to input the initial screening image and the fifth question text into the multimodal question response model, perform response result prediction processing, and obtain the secondary response result corresponding to the initial screening image;
[0062] The target screening subunit is used to identify the initial screening image, which represents the preset object as standing, as the target wearing image by the secondary response result.
[0063] In one exemplary embodiment, both the target item image and the target clothing image are at least two, and the training dataset construction unit includes:
[0064] An image feature extraction subunit is used to extract the first image features of the virtual wearable item in each target item image and to extract the second image features of the virtual wearable item in each target wearable image;
[0065] The training pair construction subunit is used to calculate the similarity between each first image feature and each second image feature, and to construct training image pairs based on the similarity calculation results; the training image pairs include training item images and training object wearing images; the similarity between the first image feature corresponding to the training item image and the second image feature corresponding to the training object wearing image satisfies the target condition;
[0066] The dataset construction subunit is used to construct the training dataset of the virtual wearable item based on the training image pairs.
[0067] In one exemplary embodiment, the dataset construction subunit includes:
[0068] The training label annotation subunit is used to annotate the training item images in the training image pairs with training object wearing image labels;
[0069] The training image extraction subunit is used to extract the object image of the preset object from the training object's wear image to obtain the training object image;
[0070] The training data construction subunit is used to construct the training dataset of the virtual wearable item based on the training item image, the training object image, and the training object wearing image label;
[0071] In one exemplary embodiment, the apparatus further includes:
[0072] The wearing result determination module is used to input the training item image and the training object image into a preset wearing model to predict the item wearing, and to obtain the predicted object wearing result after the preset object wears the virtual wearable item corresponding to the training item image;
[0073] The item wearing model determination module is used to train the preset wearing model based on the difference between the predicted object wearing result and the training object wearing image label to obtain the item wearing model.
[0074] In one exemplary embodiment, the apparatus further includes:
[0075] The test image acquisition module is used to acquire images of the test wearable items and the test object;
[0076] The test subject wearing image determination module is used to input the test wearable item image and the test subject image into the item wearing model to perform item wearing prediction, and obtain the test subject wearing image after the test subject wears the test wearable item corresponding to the test wearable item image.
[0077] The first result determination module is used to determine whether there is a defect in the wearable item in the test object's wear image based on the comparison result between the test object's wear image and the test wearable item image, and to obtain a first defect detection result;
[0078] The second result determination module is used to detect whether there are defects in the limbs of the test object in the wearing image based on the object detection algorithm, and to obtain a second defect detection result.
[0079] The model correction module is used to correct the item wearing model based on the first defect detection result and the second defect detection result to obtain a corrected item wearing model.
[0080] In an exemplary embodiment, the first result determining module includes:
[0081] A real-time image acquisition unit is used to extract a real-time image of the wearable item corresponding to the test wearable item from the image of the test object wearing the item;
[0082] The wearable item feature extraction unit is used to extract a first wearable item feature from the real-time wearable item image and a second wearable item feature from the test wearable item image; the first wearable item feature includes at least one of color feature and texture feature, and the second wearable item feature is a feature of the same type as the first wearable item feature;
[0083] The first defect determination unit is used to determine that there is a defect in the wearable item in the test object's wear image when the similarity between the first wearable item feature and the second wearable item feature is less than a preset similarity threshold.
[0084] In one exemplary embodiment, the apparatus further includes:
[0085] A mask image extraction module is used to extract a first mask image of the original wearable item area and a second mask image of the wearable item from the test object image;
[0086] The transformation parameter determination module is used to determine the transformation parameters for transforming the first mask image into the second mask image;
[0087] The image fusion module is used to fuse the image of the test wearable item with the image of the test object based on the transformation parameters, so as to obtain the shape image of the wearable item after the test object wears the test wearable item;
[0088] The edge information determination module is used to calculate the first edge information of the shape image of the wearable item and the second edge information of the wearable item in the image of the test object based on the edge detection algorithm.
[0089] The defect determination module is used to determine that there is a defect in the wearable item in the test object's wear image if the difference between the first edge information and the second edge information is greater than a preset difference threshold.
[0090] In one exemplary embodiment, the model correction module includes:
[0091] The test subject wearing image determination unit is used to acquire a first test subject wearing image in which the wearable item has a defect in the first defect detection result and a second test subject wearing image in which the limbs of the test subject have a defect in the second defect detection result.
[0092] The first corrected sample construction unit is used to construct a first corrected sample based on the test wearable item image corresponding to the test object's wearable image and the test object image; the first corrected sample is labeled with the first test object's wearable image tag.
[0093] The second corrected sample construction unit is used to construct a second corrected sample based on the test wear item image corresponding to the test object wear image and the test object image; the second corrected sample is labeled with the second test object wear image tag.
[0094] The correction training unit is used to construct correction sample data based on the first correction sample and the second correction sample, and to perform correction training on the item wearing model based on the correction sample data to obtain the corrected item wearing model.
[0095] On the other hand, an electronic device is provided, the device including a processor and a memory, the memory storing at least one instruction or at least one program, the at least one instruction or the at least one program being loaded and executed by the processor to implement the model training data construction method as described above.
[0096] On the other hand, a computer storage medium is provided that stores at least one instruction or at least one program, which is loaded and executed by a processor to implement the model training data construction method as described above.
[0097] On the other hand, a computer program product is provided, including a computer program that is loaded and executed by a processor to implement the model training data construction method as described above.
[0098] The model training data construction method and apparatus provided in this application have the following technical effects:
[0099] This application obtains the original item image set and the object wearing image set corresponding to virtual wearable items in a virtual item display platform; the original item image set is the image set of the virtual wearable items, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable items; thereby extracting the original item image corresponding to the same virtual wearable item and the image of the preset object wearing the item; then extracting the object key point information corresponding to each object wearing image in the object wearing image set, and the wearable item category corresponding to each object wearing image; determining the matching result between the object key point information corresponding to each object wearing image and the wearable item category; and identifying the object wearing images in the object wearing image set whose matching results represent successful matching results as the initial screening wearable images; thereby achieving... The process involves rapidly and accurately extracting initial images of virtual wearable items from a set of object-wearing images; then selecting images that meet preset conditions from these initial images to obtain target wearable images; further filtering from the initial images to obtain target wearable images improves their image quality; and constructing a training dataset for the virtual wearable items based on the original image set and the target wearable images. This allows for the rapid and accurate construction of batch training datasets for virtual wearable items while ensuring image quality. The training dataset is used to train an item-wearing model, which can then be trained to generate a model that automatically generates high-quality images of the target object trying on the virtual wearable items. Attached Figure Description
[0100] To more clearly illustrate the technical solutions and advantages in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0101] Figure 1 This is a schematic diagram of the application environment of a model training data construction method provided in the embodiments of this specification;
[0102] Figure 2 This is a flowchart illustrating a method for constructing model training data provided in an embodiment of this specification;
[0103] Figure 3 This is a flowchart illustrating a training method for a multimodal problem response model provided in the embodiments of this specification;
[0104] Figure 4This is a flowchart illustrating a method for fine-tuning and training the aforementioned multimodal large model based on the correction results to obtain a multimodal problem response model, as provided in the embodiments of this specification.
[0105] Figure 5 This is a flowchart illustrating a method for obtaining the original item image set and the object wearing image set corresponding to virtual wearable items in a virtual item display platform, as provided in the embodiments of this specification.
[0106] Figure 6 This is a flowchart illustrating a method for constructing a training dataset for a virtual wearable item based on the original item image set and the target wearable image provided in this specification.
[0107] Figure 7 This is a flowchart illustrating a method for constructing a training dataset for a virtual wearable item based on the target item image and the target wearable image provided in the embodiments of this specification.
[0108] Figure 8 This is a schematic flowchart of a method for detecting defects in worn items in a test object's wearing image, provided in an embodiment of this specification.
[0109] Figure 9 This is a schematic diagram of the structure of a model training data construction device provided in the embodiments of this specification;
[0110] Figure 10 This is a schematic diagram of the structure of a server provided in the embodiments of this specification. Detailed Implementation
[0111] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0112] It is understood that in the specific implementation of this application, user information is involved, such as image sets and other related data obtained by a preset object wearing the virtual wearable item. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0113] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or server that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.
[0114] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0115] Please see Figure 1 , Figure 1 This is a schematic diagram of the application environment of a model training data construction method provided in an embodiment of this application. The application environment may include at least a server 100 and a terminal 200.
[0116] In an optional embodiment, server 100 can be used to construct the training dataset of the virtual wearable item. Server 100 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.
[0117] In an optional embodiment, terminal 200 can be used to display the original set of object images and the set of images of the object wearing the virtual wearable item. Specifically, terminal 200 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, smart wearable devices, in-vehicle terminals, and smart TVs; it can also be software running on the aforementioned electronic devices, such as applications and mini-programs. The operating system running on the electronic device in this embodiment can include, but is not limited to, Android, iOS, Linux, and Windows.
[0118] The following describes a method for constructing model training data according to this application. Figure 2 This is a flowchart illustrating a method for constructing model training data according to an embodiment of this specification. This specification provides the operational steps described in the embodiments or flowchart, but based on conventional or non-inventive methods, more or fewer operational steps may be included. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual system or server products, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as shown in the flowchart... Figure 2 As shown, the above method may include:
[0119] S201: Obtain the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item.
[0120] In the embodiments of this specification, the virtual item display platform may include, but is not limited to, live streaming platforms, shopping platforms, etc.; virtual wearable items may include, but are not limited to, clothing, jewelry, shoes, hats, and other items for wearing. Preset objects may include, but are not limited to, users (e.g., mannequins), animals, etc.; when the virtual wearable item is a user item, the preset object is the user; when the virtual wearable item is an animal item, the preset object is the animal. The virtual item display platform may include link information for multiple virtual wearable items, and the original image of each virtual wearable item can be obtained based on the link information to construct its corresponding original item image set. Since each virtual wearable item has images of a preset object trying it on, an image set of the preset object wearing the virtual wearable item can be obtained for each virtual wearable item. The preset objects corresponding to different virtual wearable items may be the same or different. An original item image set corresponding to a virtual wearable item may include one or more original item images, and an object wearing image set corresponding to a virtual wearable item may include one or more object wearing images. For example, if the virtual wearable item is clothing, the original item image set is a collection of images of clothing samples, and the object wearing image set is images of a mannequin trying on the clothing.
[0121] S203: Extract the object key point information corresponding to each object wearing image in the above object wearing image set, as well as the category of the wearing item corresponding to each object wearing image.
[0122] In the embodiments of this specification, after obtaining the aforementioned set of object-wearing images, object key point information of a preset object can be extracted from each object-wearing image. This object key point information can be information on at least one limb key point of the preset object. The category of the wearable item corresponding to each object-wearing image can be used to determine the wearable limb area of the preset object, thereby determining whether the wearable item in the object-wearing image is completely worn on the preset object based on the matching relationship between the wearable limb area and the object key point information. When the virtual wearable item is clothing, the wearable item category can include, but is not limited to, tops, bottoms, dresses, etc.
[0123] S205: Determine the matching result between the key point information of the object corresponding to each of the above object wearing images and the above wearing item category.
[0124] In the embodiments of this specification, for each object wearing image, its corresponding object key point information can be matched with the aforementioned wearable item category; complete object key point information corresponding to the wearable item category can be obtained, and it can be determined whether the aforementioned object key point information includes complete object key point information. If it does, it is determined that the aforementioned object key point information matches the aforementioned wearable item category; if the aforementioned object key point information does not include complete object key point information, it indicates that the aforementioned object key point information does not match the aforementioned wearable item category, and it can be determined that the object wearing image is an incomplete image. For example, when the virtual wearable item is clothing, if the wearable item category is a top, all its corresponding object key information can be obtained, including key point information of the head and the complete upper body; then, the detected object key point information is matched with all object key information to obtain a matching result.
[0125] S207: The object wearing images that represent successful matching results in the above object wearing image set are determined as the initial screening wearing images.
[0126] In the embodiments of this specification, the object wearing image representing the successful matching result in the object wearing image set can be determined as the image of complete wearing and used as the initial screening wearing image, thereby ensuring the integrity of the initial screening wearing image.
[0127] S209: Select images that meet the preset conditions from the above preliminary screening of the wearable images of the target objects to obtain the target wearable images.
[0128] In the embodiments of this specification, images that meet preset conditions can be further filtered from the initially screened images of the wearer to obtain the target wearer image. The preset conditions can be set according to actual conditions, and for example, the preset conditions may include, but are not limited to, frontal images, standing posture images, etc.
[0129] S2011: Based on the above original object image set and the above target wearing image, construct the training dataset of the above virtual wearable object; the above training dataset is used to train the object wearing model.
[0130] In the embodiments of this specification, a training dataset for virtual wearable items can be constructed based on the filtered target wearable images and the original item image set; this training dataset can be used to train the item wearing model.
[0131] In some embodiments, a multimodal problem response model can be obtained by training a pre-trained multimodal large model, wherein, for example, Figure 3 As shown, Figure 3 A flowchart illustrating a training method for a multimodal question response model is shown, including:
[0132] S301: Obtain a comprehensive image set of samples corresponding to multiple sample virtual wearable items; the comprehensive image set of samples includes the original sample item image of each sample virtual wearable item and the sample object wearing image obtained by the sample object wearing each sample virtual wearable item;
[0133] S303: Input each sample image and the corresponding sample question text in the above-mentioned sample composite image set into a multimodal large model for response result prediction processing to obtain the sample response result corresponding to each sample image; the above-mentioned sample question text includes the above-mentioned first question text.
[0134] S305: Correct the erroneous results in the above sample response results based on the manual inspection results of each sample response result, and fine-tune the above multimodal large model based on the correction results to obtain the multimodal problem response model.
[0135] In the embodiments of this specification, the aforementioned multimodal large model is obtained by pre-training the large model with multimodal training data input; the multimodal data can include, but is not limited to, data of various modalities such as images and text. The sample object and the preset object can be objects of the same type; the sample virtual wearable items can be items from one or more different virtual item display platforms; and the number of sample virtual wearable items can be set according to actual needs. Then, each sample image in the comprehensive sample image set and the sample question text corresponding to each sample image are input into the multimodal large model for response result prediction processing to obtain the sample response result corresponding to each sample image; then, the sample response result corresponding to each sample image is manually checked, and the erroneous results are filtered out; then, based on the manual check results of each sample response result, the erroneous results in the above sample response results are corrected, and the multimodal large model is fine-tuned and trained based on the correction results to obtain the multimodal question response model, thereby improving the accuracy of the multimodal question response model. Among them, the sample question text includes the aforementioned first question text, which is used to describe whether the virtual wearable item in each image of the comprehensive image set is worn on the aforementioned preset object.
[0136] In some embodiments, the sample question text may further include a second question text, a third question text, a fourth question text, and a fifth question text. The second question text is used to describe whether the virtual wearable item in each original item image in the original item image set is complete. The third question text is used to confirm the wearable item category corresponding to each filtered item image in the filtered item image set. The fourth question text is used to describe whether the preset object in the initial screened object wearable image is a frontal image. The fifth question text is used to describe whether the preset object in the initial screened image is in a standing posture.
[0137] For example, a labeling model can be used to randomly select 1000 samples for labeling, and then these 1000 responses can be manually checked and corrected for errors. The manually corrected samples can then be used to fine-tune the model to improve its accuracy. Specifically, this can be broken down into the following steps:
[0138] 1. Collect samples: First, you need to prepare 1000 samples and let the model generate the corresponding answers.
[0139] 2. Manual review: The answers generated by the model are manually reviewed to identify and correct any errors.
[0140] 3. Constructing the training set: Combine the original samples and the corrected answers to form a new training set. This can include the original question, the model's answer, and the manually corrected answer.
[0141] 4. Fine-tuning the model: Use this new training set to fine-tune the model. The fine-tuning process helps the model learn more accurate ways to answer, thereby improving its accuracy in responding to similar questions.
[0142] In some embodiments, such as Figure 4 As shown, Figure 4 The flowchart outlines a method for correcting errors in the responses of each sample based on manual review, and then fine-tuning the multimodal large-scale model using the corrected results to obtain a multimodal problem response model. The method includes:
[0143] S3051: Obtain the manual inspection results of the response results for each sample, and determine the accuracy of the above multimodal large model based on the above manual inspection results;
[0144] S3053: When the above accuracy rate is less than the preset threshold, the sample response results that are characterized as correct results by the above manual inspection are taken as the first sample result, and the sample response results that are characterized as incorrect results by the above manual inspection are corrected to obtain the second sample result.
[0145] S3055: Obtain the first sample image corresponding to the first sample result in the above sample comprehensive image set, and the second sample image corresponding to the second sample result;
[0146] S3057: Construct a training set based on the first sample image and the second sample image mentioned above;
[0147] S3059: Input the above training set into the above multimodal large model, and fine-tune the above multimodal large model based on the first sample result label annotated by the first sample image and the first sample result label annotated by the first sample image to obtain a multimodal problem response model.
[0148] In this embodiment, the prediction accuracy of the multimodal large model can be statistically analyzed based on the manual inspection results of each sample response. If the model accuracy is low, sample result labels for the composite image can be constructed based on the manual correction results, thereby reconstructing the training set and fine-tuning the multimodal large model to obtain a multimodal question response model. Specifically, a preset loss function can be constructed, and target loss data can be determined based on the difference between the first prediction result and the first sample result label of the first sample image, and the difference between the second prediction result and the second sample result label of the second sample image. Then, the model parameters of the multimodal large model are adjusted based on the target loss data until the training termination condition is met. The training termination condition can be determined based on at least one of the target loss data and the number of training iterations. During model training, corresponding images can be selected from the composite image set for training based on different question texts. This embodiment can fine-tune the multimodal large model based on the aforementioned manual inspection results, thereby improving the prediction accuracy of the multimodal large model.
[0149] In some embodiments, the above-mentioned fine-tuning training of the multimodal large model to obtain a multimodal problem response model includes:
[0150] The above multimodal large model is fine-tuned and trained, and the model at the end of training is used as the current fine-tuning model;
[0151] Obtain the current verification image set, input the current verification image set and the sample question text into the current fine-tuning model to perform response prediction processing, and obtain the current sample prediction result;
[0152] Based on the manual inspection results of the current sample predictions, the current accuracy of the current fine-tuning model is determined.
[0153] If the current accuracy is less than the preset threshold, the erroneous results in the current sample prediction results are corrected, and the current fine-tuning model is fine-tuned and trained based on the current correction results.
[0154] If the current accuracy is greater than or equal to the preset threshold, the current fine-tuned model is determined as the multimodal problem response model.
[0155] In the embodiments of this specification, the current verification image set can be combined with sample images or other combined images. After a fine-tuning training is completed, manual inspection can continue, and the current accuracy of the current fine-tuning model can be determined based on the manual inspection results of the current sample prediction results; and it can be determined whether to continue fine-tuning training based on the current accuracy, thereby ensuring that the multimodal problem response model obtained at the end of training has a high accuracy.
[0156] In this embodiment, the process of correcting erroneous results in the current sample prediction when the current accuracy is less than the preset threshold, and then fine-tuning the current fine-tuned model based on the corrected results, may include: correcting erroneous results in the current sample prediction when the current accuracy is less than the preset threshold; fine-tuning the current fine-tuned model based on the corrected results, using the fine-tuned model as the current fine-tuned model again; and then proceeding to the steps of obtaining the current verification image set, inputting the current verification image set and the sample question text into the current fine-tuned model for response result prediction processing, and obtaining the current sample prediction result, thereby entering a repetitive training mode. This embodiment allows for multiple fine-tuning training sessions until the model's accuracy reaches the preset threshold.
[0157] For example, the model fine-tuning training process may include the following steps:
[0158] (1) Data preparation
[0159] Before fine-tuning, the dataset needs to be prepared. Assume the training set is constructed as follows:
[0160]
[0161] Where x_i is the input (in this embodiment, it is the sample integrated image set and the sample problem that needs to be labeled by the model), and y_i is the target output (the corrected answer).
[0162] (2) Introduction of LoRA adapter
[0163] In the process of fine-tuning large multimodal models, LoRA (Low-Rank Adaptation) can be introduced. In LoRA, a portion of the model's weight matrix W is decomposed into low-rank matrices. Specifically, the weight matrix W is decomposed into two low-rank matrices A and B:
[0164] W′=W+ΔW=W+AB
[0165] in:
[0166] W′ is the weight adjusted by the LoRA adapter; A is a d×r matrix (low-rank matrix), where d is the dimension of the original weight matrix and r is the rank of the low-rank matrix (usually r is much smaller than d); B is an r×d matrix.
[0167] (3) Define the loss function
[0168] The goal of fine-tuning is to minimize the model's loss on the training set. A commonly used loss function is cross-entropy loss, which is defined as:
[0169]
[0170] in:
[0171] θ is the parameter of the model.
[0172] M is the number of output categories (e.g., the size of the vocabulary).
[0173] yij is the one-hot encoding of the target output (yij = 1 if yi is of class j, otherwise 0).
[0174] pij is the probability that the model predicts input xi as class j.
[0175] 3. Backpropagation
[0176] After calculating the target loss data, the backpropagation algorithm is used to update the parameters A and B of the LoRA adapter. Since W is fixed, only A and B are updated. The core of backpropagation is calculating the gradient of the loss function with respect to A and B:
[0177]
[0178] 4. Update parameters
[0179] Optimization algorithms, such as the Adaptive Moment Estimation (Adam) algorithm, are used to update the parameters A and B of the LoRA adapter. Adam is a first-order optimization algorithm that can replace the traditional stochastic gradient descent process; it iteratively updates the neural network weights based on the training data.
[0180] Taking Adam as an example, the formula for parameter update is:
[0181]
[0182] in:
[0183] α is the learning rate.
[0184] mt is the first moment estimate of the gradient.
[0185] vt is the second moment estimate of the gradient.
[0186] ∈ is a small constant used to prevent division by zero errors.
[0187] Through the above process, a fine-tuned labeling model can be obtained. Then, 1000 samples are randomly selected again using the fine-tuned model for labeling, and the accuracy is manually checked. If the accuracy is still low, the above process can be repeated to fine-tune the model again until the accuracy of the model meets the requirements.
[0188] In some embodiments, such as Figure 5 As shown, Figure 5 A flowchart illustrating a method for obtaining the original set of item images and the set of images of objects wearing virtual items in a virtual item display platform is shown, including:
[0189] S20101: Obtain the comprehensive image set corresponding to the virtual wearable item in the virtual item display platform; the comprehensive image set includes the original item image set and the object wearing image set.
[0190] S20103: Input the above-mentioned comprehensive image set and the first question text into the multimodal question response model, perform response result prediction processing, and obtain the image recognition result of each image in the above-mentioned comprehensive image set; the above-mentioned first question text is used to describe whether the virtual wearable item in each image of the above-mentioned comprehensive image set is worn on the above-mentioned preset object.
[0191] S20105: Based on the image recognition results of each image in the above-mentioned comprehensive image set, the above-mentioned comprehensive image set is divided into the above-mentioned original item image set and the above-mentioned object wearing image set.
[0192] In the embodiments of this specification, a comprehensive image set corresponding to the virtual wearable item can be obtained by triggering the link information of the virtual wearable item in the virtual item display platform. This comprehensive image set includes the original item image set and the object wearing image set corresponding to the virtual wearable item. For example, a large number of clothing-related images can be obtained from the Internet (such as e-commerce platforms). To facilitate the subsequent construction of matching data, all images under each link can be saved in a separate folder. Then, a first question text is constructed, which describes whether the virtual wearable item in each image of the comprehensive image set is worn by the preset object. For example, the virtual wearable item is clothing, and the first question text can be "Is the clothing in the picture worn by the model?". Then, the comprehensive image set and the first question text are input into a multimodal question response model for response result prediction processing. Images in which the image recognition result indicates that the virtual wearable item is worn by the preset object are identified as object wearing images, thus obtaining the object wearing image set. Images in which the image recognition result indicates that the virtual wearable item is not worn by the preset object are identified as original item images, thus obtaining the original item image set. Thus, the image recognition result of each image in the comprehensive image set is obtained. This allows for the differentiation between the original item images and the object wearing images corresponding to the virtual wearable items in the comprehensive image set. This achieves the rapid and accurate division of the comprehensive image set into the original item image set and the object wearing image set through the multimodal question response model.
[0193] In some embodiments, the sample question text includes a second question text, which describes whether the virtual wearable item in each original item image in the original item image set is complete; such as Figure 6 As shown, Figure 6 This paper demonstrates a training dataset for constructing the aforementioned virtual wearable item based on the original item image set and the target wearable image set, including:
[0194] S20111: Input the above-mentioned original item image set and the above-mentioned second question text into the above-mentioned multimodal question response model, perform response result prediction processing, and obtain the response result corresponding to each original item image in the above-mentioned original item image set;
[0195] S20113: Select complete original images of the virtual wearable items that are the responses from the above original image set to obtain the filtered image set;
[0196] S20115: Based on the above-mentioned filtered item image set, determine the target item image;
[0197] S20117: Based on the above target item image and the above target wearing image, construct the training dataset of the above virtual wearable item.
[0198] In this embodiment, a second question text can be constructed, and the original item image set determined in the previous step of the model, along with the second question text, can be input into the multimodal question response model for response result prediction processing. This allows for the selection of complete images of virtual wearable items from the original item image set. Then, based on the selected item image set, target item images that meet the training requirements are determined. Based on the selected target item images and the target wearable images, a training dataset for the virtual wearable items is constructed. For example, when the virtual wearable item is virtual clothing, the second question text describes whether the virtual clothing in each original image in the original item image set is in a completely flat state; for example, the second question text could be "Is the clothing shown in the picture completely flat?". This allows for the rapid selection of complete images of virtual wearable items through the multimodal question response model. This embodiment can select complete images of virtual wearable items from the original item image set and then determine the target item images, thereby ensuring the image quality of the target item images in the training dataset.
[0199] In some embodiments, the sample question text includes a third question text, which is used to confirm the wearable item category corresponding to each filtered item image in the filtered item image set, and to determine the target item image based on the filtered item image set, including:
[0200] The above-mentioned set of filtered item images and the above-mentioned third question text are input into the above-mentioned multimodal question response model to perform response result prediction processing, and the wearable item category of each filtered item image in the above-mentioned set of filtered item images is obtained.
[0201] Each filtered item image in the above filtered item image set is labeled with the wearable item category label to obtain the target item image.
[0202] In this embodiment, determining the target item image from the selected item image set can also be done through a multimodal question-response model. A third question text corresponding to the selection criteria is constructed, which can be used to confirm the wearable item category corresponding to each selected item image in the selected item image set. Then, the selected item image set and the third question text are input into the multimodal question-response model for response result prediction processing to obtain the wearable item category of each selected item image in the selected item image set. Based on the category recognition result, a wearable item category label is assigned to each selected item image, thus using the selected item images labeled with the wearable item category label as target item images. When there are multiple target item images, a target item image set can be constructed. This embodiment can quickly and accurately predict the wearable item category of selected item images using a multimodal question-response model.
[0203] For example, when the virtual wearable item is virtual clothing, the third question text is used to determine the wearable item category corresponding to the filtered item image of the virtual clothing; for example, the wearable item category may include, but is not limited to, tops, dresses, and bottoms; for example, the third question text may be "Is the clothing in the picture a top, bottoms, or a dress?"; the model output may be one of these, thereby determining the wearable item category label of the filtered item image.
[0204] In some embodiments, extracting object key point information corresponding to each object wearing image in the above object wearing image set, and the category of wearable items corresponding to each object wearing image, includes:
[0205] Obtain the category of the worn items corresponding to the image of each object wearing the clothing;
[0206] For each object wearing image in the above object wearing image set, the object wearing image is input into the comprehensive image extraction model, and the image feature extraction network of the comprehensive image extraction model is used to extract the image features in the object wearing image to obtain the object wearing image features;
[0207] Based on the above-mentioned comprehensive image extraction model, the key point detection network detects the object key points in the above-mentioned object wear image features and outputs the object key point heatmap. The above-mentioned object key point heatmap represents the above-mentioned object key point information.
[0208] The correlation graph generation network based on the above-mentioned comprehensive image extraction model extracts the relative positional relationships between key points of the objects in the wearable image features, and outputs a key point correlation graph based on the above relative positional relationships.
[0209] In the embodiments of this specification, the comprehensive image extraction model may include an image feature extraction network, a key point detection network, and a correlation graph generation network; when the preset object is a user, the comprehensive image extraction model may include, but is not limited to, the OpenPose model. OpenPose is a deep learning-based human pose estimation library that can accurately detect and estimate human key points and pose information from images or videos.
[0210] The image feature extraction network can be a convolutional neural network, such as VGG or ResNet. The VGG (Visual Geometry Group) model is one of the deep learning models proposed by the Visual Geometry Group at Oxford University. The ResNet model, or Residual Neural Network, effectively alleviates the vanishing and exploding gradient problems in deep network training by introducing residual blocks and skip connections. A residual block consists of one or more convolutional layers and a skip connection. Skip connections allow the input signal to be directly passed to the output of the convolutional layer and added to the processed signal, so that the network learns the difference between the input and output, rather than the entire mapping function. This image feature extraction network is used to extract features from the input image and convert the original object-wearing image into a high-dimensional feature map.
[0211] A keypoint detection network can include a series of convolutional layers to predict a heatmap of each keypoint in the features of an object's clothing image; each heatmap corresponds to a specific keypoint, and each pixel value in the heatmap represents the probability that the location is that keypoint.
[0212] Part Affinity Fields (PAFs) generative networks can be used to generate a set of affinity maps based on the features of an object's clothing image, representing the connections between different keypoints. These maps help the model understand the spatial relationships between keypoints, thereby improving pose estimation.
[0213] Therefore, the integrated image extraction model can be used to output a keypoint heatmap and a keypoint correlation map based on an input image of an object wearing clothing. The heatmap for each keypoint is typically a two-dimensional array representing the probability that each pixel is that keypoint. The keypoint correlation map represents the connections between keypoints and is usually a two-dimensional vector field indicating the relative positions between each pair of keypoints. Thus, the limb pose of a preset object can be accurately predicted based on the keypoint heatmap and the keypoint correlation map.
[0214] In some embodiments, determining the matching result between the object key point information corresponding to each object wearing image and the wearable item category includes:
[0215] For each object wearing image in the above object wearing image set, the location information of each object key point is extracted based on the above object key point heatmap of each object key point;
[0216] Based on the above key point correlation diagram and the position information of each object's key points, the pose of the above preset object in the above object wearing image is determined.
[0217] Determine the matching result between the pose of the preset object in the image of the object wearing the above and the category of the worn items.
[0218] In the embodiments of this specification, for each object wearing image in the aforementioned object wearing image set, the position information of each object key point can be extracted based on the aforementioned object key point heatmap. For example, the specific position of each key point can be determined by finding the maximum value in the heatmap. Then, based on the aforementioned key point correlation graph and the position information of each object key point, the posture of the aforementioned preset object in the aforementioned object wearing image is determined. That is, using the correlation graph, the key points with detected position information are connected to form a complete object posture (e.g., human posture). Finally, the posture of the preset object is matched with the aforementioned wearing item category to obtain a matching result. For example, if the wearing item category is a top, then the posture of the preset object should be a complete upper body posture, and the two match. If the posture of the preset object is an incomplete upper body posture, then the matching result is determined to be a mismatch. The object wearing images in the aforementioned object wearing image set that represent successful matching results are then determined as preliminary screening wearing images. Images that meet the preset conditions are selected from the aforementioned preliminary screening object wearing images to obtain target wearing images. This embodiment improves the accuracy of matching the pose of the preset object with the above-mentioned categories of wearable items by accurately predicting the pose of the preset object, thereby further improving the quality of the target wearable image.
[0219] In some embodiments, the sample question text further includes a fourth question text, which describes whether the preset object in the initial screening object wearing image is a frontal image. The process of selecting images that meet preset conditions from the initial screening object wearing images to obtain the target wearing image includes:
[0220] The images of the objects being screened and the text of the fourth question are input into the multimodal question response model to perform response result prediction processing, thereby obtaining a response result corresponding to the images of the objects being screened.
[0221] The image of the object being worn by the initial screening object, which is characterized by the above-mentioned response result as a frontal image of the preset object, is determined as the initial screening image that meets the above-mentioned preset conditions.
[0222] Based on the initial screening images described above, the target wearable image is determined.
[0223] In the embodiments of this specification, the fourth question text can be "Is the person in the picture facing the camera?". The initial screening object's wearing image and the fourth question text can be input into the aforementioned multimodal question response model for further image filtering. Based on the response results output by the multimodal question response model, it can be determined whether the initial screening object's wearing image is a frontal image, and initial screening images that are frontal images can be selected from the initial screening object's wearing images, and the target wearing image can be further determined. This ensures that the target wearing images are all frontal wearing images of the preset object, further improving the quality of the target wearing images.
[0224] In some embodiments, the sample question text further includes a fifth question text, which describes whether the preset object in the initial screening image is in a standing posture. Based on the initial screening image, determining the target wearing image includes:
[0225] The initial screened image and the text of the fifth question are input into the multimodal question response model to perform response result prediction processing, and the secondary response result corresponding to the initial screened image is obtained.
[0226] The above-mentioned secondary response results characterize the above-mentioned preset object as the initial screening image in a standing posture, and determine it as the above-mentioned target wearing image.
[0227] In the embodiments of this specification, the fifth question text can be "Is the person in the picture standing?". There can be multiple initial screening images. Each initial screening image and the aforementioned fifth question text can be input into the aforementioned multimodal question response model to determine whether each initial screening image is in a standing posture. This allows for further filtering of images that maintain a standing posture from the initial screening images. When the worn item is clothing, images in a standing posture can more completely display the full appearance of the clothing.
[0228] In some embodiments, there are at least two images of the target item and at least two images of the target clothing, such as... Figure 7 As shown, Figure 7 A flowchart illustrating a method for constructing a training dataset for the aforementioned virtual wearable item based on the aforementioned target item image and the aforementioned target wearable image is shown, including:
[0229] S701: Extract the first image features of the virtual wearable item in each of the above target item images, and extract the second image features of the virtual wearable item in each of the above target wearable images;
[0230] S703: Calculate the similarity between each first image feature and each second image feature, and construct training image pairs based on the similarity calculation results; the training image pairs include training item images and training object wearing images; the similarity between the first image features corresponding to the training item images and the second image features corresponding to the training object wearing images satisfies the target condition;
[0231] S705: Based on the above training image pairs, construct the training dataset for the above virtual wearable items.
[0232] In the embodiments of this specification, since the same virtual wearable item may include multiple colors, styles, etc., one virtual wearable item may correspond to multiple target item images and multiple target wear images. First image features of the virtual wearable item in each of the aforementioned target item images and second image features of the virtual wearable item in each of the aforementioned target wear images are extracted. Then, the similarity between the two features is calculated, and a set of first image features and second image features whose similarity satisfies the target condition is determined. The image corresponding to the first image feature is determined as the training item image, and the image corresponding to the second image feature is determined as the training object wear image, thereby constructing a training image pair. Specifically, for any first target image feature of the virtual wearable item in any target item image, the similarity between the first target image feature and each second image feature can be calculated, and the second image feature with the highest similarity to the first target image feature among all second image features is determined as the second target image feature that satisfies the target condition. Thus, the first target image feature and the second target image feature are used as a training image pair, achieving rapid and accurate construction of a training dataset for virtual wearable items. Both the first and second image features can be derived from pre-trained models (such as the CLIP model). The CLIP model (Contrastive Language-Image Pre-Training) is a multimodal pre-trained neural network model designed to learn the alignment relationships between images and text through training on a large amount of image and text pairing data. For example, a model pre-trained on a large amount of data can be used to extract features from model images and clothing images in the same product catalog, calculate product similarity between each pair, and retain image pairs with high similarity as a matching pair of clothing flat lay image and model upper body image.
[0233] In some embodiments, constructing the training dataset for the virtual wearable item based on the training image pairs includes:
[0234] Label the training object images with training object wearing images in the above training image pairs;
[0235] Extract the object image of the preset object from the above training object wearing image to obtain the training object image;
[0236] Based on the above-mentioned training item images, training object images, and training object wearing image labels, the above-mentioned training dataset for the virtual wearable items is constructed.
[0237] In the embodiments of this specification, the training object image in the same image pair can be labeled with the corresponding training object wearing image based on the training object image and the training object wearing image in the training image pair; at the same time, the object image of the preset object can be extracted from the training object wearing image. For example, if the training object wearing image is a model wearing image, the model image can be extracted from the model wearing image, that is, the wearing items in the model wearing image can be set to transparent; thus obtaining the training object image; then the training object image, the training object image and the training object wearing image label are used as training data, thereby realizing the intelligent and batch construction of the above-mentioned training dataset of virtual wearable items.
[0238] In some embodiments, the above method further includes:
[0239] The training item image and the training object image are input into a preset wearable model to predict the item wearing, and the predicted object wearing result is obtained after the preset object wears the virtual wearable item corresponding to the training item image.
[0240] Based on the difference between the predicted object's wearing results and the training object's wearing image labels, the preset wearing model is trained to obtain the item wearing model.
[0241] In this embodiment, after constructing the training dataset, the training item image and the training object image from the same image pair can be used to train a preset wearing model to obtain the predicted object wearing result after the preset object wears the virtual wearing item corresponding to the training item image; then, based on the difference between the predicted object wearing result and the label of the training object wearing image, the target loss information is determined; the model parameters of the preset wearing model are adjusted according to the target loss information until the training termination condition is met; and the preset wearing model at the end of training is determined as the item wearing model. The training termination condition can be determined based on the target loss information and the number of training iterations. This embodiment can quickly train a highly accurate item wearing model based on an intelligently constructed training dataset.
[0242] In some embodiments, the above method further includes:
[0243] Acquire images of the item to be worn and the target object;
[0244] The above-mentioned image of the item to be worn and the above-mentioned image of the target object are input into the above-mentioned item wearing model to predict the item wearing, and the target object wearing image is obtained after the target object corresponding to the above-mentioned image of the target object wears the item corresponding to the above-mentioned image of the item to be worn.
[0245] The item wearing model obtained in this embodiment can be used to automatically generate a target object wearing image of the target object trying on the item to be worn, based on the input target object image and the image of the item to be worn, thereby improving the generation efficiency of the object wearing image.
[0246] In some embodiments, in order to improve the accuracy of the item wearing model, the item wearing model can be tested based on test data, and the model can be corrected based on the test results, so as to improve the quality and accuracy of the model's predicted images.
[0247] In some embodiments, the above method further includes:
[0248] Acquire images of the test wearable items and the test object;
[0249] The above-mentioned test wearable item image and test object image are input into the above-mentioned item wearable model to predict item wearability, and the test object wearable image after the test object image is wearing the test wearable item corresponding to the above-mentioned test wearable item image is obtained.
[0250] Based on the comparison results between the above-mentioned test subject wearing image and the above-mentioned test wearable item image, it is determined whether there is a defect in the wearable item in the above-mentioned test subject wearing image, and the first defect detection result is obtained;
[0251] Based on the object detection algorithm, the second defect detection result is obtained by detecting whether there are defects in the limbs of the test object in the above test object wearing image;
[0252] Based on the results of the first and second defect detections, the above-mentioned item wearing model is modified to obtain the modified item wearing model.
[0253] In the embodiments of this specification, a test dataset can be constructed. Each set of test data may include images of test wearable items and test objects. The test object, the preset object, and the target object can all be objects of the same type. The test wearable item images and the test object images can be input into the aforementioned item wearing model to obtain a test object wearing image after the test object wears the test wearable items. The image corresponding to the test object is the test object image, and the test wearable item image can be an image of the test wearable item. After predicting the test object wearing image, it can be checked for defects. Defects in the test object wearing image can be determined through manual inspection; alternatively, test pairs labeled with the aforementioned test wearable item images can be obtained. The model first identifies the labels on the test subject's clothing images and then compares the differences between the test subject's clothing images and their labels to determine if the predicted test subject clothing images have defects. If the difference between the test subject's clothing images and their labels is less than a preset difference value, the test subject's clothing images are determined to have defects. If the difference between the test subject's clothing images and their labels is greater than or equal to a preset difference value, the test subject's clothing images are determined to have no defects. Then, based on the defects of each test clothing image in the test dataset, the model's prediction accuracy is determined. If the prediction accuracy is less than a preset value, the above clothing model is corrected to obtain a corrected clothing model.
[0254] In the embodiments of this specification, based on the comparison results between the above-mentioned test object wearing image and the above-mentioned test wearing item image, it can be determined whether there is a defect in the wearing item in the above-mentioned test object wearing image, and a first defect detection result can be obtained; based on the object detection algorithm, it can be detected whether there is a defect in the object limb in the above-mentioned test object wearing image, and a second defect detection result can be obtained; thereby realizing the determination of whether there is a defect in the test object wearing image from two dimensions: wearing item and test object limb, and improving the accuracy of defect determination.
[0255] For example, when the test item is clothing and the test subject is a model, the defects in the generated image of the test subject wearing the clothing can be divided into three main categories: clothing color or texture defects, clothing shape defects, and model limb defects. Different defect detection algorithms can be used to detect these three different types of defects; the first defect detection result can include the detection results of clothing color or texture defects and clothing shape defects, while the second defect detection result can include the detection results of model limb defects. Based on at least one of the first and second defect detection results, the clothing model is modified to obtain a modified clothing model; thereby improving the image quality and accuracy of the predicted image from the modified clothing model.
[0256] In some embodiments, determining whether the wearable item in the test object's image has a defect based on the comparison result between the test object's wearing image and the test wearable item image includes:
[0257] Extract the real-time image of the wearable item corresponding to the test wearable item from the above test object wearable image;
[0258] Extract the first wearable item features from the real-time wearable item image and the second wearable item features from the test wearable item image; the first wearable item features include at least one of color features and texture features, and the second wearable item features are of the same type as the first wearable item features.
[0259] If the similarity between the first wearable item feature and the second wearable item feature is less than a preset similarity threshold, it is determined that the wearable item in the test subject's wear image has a defect.
[0260] In the embodiments of this specification, image features of the worn items in the test object's wearing image and the test worn item image can be extracted respectively. These features may include at least one of color features and texture features, thereby obtaining first worn item features and second worn item features. Then, based on the feature similarity between the two, it is determined whether the worn item in the test object's wearing image has a defect. If the similarity between the two is less than a preset similarity threshold, it is determined that the worn item in the test object's wearing image has a defect. The image feature extraction and similarity comparison can both be implemented through a model.
[0261] When the test item is clothing, defects in the generated image of the test subject wearing the clothing can be categorized into defects in clothing color or texture, defects in clothing design, etc. To address the potential issue of subtle differences between the clothing in the changed-up image and the provided clothing image, such as inconsistencies in patterns or colors, a parsing model can be used to first extract the clothing from the changed-up image and then compare it with the original clothing image. Figure 1 The images are fed into a feature extraction network to extract their respective features, and then the similarity between the two features is calculated. If the clothes used in the image differ significantly from the provided clothing image in texture or color, the feature similarity will be low, and the generated image is considered to have a defect.
[0262] In some embodiments, the above method further includes:
[0263] Extract the first mask image of the original wearable item area and the second mask image of the wearable item from the above test object image;
[0264] Determine the transformation parameters for transforming the first mask image into the second mask image;
[0265] Based on the above transformation parameters, the above test wearable item image and the above test object image are fused to obtain the wearable item shape image after the above test object wears the above test wearable item;
[0266] The first edge information of the shape image of the wearable item and the second edge information of the wearable item in the image of the test subject are calculated based on the edge detection algorithm.
[0267] If the difference between the first edge information and the second edge information is greater than a preset difference threshold, it is determined that the wearable items in the test object's wear image have defects.
[0268] In this embodiment, edge information of the worn items in a test subject's wearing image can be detected to determine whether defects exist. First, a first mask image of the original worn item region and a second mask image of the test worn item are extracted from the test subject image. Then, transformation parameters for transforming the first mask image into the second mask image are calculated according to a preset algorithm. The preset algorithm may include, but is not limited to, a thin plate spline (TPS) transformation algorithm (hereinafter referred to as the TPS transformation algorithm). Then, based on the transformation parameters, a shape image of the worn item after the test subject wears the test worn item is generated. Next, the difference between the first edge information of the worn item shape image and the second edge information of the worn item in the test subject's wearing image is compared using an edge detection algorithm. If the difference is significant, it is determined that the worn item in the test subject's wearing image has a defect. The edge detection algorithm may include, but is not limited to, the Canny algorithm. The Canny algorithm is one of the commonly used edge detection algorithms in computer vision, and its main characteristic is that it can accurately detect edges in an image while having good resistance to noise. This embodiment can further detect test subject wearing images with item shape defects through an edge detection algorithm, thereby improving the comprehensiveness of defect image detection.
[0269] When the test item is clothing and the test subject is a model, the defects in the generated image of the test subject wearing the clothing may include defects in the clothing design. For these defects, edge detection algorithms can be used for defect detection. Based on the difference between the first edge information and the aforementioned second edge information, it can be determined whether a clothing design defect exists. Figure 8 As shown, Figure 8This is a flowchart of a method for detecting defects in clothing items in a test object's wearing image. First, using the TPS transform algorithm, the parameters required for the TPS transform are calculated by calculating the first mask image (mask M) of the clothing area worn by the original model and the first mask image (mask c) of the clothing to be changed, and the clothing is transformed to the shape when worn by the model. Then, the edge information of the transformed clothing and the clothing on the model in the changing result is calculated by the edge detection algorithm and compared. If the edge difference between the two is too large, the clothing pattern of the generated result is considered to have defects. Among them, TPS Warp (thin plate spline interpolation) is an effective method for processing small-range deformation in two-dimensional or three-dimensional space, especially suitable for problems such as image registration, geometric modeling and data interpolation. The basic idea of TPS is to define the transformation by a set of control points (landmarks). Given a set of source points and target points, TPS calculates the transformation by minimizing the energy function, so that the transformed points are as close as possible to the target points, while maintaining the smoothness of the deformation. The energy function of TPS consists of two parts: (1) data item: ensure that the transformed points are as close as possible to the target points. (2) Smoothing term: Ensure that the transformation is smooth and avoid excessive distortion.
[0270] In some embodiments, the modification of the item wearing model based on the first defect detection result and the second defect detection result to obtain a modified item wearing model includes:
[0271] Acquire images of the first test subject whose wearable items are defective in the first defect detection results and images of the second test subject whose limbs are defective in the second defect detection results.
[0272] A first corrected sample is constructed based on the test wear item image corresponding to the test object image and the test object image; the first corrected sample is labeled with the first test object wear image.
[0273] A second corrected sample is constructed based on the test wear item image corresponding to the test object image and the test object image; the second corrected sample is labeled with the tags of the second test object image.
[0274] Based on the first and second corrected samples, corrected sample data is constructed, and the item wearing model is trained and corrected based on the corrected sample data to obtain the corrected item wearing model.
[0275] In the embodiments of this specification, test images containing defects in the two defect detection results can be corrected to construct corrected samples. The first defect detection result can include defect results determined based on color features, texture features, and pattern. Then, corrected sample data is obtained based on the corrected samples constructed from the two defects. Finally, the above-mentioned item-wearing model is trained using the corrected sample data. Specifically, the test-wearing item image and the test object image from the corrected sample data can be input into the item-wearing model for corrected training to obtain the predicted object-wearing image result after the test object wears the test-wearing item. Then, based on the difference between the predicted object-wearing image result and the corrected object-wearing image label, a target loss value is determined. The model parameters of the item-wearing model are then adjusted based on the target loss value until the training termination condition is met, and the item-wearing model at the end of training is determined as the corrected item-wearing model. The corrected object-wearing image label includes the first test object-wearing image label and the first test object-wearing image label, thereby reducing the probability of defects in the predicted image result of the corrected item-wearing model and improving the quality of the predicted image.
[0276] In some embodiments, the above method further includes:
[0277] Acquire images of the item to be worn and the target object;
[0278] The above-mentioned image of the item to be worn and the above-mentioned image of the target object are input into the above-mentioned modified item wearing model to predict the item wearing, so as to obtain the target object wearing image after the target object corresponding to the above-mentioned image of the target object wears the item to be worn corresponding to the above-mentioned image of the item to be worn.
[0279] In the embodiments of this specification, by modifying the item wearing model to predict item wearing, the image quality and prediction accuracy of the target object wearing image obtained by the target object trying on the item to be worn can be improved.
[0280] As can be seen from the technical solutions provided in the embodiments of this specification above, the embodiments of this specification obtain the original item image set and the object wearing image set corresponding to the virtual wearable items in the virtual item display platform; the original item image set is the image set of the virtual wearable items, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable items; thereby extracting the original item image corresponding to the same virtual wearable item and the image of the preset object wearing the item; then extracting the object key point information corresponding to each object wearing image in the object wearing image set, and the wearable item category corresponding to each object wearing image; determining the matching result between the object key point information corresponding to each object wearing image and the wearable item category; and representing the object wearing image with the matching result in the object wearing image set as a successfully matched result, thus determining... The process involves initial screening of virtual wearable images to quickly and accurately extract images of objects fully wearing virtual items from a set of object wearable images. Images meeting preset conditions are then selected from these initial screening images to obtain target wearable images. This further refines the selection of target wearable images by selecting images that meet preset conditions, improving the image quality of the target wearable images. Based on the original item image set and the target wearable images, a training dataset for the virtual wearable items is constructed. This allows for the rapid and accurate construction of batch training datasets for virtual wearable items while ensuring image quality. The training dataset is used to train an item wearable model, enabling the model to be trained and automatically generated using this model. High-quality images of the target object trying on the virtual wearable items can then be automatically generated using this model.
[0281] This specification also provides an apparatus for constructing model training data, such as... Figure 9 As shown, the above-mentioned device includes:
[0282] The image set acquisition module 910 is used to acquire the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item.
[0283] The item category determination module 920 is used to extract the key point information of each object wearing image in the above object wearing image set, as well as the item category of each object wearing image.
[0284] The matching result determination module 930 is used to determine the matching result between the key point information of the object corresponding to each of the above object wearing images and the above wearing item category;
[0285] The initial screening wearable image determination module 940 is used to determine the wearable image representing the successful matching result in the above-mentioned object wearable image set as the initial screening wearable image.
[0286] The target wearable image determination module 950 is used to select images that meet preset conditions from the above-mentioned preliminary screened object wearable images to obtain the target wearable image;
[0287] The training data construction module 960 is used to construct the training dataset of the virtual wearable item based on the original item image set and the target wearing image; the training dataset is used to train the item wearing model.
[0288] In some embodiments, the image set acquisition module includes:
[0289] The comprehensive image set acquisition unit is used to acquire the comprehensive image set corresponding to the virtual wearable item in the virtual item display platform; the comprehensive image set includes the original item image set and the object wearing image set.
[0290] The image recognition result prediction unit is used to input the above-mentioned comprehensive image set and the first question text into the multimodal question response model, perform response result prediction processing, and obtain the image recognition result of each image in the above-mentioned comprehensive image set; the above-mentioned first question text is used to describe whether the virtual wearable item in each image of the above-mentioned comprehensive image set is worn on the above-mentioned preset object.
[0291] The image set division unit is used to divide the comprehensive image set into the original item image set and the object wearing image set based on the image recognition results of each image in the comprehensive image set.
[0292] In some embodiments, the above-described apparatus further includes:
[0293] The sample image acquisition module is used to acquire a comprehensive image set of samples corresponding to multiple sample virtual wearable items; the comprehensive image set of samples includes the original sample item image of each sample virtual wearable item and the sample object wearing image obtained by the sample object wearing each sample virtual wearable item;
[0294] The sample result determination module is used to input each sample image and the corresponding sample question text from the above-mentioned comprehensive sample image set into a multimodal large model for response result prediction processing, so as to obtain the sample response result corresponding to each sample image; the above-mentioned sample question text includes the above-mentioned first question text.
[0295] The correction module is used to correct errors in the above sample response results based on the manual inspection results of each sample response result, and to fine-tune the above multimodal large model based on the correction results to obtain the multimodal problem response model.
[0296] In some embodiments, the above-mentioned correction module includes:
[0297] The accuracy determination unit is used to obtain the manual inspection results of the response results for each sample, and to determine the accuracy of the multimodal large model based on the above manual inspection results.
[0298] The sample result correction unit is used to, when the accuracy rate is less than a preset threshold, take the sample response results that are characterized as correct results by the manual inspection as the first sample result, and correct the sample response results that are characterized as incorrect results by the manual inspection to obtain the second sample result.
[0299] The sample image acquisition unit is used to acquire the first sample image corresponding to the first sample result in the above-mentioned sample comprehensive image set, and the second sample image corresponding to the second sample result.
[0300] The training set construction unit is used to construct a training set based on the first sample image and the second sample image mentioned above.
[0301] The model fine-tuning unit is used to input the training set into the multimodal large model, and fine-tune the multimodal large model based on the first sample result label annotated by the first sample image and the first sample result label annotated by the first sample image to obtain a multimodal problem response model.
[0302] In some embodiments, the model fine-tuning unit includes:
[0303] The sub-unit is used to fine-tune the above-mentioned multimodal large model and uses the model at the end of training as the current fine-tuning model;
[0304] The verification image set acquisition subunit is used to acquire the current verification image set. The current verification image set and the sample question text are input into the current fine-tuning model to perform response result prediction processing and obtain the current sample prediction result.
[0305] The current accuracy determination subunit is used to determine the current accuracy of the current fine-tuning model based on the manual inspection results of the current sample prediction results.
[0306] The fine-tuning training subunit is used to correct the erroneous results in the current sample prediction results when the current accuracy is less than the preset threshold, and to fine-tune the current fine-tuning model based on the current correction results.
[0307] The model determination sub-unit is used to determine the current fine-tuned model as the multimodal problem response model when the current accuracy is greater than or equal to the preset threshold.
[0308] In some embodiments, the sample question text includes a second question text, which describes whether the virtual wearable item in each original item image in the original item image set is complete; the training data construction module includes:
[0309] The response prediction unit is used to input the above-mentioned original item image set and the above-mentioned second question text into the above-mentioned multimodal question response model, perform response result prediction processing, and obtain the response result corresponding to each original item image in the above-mentioned original item image set;
[0310] The image filtering unit is used to filter out complete original images of the virtual wearable items from the original image set, thereby obtaining a filtered image set.
[0311] The target item image determination unit is used to determine the target item image based on the above-mentioned filtered item image set;
[0312] The training dataset construction unit is used to construct the training dataset of the virtual wearable item based on the target item image and the target wearable image.
[0313] In some embodiments, the sample question text includes a third question text, which is used to confirm the wearable item category corresponding to each filtered item image in the filtered item image set. The target item image determination unit includes:
[0314] The response prediction subunit is used to input the above-mentioned set of filtered item images and the above-mentioned third question text into the above-mentioned multimodal question response model, perform response result prediction processing, and obtain the wearable item category of each filtered item image in the above-mentioned set of filtered item images;
[0315] The labeling subunit is used to label each of the filtered item images in the above-mentioned filtered item image set with the wearable item category label of each of the above-mentioned filtered item images, so as to obtain the above-mentioned target item image.
[0316] In some embodiments, the wearable item category determination module includes:
[0317] The item category acquisition unit is used to acquire the item category corresponding to each object's item image.
[0318] The object wearable image feature extraction unit is used to input the object wearable image into the comprehensive image extraction model for each object wearable image in the above object wearable image set, and extract the image features in the object wearable image based on the image feature extraction network of the comprehensive image extraction model to obtain the object wearable image features;
[0319] The key point heatmap output unit is used to detect object key points in the object wear image features based on the key point detection network of the above-mentioned comprehensive image extraction model and output the object key point heatmap, which represents the object key point information.
[0320] The key point correlation graph output unit is used to extract the relative positional relationship between key points of the object in the above-mentioned object wearing image features based on the correlation graph generation network of the above-mentioned comprehensive image extraction model, and output the key point correlation graph according to the above-mentioned relative positional relationship.
[0321] In some embodiments, the matching result determination module includes:
[0322] The location information determination unit is used to extract the location information of each object key point based on the object key point heatmap of each object key point in the above object wear image set.
[0323] The pose determination unit is used to determine the pose of the preset object in the object wearing image based on the above key point correlation diagram and the position information of each object key point.
[0324] The matching result determination unit is used to determine the matching result between the pose of the preset object and the category of the wearable item in the image of the object being worn.
[0325] In some embodiments, the sample question text further includes a fourth question text, which describes whether the preset object in the initial screening object's wearing image is a frontal image. The target wearing image determination module includes:
[0326] The first-result determination unit is used to input the above-mentioned preliminary screening object wearing image and the above-mentioned fourth question text into the above-mentioned multimodal question response model, perform response result prediction processing, and obtain the first-response result corresponding to the above-mentioned preliminary screening object wearing image;
[0327] The initial screening unit is used to determine the image of the object being worn by the object, which is characterized by the first response result as a frontal image of the preset object, as an initial screening image that meets the preset conditions.
[0328] The target image determination unit is used to determine the target wearable image based on the initial screened image.
[0329] In some embodiments, the sample question text further includes a fifth question text, which describes whether the preset object in the initial screening image is in a standing posture. The target image determination unit includes:
[0330] The secondary result determination subunit is used to input the above-mentioned initial screening image and the above-mentioned fifth question text into the above-mentioned multimodal question response model, perform response result prediction processing, and obtain the secondary response result corresponding to the above-mentioned initial screening image;
[0331] The target screening subunit is used to identify the initial screening image, which represents the preset object as standing, as the target wearing image.
[0332] In some embodiments, there are at least two target object images and at least two target clothing images, and the training dataset construction unit includes:
[0333] The image feature extraction subunit is used to extract the first image features of the virtual wearable item in each of the above target item images, and to extract the second image features of the virtual wearable item in each of the above target wearable images;
[0334] The training pair construction sub-unit is used to calculate the similarity between each first image feature and each second image feature, and to construct training image pairs based on the similarity calculation results; the training image pairs include training item images and training object wearing images; the similarity between the first image features corresponding to the training item images and the second image features corresponding to the training object wearing images satisfies the target condition;
[0335] The dataset construction subunit is used to construct the training dataset for the aforementioned virtual wearable items based on the aforementioned training image pairs.
[0336] In some embodiments, the above-mentioned dataset construction subunit includes:
[0337] The training label annotation subunit is used to annotate the training object wearing image label in the above training image pair;
[0338] The training image extraction subunit is used to extract the object image of the preset object from the training object wearing image to obtain the training object image.
[0339] The training data construction sub-unit is used to construct the training dataset of the virtual wearable item based on the training item image, the training object image, and the training object wearing image label.
[0340] In some embodiments, the above-described apparatus further includes:
[0341] The wearing result determination module is used to input the above-mentioned training item image and the above-mentioned training object image into a preset wearing model to predict the wearing of the item, and to obtain the predicted wearing result of the preset object after wearing the above-mentioned virtual wearable item corresponding to the above-mentioned training item image.
[0342] The item wearing model determination module is used to train the preset wearing model based on the difference between the predicted object wearing result and the training object wearing image label to obtain the item wearing model.
[0343] In some embodiments, the above-described apparatus further includes:
[0344] The test image acquisition module is used to acquire images of the test wearable items and the test object;
[0345] The test object wearing image determination module is used to input the above-mentioned test wearable item image and test object image into the above-mentioned item wearing model to perform item wearing prediction, and obtain the test object wearing image after the test object image wears the test wearable item corresponding to the above-mentioned test wearable item image.
[0346] The first result determination module is used to determine whether there is a defect in the wearable item in the test object's wear image based on the comparison result between the test object's wear image and the test wearable item image, and to obtain the first defect detection result.
[0347] The second result determination module is used to detect whether there are defects in the limbs of the test object in the above-mentioned test object wearing image based on the object detection algorithm, and to obtain the second defect detection result.
[0348] The model correction module is used to correct the above-mentioned item wearing model based on the first defect detection result and the second defect detection result to obtain the corrected item wearing model.
[0349] In some embodiments, the first result determination module includes:
[0350] The real-time image acquisition unit is used to extract the real-time wearable item image corresponding to the test wearable item from the test object wearable image;
[0351] The wearable item feature extraction unit is used to extract the first wearable item feature from the real-time wearable item image and the second wearable item feature from the test wearable item image; the first wearable item feature includes at least one of color feature and texture feature, and the second wearable item feature is a feature of the same type as the first wearable item feature.
[0352] The first defect determination unit is used to determine that there is a defect in the wearable item in the test object's wear image when the similarity between the first wearable item feature and the second wearable item feature is less than a preset similarity threshold.
[0353] In some embodiments, the above-described apparatus further includes:
[0354] The mask image extraction module is used to extract the first mask image of the original wearable item area and the second mask image of the test wearable item from the above-mentioned test object image;
[0355] The transformation parameter determination module is used to determine the transformation parameters for transforming the first mask image into the second mask image.
[0356] The image fusion module is used to fuse the image of the test wearable item with the image of the test object based on the above transformation parameters, so as to obtain the shape image of the wearable item after the test object wears the test wearable item.
[0357] The edge information determination module is used to calculate the first edge information of the shape image of the wearable item and the second edge information of the wearable item in the image of the test object based on the edge detection algorithm.
[0358] The defect determination module is used to determine that there is a defect in the wearable item in the test object's wear image if the difference between the first edge information and the second edge information is greater than a preset difference threshold.
[0359] In some embodiments, the above-mentioned model correction module includes:
[0360] The test subject wearing image determination unit is used to acquire the first test subject wearing image in which the wearing item has a defect in the first defect detection result and the second test subject wearing image in which the limb of the object has a defect in the second defect detection result.
[0361] The first corrected sample construction unit is used to construct a first corrected sample based on the test wearable item image corresponding to the first test object wearable image and the test object image; the first corrected sample is labeled with the first test object wearable image tag.
[0362] The second modified sample construction unit is used to construct a second modified sample based on the test wearable item image corresponding to the test object image and the test object image; the second modified sample is labeled with the tag of the second test object wearable image.
[0363] The correction training unit is used to construct correction sample data based on the first correction sample and the second correction sample, and to perform correction training on the item wearing model based on the correction sample data to obtain the correction item wearing model.
[0364] The apparatus and method embodiments described herein are based on the same inventive concept.
[0365] This specification provides an electronic device including a processor and a memory. The memory stores at least one instruction or at least one program, which is loaded and executed by the processor to implement the model training data construction method provided in the above method embodiments.
[0366] Embodiments of this application also provide a computer storage medium, which can be disposed in a terminal to store at least one instruction or at least one program related to implementing a model training data construction method in the method embodiment. The at least one instruction or at least one program is loaded and executed by the processor to implement the model training data construction method provided in the above method embodiment.
[0367] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the model training data construction method provided in the above-described method embodiments.
[0368] Optionally, in the embodiments of this specification, the storage medium may be located at at least one of the multiple network servers in a computer network. Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0369] The memory described in the embodiments of this specification can be used to store software programs and modules. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory. The memory may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for the functions, etc.; the data storage area may store data created according to the use of the device, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory may also include a memory controller to provide the processor with access to the memory.
[0370] The model training data construction method embodiments provided in this specification can be executed on mobile terminals, computer terminals, servers, or similar computing devices. Taking running on a server as an example, Figure 10 This is a hardware structure block diagram of a server for a model training data construction method provided in the embodiments of this specification. For example... Figure 10 As shown, the server 1000 can vary significantly due to different configurations or performance. It may include one or more Central Processing Units (CPUs) 1010 (CPUs 1010 may include, but are not limited to, microprocessors (MCUs) or programmable logic devices (FPGAs), a memory 1030 for storing data, and one or more storage media 1020 (e.g., one or more mass storage devices) for storing application programs 1023 or data 1022. The memory 1030 and storage media 1020 may be temporary or persistent storage. The program stored in the storage media 1020 may include one or more modules, each module may include a series of instruction operations on the server. Furthermore, the CPU 1010 may be configured to communicate with the storage media 1020 and execute the series of instruction operations in the storage media 1020 on the server 1000. Server 1000 may also include one or more power supplies 1060, one or more wired or wireless network interfaces 1050, one or more input / output interfaces 1040, and / or one or more operating systems 1021, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0371] The input / output interface 1040 can be used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of server 1000. In one example, the input / output interface 1040 includes a network interface controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the input / output interface 1040 may be a radio frequency (RF) module for wireless communication with the Internet.
[0372] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, server 1000 may also include... Figure 10 The more or fewer components shown, or having the same Figure 10 The different configurations shown.
[0373] As can be seen from the embodiments of the model training data construction method, apparatus, device, or storage medium provided in this application, this application obtains the original item image set and the object wearing image set corresponding to virtual wearable items in a virtual item display platform; the original item image set is the image set of the virtual wearable items, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable items; thereby extracting the original item image corresponding to the same virtual wearable item and the image of the preset object wearing the item; then extracting the object key point information corresponding to each object wearing image in the object wearing image set, and the wearable item category corresponding to each object wearing image; determining the matching result between the object key point information corresponding to each object wearing image and the wearable item category; and representing the successfully matched object in the object wearing image set with the matching result. Wearing images are identified as initial screening images; this allows for the rapid and accurate extraction of initial screening images showing the complete virtual wearable item from the set of object wearing images; images meeting preset conditions are then selected from the initial screening object wearing images to obtain target wearing images; this further refines the selection of images meeting preset conditions from the initial screening object wearing images, improving the image quality of the target wearing images; based on the original item image set and the target wearing images, a training dataset for the virtual wearable item is constructed; this enables the rapid and accurate construction of batch training datasets for virtual wearable items while ensuring the image quality of the training dataset; the training dataset is used to train the item wearing model; thus, the item wearing model can be trained based on the training dataset; and high-quality images of the target object trying on the virtual wearable item can be automatically generated using the item wearing model.
[0374] It should be noted that the order of the embodiments described above is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, specific embodiments of this specification have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0375] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments of apparatus, devices, and storage media are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0376] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer storage medium, such as a read-only memory, a disk, or an optical disk.
[0377] The above description is only a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for constructing model training data, characterized in that, The method includes: Obtain the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item. Extract the object key point information corresponding to each object wearing image in the object wearing image set, as well as the category of the wearing item corresponding to each object wearing image; Determine the matching result between the object key point information corresponding to each object's wearing image and the category of the worn item; The object wearing images in the object wearing image set whose matching results represent successful matching results are determined as the initial screening wearing images; Images that meet preset conditions are selected from the initial screened images of the wearable objects to obtain target wearable images; Based on the original set of object images and the target wearing image, a training dataset for the virtual wearable object is constructed; the training dataset is used to train the object wearing model.
2. The method according to claim 1, characterized in that, The acquisition of the original item image set and the object wearing image set corresponding to the virtual wearable items in the virtual item display platform includes: Obtain a comprehensive image set corresponding to the virtual wearable item in the virtual item display platform; the comprehensive image set includes the original item image set and the object wearing image set. The comprehensive image set and the first question text are input into a multimodal question response model for response result prediction processing to obtain the image recognition result of each image in the comprehensive image set; the first question text is used to describe whether the virtual wearable item in each image of the comprehensive image set is worn on the preset object; Based on the image recognition results of each image in the comprehensive image set, the comprehensive image set is divided into the original item image set and the object wearing image set.
3. The method according to claim 2, characterized in that, The training method for the multimodal question response model includes: Obtain a comprehensive image set of samples corresponding to multiple sample virtual wearable items; the comprehensive image set includes the original sample item image of each sample virtual wearable item and the sample object wearing image obtained by the sample object wearing each sample virtual wearable item; Each sample image and its corresponding question text in the comprehensive sample image set are input into a multimodal large model for response prediction to obtain the sample response result for each sample image; the sample question text includes the first question text. The errors in the sample responses are corrected based on the manual inspection results of each sample response, and the multimodal large model is fine-tuned and trained based on the correction results to obtain the multimodal problem response model.
4. The method according to claim 3, characterized in that, The step of correcting erroneous results in the sample responses based on the manual inspection results of each sample response, and fine-tuning the multimodal large model based on the correction results to obtain a multimodal question response model includes: Obtain the manual inspection results of each sample response, and determine the accuracy of the multimodal large model based on the manual inspection results; If the accuracy rate is less than a preset threshold, the sample response results that are characterized as correct results by the manual inspection are taken as the first sample result, and the sample response results that are characterized as incorrect results by the manual inspection are corrected to obtain the second sample result. Obtain the first sample image corresponding to the first sample result and the second sample image corresponding to the second sample result from the sample composite image set; A training set is constructed based on the first sample image and the second sample image; The training set is input into the multimodal large model, and the multimodal large model is fine-tuned based on the first sample result label annotated with the first sample image and the first sample result label annotated with the first sample image to obtain a multimodal question response model.
5. The method according to claim 4, characterized in that, The fine-tuning and training of the multimodal large model to obtain a multimodal question response model includes: The multimodal large model is fine-tuned and trained, and the model at the end of training is used as the current fine-tuned model; Obtain the current verification image set, input the current verification image set and the sample question text into the current fine-tuning model to perform response result prediction processing, and obtain the current sample prediction result; Based on the results of manual inspection of the current sample prediction results, the current accuracy of the current fine-tuning model is determined. If the current accuracy is less than the preset threshold, the erroneous results in the current sample prediction results are corrected, and the current fine-tuning model is fine-tuned and trained based on the current correction results. If the current accuracy is greater than or equal to the preset threshold, the current fine-tuned model is determined as the multimodal problem response model.
6. The method according to claim 3, characterized in that, The sample question text includes a second question text, which describes whether the virtual wearable item in each original item image in the original item image set is complete; the step of constructing a training dataset for the virtual wearable item based on the original item image set and the target wearable image includes: The original set of item images and the second question text are input into the multimodal question response model to perform response result prediction processing, thereby obtaining the response result corresponding to each original item image in the original set of item images; The original item image set is obtained by filtering out complete original item images of the virtual wearable item whose response result is selected from the original item image set; Based on the filtered set of item images, the target item image is determined; A training dataset for the virtual wearable item is constructed based on the target item image and the target wearing image.
7. The method according to claim 6, characterized in that, The sample question text includes a third question text, which is used to confirm the wearable item category corresponding to each filtered item image in the filtered item image set. The step of determining the target item image based on the filtered item image set includes: The set of filtered item images and the text of the third question are input into the multimodal question response model to perform response result prediction processing, thereby obtaining the wearable item category of each filtered item image in the set of filtered item images; Each filtered item image in the filtered item image set is labeled with a wearable item category tag to obtain the target item image.
8. The method according to any one of claims 2-7, characterized in that, The step of extracting the object key point information corresponding to each object wearing image in the object wearing image set, and the wearing item category corresponding to each object wearing image, includes: Obtain the category of the worn items corresponding to the image of each object wearing the clothing; For each object wearing image in the object wearing image set, the object wearing image is input into the comprehensive image extraction model, and the image feature extraction network of the comprehensive image extraction model is used to extract the image features in the object wearing image to obtain the object wearing image features; Based on the key point detection network of the comprehensive image extraction model, the key points of the object in the wearable image feature are detected and the object key point heatmap is output. The object key point heatmap represents the object key point information. The correlation graph generation network based on the comprehensive image extraction model extracts the relative positional relationships between key points of the object in the wearable image features, and outputs a key point correlation graph based on the relative positional relationships.
9. The method according to claim 8, characterized in that, Determining the matching result between the object key point information corresponding to each object's wearing image and the wearing item category includes: For each object wearing image in the object wearing image set, the location information of each object key point is extracted based on the object key point heatmap of each object key point; Based on the key point correlation diagram and the position information of each object's key points, the pose of the preset object in the object wearing image is determined; Determine the matching result between the pose of the preset object in the image of the object being worn and the category of the worn item.
10. The method according to claim 9, characterized in that, The sample question text also includes a fourth question text, which describes whether the preset object in the initial screening object's wearing image is a frontal image. The step of selecting images that meet the preset conditions from the initial screening object's wearing image to obtain the target wearing image includes: The image of the initially screened object wearing the clothing and the text of the fourth question are input into the multimodal question response model to perform response result prediction processing, thereby obtaining a response result corresponding to the image of the initially screened object wearing the clothing; The image of the initially screened object wearing the device, which represents the preset object as a frontal image, is determined as the initial screened image that meets the preset conditions; Based on the initial screened images, the target wearable image is determined; The sample question text also includes a fifth question text, which describes whether the preset object in the initial screening image is in a standing posture. The step of determining the target wearing image based on the initial screening image includes: The initial screened image and the fifth question text are input into the multimodal question response model to perform response result prediction processing, thereby obtaining the secondary response result corresponding to the initial screened image; The secondary response result represents the initial screened image of the preset object as a standing posture, and is determined as the target wearing image.
11. The method according to claim 6, characterized in that, There are at least two target item images and at least two target wearable images. The step of constructing a training dataset for the virtual wearable item based on the target item images and the target wearable images includes: Extract the first image features of the virtual wearable item from each of the target item images, and extract the second image features of the virtual wearable item from each of the target wearable images; Calculate the similarity between each first image feature and each second image feature, and construct training image pairs based on the similarity calculation results; the training image pairs include training item images and training object wearing images; the similarity between the first image feature corresponding to the training item image and the second image feature corresponding to the training object wearing image satisfies the target condition; Based on the training image pairs, a training dataset for the virtual wearable item is constructed.
12. The method according to claim 11, characterized in that, The step of constructing the training dataset for the virtual wearable item based on the training image pairs includes: Label the training item images in the training image pairs with images of the training objects being worn; Extract the object image of the preset object from the training object's wearing image to obtain the training object image; The training dataset for the virtual wearable items is constructed based on the training item images, the training object images, and the training object wearing image labels. The method further includes: The training item image and the training object image are input into a preset wearable model to predict the item wearing, and the predicted object wearing result is obtained after the preset object wears the virtual wearable item corresponding to the training item image. Based on the difference between the predicted object's wearing result and the training object's wearing image label, the preset wearing model is trained to obtain the item wearing model; Acquire images of the test wearable items and the test object; The test wearable item image and the test object image are input into the item wearable model to perform item wearable prediction, and the test object wearable image is obtained after the test object image wears the test wearable item corresponding to the test wearable item image. Based on the comparison results between the test subject's wearing image and the test wearable item image, it is determined whether there are defects in the wearable items in the test subject's wearing image, and a first defect detection result is obtained; Based on the object detection algorithm, the second defect detection result is obtained by detecting whether there are defects in the limbs of the test object in the wearing image. Based on the first defect detection result and the second defect detection result, the item wearing model is corrected to obtain the corrected item wearing model.
13. The method according to claim 12, characterized in that, The step of determining whether the wearable item in the test subject's image has defects based on the comparison result between the test subject's wearing image and the test wearable item image includes: Extract the real-time image of the wearable item corresponding to the test wearable item from the image of the test subject wearing the item; Extract a first wearable item feature from the real-time wearable item image and a second wearable item feature from the test wearable item image; the first wearable item feature includes at least one of color feature and texture feature, and the second wearable item feature is a feature of the same type as the first wearable item feature; If the similarity between the first wearable item feature and the second wearable item feature is less than a preset similarity threshold, it is determined that the wearable item in the test object's wear image has a defect; The method further includes: Extract a first mask image of the original wearable item region and a second mask image of the wearable item from the image of the test object; Determine the transformation parameters for transforming the first mask image into the second mask image; Based on the transformation parameters, the image of the test wearable item and the image of the test object are fused to obtain the shape image of the wearable item after the test object wears the test wearable item; The first edge information of the shape image of the wearable item and the second edge information of the wearable item in the image of the test object are calculated based on the edge detection algorithm. If the difference between the first edge information and the second edge information is greater than a preset difference threshold, it is determined that the wearable items in the test object's wear image have defects.
14. The method according to claim 12, characterized in that, The step of correcting the item wearing model based on the first defect detection result and the second defect detection result to obtain a corrected item wearing model includes: Acquire images of a first test subject whose wearable items are defective according to the first defect detection results, and images of a second test subject whose limbs are defective according to the second defect detection results. A first corrected sample is constructed based on the test wear item image corresponding to the test object's wear image and the test object image; the first corrected sample is labeled with the first test object's wear image tag; A second corrected sample is constructed based on the test wear item image corresponding to the test object's wear image and the test object image; the second corrected sample is labeled with the second test object's wear image tag; Based on the first and second corrected samples, corrected sample data is constructed, and the item wearing model is corrected and trained based on the corrected sample data to obtain the corrected item wearing model.
15. A model training data construction apparatus, characterized in that, The device includes: The image set acquisition module is used to acquire the original item image set and the object wearing image set corresponding to the virtual wearable item in the virtual item display platform; the original item image set is the image set of the virtual wearable item, and the object wearing image set is the image set obtained by a preset object wearing the virtual wearable item. The wearable item category determination module is used to extract the object key point information corresponding to each object wearable image in the object wearable image set, as well as the wearable item category corresponding to each object wearable image; The matching result determination module is used to determine the matching result between the object key point information corresponding to each object's wearing image and the wearing item category; The initial screening wearable image determination module is used to determine the wearable images of objects in the object wearable image set whose matching results represent successful matching results as initial screening wearable images; The target wearable image determination module is used to select images that meet preset conditions from the initial screened object wearable images to obtain the target wearable image; The training data construction module is used to construct a training dataset for the virtual wearable item based on the original item image set and the target wearing image; the training dataset is used to train the item wearing model.