Method and apparatus for clothes image retrieval, device
By acquiring the color tone image of clothing and using a CNN model to extract color tone feature information for clothing matching, the problem of large workload in color calibration in existing technologies is solved, and efficient clothing image matching is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QINGDAO HAIER SMART TECH R & D CO LTD
- Filing Date
- 2020-04-08
- Publication Date
- 2026-06-02
AI Technical Summary
In existing clothing image matching and retrieval systems, color calibration is a labor-intensive process, resulting in high costs and low efficiency in the clothing matching process.
By acquiring the color tone image of the clothing to be retrieved, the activated regions of a set area of the clothing to be retrieved are extracted using a configured CNN model, and clothing matching is performed based on the color tone feature information, thereby reducing the workload of color labeling.
It enables automatic clothing matching based on color tone, reducing the cost of clothing matching, increasing the matching speed, and reducing the workload of calibration.
Smart Images

Figure CN111428738B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of smart device technology, such as methods, apparatus, and devices for retrieving images of clothing. Background Technology
[0002] With the advancement of science and technology and the development of artificial intelligence, smart homes have become closely intertwined with users. To provide users with a more intelligent experience in terms of clothing, many manufacturers have launched smart wardrobes. These wardrobes utilize sensors, remote controls, and various smart devices to achieve user-friendly functions such as automatic sterilization, dehumidification, and lighting. Furthermore, they also feature clothing recognition and recommendation functions. These functions are all based on clothing image matching and retrieval.
[0003] Currently, many clothing image matching and retrieval methods rely on clothing color features. In these Convolutional Neural Networks (CNN) algorithms for color recognition and matching, "feature extraction" is mostly based on supervised learning using color labels. This method requires highly specific labels (red, white, black, blue, yellow, etc.) during training, thus necessitating a large amount of data with corresponding color labels—a significant workload for color labeling—making clothing image matching quite labor-intensive. Summary of the Invention
[0004] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0005] This disclosure provides a method, apparatus, and device for retrieving clothing images, to solve the technical problem of large calibration workload in the clothing image matching process.
[0006] In some embodiments, the method includes:
[0007] Retrieve images of the clothing item to be searched;
[0008] Determine the color tone image corresponding to the image to be retrieved, and input the image to be retrieved into the configured first neural convolutional network (CNN) model to obtain the activation region of a set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest.
[0009] Based on the location information of the activated area, the hue feature information of the clothing to be retrieved is extracted from the hue image;
[0010] Based on the color tone feature information, retrieve matching clothing images that match the clothing to be retrieved.
[0011] In some embodiments, the device includes:
[0012] The acquisition module is configured to acquire images of the clothing to be searched, including the images of the clothing to be searched.
[0013] The determination module is configured to determine the color tone image corresponding to the image to be retrieved;
[0014] The activation module is configured to input the image to be retrieved into a configured first neural convolutional network (CNN) model to obtain an activation region of a set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest.
[0015] The extraction module is configured to extract the hue feature information of the clothing to be retrieved from the hue image based on the location information of the activated region;
[0016] The retrieval module is configured to retrieve matching clothing images that match the clothing to be retrieved based on the color tone feature information.
[0017] In some embodiments, the device includes the above-described apparatus for retrieving clothing images.
[0018] The method, apparatus, and device for clothing image retrieval provided in this disclosure can achieve the following technical effects:
[0019] The system can directly extract the hue feature information of the clothing to be retrieved from the hue image corresponding to the image to be retrieved through the configured CNN model, and perform clothing matching and retrieval based on the hue feature information. In this way, the process of automatically matching clothing based on hue is realized. Moreover, in this process, it is not necessary to perform a large number of color labeling for each sample image. Only the hue feature is extracted, thereby reducing the color data labeling when configuring the CNN model, which reduces the labeling workload in the clothing matching process and lowers the cost of clothing matching.
[0020] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description
[0021] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrations and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are shown as similar elements. The drawings are not to be scaled. And wherein:
[0022] Figure 1 This is a schematic flowchart of a method for retrieving clothing images provided in an embodiment of this disclosure;
[0023] Figure 2 This is a schematic diagram of a wardrobe structure for clothing image retrieval provided in an embodiment of this disclosure;
[0024] Figure 3 This is a schematic diagram of a clothing image retrieval system provided in an embodiment of the present disclosure;
[0025] Figure 4 This is a schematic flowchart of a method for retrieving clothing images provided in an embodiment of this disclosure;
[0026] Figure 5 This is a schematic flowchart of a method for retrieving clothing images provided in an embodiment of this disclosure;
[0027] Figure 6 This is a schematic diagram of an embodiment of the present disclosure for extracting the activation area of clothing;
[0028] Figure 7 This is a schematic diagram of the structure of a clothing image retrieval device provided in an embodiment of this disclosure;
[0029] Figure 8 This is a schematic diagram of a clothing image retrieval device provided in an embodiment of the present disclosure. Detailed Implementation
[0030] To provide a more detailed understanding of the features and technical content of the embodiments of this disclosure, the implementation of the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this disclosure. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0031] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this disclosure described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0032] Unless otherwise stated, the term "multiple" means two or more.
[0033] In this embodiment of the disclosure, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0034] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0035] Currently, in clothing image retrieval schemes, each corresponding CNN algorithm selects specific feature information to retrieve similar clothing. Among all the features of clothing, color has the most direct visual impact on users. Therefore, in this embodiment, clothing matching and retrieval can be performed based on the color tone feature information of the clothing.
[0036] In smart homes, home devices can communicate with cloud servers. This allows many devices to perform clothing image matching locally or via a server. In this embodiment, wardrobes, televisions, and other home devices can also perform clothing recognition, recommendations, or other applications based on clothing image matching. In other words, clothing image matching can be completed locally on the home device or via a server. In this embodiment, the hue feature information of the clothing to be retrieved can be extracted directly from the hue image corresponding to the image to be retrieved using a configured CNN model. Clothing matching is then performed based on this hue feature information, thus achieving automatic clothing matching based on hue. Furthermore, when configuring the CNN model, it is not necessary to perform extensive color labeling for each sample image; only type labeling or pixel score labeling is required. This reduces the data labeling workload during CNN model configuration, thereby reducing the labeling workload in the clothing matching process, lowering the cost of clothing matching, and improving the speed of clothing image matching.
[0037] Figure 1 This is a schematic flowchart illustrating a method for retrieving clothing images provided in an embodiment of this disclosure. Figure 1 As shown, the process for retrieving clothing images includes:
[0038] Step 101: Obtain the images of the clothing items to be searched.
[0039] Currently, with the development of intelligent technology, many home appliances, such as televisions, refrigerators, and projectors, are intelligent devices with image recognition capabilities. Furthermore, many pieces of furniture are gradually becoming electronic and intelligent; for example, wardrobes and bookshelves now have functions such as automatic sterilization, dehumidification, and lighting. In this embodiment, wardrobes, televisions, and other devices related to "clothing" can automatically retrieve images of clothing.
[0040] The device can acquire images containing the clothing to be searched using the image acquisition equipment configured on the device. Figure 2 This is a schematic diagram of a wardrobe structure for clothing image retrieval provided in an embodiment of this disclosure. Figure 2As shown, a camera 220 is positioned directly above the wardrobe 210, allowing the camera to capture the user's image. This results in the acquisition of a searchable image containing the clothing item to be searched. Alternatively, the searchable image, including the clothing item, can also be obtained through a human-computer interaction interface on the home appliance.
[0041] Alternatively, home appliances such as wardrobes and televisions can implement clothing image retrieval functions through a server. This allows them to send captured images to the server, whereby the executing server can receive images containing the clothing to be retrieved. In some embodiments, the system receives user video information sent by the wardrobe, which is captured by an image acquisition device when the wardrobe determines the user is in a set location. Based on the user video information, images containing the clothing to be retrieved are obtained.
[0042] Devices such as wardrobes and televisions can also control the operation of image acquisition devices. For example, a wardrobe... Figure 2 Once the user stands in the designated position on the dressing mirror while wearing the clothing to be searched, the wardrobe will activate the camera to capture the user's video information and send it to the server. The server will then receive the user's video information and identify each frame of the received video information to obtain a search image containing the clothing to be searched.
[0043] In some embodiments, devices such as wardrobes and televisions can turn off the image acquisition device when they determine that the user has left the set location, thus saving resources.
[0044] Step 102: Determine the color tone image corresponding to the image to be retrieved, and input the image to be retrieved into the configured first convolutional neural network (CNN) model to obtain the activation region of the set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest.
[0045] After obtaining the image to be retrieved, the processing can be divided into two parallel steps. One step determines the color tone of the image to be retrieved. The other step extracts activation regions of a predetermined area from the image to be retrieved, where the activation region corresponds to the convolutional feature value with the largest value.
[0046] Once the image to be searched is obtained, the RGB values of each pixel in the image can be obtained. Therefore, the RGB values need to be converted into HIS values to form a tonal image that retains only the hue (H) values.
[0047] In an HIS image, hue (H) refers to the color attribute of a pure color, determining what color it is; saturation (S) is a measure of how much a pure color is diluted by white light; higher saturation results in a more vibrant color, determining its intensity; and intensity (I) determines how bright the white light shining on the color is. Compared to RGB images, HSI images utilize color information more efficiently, making color interpretation more intuitive. Furthermore, the hue (H) in an HSI image is essentially the main information about the color characteristics; the other S and I data fluctuate with lighting and other factors. Therefore, in some embodiments of this disclosure, the acquired image to be retrieved needs to be converted from an RGB image to an HSI image, and a hue image containing only H values needs to be determined. Specifically, this may include: determining the HIS value corresponding to each pixel based on the RGB values of each pixel in the image to be retrieved; and obtaining the hue image corresponding to the image to be retrieved based on the hue (H) value in the HIS value.
[0048] The RGB values of each pixel in the image to be retrieved can be converted into the corresponding HIS value using the following formula.
[0049] in,
[0050] In addition to the two parallel steps, there is a further step of determining the activation region of a predetermined area in the image to be retrieved. In this embodiment of the disclosure, a CNN model is used to determine and extract the activation region.
[0051] Therefore, before performing activation region extraction, it is necessary to configure the corresponding CNN model for activation region extraction, namely the first convolutional neural network (CNN) model.
[0052] In some embodiments, configuring the first convolutional neural network (CNN) model may include: performing category labeling preprocessing on three sets of first clothing sample images to obtain three sets of first image feature values, namely X. a ={x a1 x a2 , ..., x aNlabel}, X p ={x p1 x p2 , ..., x pNlabel}, X n ={x n1 x n2 , ..., x nNlabel}, where the first group and the second group have the same category label, i, and the third group has a different category label, j; then, the first convolutional neural network CNN model is configured according to the loss function of formula (1).
[0053]
[0054] Where α and w are set parameters.
[0055] Here, the first CNN model is configured according to the clothing category. The image label can be 1 to N. label Clothing category labels, N label This represents the maximum value for the number of clothing categories. For example, if clothing is divided into three categories: tops, pants, and full-body clothing, then N is the maximum value for each category. label =3; Divide the clothing into: 1. Short-sleeved top, 2. Long-sleeved top, 3. Long-sleeved jacket, 4. Long-sleeved jacket, 5. Vest, 6. Tank top, 7. Shorts, 8. Pants, 9. Skirt, 10. Short-sleeved dress, 11. Long-sleeved dress, 12. Vest dress, 13. Tank top dress, then N label =13. Thus, when configuring the first CNN model, training on sample images can be performed using a Triplet approach, with preprocessing done in batches. Each batch includes two images of the same class and one image of another class. Therefore, the three sets of first clothing sample images need to undergo class labeling preprocessing. The first and second sets have the same class label, i, while the third set has a different class label, j. After this forward pass (labeling preprocessing), the feature values of the three first images are obtained, namely X... a ={x a1 x a2 , ..., x aNlabel}, X p ={x p1 x p2 , ..., x pNlabel}, X n ={x n1 x n2 , ..., x nNlabel When constructing the first CNN model, it is necessary to use a loss function to make the deep learning CNN model converge, so that the parameters in the model can fit the task of "extracting activation regions". Here, the loss function of formula (1) can be used to make the parameters in the first CNN model fit the task of "extracting activation regions".
[0056] In some embodiments, configuring the first convolutional neural network (CNN) model may include: performing preprocessing on the second clothing sample image to calibrate pixel scores in a defined region, thereby obtaining a second image feature value, X. t ={x t1 x t2 , ..., x tn}, Y t={y t1 ,,y t2 , ..., y tn Then, according to the loss function of formula (2), configure the first convolutional neural network (CNN) model;
[0057]
[0058] Where γ is a set parameter.
[0059] Here, the first CNN model can be configured using the pixel ratings of the clothing. The image to be retrieved, including the clothing to be retrieved, can be partitioned into blocks of a predetermined area. Then, each block is assigned a pixel rating. The rating principle is as follows: if the block contains all clothing pixels, the pixel rating is 1; if it does not contain any clothing pixels, the pixel rating is 0; and if it contains only a portion of the clothing pixels, the pixel rating is 0.5. For example, based on a block count of n=49, a second clothing sample image, after one forward pass (labeling preprocessing), yields the true label value of the block, i.e., the second image feature value X. t ={x t1 x t2 , ..., x t49}, Y t ={y t1 y t2 , ..., y t49 Similarly, when constructing the first CNN model, a loss function is also needed to make the deep learning CNN model converge, and the parameters in the model be able to fit the task of "extracting activation regions". Here, the loss function of formula (2) can be used to make the parameters in the first CNN model be able to fit the task of "extracting activation regions". When n=49, the loss function can be as follows:
[0060]
[0061] Of course, γ is a set parameter.
[0062] After configuring the first CNN model, the image to be retrieved is input into the first CNN model to obtain the feature map. The average value of the convolutional feature values of the specified region of the feature map can be obtained. Then, the specified region corresponding to the maximum average value is determined as the activation region, that is, the convolutional feature value corresponding to the activation region is the maximum.
[0063] Step 103: Extract the tonal feature information of the clothing to be retrieved from the tonal image based on the location information of the activated region.
[0064] Once the active region in the image to be searched is extracted, the location information of the active region can be determined, and thus, the tonal feature information of the clothing to be searched can be extracted from the tonal image.
[0065] Step 104: Based on the color tone feature information, retrieve matching clothing images that match the clothing to be retrieved.
[0066] Based directly on the color tone feature information, Euclidean distance or cosine similarity formulas can be used to search and determine the matching clothing images that match the clothing to be retrieved.
[0067] Of course, in some embodiments, the second feature information of the clothing to be retrieved can also be obtained based on the second convolutional neural network (CNN) model trained with the second label. Then, based on the hue feature information and the second feature information, a search is performed using Euclidean distance or cosine similarity formulas to determine the matching clothing image that matches the clothing to be retrieved. Clothing has many features, such as color, hue, size, type, etc. In this embodiment, the second feature information is the feature information of the clothing other than the hue feature information.
[0068] There are various ways to perform clothing matching searches based on feature information, including: Euclidean distance, cosine similarity formula, Manhattan distance, Minkowski distance, Pearson correlation coefficient, Jaccard similarity coefficient, etc., which will be described in detail here.
[0069] As can be seen, in this embodiment, the hue feature information of the clothing to be retrieved can be directly extracted from the hue image corresponding to the image to be retrieved through the configured CNN model, and clothing matching retrieval can be performed based on the hue feature information. This achieves automatic clothing matching based on hue. Furthermore, when configuring the CNN model, it is not necessary to perform extensive color labeling for each sample image; only type labeling or pixel score labeling is required. This reduces the data labeling workload during CNN model configuration, thereby reducing the labeling workload in the clothing matching process, lowering the cost of clothing matching, and improving the speed of clothing image matching.
[0070] The following describes the operation process in a specific embodiment, illustrating the image retrieval process for clothing provided by the embodiments of the present invention.
[0071] In one embodiment of this disclosure, Figure 3 This is a schematic diagram of a clothing image retrieval system provided in an embodiment of this disclosure. Figure 3 As shown, it includes: server 1000, wardrobe 2000, and terminal 3000. Server 1000 can communicate with wardrobe 2000 and terminal 3000. Server 1000 is configured with a first CNN model according to formula (1). Wardrobe 2000 can... Figure 2 As shown.
[0072] Figure 4 This is a schematic flowchart illustrating a method for retrieving clothing images provided in an embodiment of this disclosure. Combined with... Figure 4 The process for retrieving clothing images includes:
[0073] Step 401: Does the wardrobe detect the user at the designated location? If yes, proceed to step 402; otherwise, proceed to step 406.
[0074] Step 402: The wardrobe activates its camera to collect user video information.
[0075] Step 403: The wardrobe sends the collected user video information to the server, enabling the server to obtain the image of the clothing to be searched and determine the matching clothing image that matches the clothing to be searched in the image.
[0076] Step 404: The wardrobe sends a matching result request to the server.
[0077] Step 405: The wardrobe receives the matching clothing image sent by the server and performs corresponding clothing processing.
[0078] Step 406: The wardrobe confirms that the camera is off.
[0079] The wardrobe sends the user's video information to the server, which can then perform clothing image retrieval. The server has already configured the first CNN model according to formula (1).
[0080] Figure 5 This is a schematic flowchart illustrating a method for retrieving clothing images provided in an embodiment of this disclosure. Combined with... Figure 5 The process for retrieving clothing images includes:
[0081] Step 501: The server obtains the image containing the clothing item to be searched based on the user's video information.
[0082] Step 502: Input the image to be retrieved into the configured first CNN model to obtain the activation region of the set area of the clothing to be retrieved.
[0083] Figure 6 This is a schematic diagram illustrating an embodiment of the present disclosure for extracting the activation area of clothing. For example... Figure 6 As shown, the image to be retrieved is input into the configured first CNN model, resulting in a 2048×7×7 feature map. Then, the average value is taken along the 2048-dimensional vector direction in each 7×7 grid, i.e., the average value of the convolutional feature values is obtained by setting the area of each 7×7 grid. Furthermore, the top N maximum values are taken from the 7×7 values, where N=6, to obtain the activation region, as shown below. Figure 6As shown in 610.
[0084] Step 503: Determine the HIS value corresponding to each pixel based on the RGB values of each pixel in the image to be retrieved; obtain the hue image corresponding to the image to be retrieved based on the hue (H) value in the HIS value.
[0085] For example: convert the image to be retrieved into a tone image, such as... Figure 6 As shown in 620. Of course, steps 502 and 503 can be run in parallel.
[0086] Step 504: Extract the tonal feature information of the clothing to be retrieved from the tonal image based on the location information of the activated region.
[0087] For example: After obtaining a 224×224 tone image, pooling is performed according to a 7×7 grid, where stride = kernel_size = 32. The coordinates (i, j) (0 ≤ i, j < 7) of the above 6 regions are superimposed on the tone image, such as... Figure 6 As shown in Figure 630, the activated H data, i.e., hue feature information, can be obtained.
[0088] Step 505: Based on the color tone feature information, retrieve the matching clothing image that matches the clothing to be retrieved, and send the matching clothing image to the terminal.
[0089] As can be seen, in this embodiment, the server can directly extract the hue feature information of the clothing to be retrieved from the hue image corresponding to the image to be retrieved through the configured CNN model, and perform clothing matching retrieval based on the hue feature information. This achieves automatic clothing matching based on hue. Furthermore, when configuring the CNN model, it is not necessary to perform extensive color labeling for each sample image; only type labeling is required. This reduces the data labeling workload during CNN model configuration, thereby reducing the labeling workload in the clothing matching process, lowering the cost of clothing matching, and improving the speed of clothing image matching.
[0090] Based on the above process for retrieving clothing images, a device for retrieving clothing images can be constructed.
[0091] Figure 7 This is a schematic diagram of a clothing image retrieval device provided in an embodiment of this disclosure. Figure 7 As shown, the device for retrieving clothing images includes: an acquisition module 710, a determination module 720, an activation module 730, an extraction module 740, and a retrieval module 750.
[0092] The acquisition module 710 is configured to acquire images of the clothing to be searched.
[0093] The determination module 720 is configured to determine the color tone image corresponding to the image to be retrieved.
[0094] The activation module 730 is configured to input the image to be retrieved into the configured first neural convolutional network (CNN) model to obtain the activation region of a set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest.
[0095] The extraction module 740 is configured to extract the tonal feature information of the clothing to be retrieved from the tonal image based on the location information of the activated region.
[0096] The retrieval module 750 is configured to retrieve matching clothing images that match the clothing to be retrieved based on tonal feature information.
[0097] In some embodiments, the determining module 720 is specifically configured to determine the HIS value corresponding to each pixel based on the RGB values of each pixel in the image to be retrieved; and to obtain the hue image corresponding to the image to be retrieved based on the hue H value in the HIS value.
[0098] In some embodiments, the system further includes: a first configuration module configured to perform category labeling preprocessing on the three sets of first clothing sample images to obtain three sets of first image feature values, namely X... a ={x a1 x a2 , ..., x aNlabel}, X p ={x p1 x p2 , ..., x pNlabel}, X n ={x n1 x n2 , ..., x nNlabel}, where the first group and the second group have the same category label, i, and the third group has a different category label, j; according to the loss function of formula (1), configure the first convolutional neural network CNN model;
[0099]
[0100] Where α and w are set parameters.
[0101] In some embodiments, a second configuration module is further included, configured to perform preprocessing on the second clothing sample image for a defined region pixel score to obtain a second image feature value, X. t ={x t1 x t2 , ..., x tn}, Y t ={y t1 ,,yt2 , ..., y tn}; Configure the first convolutional neural network (CNN) model according to the loss function of formula (2);
[0102]
[0103] Where γ is a set parameter.
[0104] In some embodiments, the retrieval module 750 is specifically configured to obtain second feature information of the clothing to be retrieved based on a second convolutional neural network (CNN) model trained with a second label; and to search for matching clothing images that match the clothing to be retrieved based on the hue feature information and the second feature information using Euclidean distance or cosine similarity formula.
[0105] In some embodiments, the acquisition module 710 is specifically configured to receive user video information sent by the wardrobe, wherein the user video information is acquired by the wardrobe through an image acquisition device when it determines that the user is in a set position; and to obtain a searchable image containing the clothing to be searched based on the user video information.
[0106] In some embodiments, the system further includes a feedback module configured to send retrieved matching clothing images to the wardrobe for clothing recognition or recommendation.
[0107] As can be seen, in this embodiment, the clothing image retrieval device can directly extract the hue feature information of the clothing to be retrieved from the hue image corresponding to the image to be retrieved through the configured CNN model, and perform clothing matching retrieval based on the hue feature information. This achieves automatic clothing matching based on hue. Furthermore, when configuring the CNN model, it is not necessary to perform extensive color labeling for each sample image; only type labeling or pixel score labeling is required. This reduces the data labeling required during CNN model configuration, thereby reducing the labeling workload during clothing matching, lowering the cost of clothing matching, and improving the speed of clothing image matching.
[0108] This disclosure provides an apparatus for retrieving clothing images, the structure of which is as follows: Figure 8 As shown, it includes:
[0109] The processor 100 and memory 101 may further include a communication interface 102 and a bus 103. The processor 100, communication interface 102, and memory 101 can communicate with each other via the bus 103. The communication interface 102 can be used for information transmission. The processor 100 can call logical instructions stored in the memory 101 to execute the clothing image retrieval method described in the above embodiment.
[0110] Furthermore, the logic instructions in the aforementioned memory 101 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0111] The memory 101, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this disclosure. The processor 100 executes functional applications and data processing by running the program instructions / modules stored in the memory 101, that is, it implements the method for clothing image retrieval in the above method embodiments.
[0112] The memory 101 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 101 may include high-speed random access memory and may also include non-volatile memory.
[0113] This disclosure provides an apparatus comprising a wardrobe or server, including the above-described device for retrieving clothing images.
[0114] This disclosure provides a computer-readable storage medium storing computer-executable instructions configured to perform the above-described method for retrieving clothing images.
[0115] This disclosure provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions that, when executed by a computer, cause the computer to perform the above-described method for retrieving clothing images.
[0116] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0117] The technical solutions of this disclosure can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in this disclosure. The aforementioned storage medium can be a non-transitory storage medium, including: a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, and other media capable of storing program code; it can also be a transient storage medium.
[0118] The foregoing description and accompanying drawings fully illustrate embodiments of the present disclosure to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operation may vary. Parts and features of some embodiments may be included or replace parts and features of other embodiments. The scope of the embodiments of this disclosure includes the entire scope of the claims and all available equivalents of the claims. While the terms “first,” “second,” etc., may be used in this application to describe elements, these elements should not be limited by these terms. These terms are used only to distinguish one element from another. For example, a first element may be called a second element without changing the meaning of the description, and similarly, a second element may be called a first element, provided that all occurrences of “first element” are consistently renamed and all occurrences of “second element” are consistently renamed. First and second elements are both elements, but may not be the same element. Moreover, the terminology used in this application is only for describing embodiments and is not intended to limit the claims. As used in the description of the embodiments and claims, unless the context clearly indicates otherwise, the singular forms “a,” “an,” and “the” are intended to also include the plural forms. Similarly, the term “and / or” as used herein means including one or more of the associated listed elements and all possible combinations thereof. Additionally, when used herein, the terms “comprise” and its variations “comprises” and / or “comprising” refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Without further limitations, an element defined by the phrase “comprising an…” does not exclude the presence of additional identical elements in the process, method, or apparatus that includes said element. In this document, each embodiment may focus on the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, the relevant parts can be referred to the description of the method section.
[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this disclosure. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0120] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the shown or discussed units may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0121] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks may also occur in a different order than disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in a block diagram and / or flowchart, and combinations of blocks in a block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A method for retrieving images of clothing, characterized in that, include: Retrieve images of the clothing item to be searched; Determine the color tone image corresponding to the image to be retrieved, and input the image to be retrieved into the configured first convolutional neural network (CNN) model to obtain the activation region of a set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest. Based on the location information of the activated area, the hue feature information of the clothing to be retrieved is extracted from the hue image; Based on the color tone feature information, retrieve matching clothing images that match the clothing to be retrieved; Before obtaining the activated area of the set area of the garment to be retrieved, the process includes: The three sets of first clothing sample images were trained using a Triplet method. Preprocessing for category labeling was performed on a batch basis, resulting in the feature values of the three sets of first images, denoted as X. a ={ x a1 , x a2 ,..., x aNlabel }, X p ={ x p1 , x p2 ,..., x pNlabel }, X n ={ x n1 , x n2 ,..., x nNlabel }, where the first group and the second group have the same category label, i, and the third group has a different category label, j; the first convolutional neural network CNN model is configured according to the loss function of formula (1); (1) Where α and w are set parameters, a batch includes two images of the same category and one image of another category; Alternatively, the second clothing sample image can be preprocessed with pixel scoring in a defined region to obtain the second image feature value, X. t ={ x t1 , x t2 ,..., x tn }, Y t ={ y t1 , y t2 ,..., y tn }; Configure the first convolutional neural network (CNN) model according to the loss function of formula (2); (2) Where γ is a set parameter.
2. The method according to claim 1, characterized in that, Determining the color tone image corresponding to the image to be retrieved includes: Based on the RGB values of each pixel in the image to be retrieved, determine the HIS value corresponding to each pixel; Based on the hue H value in the HIS value, the hue image corresponding to the image to be retrieved is obtained.
3. The method according to claim 1, characterized in that, The process of retrieving matching clothing images that match the clothing to be retrieved includes: The second feature information of the clothing to be retrieved is obtained based on the second convolutional neural network (CNN) model trained with the second label; Based on the hue feature information and the second feature information, a search is performed using Euclidean distance or cosine similarity formula to determine the matching clothing image that matches the clothing to be retrieved.
4. The method according to claim 1, characterized in that, The acquisition of the image of the clothing item to be searched includes: The system receives user video information sent by the wardrobe, wherein the user video information is captured by the wardrobe through an image acquisition device when it determines that the user is in a set position; Based on the user's video information, a searchable image containing the clothing item to be searched is obtained.
5. The method according to claim 4, characterized in that, The method further includes: The retrieved matching clothing images are sent to the wardrobe for clothing identification or recommendation.
6. A device for retrieving images of clothing, characterized in that, include: The acquisition module is configured to acquire images of the clothing to be searched, including the images of the clothing to be searched. The determination module is configured to determine the color tone image corresponding to the image to be retrieved; The activation module is configured to input the image to be retrieved into a configured first convolutional neural network (CNN) model to obtain an activation region of a set area of the clothing to be retrieved, wherein the convolutional feature value corresponding to the activation region is the largest. The extraction module is configured to extract the hue feature information of the clothing to be retrieved from the hue image based on the location information of the activated region; The retrieval module is configured to retrieve matching clothing images that match the clothing to be retrieved based on the color tone feature information; It also includes: a first configuration module, configured to train the three sets of first clothing sample images in a Triplet manner, performing category labeling preprocessing on a batch basis to obtain three sets of first image feature values, namely X. a ={ x a1 , x a2 ,..., x aNlabel }, X p ={ x p1 , x p2 ,..., x pNlabel }, X n ={ x n1 , x n2 ,..., x nNlabel }, where the first group and the second group have the same category label, i, and the third group has a different category label, j; according to the loss function of formula (1), configure the first convolutional neural network CNN model; (1) Where α and w are set parameters, a batch includes two images of the same category and one image of another category; Alternatively, it may include a second configuration module configured to perform preprocessing on the second clothing sample image to determine pixel scores for a specified region, thereby obtaining a second image feature value, X. t ={ x t1 , x t2 ,..., x tn }, Y t ={ y t1 , y t2 ,..., y tn }; Configure the first convolutional neural network (CNN) model according to the loss function of formula (2); (2) Where γ is a set parameter.
7. An apparatus for retrieving images of clothing, the apparatus comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to perform the method for retrieving clothing images as described in any one of claims 1 to 5 when executing the program instructions.
8. A device, characterized in that, Includes the apparatus for retrieving clothing images as described in claim 6 or 7.