Electronic device and method for extracting attributes of fashion item
By converting the pose of a person in an image to a preset pose, the electronic device enhances the accuracy of fashion item attribute extraction, addressing pose-related inconsistencies in existing technologies.
Patent Information
- Application Number
- PCT/KR2025/011075
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-09
- Filing Date
- 2025-07-25
- Publication Date
- 2026-02-12
AI Technical Summary
Existing technologies struggle to accurately extract attributes of fashion items from images due to variations in human poses, leading to inconsistencies in attribute detection.
An electronic device and method that converts the pose of a person in an image to a preset pose, allowing for more accurate detection and extraction of fashion item attributes by generating a second image with a standardized pose.
Improves the accuracy of attribute extraction for fashion items by aligning the pose to a preset standard, enhancing the consistency and reliability of attribute detection.
Smart Images

Figure KR2025011075_12022026_PF_FP_ABST
Abstract
Description
Electronic device and method for extracting attributes of fashion items
[0001] An electronic device and method for extracting attributes of a fashion item from an image are disclosed. Specifically, an electronic device and method for extracting attributes of a fashion item by generating an image with a transformed pose from an input image are disclosed.
[0002] Extracting object properties from images plays a crucial role in the fields of computer vision and artificial intelligence. Extracted object properties can provide personalized experiences to users. Therefore, technologies that more accurately extract object properties from images are needed.
[0003] According to one aspect of the present disclosure, a method for extracting attributes of a fashion item from an image is disclosed. In one embodiment, the method may include obtaining a first image including a person wearing at least one fashion item. In one embodiment, the method may include obtaining a second image in which a pose of the person included in the first image is converted into a preset pose. In one embodiment, the method may include detecting at least one fashion item worn by the person converted into the preset pose based on the second image. In one embodiment, the method may include obtaining at least one attribute of at least one fashion item.
[0004] According to one aspect of the present disclosure, an electronic device for extracting attributes of a fashion item from an image is disclosed. In one embodiment, the electronic device may include a memory storing one or more instructions and at least one processor for executing one or more instructions stored in the memory. In one embodiment, by the at least one processor executing one or more instructions, the electronic device may obtain a first image including a person wearing at least one fashion item. In one embodiment, by the at least one processor executing one or more instructions, the electronic device may obtain a second image in which the pose of the person included in the first image is converted into a preset pose. In one embodiment, by the at least one processor executing one or more instructions, the electronic device may detect at least one fashion item worn by the person converted into a preset pose based on the second image. In one embodiment, by the at least one processor executing one or more instructions, the electronic device may obtain at least one attribute of at least one fashion item.
[0005] According to one aspect of the present disclosure, a computer-readable recording medium having recorded thereon a program for performing an operation of an electronic device, one of the methods described above and below.
[0006] FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.
[0007] FIG. 2 is a flowchart illustrating a method for an electronic device to obtain attributes of a fashion item from an image according to one embodiment of the present disclosure.
[0008] FIG. 3 is a diagram for explaining a method for an electronic device according to one embodiment of the present disclosure to detect a pose of a person included in an image.
[0009] FIG. 4 is a diagram illustrating a method for comparing a detected pose with a preset pose by an electronic device according to one embodiment of the present disclosure.
[0010] FIG. 5 is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to acquire a second image.
[0011] FIG. 6 is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to detect a fashion item from an image and obtain attributes of the detected fashion item.
[0012] FIG. 7 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select an image for obtaining attributes of a fashion item.
[0013] FIG. 8 is a diagram illustrating training of an image generation module according to one embodiment of the present disclosure.
[0014] FIG. 9A is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to provide image search results based on attributes of a fashion item.
[0015] FIG. 9b is a diagram for explaining a method for an electronic device according to an embodiment of the present disclosure to recommend fashion coordination based on attributes of a fashion item.
[0016] FIG. 9c is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to display an image generated based on an image generation module.
[0017] FIG. 10 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
[0018] FIG. 11 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0019] The terms used in the embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, the applicant may arbitrarily select terms, and in such cases, their meanings will be described in detail in the description of the relevant embodiments. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.
[0020] Unless the context clearly dictates otherwise, the singular forms "a," "an," and "the" are to be understood to include plural referents. Thus, for example, the description "a constituent surface" may also include reference to one or more of such surfaces.
[0021] Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by one of ordinary skill in the art described in this disclosure.
[0022] When a part of this disclosure is said to "include" a component, this does not exclude other components, unless otherwise specifically stated, but rather implies the inclusion of other components. Furthermore, terms such as "part," "module," and the like described herein refer to a unit that processes at least one function or operation, which may be implemented in hardware or software, or a combination of hardware and software.
[0023] The expression "configured to" as used herein can be used interchangeably with, for example, "suitable for", "having the capacity to", "designed to", "adapted to", "made to", or "capable of", depending on the context. The term "configured to" does not necessarily mean something is "specifically designed to" in terms of hardware. Instead, in some contexts, the expression "a system configured to" can mean that the system is "capable of" in conjunction with other devices or components. For example, the phrase "a processor configured to perform A, B, and C" can mean a dedicated processor for performing the operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or an application processor) that can perform the operations by executing one or more software programs stored in memory.
[0024] It should be understood that the blocks and combinations of flowcharts in each of the flowcharts in this disclosure can be implemented by one or more computer programs containing computer-executable instructions. The one or more computer programs may be stored entirely in a single memory, or may be divided and stored across multiple different memories.
[0025] All functions or operations described in the present disclosure may be processed by a single processor or a combination of processors. A single processor or a combination of processors may include circuitry that performs processing, such as an Application Processor (AP), a Communication Processor (CP), a Graphical Processing Unit (GPU), a Neural Processing Unit (NPU), a Microprocessor Unit (MPU), a System on Chip (SoC), or an Integrated Chip (IC).
[0026] It will be appreciated that each block of the processing flow diagrams described in the present disclosure and combinations of the flow diagrams can be performed by computer program instructions. These computer program instructions can be installed in a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing equipment, such that the instructions executed by the processor of the computer or other programmable data processing equipment create a means for performing the functions described in the flow diagram block(s). These computer program instructions can also be stored in a computer-available or computer-readable memory that can direct a computer or other programmable data processing equipment to implement the functions in a specific manner, such that the instructions stored in the computer-available or computer-readable memory can also produce an article of manufacture that includes an instruction means for performing the functions described in the flow diagram block(s). Since the computer program instructions may be installed on a computer or other programmable data processing device, a series of operational steps may be performed on the computer or other programmable data processing device to create a computer-executable process, and the instructions that cause the computer or other programmable data processing device to perform the steps for performing the functions described in the flowchart block(s) may also provide steps for performing the functions described in the flowchart block(s).
[0027] Each block described herein may represent a module, segment, or portion of code that includes one or more executable instructions for performing a specified logical function(s). It should also be noted that in some alternative implementation examples, the functions described in the blocks may occur out of order. For example, two blocks depicted in succession may actually be executed substantially concurrently, or the blocks may sometimes be executed in reverse order, depending on their respective functions.
[0028] One or more processors according to the present disclosure may include at least one of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), and a Neural Processing Unit (NPU). The one or more processors may be implemented in the form of an integrated system-on-a-chip (SoC) including one or more electronic components. Each of the one or more processors may also be implemented as separate hardware (H / W).
[0029] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by the first processor, or the first operation and the second operation may be performed by the first processor (e.g., a general-purpose processor) and the third operation may be performed by the second processor (e.g., an artificial intelligence-specific processor). However, the embodiments of the present disclosure are not limited thereto.
[0030] One or more processors according to the present disclosure may be implemented as a single core processor or as a multicore processor.
[0031] When a method according to one embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by one core or may be performed by multiple cores included in one or more processors.
[0032] In the present disclosure, a "model" or "artificial intelligence (AI) model" may refer to a set of functions or algorithms that are set to perform a desired characteristic (or purpose) by being learned using a plurality of learning data by a learning algorithm. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning. In one embodiment, the AI model may be stored in the memory of an electronic device. However, the AI model is not limited thereto, and the electronic device may transmit data input to the AI model to the server and receive data output from the AI model from the server.
[0033] In the present disclosure, a 'model' or 'artificial intelligence model' may be composed of a plurality of neural network layers. Each of the plurality of neural network layers has a plurality of weight values, and can perform neural network operations through operations between the operation results of the previous layer and the plurality of weights. The plurality of weights of the plurality of neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the plurality of weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Examples of models including a plurality of neural network layers include, but are not limited to, a deep neural network (DNN), a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), and deep Q-networks.
[0034] In the present disclosure, a "fashion item" may refer to a fashion-related object worn or used by a person. In one embodiment, a fashion item may include clothing such as tops, bottoms, pants, dresses, and outerwear; shoes such as sneakers, boots, sandals, and loafers; or accessories such as bags, hats, jewelry, glasses, and watches. However, the fashion item is not necessarily limited to the examples described above, and a fashion item may include various objects that can be recognized as being worn or used by a person included in an image.
[0035] In the present disclosure, an "attribute" may refer to information indicating a unique characteristic of an object. In one embodiment, an attribute may include at least one characteristic associated with an object that is identified based on the visual appearance of the object. In one embodiment, the attributes of a fashion item may include external characteristics of the fashion item, material characteristics, color characteristics, design pattern characteristics, style characteristics, size characteristics, and length characteristics, but are not necessarily limited to the examples described above. In one embodiment, the attributes of a fashion item may include information for identifying which category the fashion item belongs to or what visual or functional characteristics it has.
[0036] Below, with reference to the attached drawings, embodiments of the present disclosure are described in detail so that those skilled in the art can easily practice the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In the drawings, portions irrelevant to the description have been omitted for clarity of explanation, and similar reference numerals have been used throughout the specification to designate similar parts.
[0037] The present disclosure will be described below with reference to the attached drawings.
[0038] FIG. 1 is a drawing schematically illustrating the operation of electronic devices according to one embodiment of the present disclosure.
[0039] Referring to FIG. 1, an electronic device (1000) may acquire a first image (110) including a person wearing at least one fashion item. For example, the first image (110) may be an image including a person (11) wearing a top (111) and a bottom (112), which are fashion items. Examples of the top (111) may include a T-shirt, a shirt, a blouse, a sweater, etc., and examples of the bottom (112) may include pants, jeans, a skirt, leggings, etc., but are not necessarily limited to the examples described above. In one embodiment, the first image (110) may be an image captured by a camera of the electronic device (1000) and stored in a memory of the electronic device, or an image received from an external electronic device.
[0040] In one embodiment, the electronic device (1000) may acquire a second image (120) based on the first image (110). In one embodiment, the second image (120) may include an image in which the pose of the person (11) included in the first image (110) is converted into a preset pose. In other words, the person (12) included in the second image (120) may be a person who assumes a preset pose that is different from the pose of the person (11) included in the first image (110). Here, the first person (11) of the first image (110) may correspond to the second person (12) of the second image (120), and the upper garment (111) and lower garment (112) worn by the first person (11) of the first image (110) may correspond to the upper garment (121) and lower garment (122) worn by the second person (12) of the second image (120). That is, the second image (120) may include the second person (12), which is the appearance (or shape) when the pose of the person (11) included in the first image (110) is converted to a preset pose, and the upper body (111) and lower body (112) worn by the person (11) included in the first image (110) may include the upper body (121) and lower body (122), which are the appearances when the second person (12) wears the upper body (111) and lower body (112).
[0041] In one embodiment, the electronic device (1000) may obtain at least one attribute of at least one fashion item worn by a person (12) included in the second image (120) based on the second image (120). Here, obtaining at least one attribute of at least one fashion item may include the electronic device (1000) detecting at least one fashion item worn by a person (12) included in the second image (120) based on the second image (120) and extracting at least one attribute of the detected at least one fashion item. For example, the electronic device (1000) may obtain information that the attribute (131) of the top (121) worn by the person (12) converted into a preset pose based on the second image (120) is 'the neckline is a collar, the color is pink, and the sleeve length is long.' In addition, the electronic device (1000) can obtain information that the attributes (132) of the lower garment (122) worn by the person (12) converted to a preset pose based on the second image (120) are 'the color is blue and the length is long.' However, this is not necessarily limited to the example described above, and various attributes related to fashion items can be further obtained for each of the upper garment (121) and the lower garment (122).
[0042] In one embodiment, the electronic device (1000) may provide a result of an image retrieval based on at least one attribute of at least one acquired fashion item. In one embodiment, the attribute of at least one fashion item may be stored in the electronic device (1000). In one embodiment, the electronic device (1000) may provide information corresponding to a search command acquired from a user based on the attribute of at least one fashion item stored in the electronic device (1000). For example, in response to a request from a user to find a pink top (Top), the electronic device (1000) may identify an attribute (131) of a top (121) stored in the electronic device (1000) that includes information that the color is pink, and display a first image (110) corresponding to a second image (120) acquired by the attribute (131) of the top (121) on the display of the electronic device (1000). However, it is not necessarily limited to the above-described example, and the electronic device (1000) can display various UIs based on the properties of at least one stored fashion item.
[0043] In one embodiment, the attribute of at least one fashion item acquired from the second image (120) may have a higher consistency with the attribute of the actual fashion item than the attribute of at least one fashion item acquired from the first image (110). For example, the pose of a person (11) included in the first image (110) may be in a state where both arms are bent. In this case, the attribute of the sleeve length of the upper garment (111) detected from the first image (110) may be extracted as medium. On the other hand, the pose of a person (12) included in the second image (120) may be in a state where both arms are extended. In this case, the attribute of the sleeve length of the upper garment (121) detected from the second image (120) may be extracted as long. In this way, when extracting the attribute of a fashion item from an image including a person wearing a fashion item, the pose of the person in the image may affect the accuracy of the attribute of the extracted fashion item, i.e., how accurately the attribute of the actual fashion item can be acquired.
[0044] An electronic device (1000) according to one embodiment of the present disclosure may acquire attributes of a fashion item from an image by converting a pose of a person in the image into a preset pose, and may then acquire attributes of the fashion item from the generated image. In this case, the preset pose may include a pose more suitable for acquiring attributes of the fashion item. Accordingly, the accuracy with which the electronic device (1000) extracts attributes of the fashion item may be improved.
[0045] FIG. 2 is a flowchart illustrating a method for an electronic device to obtain attributes of a fashion item from an image according to one embodiment of the present disclosure.
[0046] Referring to FIG. 2, the operation of the electronic device (1000) will be schematically described, and a detailed description of each operation will be described with reference to the drawings that follow. In addition, the operations of the electronic device (1000) described in the present disclosure may be understood as the operations of the electronic device (1000) and the processor (1700) of the electronic device (1000) illustrated in FIG. 10, and the server (2000) and the processor (2300) of the server (2000) illustrated in FIG. 11. In one embodiment, some of the steps illustrated in FIG. 2 may be omitted or other steps may be added.
[0047] In step S210, the electronic device (1000) may acquire a first image including a person wearing at least one fashion item. In one embodiment, the electronic device (1000) may store multiple images including a person wearing at least one fashion item in the memory of the electronic device (1000). The electronic device (1000) may acquire at least one of the multiple images stored in the memory as the first image.
[0048] In one embodiment, the first image may include an image acquired through a camera of the electronic device (1000). For example, the electronic device (1000) may acquire an image including a user by photographing a user wearing at least one fashion item through the camera. In one embodiment, the electronic device (1000) may store the image acquired through the camera in the electronic device (1000) and acquire the first image based on a user input selecting an image stored in the electronic device (1000).
[0049] In one embodiment, the first image may include an image acquired from an external electronic device via a communication interface of the electronic device (1000). In one embodiment, the image acquired from the external electronic device may include an image acquired by photographing a person other than the user of the electronic device (1000). For example, the first image may be captured by a camera of the external electronic device, wherein the user of the external electronic device is wearing at least one fashion item, and may be stored in the external electronic device or transmitted to a server and stored in the server. The electronic device (1000) may receive the first image from the external electronic device or server via the communication interface and store the first image in the electronic device (1000).
[0050] In step S220, the electronic device (1000) may acquire a second image in which the pose of the person included in the first image is converted into a preset pose. In one embodiment, the electronic device (1000) acquiring the second image may include the electronic device (1000) generating the second image based on the first image.
[0051] In one embodiment, the electronic device (1000) can detect a pose of a person included in a first image. In one embodiment, the electronic device (1000) detecting (extracting) the pose of the person may include the electronic device (1000) estimating the pose of the person included in the first image. In other words, in the present disclosure, the operation of the electronic device (1000) detecting specific data may be understood as estimating the specific data, and conversely, the operation of the electronic device (1000) estimating the specific data may be understood as detecting the specific data. In one embodiment, the electronic device (1000) can detect a person included in the first image based on the first image and estimate the pose of the detected person.
[0052] In one embodiment, the electronic device (1000) may acquire skeleton data including a plurality of key points corresponding to joints and preset body parts of a person included in a first image as a pose of the detected person. For example, the skeleton data may include information about the positions of a plurality of key points corresponding to the nose, neck, shoulders, elbows, wrists, hips, knees, ankles, eyes, and ears of the person, and information about links or edges connecting the plurality of key points. However, the joints and preset body parts indicated by the plurality of key points of the skeleton data are not necessarily limited to the examples described above.
[0053] The specific details of how the electronic device (1000) detects a person's pose based on the first image will be described again below with reference to FIG. 3.
[0054] In one embodiment, the electronic device (1000) can compare the pose of the detected person with a preset pose. In one embodiment, data regarding the preset pose may be stored in advance in the electronic device (1000). When the pose of the person included in the first image is detected based on the first image, the electronic device (1000) can compare the pose of the detected person with the preset pose stored in advance in the electronic device (1000).
[0055] In one embodiment, the preset pose may include a preset pose of a person. In other words, comparing the detected pose of a person with the preset pose may mean comparing the detected pose of a person with the preset pose of a person. In one embodiment, the detected pose of a person and the preset pose of a person may have the same data type. For example, if the pose of a person detected based on the first image includes skeleton data, the preset pose of a person may also include skeleton data. In one embodiment, the electronic device (1000) may compare the detected pose of a person with the preset pose of a person by calculating the similarity between the skeleton data of the detected pose of a person and the skeleton data of the preset pose of a person. In one embodiment, the electronic device (1000) may calculate the similarity between the detected pose of a person and the preset pose of a person based on the difference in the positions of corresponding key points of the skeleton data. In one embodiment, the electronic device (1000) can calculate the similarity between the detected human pose and the preset human pose based on the difference in angles formed by three corresponding key points of the skeleton data.
[0056] In one embodiment, the electronic device (1000) can identify whether the pose of the detected person corresponds to the preset pose by comparing the pose of the detected person with a preset pose. In one embodiment, the electronic device (1000) can identify that the pose of the detected person corresponds to the preset pose if the similarity between the pose of the detected person and the preset pose is greater than or equal to a threshold value. In one embodiment, the electronic device (1000) can identify that the pose of the detected person does not correspond to the preset pose if the similarity between the pose of the detected person and the preset pose is less than the threshold value.
[0057] The specific details of how the electronic device (1000) compares the detected pose of a person with a preset pose will be described again below with reference to FIG. 4.
[0058] In one embodiment, the electronic device (1000) may generate a second image including a person whose detected pose has been converted to a preset pose based on the comparison result. Here, the comparison result may refer to a comparison result between the detected person's pose and the preset pose. In one embodiment, the electronic device (1000) may acquire the second image based on the identification that the detected person's pose does not correspond to the preset pose.
[0059] In one embodiment, a person and at least one fashion item included in the second image may correspond to a person and at least one fashion item included in the first image. In other words, the second image may include a person whose appearance is that of a person included in the first image when the pose of the person is converted to a preset pose, and at least one fashion item worn by the person included in the first image, which is the appearance of a person worn by the person included in the second image.
[0060] In one embodiment, the electronic device (1000) can detect the pose of the camera that captured the first image based on the acquired first image.
[0061] In one embodiment, the electronic device (1000) may acquire 6DoF (6 Degrees of Freedom) data including position information and rotation information of a camera that captured a first image as a pose of the detected camera. For example, the pose of the camera may include three-dimensional position coordinate values (x, y, z) which are position information of the camera and rotation angle values (roll, pitch, yaw) which are rotation information of the camera. However, the position information and rotation information of the camera are not necessarily limited to the examples described above.
[0062] In one embodiment, the electronic device (1000) may compare the detected pose of the camera with a preset pose. In one embodiment, when the pose of the camera that captured the first image is detected based on the first image, the electronic device (1000) may compare the detected pose of the camera with a preset pose pre-stored in the electronic device (1000).
[0063] In one embodiment, the preset pose may include the preset pose of the camera. In other words, comparing the detected pose of the camera with the preset pose may mean comparing the detected pose of the camera with the preset pose of the camera. In one embodiment, the detected pose of the camera and the preset pose of the camera may have the same data type. For example, if the pose of the camera detected based on the first image includes 6Dof data, the preset pose of the camera may also include 6Dof data. In one embodiment, the electronic device (1000) may calculate the similarity between the 6Dof data of the detected pose of the camera and the 6Dof data of the preset pose of the camera to compare the detected pose of the camera with the preset pose of the camera. In one embodiment, the electronic device (1000) may calculate the similarity for the position of the camera based on the difference between the 3D position coordinate value of the detected pose of the camera and the 3D position coordinate value of the preset pose of the camera. In one embodiment, the electronic device (1000) can calculate a similarity for the rotation of the camera based on the difference between the rotation angle value of the detected camera pose and the rotation angle value of the preset camera pose.
[0064] In one embodiment, the electronic device (1000) may compare the pose of the detected person and the pose of the detected camera, respectively, with the pose of the preset person and the pose of the preset camera, respectively, included in the preset poses. In other words, when the pose of the camera that captured the first image is detected based on the first image, the electronic device (1000) may compare the pose of the detected person and the pose of the detected camera together with the preset poses.
[0065] In one embodiment, the electronic device (1000) can identify whether the pose of the detected camera corresponds to the preset pose by comparing the pose of the detected camera with a preset pose. In one embodiment, the electronic device (1000) can identify that the pose of the detected camera corresponds to the preset pose if the similarity between the pose of the detected camera and the preset pose is greater than or equal to a threshold value. In one embodiment, the electronic device (1000) can identify that the pose of the detected camera does not correspond to the preset pose if the similarity between the pose of the detected camera and the preset pose is less than the threshold value.
[0066] The specific details of how the electronic device (1000) compares the pose of the detected camera with the preset pose will be described again below with reference to FIG. 4.
[0067] In one embodiment, the electronic device (1000) may obtain a second image that is captured with the preset camera pose and in which the pose of a person included in the first image is converted into the preset camera pose based on the identification that the detected camera pose does not correspond to the preset camera pose. In other words, when the camera pose and the preset camera pose are compared, the electronic device (1000) may obtain a second image that is in which the camera pose of the first image is converted into the preset camera pose. In this case, the second image may be an image captured with the preset camera pose and may include a person in the preset pose and at least one fashion item worn by the person in the preset pose.
[0068] In one embodiment, the electronic device (1000) may generate a second image captured with a preset camera pose while maintaining the pose of the detected person based on the identification that the pose of the camera does not correspond to the preset camera pose. In one embodiment, if the electronic device (1000) does not detect the pose of the person or does not compare the pose of the detected person with the preset camera pose, only the pose of the detected camera may be compared with the preset camera pose. In this case, the electronic device (1000) may obtain a second image in which only the pose of the camera that captured the image is changed from the detected camera pose to the preset camera pose while maintaining the pose of the person included in the first image. However, the present invention is not necessarily limited thereto, and even if the electronic device (1000) does not compare the pose of the detected person or the pose of the detected camera with the preset pose, the second image may be captured with a preset camera pose based on the first image and the pose of the person included in the first image is converted to the preset camera pose.
[0069] The specific details of how the electronic device (1000) acquires the second image will be described again below with reference to FIG. 5.
[0070] In one embodiment, the electronic device (1000) can obtain a second image as an output of the image generation model by inputting a first image and a preset pose into an image generation model. In one embodiment, the image generation model may be an artificial intelligence model trained to output a second training image in which the pose of a person included in the first training image is converted into a training pose based on the first training image and the training pose.
[0071] The specific method by which the electronic device (1000) trains the image generation model will be described again below with reference to FIG. 8.
[0072] At step S230, the electronic device (1000) can detect at least one fashion item worn by a person converted into a preset pose based on the second image.
[0073] In one embodiment, the electronic device (1000) may detect a bounding box corresponding to at least one fashion item included in a second image. Here, detecting the bounding box may include obtaining information about the location of the bounding box surrounding at least one object and information about the label (or class) of the fashion item included in the bounding box. For example, the label of the fashion item may include a category or type of the fashion item, such as 'Top, Pants, Hat, Shoes, Bag, Accessories,' but is not necessarily limited to the above-described examples.
[0074] In one embodiment, the electronic device (1000) may display a second image through a display of the electronic device (1000). In one embodiment, when the electronic device (1000) acquires the second image in step S220, the electronic device (1000) may display the second image through a UI for acquiring a user input requesting acquisition of at least one attribute from the second image before performing step S230. In one embodiment, the electronic device (1000) may perform step S230 in response to acquiring a user input requesting acquisition of at least one attribute from the second image from the user through the UI.
[0075] In step S240, the electronic device (1000) may acquire at least one attribute of at least one detected fashion item. In one embodiment, the electronic device (1000) acquiring at least one attribute may include the electronic device (1000) extracting at least one attribute of each of at least one fashion item based on the at least one detected fashion item.
[0076] In one embodiment, the electronic device (1000) may obtain an image of a bounding box corresponding to each of the at least one detected fashion item and information about a label of the bounding box based on the detection result of at least one fashion item included in the second image. In one embodiment, the electronic device (1000) may extract at least one attribute of each of the at least one fashion item based on the information about the bounding box and label of each of the at least one detected fashion item. In one embodiment, at least one attribute that can be extracted for each label of the fashion item may be preset. For example, in the case of a top among fashion items, an attribute for a neckline may be extracted, and in the case of a hat among fashion items, an attribute for a shape may be extracted. However, the at least one attribute that can be extracted for each label of the fashion item is not necessarily limited to the above-described examples.
[0077] In one embodiment, the electronic device (1000) may detect at least one fashion item worn by a person included in the first image based on the identification that the pose of the detected person corresponds to a preset pose. In other words, if the electronic device (1000) identifies that the pose of the detected person corresponds to a preset pose, the electronic device (1000) may detect at least one fashion item included in the first image without performing steps S220 and S230. In one embodiment, if at least one fashion item worn by a person included in the first image is detected, the electronic device (1000) may acquire at least one attribute of the detected at least one fashion item.
[0078] The specific details of a method for an electronic device (1000) to detect at least one fashion item from an image and to obtain attributes of at least one detected fashion item will be described again below with reference to FIG. 6.
[0079] In one embodiment, the electronic device (1000) may provide results of an image retrieval based on at least one attribute of at least one acquired fashion item.
[0080] In one embodiment, the electronic device (1000) may store at least one attribute of at least one acquired fashion item in the memory. For example, the electronic device (1000) may acquire at least one attribute of at least one fashion item from each of a plurality of pre-stored images and store the acquired at least one attribute in the memory. In one embodiment, the electronic device (1000) may acquire a search word (or query) related to an attribute of a fashion item based on a user input. In one embodiment, the electronic device (1000) may identify an attribute corresponding to the search word related to the attribute of the fashion item among the pre-stored attributes of the fashion item. In one embodiment, the electronic device (1000) may acquire (or identify an image corresponding to) an identified attribute among a plurality of pre-stored images. In one embodiment, the electronic device (1000) may provide a result of an image search by outputting information about the identified image and the identified attribute.
[0081] Specific details of a method for providing image search results based on at least one attribute of at least one fashion item acquired by an electronic device (1000) will be described again below with reference to FIGS. 9a, 9b and 9c.
[0082] FIG. 3 is a diagram for explaining a method for an electronic device according to one embodiment of the present disclosure to detect a pose of a person included in an image.
[0083] Referring to FIG. 3, the pose estimation module (1110) can estimate the pose (310) of a person included in the first image (110) and the pose (320) of the camera that captured the first image (110) based on the first image (110). In one embodiment, the pose estimation module (1110) refers to a unit that processes one function or operation of the electronic device (1000), which may be implemented as hardware included in the electronic device (1000) or software stored in the memory of the electronic device (1000) or as a combination of hardware and software. That is, the operation and function of the pose estimation module (1110) can be understood as the function and operation of the electronic device (1000).
[0084] In one embodiment, a first image (110) including a person wearing at least one fashion item may be input to a pose estimation module (1110). In one embodiment, the first image may be an image stored in the electronic device (1000) or an image received from an external electronic device. However, the present invention is not necessarily limited to the above-described example, and an area corresponding to a person wearing at least one fashion item in the first image (110) may be cropped, and the cropped area may be input to the pose estimation module (1110), thereby performing operations and functions of the pose estimation module (1110) described below.
[0085] In one embodiment, the pose estimation module (1110) may output a pose (310) of a person included in the first image (110) based on the input first image (110). In one embodiment, the pose (310) of the person included in the first image (110) may include skeleton data including a plurality of key points corresponding to joints and preset body parts of the person included in the first image (110). For example, as shown in FIG. 3, the skeleton data may include information about the positions of 18 key points each corresponding to the nose, neck, both shoulders, both elbows, both wrists, both hips, both knees, both ankles, both eyes, and both ears of the person included in the first image (110), and information about links or edges connecting the key points. However, the joints and preset body parts of the person corresponding to the key points are not necessarily limited to the examples described above.
[0086] In one embodiment, skeleton data can be normalized based on the positions of specific key points and the lengths of edges connecting the key points. For example, the skeleton data can be normalized by setting the position of a key point corresponding to a human neck to 0,0,0 and the positions of other key points to 0,0,0. Alternatively, the length of an edge between a key point corresponding to a human nose and neck can be normalized by setting the length of other edges to 1. In one embodiment, normalization of skeleton data can be performed for each of an estimated human pose and a preset human pose. In one embodiment, when the electronic device (1000) performs normalization on the skeleton data, the operations and functions of the electronic device (1000) related to the skeleton data can be understood as operations and functions for the normalized skeleton data.
[0087] In one embodiment, the pose estimation module (1110) may output a pose (320) of a camera that captured the first image (110) based on the input first image (110). In one embodiment, the pose (320) of the camera that captured the first image (110) may include 6Dof data including position information and rotation information of the camera that captured the first image (110). For example, as shown in FIG. 3, the pose (320) of the camera may include position information indicating that the position of the camera is 0 m away from a reference position in the x-axis (left-right direction), +1.7 m away from a reference position in the y-axis (up-down direction), and +1 m away from a reference position in the z-axis (front-back direction). Here, the reference position may be the center of the subject. In addition, as shown in FIG. 3, the pose (320) of the camera may include rotation information indicating that the camera roll (rotation around the z-axis) is 0 degrees, the pitch (rotation around the x-axis) is -30 degrees, and the yaw (rotation around the y-axis) is 0 degrees from the reference direction. Here, the reference direction may be a direction looking straight ahead at the center of the subject. However, the data format of the camera position information and rotation information included in the pose of the camera is not necessarily limited to the above-described example.
[0088] In one embodiment, the pose estimation module (1110) may include an artificial intelligence model that receives an image as input and outputs at least one of a pose of a person included in the image and a pose of a camera that captured the image. In one embodiment, the pose estimation module (1110) may be an artificial intelligence model trained to output at least one of a pose of a person included in the training image and a pose of a camera that captured the training image based on a training image. In one embodiment, the pose estimation module (1110) may include a convolutional layer for extracting a feature map from the input image. In one embodiment, the pose estimation module (1110) may estimate a pose by outputting a probability of where a keypoint is located for each of a plurality of locations based on the extracted feature map, and identifying the location with the highest probability as the location of the keypoint. In one embodiment, the pose estimation module (1110) may estimate a pose of a camera by outputting a 6Dof representing a pose of a camera of an input image based on the extracted feature map.
[0089] In one embodiment, the pose estimation module (1110) may be trained based on a training data set for training the pose estimation module (1110). In one embodiment, the training data set for training the pose estimation module may include a plurality of training images, an actual pose of a person included in each of the plurality of training images, and an actual pose of a camera that captured each of the plurality of training images. In one embodiment, if the pose estimation module (1110) is a model trained to output only one of the person's pose or the camera's pose, the training data set may not include data that does not correspond to data output from either the person's actual pose or the camera's actual pose.
[0090] In one embodiment, the pose estimation module (1110) may be trained by calculating the difference between the estimated pose of a person and the pose of a camera based on training images during the training process and the actual pose of the person and the actual pose of the camera included in the training data set, and using a loss function based on the calculated difference. For example, the loss function may include a Mean Squared Error (MSE) loss function calculated based on the square of the difference in the positions of keypoints or the difference in the positions of cameras, or a Quaternion loss function calculated based on the difference in rotation, but is not necessarily limited to the examples described above.
[0091] FIG. 4 is a diagram illustrating a method for comparing a detected pose with a preset pose by an electronic device according to one embodiment of the present disclosure.
[0092] Referring to FIG. 4, the pose comparison module (1120) may compare at least one of the estimated human pose (310) and the estimated camera pose (320) with a preset pose. In one embodiment, the preset pose may include a preset human pose (410) compared with the estimated human pose (310) and a preset camera pose (420) compared with the estimated camera pose (320). In one embodiment, the pose comparison module (1120) refers to a unit that processes one function or operation of the electronic device (1000), which may be implemented as hardware included in the electronic device (1000) or software stored in a memory of the electronic device (1000) or as a combination of hardware and software. That is, the operation and function of the pose estimation module (1110) may be understood as the function and operation of the electronic device (1000).
[0093] In one embodiment, the estimated human pose (310) and the estimated camera pose (320) may be data acquired through the pose estimation module (1110) of FIG. 3. In one embodiment, when the pose estimation module (1110) outputs only one of the estimated human pose (310) and the estimated camera pose (320), the pose comparison module (1120) may compare only the preset pose corresponding to the output data, and may not compare the data that is not output with the preset pose.
[0094] In one embodiment, the estimated human pose (310) may be input to the pose comparison module (1120). In one embodiment, the pose comparison module (1120) may compare the input estimated human pose (310) with a preset human pose (410).
[0095] In one embodiment, the preset human pose (410) may include an upright posture of a human. The upright posture may include a posture in which the head and cervical spine of the human are in a straight line, both shoulders are balanced, and the pelvis is in a neutral position. In addition, the upright posture may include a posture in which both arms and both legs are straight. However, the present invention is not necessarily limited to the above-described examples, and the preset human pose (410) may include a pose of a human with the highest accuracy in extracting attributes of a fashion item among a plurality of preset human pose candidates. In one embodiment, the data type of the preset human pose (410) may be the same as the data type of the estimated human pose (310). For example, if the estimated human pose (310) includes skeleton data, the preset human pose (410) may also include skeleton data.
[0096] In one embodiment, the pose comparison module (1120) can calculate the similarity based on the difference between the positions of key points of the estimated human pose (310) and the positions of key points of the corresponding preset human pose (410). For example, the pose comparison module (1120) can calculate the Euclidean distance based on the difference between the positions of the first key point (310-1) corresponding to the right eye of the estimated human pose (310) and the positions of the first key point (410-1) corresponding to the right eye of the preset human pose (410). The pose estimation module (1110) can calculate the Euclidean distance between the corresponding key points for other key points in the same manner, and calculate the similarity between the estimated human pose (310) and the preset human pose (410) based on the average or sum of the calculated Euclidean distances. In this case, the larger the average or sum of the Euclidean distances, the lower the similarity is calculated, and conversely, the smaller the average or sum of the Euclidean distances, the higher the similarity is calculated.
[0097] In one embodiment, the pose comparison module (1120) can calculate the similarity based on the difference between the angles formed by three key points of the estimated human pose (310) and the angles formed by three key points of the corresponding preset human pose (410). For example, the pose comparison module (1120) can calculate the difference between the angles formed by the second key point (310-2) corresponding to the right hand of the estimated human pose (310), the third key point (310-3) corresponding to the right elbow, and the fourth key point (310-4) corresponding to the right shoulder, and the angles formed by the second key point (410-2) corresponding to the right hand of the preset human pose (410), the third key point (410-3) corresponding to the right elbow, and the fourth key point (410-4) corresponding to the right shoulder. The pose estimation module (1110) can calculate the difference in angles formed by three corresponding key points for other key points in the same manner, and calculate the similarity between the estimated human pose (310) and the preset human pose (410) based on the average or sum of the calculated angle differences.
[0098] In one embodiment, the pose comparison module (1120) may apply different weights to each of the plurality of keypoints, and calculate the similarity based on the difference in the positions of the weighted keypoints or the difference in the angles formed by three weighted keypoints. In one embodiment, different weights may be preset for each of the plurality of keypoints. For example, among the plurality of keypoints, the first keypoint (310-1) of the estimated human pose (310) and the first keypoint (410-1) of the preset human pose (410) may have low weights preset because they do not have a significant effect on accurately extracting the attributes of the fashion item corresponding to the upper body even if the positions are far apart from each other. On the other hand, since the difference in the positions of the second key point (310-2), the third key point (310-3), and the fourth key point (310-4) of the estimated human pose (310) among the plurality of key points and the difference in the angles formed by the second key point (410-2), the third key point (410-3), and the fourth key point (410-4) of the preset human pose (410) can have a great influence on accurately extracting the properties of the fashion item corresponding to the upper part, the weights for them can be preset high.
[0099] In one embodiment, the estimated camera pose (3200) may be input to the pose comparison module (1120). In one embodiment, the pose comparison module (1120) may compare the input estimated camera pose (320) with a preset camera pose (420).
[0100] In one embodiment, the preset camera pose (420) may include a frontal camera pose. The frontal camera pose may refer to a pose in which the camera captures a subject directly from the front. The frontal camera pose may refer to a pose in which the camera is positioned at a certain distance from the subject while being aligned horizontally and vertically without rotation or tilt. However, the present invention is not necessarily limited to the above-described example, and the preset camera pose (420) may include a pose of the camera with the highest accuracy in extracting attributes of a fashion item among a plurality of preset camera pose candidates. In one embodiment, the data type of the preset human pose (410) may be the same as the data type of the estimated human pose (310). For example, if the estimated camera pose (320) includes 6Dof data including camera position information and rotation information, the preset camera pose (420) may also include 6Dof data including camera position information and rotation information.
[0101] In one embodiment, the pose comparison module (1120) may calculate the similarity for the camera position based on the difference between the 3D position coordinate value of the estimated camera pose (320) and the 3D position coordinate value of the preset camera pose (420). For example, the pose comparison module (1120) may calculate the Euclidean distance between the 3D position coordinate value (0, 1.7, 1) of the estimated camera pose (320) and the 3D position coordinate value (0, 1, 1) of the preset camera pose (420), and may calculate the similarity based on the calculated Euclidean distance. In this case, the larger the Euclidean distance, the smaller the similarity for the camera position is calculated, and the smaller the Euclidean distance, the larger the similarity for the camera position can be calculated.
[0102] In one embodiment, the pose comparison module (1120) can calculate the similarity for the rotation of the camera based on the difference between the rotation angle value of the estimated pose (320) of the camera and the rotation angle value of the preset pose (420) of the camera. For example, the pose comparison module (1120) can calculate the difference in the quaternion angle or the Euler angle between the rotation angle value (0, -30, 0) of the pose (320) of the camera and the rotation angle value (0, 0, 0) of the preset pose (420) of the camera, and can calculate the similarity based on the difference in the calculated angles. In this case, the larger the difference in angles, the smaller the similarity for the rotation of the camera is calculated, and the smaller the difference in angles, the larger the similarity for the rotation of the camera can be calculated.
[0103] In one embodiment, the pose comparison module (1120) may output a comparison result (430) based on the calculated similarity. In one embodiment, the comparison result (430) may include at least one of whether the estimated human pose (310) corresponds to the preset human pose (310) and whether the estimated camera pose (320) corresponds to the preset camera pose (420). For example, if the similarity between the estimated human pose (310) and the preset human pose (410) is less than a threshold, the pose comparison module (1120) may identify that the estimated human pose (310) does not correspond to the preset human pose (410) and output data indicating that the human pose is 'mismatched' as the comparison result (430). In addition, if the similarity between the estimated pose (320) of the camera and the preset pose (420) of the camera estimated by the comparison result (430) is less than a threshold value, the pose comparison module (1120) can identify that the estimated pose (320) of the camera does not correspond to the preset pose (420) of the camera and output data indicating that the pose of the camera is ‘inconsistent’ by the comparison result (430).
[0104] FIG. 5 is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to acquire a second image.
[0105] Referring to FIG. 5, the image generation module (1130) may generate a second image (120) based on the first image (110). In one embodiment, a person and at least one fashion item included in the second image (120) may correspond to a person and at least one fashion item included in the first image (110). In other words, the electronic device (1000) may generate a second image (120) in which at least one of a pose of a person included in the first image (110) and a pose of a camera that captured the first image (110) is changed, while maintaining the identity of the person and at least one fashion item included in the first image (110). In one embodiment, the image generation module (1130) refers to a unit that processes one function or operation of the electronic device (1000), which may be implemented as hardware included in the electronic device (1000) or software stored in a memory of the electronic device (1000) or as a combination of hardware and software. That is, the operation and function of the image generation module (1130) can be understood as the function and operation of the electronic device (1000).
[0106] In one embodiment, the image generation module (1130) can generate the second image (120) based on the comparison result (430) output from the pose comparison module (1120) in FIG. 4. In one embodiment, the comparison result (430) can be input to the image generation module (1130), or a signal for activating an operation or function of the image generation module (1130) based on the comparison result (430) can be input to the image generation module (1130).
[0107] In one embodiment, the image generation module (1130) may generate a second image (120) including a person whose estimated human pose (310) is converted into a preset human pose (410) based on the comparison result (430). For example, if the image generation module (1130) determines that the estimated human pose (310) does not correspond to the preset human pose (410), the image generation module (1130) may generate a second image (120) based on the first image (110) and the preset human pose (410). In this case, the second image (120) may include a person included in the first image (110) in the preset human pose (410) and at least one fashion item included in the first image (110) worn by a person included in the first image (110) in the preset human pose (410).
[0108] In one embodiment, the image generation module (1130) may generate a second image (120) captured with a preset camera pose (420) based on the comparison result (430). For example, if the image generation module (1130) determines that the estimated camera pose (320) does not correspond to the preset camera pose (420), the image generation module (1130) may generate a second image (120) based on the first image (110) and the preset camera pose (420). In this case, the second image (120) may include a person included in the first image (110) captured with the preset camera pose (420) and at least one fashion item worn by the person included in the first image (110) captured with the preset camera pose (420).
[0109] In one embodiment, the image generation module (1130) may generate a second image (120) including a person captured with a preset camera pose (420) based on the input comparison result and whose estimated person pose (310) is converted to the preset person pose (410). For example, if the image generation module (1130) identifies that the estimated person pose (310) does not correspond to the preset person pose (410) and the estimated camera pose (320) does not correspond to the preset camera pose (420), the image generation module (1130) may generate a second image (120) based on the first image (110), the preset person pose (410), and the preset camera pose (420). In this case, the second image (120) may be captured with a preset camera pose (420) and may include a person included in the first image (110) in a preset person pose (410) and at least one fashion item included in the first image (110) worn by a person included in the first image (110) in a preset person pose (410) and captured with a preset camera pose (420).
[0110] In one embodiment, the image generation module (1130) may include an artificial intelligence model trained to output a second image (120) by inputting a first image (110) and a preset pose including at least one of a preset human pose (410) and a preset camera pose (420). Accordingly, the image generation module (1130) may be referred to as an image generation model. In one embodiment, the image generation model may be an artificial intelligence model trained to output a second training image including a human whose pose included in the first training image is converted into the pose of the training human based on the first training image and the pose of the training human. In one embodiment, the image generation model may be an artificial intelligence model trained to output a second training image captured with the pose of the training camera based on the first training image and the pose of the training camera. In one embodiment, the image generation model may be an artificial intelligence model configured to output a second training image, which is captured in the pose of the training camera based on a first training image, a pose of the training person, and a pose of the training camera, and which includes a person whose pose included in the first training image is converted to the pose of the training person.
[0111] In one embodiment, the image generation module (1130) may be trained based on a training data set for training the image generation module (1130). Details regarding this will be further described below with reference to FIG. 8.
[0112] In this way, the electronic device (1000) according to one embodiment of the present disclosure can generate a second image (120) based on at least one of a first image (110), a preset human pose (410), and a preset camera pose (420). In this case, the second image (120) may include a human pose that is more suitable for extracting attributes of a fashion item than the first image (110), or may be an image captured with a camera pose.
[0113] FIG. 6 is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to detect a fashion item from an image and obtain attributes of the detected fashion item.
[0114] Referring to FIG. 6, the fashion item detection module (1140) can detect at least one fashion item included in the second image (120) based on the second image (120). In addition, the fashion item attribute extraction module (1150) can extract at least one attribute of each of the at least one detected fashion items. In one embodiment, the fashion item detection module (1140) and the fashion item attribute extraction module (1150) refer to a unit that processes one function or operation of the electronic device (1000), which may be implemented as hardware included in the electronic device (1000) or software stored in a memory of the electronic device (1000) or as a combination of hardware and software. That is, the operation and function of the fashion item detection module (1140) and the fashion item attribute extraction module (1150) may be understood as the function and operation of the electronic device (1000).
[0115] In one embodiment, the second image (120) input to the fashion item detection module (1140) may be an image acquired through the image generation module (1130) of FIG. 4. In one embodiment, if the pose comparison module (1120) of FIG. 4 identifies that the estimated human pose (310) corresponds to the preset human pose (410), the first image (110) may be input to the fashion item detection module (1140) instead of the second image (120). In one embodiment, if the pose comparison module (1120) identifies that the estimated camera pose (320) corresponds to the preset camera pose (420), the first image (110) may be input to the fashion item detection module (1140) instead of the second image (120). The fashion item detection module (1140) can detect at least one fashion item included in the input image based on the input first image (110) or second image (120). For the convenience of explaining the invention, the following description will be made based on the operation based on the second image (120).
[0116] In one embodiment, the fashion item detection module (1140) may extract a bounding box corresponding to at least one fashion item included in the second image (120) based on the second image (120). In one embodiment, the fashion item detection module (1140) may obtain information about the location of the bounding box corresponding to the at least one extracted fashion item and information about the label (or class) of the fashion item included in the bounding box. Here, the label may mean the category of the fashion item. For example, the fashion item detection module (1140) may obtain a first bounding box (611) including the upper garment and a second bounding box (612) including the lower garment in the second image (120) based on the second image (120). In one embodiment, the fashion item detection module (1140) can obtain information that the label of the fashion item included in the first bounding box (611) is 'top' and the label of the fashion item included in the second bounding box (612) is 'pants'.
[0117] In one embodiment, the fashion item attribute extraction module (1150) may extract attributes corresponding to each of at least one fashion item based on a bounding box corresponding to at least one fashion item included in the second image (120). In one embodiment, an image of a bounding box including at least one fashion item and information about a label of the bounding box may be input to the fashion item attribute extraction module (1150). In this case, the image of the bounding box may be obtained by cropping an area corresponding to the bounding box from the second image (120) based on information about the location of the bounding box. For example, the fashion item attribute extraction module (1150) may obtain attributes (621) of an upper garment included in a first bounding box (611) based on the first bounding box (611), and may obtain attributes (622) of a lower garment included in a second bounding box (612) based on the second bounding box (612). In one embodiment, the attributes obtained from the fashion item attribute extraction module (1150) may include label information of the detected fashion item.
[0118] In one embodiment, the fashion item detection module (1140) may include an artificial intelligence model trained to input an image and output a bounding box surrounding a fashion item included in the image. In one embodiment, the fashion item attribute extraction module (1150) may include an artificial intelligence model trained to input a bounding box and output attributes of a fashion item included in the bounding box. In one embodiment, the fashion item detection module (1140) and the fashion item attribute extraction module (1150) may include a convolutional layer for extracting a feature map from the input image. In one embodiment, the fashion item detection module (1140) may identify the location of the bounding box and the label of the fashion item included in the bounding box based on the extracted feature map. In one embodiment, the fashion item attribute extraction module (1150) may identify attributes of the fashion item included in the bounding box based on the extracted feature map.
[0119] In one embodiment, the fashion item detection module (1140) and the fashion item attribute extraction module (1150) may be implemented as a single integrated module. In this case, the integrated module may include an artificial intelligence model trained to input an image and output labels and attributes of fashion items contained in the image. For example, the integrated module may input a second image (120) and output attributes of the upper garment (621) and attributes of the lower garment (622).
[0120] In one embodiment, the fashion item detection module (1140) and the fashion item attribute extraction module (1150) may be trained based on a training data set for training each module. In one embodiment, the training data set for training the fashion item detection module (1140) may include a plurality of training images and information about bounding boxes and labels surrounding the fashion items included in each of the plurality of training images. In one embodiment, the training data set for training the fashion item attribute extraction module (1150) may include images of each of the plurality of fashion items and information about attributes of the plurality of fashion items.
[0121] In one embodiment, the fashion item detection module (1140) and the fashion item attribute extraction module (1150) may be trained using a loss function calculated based on the difference between the positions, labels, and attributes of the bounding boxes output during the training process and the positions, labels, and attributes of the bounding boxes included in the training data set. For example, the loss function may include an L1 loss function, a cross-entropy loss function, etc., but is not necessarily limited to the examples described above.
[0122] FIG. 7 is a flowchart illustrating a method for an electronic device according to one embodiment of the present disclosure to select an image for obtaining attributes of a fashion item.
[0123] In step S710, the electronic device (1000) can estimate at least one of a pose of a person wearing at least one fashion item included in the first image (110) and a pose of a camera that captured the first image (110). Since the operation of the electronic device (1000) estimating the pose of the person and the pose of the camera based on the first image (110) corresponds to the operation of the pose estimation module (1110) of FIG. 3, a redundant description will be omitted.
[0124] In step S720, the electronic device (1000) can identify whether the estimated pose matches a preset pose. In one embodiment, the electronic device (1000) can identify whether the estimated human pose corresponds to a preset human pose among the preset poses, and can identify whether the estimated camera pose corresponds to a preset camera pose among the preset poses. In one embodiment, the electronic device (1000) can identify that the estimated pose does not match the preset pose if it is identified that the estimated human pose does not correspond to the preset human pose or if it is identified that the estimated camera pose does not correspond to the preset camera pose. However, the present invention is not necessarily limited to the above-described example, and the electronic device (1000) can also identify that the estimated pose does not match the preset pose if it is identified that both the estimated human pose and the estimated camera pose do not correspond to the preset human pose and the preset camera pose, respectively. Since the operation of the electronic device (1000) comparing the estimated pose with the preset pose corresponds to the operation of the pose comparison module (1120) of FIG. 4, a redundant description will be omitted.
[0125] At step S730, if the electronic device (1000) determines that the estimated pose does not match the preset pose (S720-N), it can generate a second image (120). Since the operation of the electronic device (1000) generating the second image (120) corresponds to the operation of the image generation module (1130) of FIG. 5 generating the second image (120), a redundant description will be omitted.
[0126] At step S740, if the electronic device (1000) identifies that the estimated pose matches the preset pose (S720-Y), the electronic device (1000) can detect at least one fashion item included in the first image (110). At step S740, if the second image (120) is generated through step S730, the electronic device (1000) can detect at least one fashion item included in the second image (120). Since the operation of the electronic device (1000) detecting at least one fashion item included in the image corresponds to the operation of the fashion item detection module (1140) of FIG. 6, a redundant description thereof will be omitted.
[0127] At step S750, the electronic device (1000) can extract attributes of at least one detected fashion item. Since the operation of the electronic device (1000) to extract attributes of a fashion item corresponds to the operation of the fashion item attribute extraction module (1150) of FIG. 6 , a redundant description will be omitted.
[0128] In this way, the electronic device according to one embodiment of the present disclosure can generate the second image (120) based on whether the estimated pose matches the preset pose. In other words, if the pose of the person included in the first image (110) or the pose of the camera that captured the first image (110) matches the preset pose and is suitable for extracting the attributes of the fashion item, the electronic device (1000) can save system resources by not generating the second image (120). On the other hand, if the pose of the person included in the first image (110) or the pose of the camera that captured the first image (110) does not match the preset pose and is therefore unsuitable for extracting the attributes of the fashion item, the electronic device (1000) can generate the second image (120) and accurately extract the attributes of the fashion item.
[0129] FIG. 8 is a diagram illustrating training of an image generation module according to one embodiment of the present disclosure.
[0130] Referring to FIG. 8, the image generation module (1130) may be trained based on a training data set for training the image generation module (1130). In one embodiment, the training data set may include a plurality of training input images (810) and a plurality of training input attributes (820). The plurality of training input attributes (820) may include attributes of at least one fashion item worn by a person included in each of the plurality of training input images (810). In other words, the plurality of training input attributes (820) may mean ground-truth data of attributes of at least one fashion item worn by a person included in each of the plurality of training input images (810). For example, a first training attribute (820-1), which is one of the plurality of training input attributes (820), may include attributes of a top and pants worn by a person included in the first training input image (810-1), which is one of the plurality of training input images (810).
[0131] In one embodiment, the image generation module (1130) may output a plurality of training output images (830) based on each of a plurality of training images (810) during the training process. For example, the image generation module (1130) may receive a first training input image (810-1) and output a first training output image (830-1). Since the plurality of training input images (810), the plurality of training output images (830), and the plurality of training human poses correspond to the first image (110) input to the image generation module (1130) of FIG. 5 and the second image (120) output from the image generation module (1130), a redundant description thereof will be omitted.
[0132] In one embodiment, a plurality of training output attributes (840) may be extracted based on a plurality of training output images (830), each of which includes attributes of at least one fashion item worn by a person included in each of the plurality of training output images (830). For example, a first training output attribute (840-1), which is one of the plurality of training output attributes (840), may include attributes of a top and pants worn by a person included in the first training output image (830-1) extracted from the first training output image (830-1), which is one of the plurality of training output images (830).
[0133] In one embodiment, the plurality of training output attributes (840) may be obtained by the fashion item detection module (1140) and the fashion item attribute extraction module (1150) of FIG. 6, or may be obtained by a separate artificial intelligence model trained to perform the same operations and functions as the fashion item detection module (1140) and the fashion item attribute extraction module (1150) of FIG. 6.
[0134] In one embodiment, the image generation module (1130) may be an artificial intelligence model trained using a loss function calculated based on the difference between a plurality of training input attributes (820) and a plurality of training output attributes (840). For example, the image generation module (1130) may be trained using a loss function (850) calculated based on the difference between a first training input attribute (820-1) and a first training output attribute (820-1). In one embodiment, the loss function (850) may include, but is not necessarily limited to, a cross-entropy loss function, and may further include various loss functions calculated based on the difference between the input attribute and the inferred attribute.
[0135] In one embodiment, the training data set may include a plurality of training person poses including a pose of a person included in each of a plurality of training input images (810) and a plurality of training camera poses including a pose of a camera that captured each of the plurality of training images (810). In one embodiment, the image generation module (1130) may output a plurality of training output images (830) based on at least one of the plurality of training images (810), the plurality of training person poses, and the plurality of training camera poses. In one embodiment, the image generation module (1130) may estimate the poses of the plurality of people included in each of the plurality of training output images (830) and the poses of the plurality of cameras that captured each of the plurality of training output images (830) based on the plurality of training output images (830).
[0136] In one embodiment, poses of multiple people and poses of multiple cameras may be acquired by the pose estimation module (1110) of FIG. 3, or may be acquired by a separate artificial intelligence model trained to perform the same actions and functions as the pose estimation module (1110).
[0137] In one embodiment, the image generation module (1130) may be trained based on at least one of a loss function calculated based on the differences between the estimated poses of a plurality of people and the poses of a plurality of training people, and a loss function calculated based on the differences between the estimated poses of a plurality of cameras and the poses of a plurality of training cameras. Since this may correspond to the loss function used in the training of the pose estimation module (1110) of FIG. 3, a redundant description thereof will be omitted.
[0138] In this way, when training the image generation module (1130), the electronic device (1000) according to one embodiment of the present disclosure can train the image generation module (1130) using a loss function calculated based on the difference between the attributes of an actual fashion item included in an input image and the attributes of the fashion item inferred from the generated image. Accordingly, the image generation module (1130) can be trained to generate an image from which the attributes of the fashion item can be more accurately extracted, which can improve the accuracy of the electronic device (1000) in extracting attributes of the fashion item in the inference step.
[0139] FIG. 9A is a diagram illustrating a method for an electronic device according to an embodiment of the present disclosure to provide image search results based on attributes of a fashion item.
[0140] Referring to FIG. 9A, the electronic device (1000) may extract a plurality of attributes (920) corresponding to each of the plurality of images (910) from the plurality of stored images (910). For example, the electronic device (1000) may extract a first attribute (920-1) including at least attributes of upper and lower garments included in the first stored image (910-1), based on a first stored image (910-1) that is one of the plurality of stored images (910). In one embodiment, the plurality of attributes (920) corresponding to each of the plurality of stored images (910) may include attributes extracted by acquiring each of the plurality of stored images (910) as a first image in step S210 of FIG. 2 and performing steps S210 to S240, and a label of a fashion item from which the attributes are extracted. Accordingly, a detailed description of a specific method by which the electronic device (1000) extracts attributes from an image will be omitted.
[0141] In one embodiment, the electronic device (1000) may store the extracted plurality of attributes (920) in a database (930) of the electronic device (1000). In one embodiment, the database (930) may be implemented in a memory of the electronic device (1000) or in a storage space connected to the electronic device (1000). In one embodiment, the electronic device (1000) may map or associate each of the extracted plurality of attributes (920) to a corresponding plurality of stored images (910) and store the extracted plurality of attributes (920) together with the plurality of stored images (910) in the database (930). For example, the extracted plurality of attributes (920) may be stored in the database (930) together with the plurality of stored images (910) in the form of metadata of the plurality of stored images (910).
[0142] In one embodiment, the electronic device (1000) may obtain a search word (941) related to an attribute of a fashion item. The search word (941) related to an attribute of a fashion item may be directly input through a user input unit of the electronic device (1000) or may be obtained as a result of voice recognition obtained based on a voice signal of the user. In one embodiment, the search word (941) related to an attribute of a fashion item may include at least one of a label of a fashion item that can be detected by the electronic device (1000) and attributes of the fashion item that can be extracted from the electronic device (1000). In one embodiment, the label of the fashion item that can be detected by the electronic device (1000) may correspond to a label of a fashion item included in a training data set of a fashion item detection module (1140), and the attribute of the fashion item that can be extracted may correspond to an attribute of a fashion item included in a training data set of a fashion item attribute extraction module (1150). If a word or sentence included in a search term (941) related to an attribute of a fashion item does not match a label or attribute of a fashion item that can be extracted from the electronic device (1000), the electronic device (1000) can convert the word or sentence into a label or attribute of a fashion item that can be extracted from the electronic device (1000) through query expansion or query correction.
[0143] In one embodiment, the electronic device (1000) may obtain an image and an attribute corresponding to a search word (941) related to an attribute of an acquired fashion item from a database (930). In one embodiment, the electronic device (1000) may identify an attribute matching the attribute of the fashion item included in the search word (941) among a plurality of attributes (920) stored in the database (930), and may identify an image corresponding to the identified attribute among a plurality of images (910) stored in the database (930). In one embodiment, when there are multiple attributes matching the attribute included in the search word (941) among the plurality of attributes (920), an attribute having the highest degree of matching may be identified. For example, if the search word (941) related to the attribute of the fashion item includes 'pink', 'collar', and 'top', the electronic device (1000) can identify a fashion item whose label is 'top' that includes all of the search words (941) related to the attribute of the fashion item among the plurality of attributes (920) stored in the database (930), a first attribute (920-1) in which the color is 'pink' and the neckline is 'collar' as attributes of the fashion item of the top, and a first stored image (910-1) corresponding to the first attribute (920-1). However, the present invention is not necessarily limited to the above-described example, and if there is no attribute among the plurality of attributes (920) that includes all of the search words (941) related to the attribute of the fashion item, an attribute that includes only some of the attributes and an image corresponding to the attribute may be identified.
[0144] In one embodiment, the electronic device (1000) may output information about the identified image and attributes based on a search term (941) related to attributes of a fashion item. In one embodiment, the electronic device (1000) may display information about the identified image and attributes through an image search result UI (950). In one embodiment, the image search result UI (950) may include text (951) indicating information about the first stored image (910-1), which is the identified image, and the first stored attribute (920-1), which is the identified attribute.
[0145] In one embodiment, the image search result UI (950) may include a button (952) for obtaining a request for display of another image from a user. In one embodiment, in response to obtaining an input for selecting the button (952) from the user, the electronic device (1000) may identify an attribute matching an attribute included in a search term (941) related to an attribute of a fashion item and an image corresponding to the matching attribute among the remaining attributes except for the first attribute (920-1) among the plurality of attributes (920) stored in the database (930), and display information about the identified image and attribute through the image search result UI (950).
[0146] FIG. 9b is a diagram for explaining a method for an electronic device according to an embodiment of the present disclosure to recommend fashion coordination based on attributes of a fashion item.
[0147] Referring to FIG. 9B, the electronic device (1000) can extract a second attribute (920-2), a third attribute (920-3), and a fourth attribute (920-4) from the second stored image (910-2), the third stored image (910-3), and the fourth stored image (910-4), respectively, and store them in the database (930). Since the second stored image (910-2), the third stored image (910-3), and the fourth stored image (910-4) are included in the plurality of stored images (910) of FIG. 9A, and the second attribute (920-2), the third attribute (920-3), and the fourth attribute (920-4) are included in the plurality of attributes (920) of FIG. 9A, a redundant description thereof will be omitted.
[0148] In one embodiment, the electronic device (1000) may obtain a coordination recommendation command (942). In one embodiment, the electronic device (1000) may obtain an image and an attribute corresponding to the obtained coordination recommendation command (942) from the database (930). In one embodiment, the coordination recommendation command (942) may be directly input through a user input unit of the electronic device (1000) or may be obtained as a result of voice recognition obtained based on a user's voice signal. In one embodiment, the coordination recommendation command (942) may include at least one of a query for activating a coordination recommendation function, a label of a fashion item that can be detected by the electronic device (1000), and attributes of the fashion item that can be extracted by the electronic device (1000). In one embodiment, when the electronic device (1000) identifies that the coordination recommendation command (942) includes a query that activates a coordination recommendation function, the electronic device (1000) may identify an attribute that matches the attribute included in the coordination recommendation command (942) among a plurality of attributes (920) stored in the database (930), and may identify an image corresponding to the identified attribute. For example, when the coordination recommendation command (942) includes the sentence “Recommend a black short-sleeved T-shirt,” the electronic device (1000) may perform query analysis to extract the attributes of “coordination recommendation,” which is a query that activates the coordination recommendation function, and “top,” which is a label of a fashion item, and a fashion item whose color is “black” and whose sleeve length is “short.”
[0149] In one embodiment, the electronic device (1000) may identify an attribute including an attribute of the extracted fashion item among a plurality of attributes (920) stored in the database (930) based on a query extracted based on a coordination recommendation command (942) and an attribute of the fashion item, and may identify an image corresponding to the identified attribute. For example, the electronic device (1000) may identify a second attribute (920-2), a third attribute (930-3), and a fourth attribute (940-4) including an attribute that the label of the fashion item included in the coordination recommendation command (942) is 'top', the color is 'black', and the sleeve length is 'short' among the plurality of attributes (920), and may identify a second stored image (910-2), a third stored image (910-3), and a fourth stored image (910-4) corresponding to the identified attributes among a plurality of images (910) stored in the database (930).
[0150] In one embodiment, the electronic device (1000) may output information about the identified image and attribute based on the coordination recommendation command (942). In one embodiment, the electronic device (1000) may display information about the identified image and attribute through the coordination recommendation UI (960). In one embodiment, the coordination recommendation UI (960) may include text indicating information about the second stored image (910-2), which is the identified image, and the second attribute (920-2), which is the identified attribute. In one embodiment, the text indicating information about the second attribute (920-2), which is the identified attribute, may include other labels and attributes for other labels excluding the labels included in the coordination recommendation command (942) in the second attribute (920-2). For example, text representing information about the second attribute (920-2) may include attributes (961) for 'pants' and 'pants', and attributes (962) for 'hat' and 'hat', which are labels of fashion items included in the coordination recommendation command (942) in the second attribute (920-2) excluding 'top'.
[0151] In one embodiment, the coordination recommendation UI (960) may include a button (963) for obtaining a request for another coordination from the user. In one embodiment, the electronic device (1000) may, in response to obtaining a user input of selecting the button (963) from the user, display information about identified images and attributes other than the images and attributes included in the coordination recommendation UI (960) through the coordination recommendation UI (960). For example, in response to obtaining a user input of selecting the button (963) from the user, the electronic device (1000) may display a third stored image (910-3) and a third attribute (920-3) through the coordination recommendation UI (960), or display a fourth stored image (910-4) and a fourth attribute (920-4) through the coordination recommendation UI (960).
[0152] FIG. 9c is a diagram illustrating a method for an electronic device according to one embodiment of the present disclosure to display an image generated based on an image generation module.
[0153] Referring to FIG. 9C, the electronic device (1000) can obtain a user input for determining a first image through an image selection UI (970). In one embodiment, the image selection UI (970) can include a first button (971) for obtaining a user input for requesting extraction of the attributes of the fashion item from a first stored image (910-1) from which the attributes of the fashion item are to be extracted, and a second button (972) for obtaining a user input for requesting generation of an image in which a pose of a person included in the first stored image (910-1) is converted into a preset pose based on the first stored image (910-1).
[0154] In one embodiment, the electronic device (1000) may detect at least one fashion item worn by a person included in the first stored image (910-1) based on the first stored image (910-1) in response to obtaining a user input of selecting the first button (971), and extract attributes of the detected item. Here, the operation of detecting the fashion item based on the first stored image (910-1) and extracting attributes of the detected item may correspond to the operation of detecting at least one fashion item worn by a person included in the first image based on the first image in the description with respect to FIG. 3.
[0155] In one embodiment, the electronic device (1000) may obtain a first generated image (911) based on the first stored image (910-1) in response to obtaining a user input selecting the second button (972). In one embodiment, the electronic device (1000) may display the first generated image (911) through an image output UI (980). In one embodiment, the image output UI (980) may include a third button (981) for obtaining a user input requesting extraction of an attribute of a fashion item and a fourth button (982) for obtaining a user input requesting generation of a new image. Here, since the operation of the electronic device (1000) obtaining the first generated image (911) based on the first stored image (910-1) may correspond to the operation of the image generation module (1130) of FIG. 5, a redundant description thereof will be omitted.
[0156] In one embodiment, the electronic device (1000) may detect at least one fashion item included in the first generated image (911) based on the first generated image (911) in response to obtaining a user input selecting the third button (981), and extract attributes of the detected fashion item. In one embodiment, the electronic device (1000) may obtain a second generated image (912) based on the first stored image (910-1) in response to obtaining a user input selecting the fourth button (982). In one embodiment, the electronic device (1000) may display the second generated image (912) through the image output UI (980). Here, the second generated image (912) may be a different image from the first generated image (911). For example, if the image generation module (1130) of FIG. 5 includes an artificial intelligence model, output data may vary for the same input data if the seed value that determines the initial state of the random number generator of the artificial intelligence model varies. That is, the electronic device (1000) may regenerate an image suitable for extracting attributes of a fashion item in response to obtaining a user input for selecting the fourth button (982).
[0157] In this way, the electronic device (1000) according to one embodiment of the present disclosure can generate a new pose-converted image based on a user input requesting the reproduction of a pose-converted image. Accordingly, among the images output by the image generation module (1130), an image suitable for extracting attributes of a fashion item is selected, and attributes of the fashion item can be extracted based on the selected image.
[0158] FIG. 10 is a detailed configuration diagram of an electronic device according to one embodiment of the present disclosure.
[0159] Referring to FIG. 10, the electronic device (1000) may include a memory (1100), a display (1200), a communication interface (1300), an input interface (1400), an output interface (1500), a camera (1600), and a processor (1700). The memory (1100), the display (1200), the communication interface (1300), the input interface (1400), the output interface (1500), the camera (1600), and the processor (1700) may each be electrically and / or physically connected to each other.
[0160] The components illustrated in FIG. 10 are merely according to one embodiment of the present disclosure, and the components included in the electronic device (1000) are not limited to those illustrated in FIG. 10. The electronic device (1000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 10, and may further include components not illustrated in FIG. 10.
[0161] The memory (1100) may store instructions or program codes for performing functions or operations of the electronic device (1000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (1100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or an assembler.
[0162] In one embodiment, the memory (1100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a RAM (Random Access Memory), a SRAM (Static Random Access Memory), a ROM (Read-Only Memory), an EEPROM (Electrically Erasable Programmable Read-Only Memory), a PROM (Programmable Read-Only Memory), a Mask ROM, a Flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD). The memory (1100) may not exist separately and may be configured to be included in the processor (1700). The memory (1100) may be configured as a volatile memory, a nonvolatile memory, or a combination of a volatile memory and a nonvolatile memory. A program or at least one instruction for performing operations according to embodiments described below may be stored in the memory (1100). The memory (1100) may also provide stored data to the processor (1700) at the request of the processor (1700).
[0163] In one embodiment, the memory (1100) may include at least one of a pose estimation module (1110), a pose comparison module (1120), an image generation module (1130), a fashion item detection module (1140), and a fashion item attribute extraction module (1150). In one embodiment, the memory (1100) may include a preset human pose (410) and a preset camera pose (420). In one embodiment, the memory (1100) may include a plurality of stored images (910) and a plurality of attributes (920). In one embodiment, the memory (1100) may include a training data set for training at least one of the pose estimation module (1110), the pose comparison module (1120), the image generation module (1130), the fashion item detection module (1140), and the fashion item attribute extraction module (1150). However, it is not necessarily limited to the above-described examples, and the memory (1100) may further include various data necessary to perform the operations and functions of the electronic device (1000) described in the present disclosure.
[0164] The display (1200) is a component for displaying images and / or videos. In one embodiment, the display (1200) may be configured as a physical device including at least one of a liquid crystal display, a thin film transistor-liquid crystal display, an organic light-emitting diode (OLED), a flexible display, a 3D display, and an electrophoretic display. In one embodiment, the display (1200) may display at least one of an image search result UI (950), a coordination recommendation UI (960), an image selection UI (970), and an image output UI (980) based on a signal received from the processor (1700). However, the present invention is not limited to the above-described example, and the display (1200) may display various UIs and data necessary for performing operations and functions of the electronic device (1000) described in the present disclosure based on a signal received from the processor (1700).
[0165] The communication interface (1300) is a component for the electronic device (1000) to communicate with an external electronic device. In one embodiment, the communication interface (1300) may perform data communication between the electronic device (1000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.
[0166] In one embodiment, the communication interface (1300) may receive an image including a person wearing at least one fashion item from an external electronic device or transmit the image to the external electronic device. In one embodiment, the communication interface (1300) may receive a preset pose of a person (410) and a preset pose of a camera (420) from the external electronic device or transmit the image to the external electronic device. In one embodiment, the communication interface (1300) may receive data input to or output from at least one of a pose estimation module (1110), a pose comparison module (1120), an image generation module (1130), a fashion item detection module (1140), and a fashion item attribute extraction module (1150) from the external electronic device or transmit the data to the external electronic device. In one embodiment, the communication interface (1300) may receive an attribute of a fashion item extracted from an image from the external electronic device or transmit the attribute to the external electronic device. However, the present invention is not limited to the examples described above, and the communication interface (1300) may receive or transmit various data necessary to perform the operations and functions of the electronic device (1000) described in the present disclosure from or to an external electronic device.
[0167] The input interface (1400) is a component for receiving various user inputs. In one embodiment, the input interface (1400) may include a touch panel, a physical button, a microphone, etc. In one embodiment, information input through the input interface (1400) may be provided to the processor (1700). In one embodiment, a search term (941) related to the attributes of a fashion item may be obtained through the input interface (1400). In one embodiment, a coordination recommendation command (942) may be obtained through the input interface (1400). In one embodiment, a user input for selecting at least one of a button (952) of an image search result UI (950), a button (963) of a coordination recommendation UI (960), a first button (971) and a second button (972) of an image selection UI (970), a third button (981) and a fourth button (982) of an image output UI (980) may be obtained through the input interface (1400). However, the present invention is not limited to the examples described above, and the input interface (1400) can obtain various data necessary to perform the operations and functions of the electronic device (1000) described in the present disclosure.
[0168] The output interface (1500) is a component for the electronic device (1000) to provide various information to the user. In one embodiment, the electronic device (1000) may include a speaker, which is a component that outputs sound. In one embodiment, the output interface (1500) may output voice corresponding to text displayed through the user interface based on a signal received from the processor (1700). However, the present invention is not necessarily limited to the above-described example, and the output interface (1500) may output various voices or information for performing the operations and functions of the electronic device (1000) described in the present disclosure.
[0169] The camera (1600) is a component that can acquire information about the real environment by photographing the real environment of the electronic device (1000). In one embodiment, the camera (1600) may include a single camera or three or more multi-cameras. In one embodiment, the camera (1600) may include a lens module, an image sensor, and an image processing module. The camera (1600) can acquire still images or moving images obtained by an image sensor (e.g., CMOS or CCD). In one embodiment, the image processing module can process the still images or moving images acquired through the image sensor to extract necessary information. In one embodiment, information acquired through the camera (1600) can be provided to the processor (1700).
[0170] In one embodiment, the electronic device (1000) can obtain an image including a user wearing at least one fashion item by photographing the user wearing at least one fashion item through the camera (1600). In one embodiment, the electronic device (1000) can store the image obtained through the camera (1600) in the memory (1100) of the electronic device (1000), and obtain the selected image as a first image based on obtaining a user input for selecting an image stored in the memory (1100) through the input interface (1400). However, the present invention is not limited to the above-described example, and the electronic device (1000) can obtain various images or videos for performing the operations and functions of the electronic device (1000) described in the present disclosure through the camera (1600).
[0171] The processor (1700) can control the overall operations of the electronic device (1000). In one embodiment, the processor (1700) can include multiple processors. In one embodiment, at least one processor (1700) can perform the operations and functions of the electronic device (1000) described in the present disclosure by executing one or more instructions of a program stored in the memory (1100).
[0172] The processor (1700) may be configured as at least one of, for example, a Central Processing Unit, a microprocessor, a Graphic Processing Unit, Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), an Application Processor, a Neural Processing Unit, or an artificial intelligence processor designed with a hardware structure specialized for processing artificial intelligence models, but is not limited thereto.
[0173] When a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single processor or by multiple processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first and second operations may be performed by a first processor and the third operation may be performed by a second processor. However, the embodiments of the present disclosure are not limited thereto.
[0174] One or more processors according to the present disclosure may be implemented as a single-core processor or a multi-core processor. If a method according to an embodiment of the present disclosure includes multiple operations, the multiple operations may be performed by a single core or by multiple cores included in one or more processors.
[0175] In one embodiment, at least one processor (1700) may obtain a first image including a person wearing at least one fashion item by executing at least one command. In one embodiment, at least one processor (1700) may obtain a second image in which a pose of a person included in the first image is converted into a preset pose by executing at least one command. In one embodiment, at least one processor (1700) may detect at least one fashion item worn by a person converted into a preset pose based on the second image by executing at least one command. In one embodiment, at least one processor (1700) may extract at least one attribute of the detected at least one fashion item by executing at least one command.
[0176] In one embodiment, at least one processor (1700) can detect a pose of a person included in a first image by executing at least one command. In one embodiment, at least one processor (1700) can compare the pose of the detected person with a preset pose by executing at least one command. In one embodiment, at least one processor (1700) can obtain a second image based on the comparison result by executing at least one command.
[0177] In one embodiment, at least one processor (1700) may acquire a second image based on identifying that the pose of the detected person does not correspond to a preset pose by executing at least one command.
[0178] In one embodiment, at least one processor (1700) can acquire skeleton data including a plurality of key points corresponding to joints and preset body parts of a person as a detected pose of the person by executing at least one command.
[0179] In one embodiment, at least one processor (1700) may detect a pose of a camera that captured a first image based on the acquired first image by executing at least one command. In one embodiment, at least one processor (1700) may compare each of a detected person's pose and a detected camera's pose with each of a preset person's pose and a preset camera's pose included in preset poses by executing at least one command.
[0180] In one embodiment, at least one processor (1700) may execute at least one command to obtain a second image including a person captured in a pose of a preset camera and a pose of a detected person converted into a pose of a preset person, based on the identification that the pose of the detected camera does not correspond to a pose of a preset camera.
[0181] In one embodiment, at least one processor (1700) can acquire 6DoF (6 Degrees of Freedom) data including 3D coordinate information and rotation information of a camera that captured a first image as a pose of the detected camera by executing at least one command.
[0182] In one embodiment, at least one processor (1700) can obtain a second image as an output of the image generation model by inputting the acquired first image and the preset pose into the image generation model by executing at least one command.
[0183] In one embodiment, the image generation model may be an artificial intelligence model trained to output a second training image that transforms the pose of the person included in the first training image into the pose of the training person based on the first training image and the pose of the training person.
[0184] In one embodiment, at least one processor (1700) may detect at least one fashion item worn by a person included in a first image based on identifying that the pose of the detected person corresponds to a preset pose by executing at least one command. In one embodiment, at least one processor (1700) may extract at least one attribute of each of at least one fashion item worn by a person included in the detected first image by executing at least one command.
[0185] In one embodiment, at least one processor (1700) may display a second image acquired through a display by executing at least one command. In one embodiment, at least one processor (1700) may detect at least one fashion item in response to obtaining a user input requesting acquisition of at least one attribute from the second image by executing at least one command.
[0186] In one embodiment, at least one processor (1700) may detect at least one fashion item worn by a person included in a first image based on identifying that the pose of the detected person corresponds to a preset pose by executing at least one command. In one embodiment, at least one processor (1700) may extract at least one attribute of each of at least one fashion item worn by a person included in the detected first image by executing at least one command.
[0187] However, the present invention is not necessarily limited to the above-described examples, and at least one processor (1700) may perform the operations and functions of the electronic device (1000) described in the present disclosure by executing at least one command.
[0188] FIG. 11 is a detailed configuration diagram of a server according to one embodiment of the present disclosure.
[0189] Referring to FIG. 11, the server (2000) may include a memory (2100), a communication interface (2200), and a processor (2300). Each may be electrically and / or physically connected to each other.
[0190] The components illustrated in FIG. 11 are merely in accordance with one embodiment of the present disclosure, and the components included in the server (2000) are not limited to those illustrated in FIG. 11. The server (2000) according to one embodiment of the present disclosure may not include some of the components illustrated in FIG. 11, and may further include components not illustrated in FIG. 11.
[0191] The memory (2100) may store instructions or program codes for performing functions or operations of the server (2000). In one embodiment, at least one instruction, algorithm, data structure, program code, and application program stored in the memory (2100) may be implemented in a programming or scripting language such as, for example, C, C++, Java, or assembler.
[0192] In one embodiment, the memory (2100) may include at least one of a flash memory type, a hard disk type, a multimedia card micro type, a card type memory (e.g., SD or XD memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a mask ROM, a flash ROM, etc.), a hard disk drive (HDD), or a solid state drive (SSD).
[0193] In one embodiment, the data stored in the memory (2100) may correspond to the information stored in the memory (2100) of the electronic device (1000) of FIG. 10, so redundant description will be omitted.
[0194] The communication interface (2200) is a component for the server (2000) to communicate with an external electronic device. In one embodiment, the communication interface (2200) may perform data communication between the server (2000) and the external electronic device using at least one of data communication methods including wired LAN, wireless LAN, Wi-Fi, Bluetooth, zigbee, Wi-Fi Direct (WFD), infrared Data Association (IrDA), Bluetooth Low Energy (BLE), Near Field Communication (NFC), Wireless Broadband Internet (Wibro), World Interoperability for Microwave Access (WiMAX), Shared Wireless Access Protocol (SWAP), Wireless Gigabit Alliance (WiGig), and RF communication.
[0195] In one embodiment, the operation and function of the communication interface (2200) may correspond to the operation and function of the communication interface (1300) of the electronic device (1000) of FIG. 10, so redundant description will be omitted.
[0196] The processor (2300) can control the overall operations of the server (2000). In one embodiment, the processor (2300) may include multiple processors. In one embodiment, at least one processor (2300) can perform the operations and functions of the electronic device (1000) described in the present disclosure by executing one or more instructions of a program stored in the memory (2100).
[0197] In one embodiment, the operations and functions performed by at least one processor (2300) by executing at least one instruction stored in the memory (2100) may correspond to the operations and functions of the processor (1700) of the electronic device (1000) of FIG. 10, so redundant descriptions will be omitted.
[0198] Meanwhile, embodiments of the present disclosure may also be implemented in the form of a recording medium containing computer-executable instructions, such as program modules, executed by a computer. Computer-readable media may be any available media that can be accessed by a computer, and include both volatile and nonvolatile media, removable and non-removable media. Furthermore, computer-readable media may include computer storage media and communication media. Computer storage media include both volatile and nonvolatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Communication media may typically include computer-readable instructions, data structures, or other data in a modulated data signal, such as program modules.
[0199] Additionally, a computer-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.
[0200] The above description of the present disclosure is provided for illustrative purposes only, and those skilled in the art will readily appreciate that modifications to other specific forms can be made without altering the technical spirit or essential features of the present disclosure. Therefore, the embodiments described above should be understood as illustrative in all respects and not restrictive. For example, components described as being single may be implemented in a distributed manner, and similarly, components described as being distributed may be implemented in a combined manner.
[0201] The scope of the present disclosure is indicated by the claims described below rather than the detailed description above, and all changes or modifications derived from the meaning and scope of the claims and their equivalent concepts should be interpreted as being included in the scope of the present disclosure.
Claims
1. A method for obtaining attributes of a fashion item from an image, A step (S210) of acquiring a first image including a person wearing at least one fashion item; A step (S220) of obtaining a second image in which the pose of a person included in the first image is converted into a preset pose; A step (S230) of detecting at least one fashion item worn by a person converted into the preset pose based on the second image; and A method comprising a step (S240) of acquiring at least one attribute of at least one fashion item.
2. In paragraph 1, The step of obtaining the second image is as follows: A step of detecting a pose of a person included in the first image; A step of comparing the pose of the detected person with the preset pose; and A method comprising: a step of obtaining the second image based on the comparison result; 3. In paragraph 2, The step of detecting the pose of the person above is: A method comprising: acquiring skeleton data including a plurality of key points corresponding to joints and preset body parts of a human, as the detected human pose.
4. In any one of paragraphs 2 to 3, The above method, Further comprising a step of detecting a pose of a camera that captured the first image based on the first image; The step of comparing the pose of the detected person with the preset pose is as follows: A method comprising: a step of comparing each of the detected pose of the person and the detected pose of the camera with each of the preset poses of the person and the preset pose of the camera included in the preset poses.
5. In paragraph 4, The step of obtaining the second image is as follows: A method comprising: obtaining a second image in which the pose of the person included in the first image is converted into the pose of the person, based on the detection that the pose of the camera does not correspond to the pose of the preset camera; 6. In any one of paragraphs 4 to 5, The step of detecting the pose of the above camera is: A method comprising: a step of acquiring 6DoF (6 Degrees of Freedom) data including position information and rotation information of a camera that captured the first image as a pose of the detected camera.
7. In any one of paragraphs 1 to 6, The step of obtaining the second image is as follows: A step of obtaining the second image as an output of the image generation model by inputting the first image and the preset pose into the image generation model; The above image generation model is, A method, wherein the artificial intelligence model is trained to output a second training image in which the pose of the person included in the first training image is converted into the pose of the training person based on the first training image and the pose of the training person.
8. In an electronic device (1000) for obtaining attributes of a fashion item from an image, Memory (1100) for storing instructions; and At least one processor (1700) operably coupled to the memory (1000) and including a processing circuit; The electronic device (1000) executes the instructions by at least one processor (1700) alone or in cooperation with each other, Obtain a first image including a person wearing at least one fashion item, Generate a second image in which the pose of a person is converted into a preset pose based on the first image, Detecting at least one fashion item worn by a person converted to the preset pose based on the second image, An electronic device that acquires at least one attribute of at least one fashion item.
9. In paragraph 8, The electronic device, wherein the at least one processor executes one or more instructions, Detecting the pose of a person included in the first image above, Compare the pose of the detected person with the preset pose, An electronic device that obtains the second image based on the comparison result.
10. In paragraph 9, The electronic device, wherein the at least one processor executes one or more instructions, An electronic device that acquires skeleton data including a plurality of key points corresponding to joints and preset body parts of a human being as the detected human pose.
11. In any one of paragraphs 9 to 10, The electronic device, wherein the at least one processor executes one or more instructions, Detecting the pose of the camera that captured the first image based on the first image, An electronic device that compares each of the detected pose of the person and the detected pose of the camera with each of the preset poses of the person and the preset pose of the camera included in the preset poses.
12. In paragraph 11, The electronic device, wherein the at least one processor executes one or more instructions, An electronic device that obtains a second image in which the pose of the person included in the first image is converted into the pose of the person, based on the detection that the pose of the camera does not correspond to the pose of the preset camera.
13. In any one of paragraphs 11 to 12, The electronic device, wherein the at least one processor executes one or more instructions, An electronic device that acquires 6DoF (6 Degrees of Freedom) data including 3D coordinate information and rotation information of the camera that captured the first image as the pose of the detected camera.
14. In any one of paragraphs 8 to 13, The electronic device, wherein the at least one processor executes one or more instructions, By inputting the first image and the preset pose into an image generation model, the second image is obtained as an output of the image generation model, The above image generation model is, An electronic device, which is an artificial intelligence model trained to output a second training image in which the pose of a person included in the first training image is converted into the pose of the training person based on a first training image and a pose of the training person.
15. A computer-readable recording medium having recorded thereon a program for performing the method of any one of clauses 1 to 7 on a computer.
Citation Information
Patent Citations
A device and method for determining a pose of a camera
KR1020180112090A
Driving device for vehicle
KR1020260004041A
Apparatus and method for virtual assembly simulation of propeller and shaft for ship
KR102932037B1
KR20210030239A