Image processing methods, devices, chip systems, storage media and software products
By identifying human feature points through a neural network model, outfit recommendations that match the user's physical characteristics are generated, solving the problem of tedious manual input and improving the accuracy of outfit recommendations and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-06
- Publication Date
- 2026-04-07
AI Technical Summary
In existing technologies, the process of users manually inputting clothing information, physical characteristics, and personalized needs is cumbersome and time-consuming, resulting in poor accuracy of outfit recommendations and a poor user experience.
By using a neural network model to identify the location information of human feature points, the system determines a person's face shape, body proportions, skin color, gender, hairstyle, etc., and generates clothing recommendations that match the user's physical characteristics.
It improved the accuracy of outfit recommendations and user experience, enhancing the user's visual experience and satisfaction.
Smart Images

Figure CN119091232B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of terminal, and in particular, to an image processing method and device, a chip system, a storage medium and a program product. BACKGROUND
[0002] In order to improve user experience, some electronic devices can provide a dressing recommendation function.
[0003] At present, the user can manually input desired clothing information, physical features or personalized needs in the electronic device. After the electronic device receives the information input by the user, the corresponding dressing recommendation can be obtained based on the information.
[0004] However, manually inputting clothing information, physical features and personalized needs by the user can be tedious and time-consuming, which affects the user experience. SUMMARY
[0005] The embodiments of the present application provide an image processing method, device, chip system, storage medium and program product, which are applied to the technical field of terminal, and can make the electronic device recommend appropriate dressing to the user according to the recognized character information, improve the accuracy of dressing recommendation, and thus improve the user experience.
[0006] In a first aspect, an image processing method is provided. The method includes: obtaining a target image;
[0007] obtaining first information of the target image based on a first model; the first model is used to obtain position information of a human body part in an image input into the first model, and the first information includes the position information of the human body part in the target image; obtaining second information of the target image based on a second model; the second model is used to obtain gender information, skin color information and / or hairstyle information of a character in an image input into the second model, and the second information includes the gender information, the skin color information and / or the hairstyle information of the character in the target image; and obtaining a dressing image of the character in the target image based on the target image, the first information, the second information and a third model; the third model is used to obtain physical features of the character in an image input into the third model based on output information of the first model and output information of the second model, and determine a dressing image of the character based on the physical features of the character. The electronic device can obtain the physical features of the character based on the target image input by the user, and the first model, the second model and the third model, and can also generate a dressing picture conforming to the physical features of the character in the target image based on the third model. In this way, the user can more intuitively show the actual effect of the dressing recommendation on himself, improve the visual experience of the user, and improve the user satisfaction.
[0008] In conjunction with the first aspect, in certain implementations of the first aspect, obtaining the first information of the target image based on the first model includes: inputting the target image into the first model; the first model includes one or more first fusion convolutional modules; extracting features from the person in the target image based on the one or more first fusion convolutional modules to obtain a first feature image output by the first target fusion convolutional module; the first target fusion convolutional module is any one of the one or more first fusion convolutional modules; wherein the size of the first feature image output by the first target fusion convolutional module is smaller than the size of the feature image input to the first target fusion convolutional module, and the number of channels of the first feature image output by the first target fusion convolutional module is greater than the number of channels of the feature image input to the first target fusion convolutional module; performing convolution, pooling, and fully connected operations on the first feature image output by the one or more first fusion convolutional modules to obtain a second feature image; the second feature image includes the first information. After inputting the target image into the first model, the first information of the target image can be obtained, that is, the positional information of each body part of the person in the target image, which lays the foundation for subsequent calculation of the person's facial features, body proportions, and weight based on the positional information of the body parts.
[0009] In conjunction with the first aspect, in some implementations of the first aspect, obtaining the second information of the target image based on the second model includes: inputting the target image into the second model; the second model includes one or more second fusion convolutional modules; extracting features from the person in the target image based on the one or more second fusion convolutional modules to obtain a third feature image output by the second target fusion convolutional module; the second target fusion convolutional module is any one of the one or more second fusion convolutional modules; performing convolution, pooling, and fully connected operations on the third feature image output by the one or more second fusion convolutional modules to obtain a fourth feature image; the fourth feature image includes the second information; wherein, the fully connected operation includes a first fully connected layer, a second fully connected layer, and a third fully connected layer, the first fully connected layer being used to output information related to the person's gender in the target image; the second fully connected layer being used to output information related to the person's skin color in the target image; and the third fully connected layer being used to output information related to the person's hairstyle in the target image. Based on the second model, the second information of the person in the target image, namely the person's gender information, skin color information, and hairstyle information, can be identified. This can improve the accuracy of electronic devices in acquiring physical features such as gender, skin color, and hairstyle of a person, and make the clothing images output by the third model more consistent with the person's physical features, thereby improving user satisfaction.
[0010] In conjunction with the first aspect, in certain implementations of the first aspect, an image of the clothing of a person in the target image is obtained based on the target image, first information, second information, and a third model. This includes: obtaining facial features, body proportions, and / or weight information of the person in the target image based on the first information; inputting the target image, the facial features, body proportions, and / or weight information of the person in the target image, as well as the gender, skin color, and / or hairstyle information of the person in the target image into the third model; and obtaining the clothing image of the person in the target image based on the third model. The clothing in the clothing image of the person in the target image obtained based on the third model can be clothing that highly matches the person's physical characteristics. This makes the clothing recommendation results more accurate and improves user satisfaction.
[0011] In conjunction with the first aspect, in some implementations of the first aspect, the first information includes the coordinates of the person's forehead, chin, cheeks, shoulders, waist, knees, and ankles in the target image. Obtaining relatively detailed and accurate coordinate information for each body part in the target image helps in analyzing the person's facial features, body proportions, and weight characteristics. It also helps improve the accuracy of clothing recommendations and enhances the user experience.
[0012] In conjunction with the first aspect, in certain implementations of the first aspect, facial shape information, body proportion information, and / or weight information of a person in the target image are obtained based on the first information, including: determining the face length information of the person based on the coordinate information of the person's forehead and the chin in the target image; determining the face width information of the person based on the coordinate information of both sides of the person's cheeks in the target image; calculating the ratio of the face length information to the face width information to determine the face shape information of the person in the target image; determining the height proportion information of the person in the target image based on the coordinate information of both sides of the person's shoulders and the coordinate information of both sides of the person's ankles in the target image; determining the weight proportion information of the person in the target image based on the coordinate information of both sides of the person's waist and the coordinate information of both sides of the person's knees in the target image; and calculating the ratio of the height proportion information to the weight proportion information to determine the body proportion information of the person in the target image. Based on the coordinate information of various body parts of the person, facial shape information, body proportion information, and other physical features can be determined. This allows electronic devices to generate outfit images that better match a person's face shape and body proportions, improving the accuracy of outfit recommendations and enhancing the user experience.
[0013] Secondly, embodiments of this application provide an image processing apparatus. The image processing apparatus can be an electronic device, or a chip or chip system within an electronic device. The image processing apparatus may include a processing unit, which may be a processor. The apparatus may also include a storage unit, which may be a memory. The storage unit stores instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement an image processing method described in the first aspect or any possible implementation of the first aspect. When the apparatus is a chip or chip system within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to cause the electronic device to implement an image processing method described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (e.g., a register, cache, etc.), or a storage unit located outside the chip within the electronic device (e.g., a read-only memory, random access memory, etc.).
[0014] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, the memory for storing code instructions, and the processor for running the code instructions to perform the methods described in the first aspect or any possible implementation of the first aspect.
[0015] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program or instructions that, when executed on a computer, cause the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0016] Fifthly, embodiments of this application provide a computer program product including a computer program, which, when run on a computer, causes the computer to perform the methods described in the first aspect or any possible implementation thereof.
[0017] Sixthly, this application provides a chip or chip system including at least one processor and a communication interface. The communication interface and the at least one processor are interconnected via a circuit. The at least one processor is used to run computer programs or instructions to perform the methods described in the first aspect or any possible implementation of the first aspect. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.
[0018] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).
[0019] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0021] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;
[0022] Figure 3 A schematic diagram of the overall structure of an outfit recommendation network model provided in the application embodiment;
[0023] Figure 4 A schematic diagram of a human feature model provided in an embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the structure of a feature point recognition network model provided in an embodiment of this application;
[0025] Figure 6 This is a schematic diagram of the structure of a fusion convolution module provided in an embodiment of this application;
[0026] Figure 7 This is a schematic diagram of the structure of the multi-classification network model provided in the embodiments of this application;
[0027] Figure 8 A schematic flowchart of an image processing method provided in an embodiment of this application;
[0028] Figure 9 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation
[0029] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:
[0030] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0031] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0032] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.
[0033] With the continuous development of terminal technology and artificial intelligence, users' expectations for electronic devices are also constantly increasing. Users expect these electronic devices to not only meet basic communication and entertainment needs, but also provide more intelligent and personalized services. For example, to improve user experience, some electronic devices can provide outfit recommendation functions.
[0034] In one implementation, the user can manually input desired clothing information, physical characteristics, or personalized needs into an electronic device. For example, clothing information can include the user's specific requirements for clothing, such as color and style. Physical characteristics may include the user's height and weight. Personalized needs may include the current season and the shooting scenario. After receiving the user's input, the electronic device can match this information with clothing information in a clothing database to filter out outfits that meet the user's requirements.
[0035] However, requiring users to manually input clothing information, physical characteristics, and personalized needs can be cumbersome and time-consuming. For users unfamiliar with the process, this method may increase the difficulty of use and reduce the user experience. Secondly, since clothing recommendations are highly dependent on the information input by the user, inaccurate or incomplete information, such as incorrect height or weight, or unclear clothing requirements, will affect the system's matching results, leading to inaccurate recommendations and a reduced user experience.
[0036] To address the aforementioned technical problems, this application provides an image processing method that can identify the location information of feature points such as the forehead, cheeks, and chin of a human body based on a neural network model, thereby determining the person's face shape, height, weight, and other body proportions. It can also identify information such as skin color, gender, and hairstyle based on the neural network model. In this way, suitable clothing recommendations can be given to users based on the identified person information, improving the accuracy of clothing recommendations and thus enhancing the user experience.
[0037] It is understood that the electronic devices in the embodiments of this application can be any form of terminal device. For example, electronic devices may include: mobile phones, tablet computers, handheld computers, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, electronic devices in 5G networks, or future evolved public land mobile communication networks (PLTs). The embodiments of this application do not limit the scope of electronic devices in a mobile network (PLMN).
[0038] Furthermore, in this embodiment of the application, the electronic device can also be an electronic device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.
[0039] The electronic devices in the embodiments of this application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.
[0040] In this embodiment, the electronic device includes a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and main memory. The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.
[0041] For example, Figure 1 A schematic diagram of the electronic device is shown.
[0042] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0043] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may include hardware, software, or a combination of software and hardware.
[0044] Processor 110 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.
[0045] For example, in this embodiment of the application, the processor 110 can be used to extract the location information of feature points in a user's portrait photo, or the processor 110 can also be used to classify features such as skin color or hairstyle in the portrait photo to obtain classification information.
[0046] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or is reusing. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the aforementioned memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. For example, in the embodiments of this application, the processor 110 can be used to store the location information of feature points in a portrait photograph or the classification information of skin color, hairstyle, etc., without specific limitations.
[0047] Internal memory 121 can be used to store computer executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device, etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the electronic device by running instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor. For example, in this embodiment, internal memory 121 can be used to store trained outfit recommendation network models. For example, it may include parameters and structures of feature point recognition network models, multi-classification network models, and conditional generative adversarial network models, etc.
[0048] In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. The cameras 193 can be used to capture still images or videos. For example, in this embodiment, the electronic device can use the cameras 193 to collect portrait photos of the user.
[0049] Figure 2 This is a software structure block diagram of an electronic device according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, the hardware adaptation layer (HAL), and the kernel layer.
[0050] The application layer, also known as the application layer, can include a series of application packages. For example... Figure 2 As shown, the application package can include applications such as phone, music, calendar, camera, games, notes, and video. Applications can include system applications and third-party applications.
[0051] Optionally, the application layer may also include an outfit recommendation app, which can provide outfit recommendations to users based on their portrait photos.
[0052] The application framework layer provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The application framework layer can include some predefined functions.
[0053] like Figure 2 As shown, the application framework layer may include an activity manager, window manager, resource manager, notification manager, content provider, and view system, etc.
[0054] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the control and management of the Android system.
[0055] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.
[0056] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of this application, the virtual machine can be used to extract the location information of feature points in a user's portrait photo, or it can also be used to classify features such as skin color or hairstyle in the portrait photo to obtain classification information.
[0057] System libraries, also known as the native layer, can include multiple functional modules. Examples include media libraries, function libraries, and graphics processing libraries (such as OpenGL ES).
[0058] The Hardware Abstraction Layer (HAL) is an abstraction layer situated between the kernel layer and the Android runtime. The HAL can be a wrapper around hardware drivers, providing a unified interface for calls from upper-layer applications.
[0059] The kernel layer is the layer between hardware and software. The kernel layer can include display drivers, camera drivers, audio drivers, battery drivers, Bluetooth drivers, CPU drivers, USB drivers, etc.
[0060] It should be noted that the embodiments of this application are only illustrated using the Android system. In other operating systems (such as Windows, iOS, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of this application, the solution of this application can also be implemented.
[0061] The methods of this application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be described again in some embodiments.
[0062] It is understood that in this embodiment, the training process of the outfit recommendation network model can be implemented on a cloud server or on the user's electronic device. The user's electronic device may include, for example, a mobile phone, watch, tablet, or computer. The cloud server can also be simply referred to as a cloud server or a server. When the outfit recommendation network model training process is implemented on a cloud server, the computational load on the user's electronic device can be reduced, improving performance and reducing lag. For ease of description, the following explanation will use training the outfit recommendation network model on a cloud server as an example.
[0063] Let's combine the following... Figure 3 This section describes the overall structure of the outfit recommendation network model provided in the application embodiments. For example... Figure 3 As shown, the outfit recommendation network model provided in this application embodiment may include a feature point recognition network model, a multi-classification network model, and a conditional generative adversarial network model.
[0064] It is understandable that the input to the outfit recommendation network model can be a user's portrait photo. Optionally, the portrait photo can be a real-time captured photo of the user, or a portrait photo stored on an electronic device; there is no limitation here.
[0065] When a user inputs a portrait photo into the clothing recommendation network model, the feature point recognition network model can extract multiple feature point information of the human body in the portrait photo. Optionally, the feature point information of the human body can include the location information of key points of the human body, such as the forehead, chin, both sides of the cheeks, both sides of the shoulders, both sides of the waist, both knees, and both ankles.
[0066] It is understandable that by using the location information of these key points on the human body, one can calculate a person's facial features, body proportions, weight, and other physical characteristics.
[0067] For the sake of brevity, the location information of key points on the human body included in the feature point information will be referred to as the location information of the feature points in the following text.
[0068] Multi-classification network models can obtain classification information such as skin color, hairstyle, or gender from a user's portrait photo. After the feature point recognition network model obtains feature point information, and the multi-classification network model obtains classification information for skin color, hairstyle, or gender, the feature point information, the classification information for skin color, hairstyle, or gender, and the user's portrait photo can be input into a conditional generative adversarial network model.
[0069] Furthermore, the conditional generative adversarial network model can generate one or more sets of outfit recommendations based on the user's portrait photo and the information output by the feature point recognition network model and the multi-classification network model. Optionally, the outfit recommendations can be described in text form.
[0070] For example, the textual description of outfit recommendations may include detailed clothing types and corresponding outfit recommendations for each clothing type. Clothing types may include, for example, tops, bottoms, outerwear, and shoes. Outfit recommendations for each clothing type could be: "Top: White shirt"; "Bottoms: Black long skirt"; "Outerwear: Light-colored trench coat"; "Shoes: White sneakers".
[0071] Alternatively, outfit recommendations can include images of clothing in various styles suitable for the user. These images could include, for example, images of a white shirt, a black long skirt, a light-colored trench coat, and white sneakers. These images of different clothing styles can be presented as a set of images for outfit recommendations.
[0072] Furthermore, the outfit recommendation network model can generate an outfit list based on one or more sets of outfit recommendations. Optionally, each set of outfit recommendations in the outfit list can include textual descriptions. Alternatively, each set of outfit recommendations in the outfit list can also include images of clothing of various types. Alternatively, the outfit list can also include outfit recommendations that combine textual descriptions and clothing images, where the textual descriptions can provide specific clothing type or color suggestions, and the clothing images can provide visual references.
[0073] It should be noted that the clothing types described above, including tops, bottoms, outerwear, and shoes, are merely illustrative descriptions and do not limit the number or specific types of clothing. In the actual application of the embodiments of this application, more clothing types may be included. For example, clothing types may also include accessories, bags, or glasses, etc. For the sake of brevity, they will not be listed here.
[0074] Optionally, the clothing recommendation network model provided in this application embodiment can also be used to process images of various clothing types generated by a conditional generative adversarial network model to obtain clothing recommendation images from user-input portrait photos. The person in the clothing recommendation image can "wear" clothing images of various clothing types generated by the conditional generative adversarial network model.
[0075] Furthermore, the outfit list can also include outfit recommendation images, allowing users to more intuitively see how the outfit recommendations actually look on them, improving the user's visual experience and increasing user satisfaction.
[0076] In a possible implementation, the clothing recommendation network model can perform image segmentation on the user-input portrait photo, identifying and separating different parts of the person, such as the head, torso, and limbs. It can also locate key points in the portrait photo, such as the shoulders, waist, knees, and ankles, based on feature point information. Furthermore, the clothing recommendation network model can perform geometric transformations on the clothing images in the clothing recommendations based on these key points to obtain deformed clothing images that fit the key points of the person.
[0077] Furthermore, the outfit recommendation network model can stitch together images of different parts of a person, such as the head, torso, and limbs, as well as geometrically transformed clothing images to obtain corresponding outfit recommendation pictures.
[0078] For example, suppose the user inputs a portrait photo in which the person's original outfit is a white short-sleeved shirt and dark blue jeans. The conditional generative adversarial network model, based on the person's facial features such as face shape, body proportions, and weight, as well as classification information such as skin color, hairstyle, or gender, can recommend an outfit suitable for the user: a white shirt and a black long skirt.
[0079] Furthermore, the clothing recommendation network model can identify and separate different parts of a person, such as the head, torso, and limbs. It can also perform geometric transformations on the clothing image of the white shirt and black skirt based on feature point information. Then, it can generate corresponding clothing recommendations based on the person's head, torso, limbs, and the geometrically transformed images of the white shirt and black skirt.
[0080] The implementation process of the image processing method of this application embodiment will be described in detail below from aspects such as (i) defining the human body feature model, (ii) recognizing human body feature points, and (iii) recognizing skin color, gender, and hairstyle of a person.
[0081] (a) Define the human body feature model.
[0082] In order to ensure that the recommended outfits are in line with the physical characteristics of the person, this application defines a human body feature model, which can be used to identify the physical characteristics of a person, such as face shape, body proportions, and weight.
[0083] For example, a human feature model may include multiple feature points of the human body, and the distribution map of these multiple feature points may be as follows: Figure 4 As shown. Feature points can include the forehead, chin, both sides of the cheeks, both sides of the shoulders, both sides of the waist, both knees, and both ankles in a portrait.
[0084] Furthermore, multiple feature points of the human body feature model can be used to calculate a person's facial features such as face shape, body proportions, and weight. For example, embodiments of this application can calculate a person's face shape based on the positional information of four feature points on the forehead, chin, and both sides of the cheeks. For instance, the electronic device can calculate a person's face length based on the positional information of the forehead and chin. The electronic device can also calculate a person's face width based on the positional information of both sides of the cheeks. Furthermore, the electronic device can calculate the face length-to-width ratio based on the person's face length and width.
[0085] After obtaining the aspect ratio of a person's face, the face shape can be determined based on this ratio. The aspect ratio can be calculated by dividing the face length by the face width.
[0086] In one possible implementation, the electronic device can preset different threshold ranges for different face shapes. For example, if a person's face shape includes round, oval, and oblong faces, the aspect ratio of a round face can correspond to threshold range a, the aspect ratio of an oval face can correspond to threshold range b, and the aspect ratio of an oblong face can correspond to threshold range c. After determining the threshold range to which the aspect ratio of a person's face belongs, the electronic device can determine the person's face shape.
[0087] It should be noted that the above classification of human face shapes into round, oval, and oblong faces is merely an illustrative description and does not limit the number of face shapes or their names. In the actual application of this application's embodiments, the number and names of face shapes can be adjusted and expanded according to specific needs. For example, electronic devices can also add more face shapes such as square, heart-shaped, and diamond-shaped faces to accommodate a wider range of users.
[0088] It should be noted that the method described above for calculating a person's face shape based on the location information of four feature points on the forehead, chin, and both sides of the cheeks is merely an illustrative description. In the actual application of this application embodiment, the calculation of a person's face shape can also be based on additional feature points to obtain more accurate facial features.
[0089] For example, additional feature points may also include the cheekbones or jaw angles on both sides. That is, the electronic device can identify more face shape categories based on the location information of the forehead, chin, and sides of the cheeks, as well as additional feature points, without limitation. Therefore, in the actual application of this application embodiment, the feature points for calculating face shape are not limited to a few fixed points, but can be flexibly adjusted according to needs to adapt to the diverse features of different users. For simplicity, the process of identifying face shape based on four feature points on the forehead, chin, and sides of the cheeks will be described in detail below.
[0090] Furthermore, embodiments of this application can also calculate a person's body proportions or physical features such as weight based on the positional information of all or some of the feature points. For example, body proportions in embodiments of this application may include waist-to-shoulder ratio, leg-to-body ratio, and upper-to-lower-body ratio. The electronic device can calculate a person's shoulder width based on the positional information of feature points on both sides of the shoulders, calculate a person's waist circumference based on the positional information of feature points on both sides of the waist, and calculate the person's waist-to-shoulder ratio based on the shoulder width and waist circumference.
[0091] Similarly, electronic devices can calculate a person's leg length based on the location of feature points at the waist and ankles, calculate a person's height based on the location of feature points at the forehead and ankles, and calculate the person's leg-to-body ratio based on leg length and height. Furthermore, electronic devices can also calculate a person's upper body length based on the location of feature points at the shoulders and waist, and calculate the person's upper-to-lower body ratio based on upper body length and leg length.
[0092] It should be noted that the above process of calculating the waist-to-shoulder ratio, leg-to-body ratio, and upper-to-lower-body ratio based on the positional information of all or some of the feature points is merely an illustrative description and does not limit the indicators of body proportions or the feature points used to calculate them. In practical applications, more body proportion indicators can be calculated based on the positional information of additional feature points.
[0093] For example, additional feature points may include feature points on both sides of the buttocks. More body proportion indicators may include, for example, the waist-to-hip ratio. The electronic device can calculate the hip circumference of a person based on the positional information of the feature points on both sides of the buttocks, and calculate the waist-to-hip ratio based on the hip circumference and waist circumference. Therefore, in the practical application of this embodiment, the feature points for calculating body proportions are not limited to a few fixed points, but can be flexibly adjusted according to needs to obtain more accurate body proportions.
[0094] One possible way to calculate a person's weight is based on their height and waist circumference. For example, after obtaining a person's height and waist circumference, their weight can be determined based on the ratio of height to waist circumference. This ratio is calculated by dividing height by waist circumference. A smaller result when dividing height by waist circumference suggests the person is overweight around the waist, while a larger result suggests a slimmer waist.
[0095] It's understandable that the feature points used to calculate a person's weight can be adjusted as needed to obtain a more accurate assessment. For example, feature points in a human body feature model can also include those on both sides of the chest, allowing for the calculation of a person's chest circumference. A larger chest circumference can be interpreted as chest obesity.
[0096] Optionally, to obtain more comprehensive information about a person's body shape and weight, in addition to obtaining the degree of fatness based on feature points, users can also input data such as height and weight. That is, the human body feature model can also include a person's height and weight, and can calculate the person's body mass index (BMI) based on height and weight. BMI can be used to indicate a person's degree of obesity; a higher BMI can be understood as a more obese person.
[0097] The above process describes how, when a human body feature model includes multiple feature points, facial features such as face shape, body proportions, and weight can be determined based on these feature points. Different physical features can correspond to different clothing recommendations. Furthermore, the human body feature model can also identify a person's skin color, hairstyle, and gender. For example, in the human body feature model provided in this embodiment, skin color can be divided into black, white, and yellow; hairstyle can be divided into long hair and short hair; and gender can be divided into male and female.
[0098] It should be noted that the above classification of skin tone and hairstyle categories is merely illustrative and does not limit the names or number of categories. In practical applications, skin tone can also be categorized as warm or cool, and hairstyle can be categorized as straight or curly. In other words, the names and number of skin tone and hairstyle categories can be adjusted and expanded according to specific needs, allowing the human feature model to better adapt to the personalized needs of different users and provide more accurate and diverse clothing recommendations.
[0099] It should be noted that the human feature model defined above does not directly participate in the training of the neural network model. In other words, the human feature model described above can be used to describe a person's feature points, skin color, hairstyle, gender, and other characteristic data. This defined human feature model helps improve the consistency and completeness of data during sample collection and preprocessing, but it is not used to train the human feature model itself in the subsequent actual training process.
[0100] (ii) Identification of human body feature points.
[0101] The embodiments of this application can identify and obtain the location information of a person's feature points based on a feature point recognition network model.
[0102] In one possible implementation, before identifying the location information of human feature points, the feature point recognition network model can be trained based on training samples, which include multiple different images. For example, the training samples can be used as input data and fed into the feature point recognition network model. Through multiple layers of deep learning, the model learns the location information of human feature points. The specific number of images included in the training samples during the model training process is not limited in this embodiment.
[0103] Optionally, the coordinates of the feature points can be used to represent their positional information. That is, for any one feature point, there can be two coordinate values: an x-coordinate and a y-coordinate. In a human feature model that includes 12 feature points—forehead, chin, both cheeks, both shoulders, both waists, both knees, and both ankles—the positional information of these 12 feature points can be represented by 24 coordinate values.
[0104] Optionally, the 24 coordinate values can be represented by data or by vectors; no restriction is placed here. Therefore, the output of the feature point recognition network model can include 24 dimensions of values, which can be used to indicate the location information of 12 feature points.
[0105] For example, a feature point recognition network model can be composed of fusion convolutional modules, which can be used for feature extraction. In this embodiment, two fusion convolutional modules can be set as one fusion convolutional module group, and five fusion convolutional module groups can be stacked in the feature point recognition network model, that is, there are a total of 10 fusion convolutional modules in the feature point recognition network model. Each stacked fusion convolutional module group can reduce the size of the feature map to half its original size and double the number of channels in the feature map.
[0106] It should be noted that the size of a feature map can include both its height and width; that is, the height and width of the feature map are each halved. For ease of description, the following text will refer to the situation where the height and width of the feature map are each halved as the feature map size being halved.
[0107] Furthermore, the feature point recognition network model can also include an output network. That is, after feature extraction from the input image, the resulting feature map can be fed into the output network. The output network can sequentially include convolutional layers, global pooling layers, and fully connected layers, so that the final output feature map has a size of 1×1×24, meaning it has 24 channels. The values output by these 24 channels can be used to indicate the positional information of 12 feature points.
[0108] Figure 5This is a schematic diagram of the feature point recognition network model provided in an embodiment of this application. Figure 5 As shown, for ease of description, the feature point recognition network model is divided into modules 501 to 505. Specifically, the structure of each part of the feature point recognition network model will be described below through modules 501 to 505. Furthermore, the structure of image data or feature map data will be described in the following text using the format H×W×C, where H represents height, W represents width, and C represents the number of channels.
[0109] (5.1) First Module 501: May include convolutional layers (conv), batch normalization (BN), and the H-Swish (Hard-Swish) activation function. Convolutional layers can be used to extract local features from the input feature map. Batch normalization can be used to normalize the output of the convolutional layers, thereby stabilizing the training process. The H-Swish activation function can be used to improve training efficiency and model performance. That is, in the first module 501, local features of the input image can be extracted through convolutional layer a to obtain a convolutional feature map, and then batch normalization and the H-Swish activation function are applied sequentially to this convolutional feature map to obtain feature map a.
[0110] For example, such as Figure 5 As shown, when the kernel size of convolutional layer a is set to 3×3, the stride of the convolution can be set to 2, thereby adjusting the size of the output feature map a of convolutional layer a to half that of the input image. Alternatively, the number of kernels in convolutional layer a can be set to 16 to adjust the number of channels in feature map a to 16. It can be understood that when the H×W×C of the input image is 448×448×3, the feature map a obtained after feature extraction by the first module 501 is 224×224×16.
[0111] It is understandable that the first module 501 can be used to perform preliminary feature extraction and preprocessing on the input image. Through a combination of 3×3 convolution, batch normalization, and H-Swish activation function, the first module 501 can extract preliminary features from the input image, while adjusting the spatial size and number of channels of the feature map to prepare for feature extraction in subsequent modules.
[0112] (5.2) Second module 502: can be composed of multiple fused convolutional modules. It can be understood that the input to the second module 502 can be feature map a output from the first module 501, and the output of the second module 502 can be feature map b. In this embodiment, two fused convolutional modules can be grouped together, and the fused convolutional module groups can be stacked 5 times in the second module 502, that is, the second module 502 can include 10 fused convolutional modules.
[0113] To facilitate understanding, the following will be combined with... Figure 6 This application introduces the fusion convolution module provided in its embodiments. Figure 6 This is a schematic diagram of the structure of the fusion convolution module provided in an embodiment of this application, where H1×W1×C1 can represent the size and number of channels of the input feature map of the fusion convolution module. In this fusion convolution module, the input feature map can first be processed as follows: Figure 6 The four parts shown are processed, and the results of these four parts can be spliced together to obtain the spliced feature map.
[0114] The specific processing steps for the four parts are as follows:
[0115] (1) 1×1 Convolution and Squeeze-and-Excitation (SE) Module: 1×1 convolution can be used for channel fusion to reduce the dimensionality of the input feature map. The SE module can enhance salient features by calculating the importance weight of each channel. That is, the SE module can adjust the weight of each channel using global information to highlight important feature regions and reduce unnecessary computation. It can be understood that this part can fuse the features between the channels of the feature map while reducing the amount of computation, and apply weights to the fused features to highlight salient features.
[0116] Optionally, the number of channels in a 1×1 convolution can be set to C1, meaning that a convolutional layer with a kernel size of 1×1 and a number of channels of C1 is used to convolve the input feature map. The number of kernels can be set according to requirements and is not limited here. Optionally, the channel vector of the SE module can also be set to 1×1×C1.
[0117] (2) Depthwise separable convolution: This can include depthwise convolution and 1×1 pointwise convolution. Optionally, the depth of the depthwise convolution and the pointwise convolution can be C1. That is, first, a convolutional layer with C1 channels is used to perform depthwise convolution on the input feature map, and then a convolutional layer with a kernel size of 1×1 and C1 channels is used to perform pointwise convolution on the input feature map. This part can fuse features within the neighborhood of the feature map channels while reducing the amount of computation.
[0118] (3) 3×3 convolution: This includes convolutional layers with a kernel size of 3×3, where the depth of the convolutional layer can be C1. This part can perform spatial and channel feature fusion simultaneously.
[0119] (4) Skip connection: The input feature map is not processed, that is, the input feature map is directly passed to the output, which can reduce the gradient vanishing phenomenon that occurs during training.
[0120] Furthermore, the fusion convolution module can concatenate the features processed from the above four parts, fusing multiple feature representations to obtain a richer and more comprehensive feature map c. The resulting features will be richer and more comprehensive.
[0121] Furthermore, the concatenated feature map c can be input into path 1 or path 2 according to specific requirements. Path 1 can include a convolutional layer with a kernel size of 3×3 and the number of channels C2. Path 2 can include a pooling layer. That is, path 1 can perform a 3×3 convolution on feature map c, and path 2 can perform pooling on feature map c.
[0122] It should be noted that, in this embodiment, when it is necessary to change the feature map size and number of channels, the stitched feature map can be sent to path 1, that is, feature extraction is performed on the stitched feature map through path 1. When it is not necessary to change the feature map size and number of channels, the stitched feature map can be sent to path 2, that is, feature extraction is performed on the stitched feature map through path 2 using a pooling operation.
[0123] In this embodiment, the second module 502 can stack five fused convolutional module groups. Each stack can halve the size of the feature map and double the number of channels. Taking the input feature map a of the second module 502 as having a size and number of channels of 224×224×16, and the output feature map b as having a size and number of channels of 7×7×512, the size and number of channels of feature map a can be changed from 224×224×16 to 112×112×32, 56×56×64, 28×28×128, 14×14×256, and 7×7×512, thus obtaining feature map b.
[0124] Therefore, for the two fusion convolutional modules in each fusion convolutional module group, the first fusion convolutional module can be configured to change the feature map size and number of channels, while the second fusion convolutional module can be configured not to change the feature map size and number of channels. That is, the feature map c in the first fusion convolutional module can be input to path 1, and the feature map c in the second fusion convolutional module can be input to path 2.
[0125] In other words, within each group of fused convolutional modules, the first fused convolutional module can halve the size of the input feature map and double the number of channels. For example, with an input feature map size and number of channels of 224×224×16, the first fused convolutional module can change the feature map size and number of channels to 112×112×32. The second fused convolutional module does not change the size and number of channels of the input feature map; that is, with an input feature map size and number of channels of 112×112×32, its output feature map size and number of channels remain 112×112×32.
[0126] It is understandable that since the size of the feature map can be halved each time it passes through the fusion convolution module corresponding to path 1, the stride of the convolution can be set to 2 during the convolution process of path 1 so that the size of the feature map can be halved.
[0127] It should be noted that, as the number of channels in the input and output feature maps differs, the number of convolutional kernels in the convolutional layer of the first fusion convolutional module in each group may vary. In other words, the number of convolutional kernels in the convolutional layer of path 1 is related to the number of channels in the output feature map. The following example uses the first fusion convolutional module group, where the input feature map size and number of channels are 224×224×16, and the output feature map size and number of channels are 112×112×32, to illustrate the setting of the number of convolutional kernels in the fusion convolutional module.
[0128] It should also be noted that the number of channels in each of the four processing steps described above is not limited in this embodiment of the application. In actual applications, it can be set according to application needs. That is, the number of channels in the feature map c obtained after stitching can be the same as or different from C1.
[0129] For the first fusion convolutional module in the first group, the feature map c obtained after concatenation is fed into path 1 for a 3×3×C2 convolution operation, where the size of C2 can be the same as the number of channels of feature map c. Since the size and number of channels of the output feature map of the first fusion convolutional module are 112×112×32, the convolutional layer in path 1 includes 32 3×3×C2 convolutional kernels.
[0130] Similarly, for the first fused convolutional module in the second group, the convolutional layer in path 1 can be configured to include 64 3×3×C2 convolutional kernels; for the first fused convolutional module in the third group, the convolutional layer in path 1 can be configured to include 128 3×3×C2 convolutional kernels; for the first fused convolutional module in the fourth group, the convolutional layer in path 1 can be configured to include 256 3×3×C2 convolutional kernels; and for the first fused convolutional module in the fifth group, the convolutional layer in path 1 can be configured to include 512 3×3×C2 convolutional kernels.
[0131] Optionally, for each fusion convolutional module in the second module 502, batch normalization and the H-Swish activation function can also be applied to the results of the fusion convolutional modules.
[0132] It is understood that the fusion convolutional model proposed in this application can enhance feature extraction capabilities by fusing adjacent feature maps. Specifically, the fusion convolutional module concatenates the inter-channel features, channel neighborhood features, combined features of channel neighborhoods and channels obtained after the four-part processing, along with the original features, to obtain a concatenated feature map. Further feature extraction can be performed on the concatenated feature map, resulting in richer extracted features and improved recognition accuracy. This allows the feature point recognition network model to capture more detailed feature information, contributing to improved overall model performance.
[0133] (5.3) Output Network: Includes a third module 503, a fourth module 504, and a fifth module 505. The third module 503 can be a 1×1 convolutional layer, the fourth module 504 can be a global pooling layer, and the fifth module 505 can be a fully connected layer. Optionally, batch normalization and the H-Swish activation function can be applied to the convolutional results of the third module 503; and batch normalization and the H-Swish activation function can be applied to the pooling results of the fourth module 504.
[0134] It can be understood that in the feature point recognition network model, the input image, after passing through the first module 501 to the fifth module 505, can output a 1×1×24 feature map. These 24 values can correspond to the horizontal and vertical coordinates of 12 feature points, indicating the positional information of each feature point in the training samples. It should be noted that when the accuracy of the feature point recognition network model in recognizing the positional information of human feature points reaches a preset standard, the feature point recognition network model can be used to perform human feature point recognition on images and can output the positional information of human feature points in portrait photos.
[0135] In this embodiment, a feature point recognition network model can identify the location information of feature points in a portrait photo. Furthermore, based on the location information of these feature points, facial features such as face shape, body proportions, and weight can be calculated. This allows the clothing recommendation network to recommend clothing items that better match the user's physical characteristics, resulting in more accurate recommendations and improved user experience.
[0136] (iii) Identification of skin color, gender and hairstyle of people.
[0137] The embodiments of this application can identify and obtain a person's skin color, gender, and hairstyle based on a multi-classification network model.
[0138] In a possible implementation, before identifying a person's skin color, gender, and hairstyle, a multi-classification network model can be trained based on training samples, which include multiple different images. For example, the training samples can be used as input data and fed into the multi-classification network model. Through multiple layers of deep learning, the model learns the person's skin color, gender, and hairstyle. The specific number of images included in the training samples during the model training process is not limited in this embodiment.
[0139] Optionally, the multi-classification network model can also be composed of fused convolutional modules. In this embodiment, similar to the feature point recognition network model, the multi-classification network model can also include fused convolutional module groups, and five fused convolutional module groups can be stacked in the multi-classification network model.
[0140] For example, Figure 7 This is a schematic diagram of the structure of a multi-classification network model provided in an embodiment of this application. Figure 7 As shown, for ease of description, the multi-class network model is divided into modules 6 (701) to 10 (705). It can be understood that the backbone network of the multi-class network model is the same as that of the feature point recognition network model; that is, module 6 (701) can be the same as module 1 (501). For details, please refer to the relevant description of module 1 (501), which will not be repeated here.
[0141] Similarly, Module 702 can be the same as Module 2 502; please refer to the relevant description of Module 2 502 for details. Module 8 703 can be the same as Module 3 503; please refer to the relevant description of Module 3 503 for details. Module 9 704 can be the same as Module 4 504; please refer to the relevant description of Module 4 504 for details, which will not be repeated here. The following focuses on Module 10 705.
[0142] In this embodiment, the multi-classification network model includes a tenth module 705 in its output network. This tenth module 705 may include three fully connected layers, which can be learning layers for skin color, gender, and hairstyle, respectively. The output of the skin color learning layer can be the probability of recognizing a person's skin color. When the clothing recommendation network model classifies a person's skin color as black, white, or yellow, the skin color learning layer can output the probability value of recognizing the person as having one of these three skin colors.
[0143] In a possible implementation, after obtaining the probability values for the three skin tones, and if the highest probability value among the three exceeds a pre-set skin tone threshold, the skin tone with the highest probability value output by the skin tone learning layer can be set as the recognition result of the multi-classification network model. For example, assuming the probability value of black is 0.8, the probability value of white is 0.2, and the probability value of yellow is 0.1, and 0.8 exceeds the preset skin tone threshold, then the recognition result of the person's skin tone in the multi-classification network model can be considered as black.
[0144] It is understandable that the process of multi-class network models recognizing hairstyles and genders is similar to that of recognizing skin color, so it will not be elaborated here.
[0145] It is understandable that in a multi-classification network model, the input image, after passing through modules 6 (701) to 10 (705), can output probability values for the recognition results of a person's skin color, hairstyle, and gender. It should be noted that when the accuracy of the multi-classification network model in recognizing a person's skin color, hairstyle, and gender reaches a preset standard, the multi-classification network model can be used to recognize portrait photos and can output the recognition results of the person's skin color, hairstyle, and gender in the portrait photo.
[0146] In this embodiment of the application, a multi-classification network model can identify physical features such as skin color, hairstyle, and gender of people in portrait photos, so that the clothing recommendation network can recommend clothing to users with a higher degree of matching with the user's physical features, making the recommendation results more accurate and thus improving the user experience.
[0147] It should be noted that the above description of identifying skin color, hairstyle, and gender using a multi-classification network model is merely an illustrative example. In practical applications of clothing recommendation network models, this multi-classification network model can also include classifications of other physical features, such as age group and whether someone wears glasses. Specific classification features can be adjusted and expanded according to actual needs to better meet the personalized needs of different user groups. This will not be elaborated upon here.
[0148] (iv) Generate outfit recommendations.
[0149] This application embodiment can generate clothing recommendations for characters based on a conditional generative adversarial network model.
[0150] In one possible implementation, before generating outfit recommendations for a person, the conditional generative adversarial network (GAN) model can be trained based on training samples, which include multiple different images. Optionally, the input data of the GAN model may include the training samples, the location information of each feature point in the images in the training samples identified by a feature point recognition network model, and physical features such as gender, skin color, and hairstyle identified by a multi-classification network model in the images in the training samples.
[0151] It should be noted that the conditional generative adversarial network (GAN) model includes a generator and a discriminator. The generator can be used to generate one or more outfit recommendations based on the location information of each feature point in the input data, as well as classification information such as skin color, hairstyle, and gender. Optionally, the outfit recommendation can be one or more recommendation images, which can be used to instruct the outfit recommendation network model to generate outfit recommendations for a person.
[0152] Correspondingly, the discriminator can output a recommendation probability value based on the real images in the training samples in the input data and the recommended images generated by the generator. This recommendation probability value can represent the probability that the recommended image is a real image.
[0153] Furthermore, the conditional generative adversarial network (GAN) model can be trained adversarially based on input data. After multiple rounds of iterative training, once the generator and discriminator in the GAN model reach a preset standard, the GAN model can be used to generate clothing recommendations for users based on images, the location information of each feature point identified by the image through a feature point recognition network model, and physical features such as gender, skin color, and hairstyle identified by the image through a multi-classification network model.
[0154] Optionally, the aforementioned preset standard can be that the generator can generate highly realistic recommended images, while the discriminator has difficulty distinguishing these recommended images from real images.
[0155] In this embodiment, a conditional generative adversarial network model can generate clothing recommendations that better match the user's physical features, making the recommendation results more accurate and thus improving the user experience.
[0156] It is understood that the clothing recommendation network model provided in this application embodiment adopts the same end-to-end neural network model. That is, based on different recognition targets, such as different body features, feature extraction and target recognition are performed separately through multiple learning links. In this way, each learning link can focus on the feature learning of a specific target, thereby improving the specificity and accuracy of recognition. As a result, the clothing recommendation network model generates clothing recommendations that are more suitable for the person's body features, resulting in better recommendation performance and improved user satisfaction.
[0157] Figure 8 An image processing method according to an embodiment of this application is illustrated. The method includes:
[0158] S801, Acquire the target image.
[0159] The target image can be a photograph containing a human figure, which can be understood as the portrait photograph mentioned above. This target image can be a portrait photograph taken in real-time by an electronic device, or a portrait photograph stored in an electronic device; there are no restrictions here.
[0160] S802. Obtain first information of the target image based on the first model; the first model is used to obtain the position information of human body parts in the image input to the first model, and the first information includes the position information of human body parts in the target image.
[0161] Here, "human body parts" can be understood as the feature points mentioned above, meaning that human body parts can include, but are not limited to, the forehead, chin, both sides of the cheeks, both sides of the shoulders, both sides of the waist, both knees, and both ankles. "First information" can be understood as the feature point information mentioned above, meaning that the first information can include the positional information of the human body parts in the target image. Optionally, the positional information can be in array form or coordinate form; there is no limitation here.
[0162] The first model can be understood as the feature point recognition network model mentioned above. The model structure and implementation process of the first model can be found in the relevant description of the feature point recognition network model above, and will not be repeated here. It can be understood that for an image input into the first model, the first model can obtain the positional information of human body parts in the image. In this embodiment, after the electronic device acquires the target image, it can input the target image into the first model and can identify the positional information of human body parts in the target image based on the first model.
[0163] S803. Obtain second information of the target image based on the second model; the second model is used to obtain the gender information, skin color information and / or hairstyle information of the person in the image input to the second model, and the second information includes the gender information, skin color information and / or hairstyle information of the person in the target image.
[0164] In this embodiment, gender can be categorized as male and female, skin color as black, white, and yellow, and hairstyle as long and short. Gender information can be understood as the classification information of the gender of a person in the target image by the second model. This classification information can be, for example, probability information, meaning the second model can identify the probability that a person in the target image is male and the probability that a person is female.
[0165] Similarly, skin color information can be understood as the second model's classification information of the skin color of the person in the target image, and hairstyle information can be understood as the second model's classification information of the hairstyle of the person in the target image. These will not be elaborated on here.
[0166] The second model can be understood as the multi-classification network model mentioned above, and its structure or implementation process can be found in the relevant description of the multi-classification network model above, which will not be repeated here. It can be understood that for an image input into the second model, the second model can obtain the image's gender information, skin color information, and / or hairstyle information. In this embodiment, after the electronic device acquires the target image, it can input the target image into the second model and obtain the person's gender information, skin color information, and / or hairstyle information based on the second model.
[0167] It should be noted that the second information in this application embodiment, including the person's gender, skin color, and / or hairstyle information, is merely an illustrative description. In the actual application of the image processing method provided in this application embodiment, the second model can also identify other features of the person. For example, the second model can also identify the person's age group and whether they wear glasses; correspondingly, the second information can also include the person's age group information, glasses wearing information, etc. Specific classification features can be adjusted and expanded according to actual needs to better meet the personalized needs of different user groups.
[0168] S804. Based on the target image, the first information, the second information, and the third model, obtain the clothing image of the person in the target image; the third model is used to obtain the physical features of the person in the image input to the third model based on the output information of the first model and the output information of the second model, and determine the clothing image of the person based on the physical features of the person.
[0169] The output information of the first model can be understood as the positional information of human body parts in the image input to the first model. Similarly, the output information of the second model can be understood as the gender, skin color, and / or hairstyle information of the image input to the second model. The third model can be understood as the conditional generative adversarial network model mentioned above, and the structure or implementation process of the third model can be found in the relevant description of the conditional generative adversarial network model mentioned above, which will not be repeated here.
[0170] The physical characteristics of a person can be referred to in the description above. These characteristics include, but are not limited to, facial features such as face shape, body proportions, and weight, as well as gender, skin color, and hairstyle. It can be understood that the third model can derive facial features such as face shape, body proportions, and weight based on the positional information of body parts. The specific implementation process can be found in the description above and will not be repeated here.
[0171] Correspondingly, the third model can also obtain physical features such as gender, skin color, and hairstyle based on a person's gender information, skin color information, and / or hairstyle information. For example, when each piece of information is represented by a probability, the third model can select the feature value with the highest probability value as the corresponding physical feature of the person. Assuming that in a person's skin color information, the probability value for black is 0.8, the probability value for white is 0.2, and the probability value for yellow is 0.1, then the third model can determine that the person's skin color is black.
[0172] Furthermore, the third model can also determine a person's clothing image based on their physical features. This clothing image can be understood as the clothing recommendation image mentioned above. For the specific implementation process, please refer to the description of generating clothing recommendation images above; it will not be repeated here. In other words, the person in this clothing image is generated based on the person in the input image.
[0173] Specifically, when the image input to the first model is the target image, the output information of the first model can be understood as the first information. Similarly, when the image input to the second model is the target image, the output information of the second model can be understood as the second information. It can be understood that by inputting the target image, the first information, and the second information into the third model, an image of the clothing worn by the person in the target image can be obtained. The clothing worn by the person in this clothing image can be understood as clothing suitable for the person's physical characteristics.
[0174] In this embodiment, the electronic device can obtain the physical features of a person based on a target image input by the user, as well as a first model, a second model, and a third model. Furthermore, it can generate an outfit image that matches the physical features of the person in the target image based on the third model. This allows users to more intuitively see the actual effect of the outfit recommendations on themselves, improving the user's visual experience and increasing user satisfaction.
[0175] Optional, in Figure 8Based on the corresponding embodiment, obtaining the first information of the target image based on the first model includes: inputting the target image into the first model; the first model includes one or more first fusion convolutional modules; extracting features of the person in the target image based on the one or more first fusion convolutional modules to obtain a first feature image output by the first target fusion convolutional module; the first target fusion convolutional module is any one of the one or more first fusion convolutional modules; wherein, the size of the first feature image output by the first target fusion convolutional module is smaller than the size of the feature image input to the first target fusion convolutional module, and the number of channels of the first feature image output by the first target fusion convolutional module is greater than the number of channels of the feature image input to the first target fusion convolutional module; performing convolution, pooling, and fully connected operations on the first feature image output by the one or more first fusion convolutional modules to obtain a second feature image; the second feature image includes the first information.
[0176] The first fusion convolution module can be understood as the fusion convolution module in the feature point recognition network model mentioned above. For details, please refer to the above text. Figure 6 The description of the fusion convolution module in the embodiment will not be repeated here. The first target fusion convolution module can be understood as the fusion convolution module corresponding to path 1 in the fusion convolution module group of the feature point recognition network model above.
[0177] That is, the size of the first feature map output by the first target fusion convolution module is half the size of the feature image input to the first target fusion convolution module, and the number of channels of the first feature map output by the first target fusion convolution module is twice the number of channels of the feature image input to the first target fusion convolution module.
[0178] The first feature image can be understood as feature map b mentioned above. Convolution operation can be performed through the third module 503 mentioned above, pooling operation can be performed through the fourth module 504 mentioned above, and fully connected operation can be performed through the fifth module 505 mentioned above. For details, please refer to the relevant descriptions above, which will not be repeated here.
[0179] That is, after performing convolution, pooling, and fully connected operations on feature map b, a second feature image can be obtained. The second feature image can be understood as the feature map output by module 505 mentioned above. This feature map can include first information, that is, the feature map can be used to indicate the positional information of human body parts.
[0180] For example, assuming that the human body parts in this embodiment are the forehead, chin, both sides of the cheeks, both sides of the shoulders, both sides of the waist, both knees, and both ankles (12 parts in total), the second feature map can be a 1×1×24 feature map. These 24 values can correspond to the horizontal and vertical coordinates of the 12 human body parts, indicating the positional information of each human body part in the target image. Similarly, if the human body parts also include other parts of the person, the number of channels in the second feature image can be twice the number of human body parts, to indicate the horizontal and vertical coordinates of each human body part respectively.
[0181] In this embodiment of the application, after inputting the target image into the first model, the first information of the target image can be obtained, that is, the position information of each human body part of the person in the target image, which lays the foundation for subsequent calculation of the person's facial features, body proportions, and weight based on the position information of the human body parts.
[0182] Optional, in Figure 8 Based on the corresponding embodiment, obtaining the second information of the target image based on the second model includes: inputting the target image into the second model; the second model includes one or more second fusion convolutional modules; extracting features from the person in the target image based on the one or more second fusion convolutional modules to obtain a third feature image output by the second target fusion convolutional module; the second target fusion convolutional module is any one of the one or more second fusion convolutional modules; performing convolution, pooling, and fully connected operations on the third feature image output by the one or more second fusion convolutional modules to obtain a fourth feature image; the fourth feature image includes the second information; wherein, the fully connected operation includes a first fully connected layer, a second fully connected layer, and a third fully connected layer, the first fully connected layer being used to output information related to the person's gender in the target image; the second fully connected layer being used to output information related to the person's skin color in the target image; and the third fully connected layer being used to output information related to the person's hairstyle in the target image.
[0183] The second fusion convolutional module can be understood as the fusion convolutional module in the multi-classification network model mentioned above. For details, please refer to the above text. Figure 6 The description of the fusion convolution module in the embodiments will not be repeated here. The second target fusion convolution module can be understood as the fusion convolution module corresponding to path 1 in the fusion convolution module group of the multi-classification network model above.
[0184] Similar to the first target fusion convolutional module, the size of the third feature map output by the second target fusion convolutional module is half the size of the feature image input to the second target fusion convolutional module, and the number of channels of the third feature map output by the second target fusion convolutional module is twice the number of channels of the feature image input to the second target fusion convolutional module.
[0185] Furthermore, the third feature image output by the second target fusion convolution module can be subjected to convolution, pooling, and fully connected operations to obtain the fourth feature image. The convolution and pooling operations can be performed by the eighth module 703 and the ninth module 704 mentioned above, respectively, and will not be elaborated further here.
[0186] The fully connected operation can be executed by module 705 in the above text. That is, the first fully connected layer can be understood as the gender learning layer in the above text, the second fully connected layer can be understood as the skin color learning layer in the above text, and the first fully connected layer can be understood as the hairstyle learning layer in the above text. For details, please refer to the relevant descriptions in the above text, which will not be repeated here.
[0187] In this embodiment, the second model can identify second information about a person in the target image, namely, the person's gender, skin color, and hairstyle. This improves the accuracy of the electronic device in acquiring physical features such as gender, skin color, and hairstyle, and allows the clothing image output by the third model to better match the person's physical features, thus enhancing user satisfaction.
[0188] Optional, in Figure 8 Based on the corresponding embodiments, an image of the clothing of a person in the target image is obtained based on the target image, first information, second information, and third model, including: obtaining facial shape information, body proportion information, and / or weight information of the person in the target image based on the first information; inputting the target image, facial shape information, body proportion information, and / or weight information of the person in the target image, as well as gender information, skin color information, and / or hairstyle information of the person in the target image into the third model; and obtaining the clothing image of the person in the target image based on the third model.
[0189] In this embodiment, the face shape of a person can be categorized as round, oval, or oblong. Body proportion information may include, for example, the waist-to-shoulder ratio, leg-to-body ratio, and upper-to-lower-body ratio. Fat / thin information may include the degree of fatness / thinness of various parts of the body, such as the degree of fatness / thinness of the waist and chest. For details, please refer to the relevant description of the degree of fat / thinness of the person above, which will not be repeated here.
[0190] In this embodiment, the clothing in the clothing image of the person in the target image obtained based on the third model can be clothing that highly matches the person's physical characteristics. This makes the clothing recommendations more accurate and improves user satisfaction.
[0191] Optional, in Figure 8Based on the corresponding embodiments, the first information includes the coordinate information of the person's forehead in the target image, the coordinate information of the person's chin in the target image, the coordinate information of the two sides of the person's cheeks in the target image, the coordinate information of the two sides of the person's shoulders in the target image, the coordinate information of the two sides of the person's waist in the target image, the coordinate information of the two sides of the person's knees in the target image, and the coordinate information of the two sides of the person's ankles in the target image.
[0192] The coordinate information can include the x-coordinate and y-coordinate of each human body part in the image. Using these coordinates, the electronic device can determine the position of each human body part within the target image.
[0193] For example, the coordinates of a person's forehead, chin, and sides of their cheeks can be used to calculate their face shape. Similarly, the positional information of all or part of a person's body parts can be used to calculate their body proportions or other physical characteristics such as weight. The specific calculation process can be found in the description above and will not be repeated here.
[0194] In this embodiment, obtaining detailed and accurate coordinate information of various body parts in the target image helps in analyzing facial features such as face shape, body proportions, and weight. It also helps improve the accuracy of clothing recommendations and enhances the user experience.
[0195] Optional, in Figure 8 Based on the corresponding embodiments, the facial shape information, body proportion information, and / or weight information of the person in the target image are obtained based on the first information, including: determining the face length information of the person based on the coordinate information of the person's forehead and the coordinate information of the person's chin in the target image; determining the face width information of the person based on the coordinate information of the two sides of the person's cheeks in the target image; calculating the ratio of the face length information to the face width information to determine the face shape information of the person in the target image; determining the height proportion information of the person in the target image based on the coordinate information of the two sides of the person's shoulders and the coordinate information of the two sides of the person's ankles in the target image; determining the weight proportion information of the person in the target image based on the coordinate information of the two sides of the person's waist and the coordinate information of the two sides of the person's knees in the target image; and calculating the ratio of the height proportion information to the weight proportion information to determine the body proportion information of the person in the target image.
[0196] The facial shape information can include the aspect ratio of the face, which can be understood as the aspect ratio mentioned above. For details on calculating the aspect ratio, please refer to the description above; it will not be repeated here. Specifically, the electronic device can determine the face length based on the difference between the vertical coordinates of the forehead and the vertical coordinates of the chin. Similarly, the electronic device can determine the face width based on the difference between the horizontal coordinates of the two sides of the face.
[0197] It should be noted that the above process of determining facial features based on the forehead, chin, and sides of the cheeks is only an illustrative description. In actual applications, the human body parts used to calculate facial features are not limited to these few fixed human body parts, but can be flexibly adjusted according to needs to adapt to the diverse characteristics of different users.
[0198] Furthermore, embodiments of this application can also calculate the body proportions of a person. Specifically, the electronic device can calculate the center position of the shoulders based on the coordinates of the two sides of the shoulders in the target image, and can calculate the center position of the ankles based on the coordinates of the two sides of the ankles in the target image. Further, the height proportions of the person are determined based on the difference between the vertical coordinates of these two center positions.
[0199] Similarly, the electronic device can determine a person's waist circumference based on the difference in the horizontal coordinates of the two sides of the waist in the target image, and the knee circumference based on the difference in the horizontal coordinates of the two knees in the target image. Furthermore, the electronic device can determine the person's body proportions based on these waist and knee circumferences.
[0200] Furthermore, electronic devices can determine a person's body proportions based on the ratio of this height ratio information to their weight ratio information.
[0201] It should be noted that the method described above for determining the body proportions of a person in a target image using height and weight ratio information is merely an exemplary description. In the actual application of this application embodiment, other body proportion information may also be included, such as waist-to-shoulder ratio, leg-to-body ratio, and upper-to-lower-body ratio. The calculation methods for these body proportions can be found in the above descriptions of the calculation of waist-to-shoulder ratio, leg-to-body ratio, and upper-to-lower-body ratio, and will not be repeated here.
[0202] In this embodiment, the coordinate information of various body parts of a person can be used to determine facial features such as face shape and body proportions. This allows the clothing images generated by the electronic device to better match the person's face shape and body proportions, improving the accuracy of clothing images recommended to the user and enhancing the user experience.
[0203] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0204] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the method steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0205] This application embodiment can divide the apparatus for implementing the method into functional modules based on the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0206] like Figure 9 The diagram shows a schematic of a chip provided in an embodiment of this application. The chip 900 includes one or more processors 901, a communication line 902, a communication interface 903, and a memory 904.
[0207] In some implementations, memory 904 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.
[0208] The methods described in the embodiments of this application can be applied to or implemented by processor 901. Processor 901 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of processor 901 or by instructions in software form. The processor 901 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. Processor 901 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0209] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 904, and processor 901 reads the information in memory 904 and, in conjunction with its hardware, completes the steps of the above method.
[0210] The processor 901, memory 904 and communication interface 903 can communicate with each other through communication line 902.
[0211] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.
[0212] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from a website site, computer, server, or data center to another website site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).
[0213] This application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium capable of transferring a computer program from one place to another. The storage medium can be any target medium accessible by a computer.
[0214] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may also include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers.
[0215] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
Claims
1. An image processing method applied to electronic devices, characterized in that, The method includes: Acquire a target image containing a human body; The target image is input into the first model, and the first model is used to extract features from the target image to obtain the position information of human body parts in the target image. The position information of human body parts includes the coordinate information of the forehead, chin, cheeks, shoulders, waist, knees, and ankles in the target image. The target image is input into the second model, and the second model performs feature extraction on the target image to obtain the gender information, skin color information and hairstyle information of the person in the target image; The target image, the location information of the human body parts, the gender information, skin color information, and hairstyle information of the person are input into the third model. The third model obtains the physical features of the person based on the location information of the human body parts, the gender information, skin color information, and hairstyle information. Based on the physical features of the person, it determines the clothing image to be recommended, performs a geometric transformation on the clothing image based on the location information of the human body parts, and obtains a recommended outfit image based on the target image and the geometrically transformed clothing image. The first model includes a first module, a second module, and a first output network. The second module includes at least one fused convolutional module group, which includes a first fused convolutional module and a second fused convolutional module. The step of inputting the target image into a first model and extracting features from the target image using the first model to obtain the location information of human body parts in the target image includes: The target image is input into the first module, and the first module performs feature extraction on the target image to obtain a first initial feature image; The first initial feature image is input into the second module, and features are extracted from the first initial feature image through the first fusion convolution module and the second fusion convolution module to obtain the first feature image; The first feature image is input into the first output network, and convolution, pooling and full connection are performed on the first feature image to obtain the position information of the human body part; The step of extracting features from the first initial feature image using the first fusion convolution module and the second fusion convolution module includes: The feature image input to the first fusion convolution module is sequentially processed through a first 1×1 convolutional layer and a compression and activation module to perform feature fusion between feature map channels, and weights are applied to the fused features to obtain a second feature map. The feature image input to the first fusion convolutional module is sequentially processed through a depthwise separable convolutional layer and a second 1×1 convolutional layer to perform feature fusion within the neighborhood of the feature map channel, resulting in a third feature map. A first 3×3 convolutional layer is used to perform spatial and channel feature fusion on the feature image input to the first fusion convolutional module to obtain a fourth feature map; The second feature map, the third feature map, the fourth feature map, and the feature image input to the first fusion convolution module are concatenated to obtain the fifth feature map; The fifth feature image is convolved by a second 3×3 convolutional layer to obtain a sixth feature image; the size of the sixth feature image is half the size of the feature image input to the first fusion convolutional module, and the number of channels of the sixth feature image is twice the number of channels of the feature image input to the first fusion convolutional module. The sixth feature image is input into the second fusion convolution module for feature extraction.
2. The method according to claim 1, characterized in that, The second model includes a sixth module, a seventh module, and a second output network. The seventh module includes at least one second fused convolutional module group, and the second output network includes a first fully connected layer, a second fully connected layer, and a third fully connected layer. The step of inputting the target image into the second model and extracting features from the target image using the second model to obtain gender information, skin color information, and hairstyle information of the person in the target image includes: The target image is input into the sixth module, and the sixth module performs feature extraction on the target image to obtain a second initial feature image; The second initial feature image is input into the seventh module, and the second initial feature image is extracted by the at least one second fusion convolution module group to obtain the third feature image; The third feature image is input into the second output network, and the third feature image is convolved and pooled through the second output network; The feature images obtained by pooling are input into the first fully connected layer, the second fully connected layer, and the third fully connected layer, respectively, to obtain the gender information, skin color information, and hairstyle information of the person in the target image.
3. The method according to claim 2, characterized in that, The process of obtaining physical characteristics of a person based on the location information of the body parts, the person's gender, skin color, and hairstyle includes: Based on the location information of the human body parts, facial features, body proportions, and / or weight information of the person in the target image are obtained; The gender, skin color, and hairstyle of the person in the target image are obtained based on the person's gender information, skin color information, and hairstyle information.
4. The method according to claim 3, characterized in that, Based on the location information of the human body parts, facial features, body proportions, and / or weight information of the person in the target image are obtained, including: Based on the coordinate information of the person's forehead in the target image and the coordinate information of the person's chin in the target image, the face length information of the person is determined; Based on the coordinate information of the two sides of the person's cheeks in the target image, the face width information of the person is determined; Calculate the ratio of the face length information to the face width information of the person to determine the face shape information of the person in the target image; Based on the coordinate information of the person's shoulders on both sides in the target image and the coordinate information of the person's ankles on both sides in the target image, the height ratio information of the person in the target image is determined; Based on the coordinate information of the waist sides of the person in the target image and the coordinate information of the knees on both sides of the person in the target image, the fat-to-thin ratio information of the person in the target image is determined; Calculate the ratio of the height ratio of the person to the fat ratio of the person to determine the body proportion information of the person in the target image.
5. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-4.
6. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-4.
8. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-4.
Citation Information
Patent Citations
Garment style matching algorithm based on face features
CN115482577A