Clothing recommendation method and device

By obtaining user vector representation and combining diffusion model and virtual fitting technology to generate clothing display videos, the problem of difficult to meet personalized needs in the existing clothing recommendation system is solved, and the success rate and user experience improvement of personalized clothing recommendations are achieved.

CN120338929APending Publication Date: 2025-07-18BEIJING WODONG TIANJUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510501076.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing clothing recommendation system cannot accurately reflect the user's personalized needs, and the recommendation results are displayed in a single form, resulting in low recommendation success rate and poor user experience.

Method used

By obtaining user vector representation, combining diffusion model and virtual fitting technology to generate clothing display videos, using noise image encoding for reverse denoising and decoding processing, personalized clothing recommendation results are generated, and clothing matching that users are interested in is displayed through dynamic videos.

Benefits of technology

It significantly improves the success rate of clothing recommendations, meets users' personalized clothing selection and matching needs, improves user experience, and provides high reference value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338929A_ABST
    Figure CN120338929A_ABST
Patent Text Reader

Abstract

The invention discloses a costume recommendation method and device, and relates to the technical field of computers. A specific embodiment of the method comprises the steps of obtaining a recommendation user identifier in response to a clothing recommendation request, and obtaining a user vector representation according to the recommendation user identifier; the noise picture codes are subjected to reverse denoising and decoding processing in combination with user vector representation, recommended costume pictures are obtained, and the noise picture codes are obtained by conducting forward diffusion on costume picture codes corresponding to all costume pictures in a costume picture library; and generating a costume display video according to the recommended costume picture, and performing recommended costume display based on the costume display video. According to the embodiment, personalized clothes selection and matching requirements of the user are met visually and greatly through dynamic video display, high reference value can be brought to the user, the recommendation success rate is remarkably increased, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and particularly to a method and device for clothing recommendation. Background Art

[0002] A recommendation system is an information filtering system designed to predict and recommend products or services that a user may be interested in. It analyzes the user's historical behavior, preferences, context information, and the characteristics of items, and uses algorithms such as collaborative filtering, content-based recommendation, and hybrid recommendation to generate a personalized recommendation list, and displays the recommendation list in the form of pictures and texts. Recommendation systems are widely used in fields such as e-commerce, social media, music, and video streaming services, aiming to improve the user experience, increase user engagement, and promote sales and content consumption.

[0003] However, in the scenario of clothing recommendation, the existing recommendation systems cannot accurately reflect the personalized needs of users, and only display the recommendation results through pictures and texts. The display form is single and cannot bring enough reference value to users, resulting in a low recommendation success rate. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a method and device for clothing recommendation, which can greatly meet the user's personalized clothing selection and matching needs visually through dynamic video display, can bring higher reference value to users, significantly improve the recommendation success rate, and improve the user experience.

[0005] To achieve the above object, according to one aspect of the embodiments of the present invention, a method for clothing recommendation is provided, including: Responding to a clothing recommendation request, obtaining a recommended user identifier, and obtaining a user vector representation according to the recommended user identifier; Combining the user vector representation to perform reverse denoising and decoding processing on a noise image encoding, to obtain a recommended clothing image, where the noise image encoding is obtained by forward diffusion of the clothing image encoding corresponding to each clothing image in a clothing image library; Generating a clothing display video according to the recommended clothing image, and performing recommended clothing display based on the clothing display video.

[0006] Optionally, the user vector representation is obtained by inputting user attribute features into the feature embedding layer of the recommendation model to extract embedding vectors; wherein, the recommendation model further includes a deep learning network layer, and the recommendation model is trained in the following manner: obtaining user attribute features, user context features, and item attribute features of multiple users within a specified time period; respectively inputting the user attribute features, user context features, and item attribute features of each user into the feature embedding layer to extract embedding vectors, obtaining the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user; inputting the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user into the deep learning network layer to train the recommendation model.

[0007] Optionally, the clothing picture code corresponding to each clothing picture in the clothing picture library is generated by the encoder of the vector diffusion model; wherein, the vector diffusion model further includes a decoder and a denoising layer; combining the user vector representation to perform reverse denoising and decoding processing on the noise picture code to obtain a recommended clothing picture, including: through the denoising layer, combining the user vector representation to perform reverse denoising processing on the noise picture code; through the decoder, performing decoding processing on the picture code after reverse denoising to obtain a recommended clothing picture.

[0008] Optionally, the vector diffusion model is trained in the following manner: obtaining clothing pictures purchased by target users with clothing purchase records within a specified time period and user vector representations; inputting the clothing pictures purchased by the target users into the encoder for image encoding to obtain the clothing picture code corresponding to the clothing pictures; inputting the user vector representation and the clothing picture code corresponding to the clothing pictures into the denoising layer for forward diffusion and reverse denoising processing to obtain the processed clothing picture code; inputting the processed clothing picture code into the decoder for image decoding to train the vector diffusion model.

[0009] Optionally, generating a clothing display video according to the recommended clothing picture, including: inputting the recommended clothing picture and a preset virtual fitting model of a fitting model character image to generate a fitting result picture; inputting the fitting result picture and a preset fitting pose into a pose video generation model to generate a clothing display video.

[0010] Optionally, the recommended clothing picture includes multiple clothing pictures corresponding to a set of matching clothing.

[0011] According to another aspect of the embodiments of the present invention, there is provided a clothing recommendation device, including: A user vector acquisition module, configured to, in response to a clothing recommendation request, acquire a recommended user identifier, and acquire a user vector representation according to the recommended user identifier; A recommended clothing determination module, which is used to perform reverse denoising and decoding processing on the noise image encoding by combining the user vector representation to obtain a recommended clothing image. The noise image encoding is obtained by performing forward diffusion on the clothing image encoding corresponding to each clothing image in the clothing image library. A recommended clothing display module, which is used to generate a clothing display video based on the recommended clothing image and perform recommended clothing display based on the clothing display video.

[0012] Optionally, the user vector representation is obtained by inputting user attribute features into the feature embedding layer of the recommendation model to extract embedding vectors. Among them, the recommendation model further includes a deep learning network layer, and the recommendation model is trained in the following manner: Obtain the user attribute features, user context features, and item attribute features of multiple users within a specified time period; respectively input the user attribute features, user context features, and item attribute features of each user into the feature embedding layer to extract embedding vectors, obtaining the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user; input the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user into the deep learning network layer to perform recommendation model training.

[0013] Optionally, the clothing image encoding corresponding to each clothing image in the clothing image library is generated by the encoder of the vector diffusion model. Among them, the vector diffusion model further includes a decoder and a denoising layer. The recommended clothing determination module is further used to: perform reverse denoising processing on the noise image encoding by combining the user vector representation through the denoising layer; perform decoding processing on the reverse denoised image encoding through the decoder to obtain a recommended clothing image.

[0014] Optionally, the vector diffusion model is trained in the following manner: Obtain the clothing images purchased by target users with clothing purchase records within a specified time period and the user vector representation; input the clothing images purchased by the target users into the encoder for image encoding to obtain the clothing image encoding corresponding to the clothing images; input the user vector representation and the clothing image encoding corresponding to the clothing images into the denoising layer for forward diffusion and reverse denoising processing to obtain the processed clothing image encoding; input the processed clothing image encoding into the decoder for image decoding to perform vector diffusion model training.

[0015] Optionally, the recommended clothing display module is further used to: input the recommended clothing image and a preset fitting model character image into a virtual fitting model to generate a fitting result image; input the fitting result image and a preset fitting pose into a pose video generation model to generate a clothing display video.

[0016] Optionally, the recommended clothing pictures include multiple clothing pictures corresponding to a set of dressed clothing.

[0017] According to another aspect of the embodiments of the present invention, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the clothing recommendation method provided by the embodiments of the present invention.

[0018] According to another aspect of the embodiments of the present invention, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the clothing recommendation method provided by the embodiments of the present invention.

[0019] According to still another aspect of the embodiments of the present invention, there is provided a computer program product, including a computer program, which when executed by a processor, implements the clothing recommendation method provided by the embodiments of the present invention.

[0020] One embodiment of the above invention has the following advantages or beneficial effects: By responding to a clothing recommendation request, obtaining a recommended user identifier, and obtaining a user vector representation according to the recommended user identifier; combining the user vector representation to perform reverse denoising and decoding processing on the noise picture encoding to obtain a recommended clothing picture, where the noise picture encoding is obtained by forward diffusion of the clothing picture encoding corresponding to each clothing picture in the clothing picture library; generating a clothing display video according to the recommended clothing picture, and performing recommended clothing display based on the clothing display video. Through the technical solution of integrating a recommendation system and image / video generation technology, using a diffusion model combined with a user vector representation to generate clothing that users are interested in, personalized recommendation results can be generated; on this basis, using virtual fitting technology to generate pictures of a model wearing clothing, and using picture-to-video technology to generate detailed videos of the model showing the clothing. Compared with the graphic display recommendations of traditional e-commerce, dynamic video display can greatly meet the user's personalized clothing selection and matching needs visually, can bring higher reference value to users, significantly improve the recommendation success rate, and improve the user experience.

[0021] The further effects of the above non-conventional optional methods will be described in combination with specific embodiments below. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them: Figure 1 is a schematic diagram of the main steps of the clothing recommendation method according to the embodiments of the present invention; Figure 2 is a schematic diagram of the structure of a recommendation model according to an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a vector diffusion model according to an embodiment of the present invention; Figure 4 It is a schematic diagram of the generation process of a clothing display video according to an embodiment of the present invention; Figure 5 It is a schematic diagram of the main modules of a clothing recommendation device according to an embodiment of the present invention; Figure 6 It is an exemplary system architecture diagram to which an embodiment of the present invention can be applied; Figure 7 It is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing an embodiment of the present invention. Detailed implementation manners

[0023] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following.

[0024] It should be noted that in the technical solutions disclosed by the present invention, in terms of the collection, gathering, update, analysis, processing, use, transmission, storage, etc. of the user's personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the security of the user's personal information, network security, and national security.

[0025] In order to solve the technical problems existing in the prior art, the present invention provides a clothing recommendation method. By integrating the recommendation system and image / video generation technology, a diffusion model is used to generate clothing that the user is interested in. When training the diffusion model, a user vector representation is added as a guiding condition to generate clothing that the user is interested in. On this basis, virtual fitting technology is used to generate pictures of models wearing clothing, and image-to-video technology is used to generate detailed videos of models showing clothing. Compared with the text and picture display recommendations of traditional e-commerce, the dynamic video display can greatly meet the user's personalized clothing selection and matching needs visually, bring higher reference value to the user, significantly improve the recommendation success rate, and improve the user experience.

[0026] Figure 1 It is a schematic diagram of the main steps of a clothing recommendation method according to an embodiment of the present invention. As Figure 1 shown, the clothing recommendation method according to an embodiment of the present invention mainly includes the following steps S101 to step S103.

[0027] Step S101: In response to a clothing recommendation request, obtain a recommended user identifier, and obtain a user vector representation based on the recommended user identifier. When a user accesses the main page of an e-commerce platform or the main page of the clothing category, a clothing recommendation request will be triggered. At this time, obtain the user name, platform ID (Identity Document), etc. corresponding to this user as the recommended user identifier, so as to obtain the user vector representation based on the recommended user identifier.

[0028] According to an embodiment of the present invention, the user vector representation is obtained by inputting user attribute features into the feature embedding layer of the recommendation model for extracting embedding vectors. Among them, the recommendation model further includes a deep learning network layer, and the recommendation model is trained in the following manner: obtain the user attribute features, user context features, and item attribute features of multiple users within a specified time period; respectively input the user attribute features, user context features, and item attribute features of each user into the feature embedding layer for extracting embedding vectors to obtain the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user; input the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user into the deep learning network layer for training the recommendation model.

[0029] In an embodiment of the present invention, in order to optimize the user vector representation to better reflect the personalized interests of the user, a recommendation model can be pre-trained to obtain model parameters that meet the requirements, and then the feature embedding layer in the recommendation model can be used to generate the user vector representation. Among them, the recommendation model mainly includes a feature embedding layer and a deep learning network layer. The feature embedding layer is used to extract feature embedding vectors for user attribute features to obtain a user vector representation; the deep learning network layer is used to implement a recommendation algorithm based on the embedding vectors extracted by the feature embedding layer to obtain a recommendation list for output.

[0030] In an embodiment of the present invention, when training the recommendation model, the training data used includes, for example, the user attribute features, user context features, and item attribute features of multiple users within a specified time period. Among them, the specified time period is, for example, the most recent year, the most recent month, etc., which can be set according to business needs; the user attribute features are, for example, user age, gender, preferred style, etc.; the user context features are, for example, the geographical address and login time when the user usually logs in to the e-commerce platform; the item attribute features are, for example, the price and operation duration of the items that the user operates through behaviors such as browsing, searching, consulting, and purchasing. It should be noted that the items here generally refer to products, videos, graphic and text posts, etc. that can be recommended by the recommendation system.

[0031] Figure 2 is a schematic structural diagram of the recommendation model according to an embodiment of the present invention. As Figure 2As shown, in an embodiment of the present invention, the recommendation model has an input layer, a feature embedding layer, a deep learning network layer, and an output layer. Among them, the input data of the input layer includes feature data such as user attribute features, user context features, and item attribute features of multiple users within a specified time period. The input feature data is subjected to feature embedding mapping processing through the feature embedding layer to convert the input feature data from a discrete high-dimensional space to a continuous low-dimensional space. As shown in the figure, after the user attribute features, user context features, and item attribute features pass through the feature embedding layer, a user attribute embedding vector, a user context embedding vector, and an item attribute embedding vector are obtained respectively.

[0032] After that, the embedding vectors output by the feature embedding layer (including the user attribute embedding vector, the user context embedding vector, and the item attribute embedding vector) are input into the deep learning network layer to learn and estimate the browsing probability of users for each item. Commonly used recommendation systems in the industry often design many features to increase the upper limit of the model, while the present invention hopes to pre-train the personalized vector representation of users, so complex feature construction is not required to ensure that the model can focus on optimizing the user vector representation. Among them, the deep learning network layer can be any model commonly used in the recommendation model, and the present invention does not specifically limit its structure.

[0033] After the embedding vectors pass through the deep learning network layer, the browsing probability of users for each item can be obtained, and the recommendation list can be determined accordingly.

[0034] In an embodiment of the present invention, when training the recommendation model, the cross-entropy loss function can be used as the objective function of the recommendation model, and the gradient descent method is used to optimize the network parameters of the recommendation model. Among them, the cross-entropy loss function is as follows: ; Here, y is the label of the user clicking on the item (that is: if the user clicks on the item, the label value is 1; if the user does not click on the item, the label value is 0), is the estimated browsing probability of the user clicking on the item for browsing.

[0035] After training the recommendation model, that is, optimizing the network parameters of the recommendation model, correspondingly, the network parameters of the feature embedding layer of the recommendation model are optimized. After that, by inputting the user attribute features into the feature embedding layer, an optimized user vector representation can be obtained, which can more accurately reflect the personalized interests of users. In the e-commerce recommendation system, the user vector representation can not only directly reflect the user's preference for clothing matching, but also express the user's preference for various categories of goods and consumption habits. These implicit information can indirectly express the user's clothing style, so as to provide personalized clothing recommendations for users.

[0036] In an embodiment of the present invention, the user vector representation can be obtained by acquiring the user attribute features of the recommended user through the recommended user identifier after generating a clothing recommendation request and inputting the user attribute features into the feature embedding layer of the recommendation model; or it can be obtained by inputting the user attribute features into the feature embedding layer of the recommendation model after the user attribute features change or the recommendation model is updated, and then saved in the database for acquisition when needed. The present invention does not limit the acquisition method of the user vector representation.

[0037] Step S102: Perform reverse denoising and decoding processing on the noisy image encoding in combination with the user vector representation to obtain a recommended clothing image. The noisy image encoding is obtained by forward diffusion of the clothing image encoding corresponding to each clothing image in the clothing image library. In an embodiment of the present invention, determining the recommended clothing image in combination with the user vector representation can better perform personalized clothing matching recommendations and meet the user's personalized clothing selection and matching needs. Among them, for example, the clothing image library stores the images of all clothing items on the e-commerce platform, and each clothing image corresponds to a clothing image encoding, so as to facilitate subsequent processing based on the clothing image encoding to determine the recommended service image.

[0038] According to an embodiment of the present invention, the recommended clothing image includes multiple clothing images corresponding to a set of matching clothing. When the existing recommendation system makes clothing recommendations, it mostly recommends single clothing items, while the recommendation system of the present invention can make matching clothing recommendations when making clothing recommendations, so as to better meet the user's matching needs.

[0039] According to an embodiment of the present invention, the clothing image encoding corresponding to each clothing image in the clothing image library is, for example, generated by the encoder of the vector diffusion model. Among them, the vector diffusion model may further include a decoder and a denoising layer. When this step S102 performs reverse denoising and decoding processing on the noisy image encoding in combination with the user vector representation to obtain a recommended clothing image, it may specifically include: performing reverse denoising processing on the noisy image encoding in combination with the user vector representation through the denoising layer; and performing decoding processing on the reverse-denoised image encoding through the decoder to obtain the recommended clothing image.

[0040] In an embodiment of the present invention, a vector diffusion model (Latent Diffusion Models) is used to generate recommended clothing pictures, which can reduce the computational complexity and accelerate the convergence speed of model training. Among them, the diffusion model is a generative model that defines a forward diffusion process, decomposing the original picture into countless tiny noise steps, adding a small amount of Gaussian noise to the picture each time, and turning the original picture into a completely random noise map. The denoising process uses a denoising network to gradually denoise the noise map until a clear picture is restored. The most important part of the diffusion model is the denoising network, which is a U-Net network (a fully convolutional network and an image segmentation model with a symmetric encoding-decoding structure). The denoising network is optimized by optimizing the objective loss function, thereby optimizing the diffusion model. Among them, the objective loss function is, for example, an optimization objective function based on predicted noise. Once the model training is completed, the model can obtain denoised data from random noise through iterative denoising.

[0041] Figure 3 It is a schematic structural diagram of the vector diffusion model according to an embodiment of the present invention. As Figure 3 shown, the vector diffusion model of the embodiment of the present invention mainly includes an encoder, a decoder, and a denoising layer. Among them, the encoder is used to encode the original clothing picture to obtain a clothing picture encoding; the denoising layer is used to perform forward diffusion on the clothing picture encoding to obtain a noise picture encoding corresponding to each clothing picture, and perform backward denoising processing on the noise picture encoding corresponding to each clothing picture in combination with the user vector representation; the decoder is used to perform decoding processing on the picture encoding after backward denoising, and finally output the dressed clothing picture.

[0042] According to an embodiment of the present invention, the vector diffusion model is, for example, trained in the following manner: obtaining the clothing pictures and user vector representations purchased by target users with clothing purchase records within a specified time period; inputting the clothing pictures purchased by the target users into the encoder for image encoding to obtain a clothing picture encoding corresponding to the clothing pictures; inputting the user vector representation and the clothing picture encoding corresponding to the clothing pictures into the denoising layer for forward diffusion and backward denoising processing to obtain a processed clothing picture encoding; inputting the processed clothing picture encoding into the decoder for image decoding to perform vector diffusion model training.

[0043] In an embodiment of the present invention, the specified time period is, for example, the last year, the last month, etc., and can be flexibly set according to business needs. It is possible to obtain target users with clothing purchase records by obtaining the user log data of the e-commerce platform during the specified time period, and then obtain the clothing pictures purchased by the target users, as well as the user vector representations corresponding to the target users, and use the clothing pictures purchased by the target users and the user vector representations as the training data of the vector diffusion model.

[0044] When training a vector diffusion model, the clothing pictures purchased by the target user can be input into an encoder for image encoding to obtain the clothing picture encoding corresponding to the clothing pictures; the user vector representation and the clothing picture encoding corresponding to the clothing pictures are input into a denoising layer for forward diffusion and reverse denoising processing to obtain the processed clothing picture encoding; the processed clothing picture encoding is input into a decoder for image decoding to train the vector diffusion model. During the model training process, the target loss function is calculated after each round of training. Finally, the model corresponding to the target loss function that meets the requirements is used as the trained vector diffusion model.

[0045] After the vector diffusion model is trained, whether the user has a clothing consumption record or not, as long as the user vector representation is provided, the vector diffusion model can generate recommended clothing pictures based on the noise picture encoding generated during model training and the user vector representation. The recommended clothing pictures include multiple clothing pictures corresponding to a set of matching clothing, so that matching clothing recommendations can be made to better meet the user's dressing needs.

[0046] Step S103: Generate a clothing display video based on the recommended clothing pictures and perform recommended clothing display based on the clothing display video.

[0047] According to an embodiment of the present invention, generating a clothing display video based on the recommended clothing pictures may specifically include: inputting the recommended clothing pictures and a preset virtual fitting model of a fitting model character image to generate a fitting result picture; inputting the fitting result picture and a preset fitting pose into a pose video generation model to generate a clothing display video. In the embodiment of the present invention, after personalized recommended clothing pictures are generated for the user, the user cannot see the specific fitting effect from the static clothing pictures and cannot judge the actual wearing situation. Based on this, the present invention adopts virtual fitting technology and pose video generation technology to generate a dynamic clothing display video, so that the user can intuitively see the fitting effect and improve the conversion rate of the recommended clothing. In specific implementation, the virtual fitting model is, for example, the OOTDiffusion model (an open-source generation model for fitting), and the pose video generation model is, for example, the AnimateAnyone model (an open-source generation model for action video generation). It should be understood that other network models can also be selected for the virtual fitting model and the pose video generation model as long as the corresponding functions can be realized.

[0048] Figure 4 is a schematic diagram of the generation process of a clothing display video according to an embodiment of the present invention. As Figure 4As shown, in one embodiment of the present invention, when generating a clothing display video, first, the recommended clothing pictures and the preset virtual fitting model of the fitting model are input into the virtual fitting model to generate a fitting result picture of the model wearing the corresponding clothing; then, the fitting result picture and the preset fitting pose are input into the pose video generation model to generate a clothing display video. Among them, the preset fitting poses can include multiple fitting poses to more completely simulate the posture of the user wearing the clothing.

[0049] The dynamic clothing display video generated by this embodiment is a clothing display video generated by using the user's personalized interest guidance vector diffusion model by mining the user's consumption behavior and browsing behavior through the recommendation system to fully capture the user's personalized interest, which can achieve the effects of personalized recommendation and visual dynamic display of clothing.

[0050] Figure 5 It is a schematic diagram of the main modules of the clothing recommendation device according to an embodiment of the present invention. As Figure 5 shown, the clothing recommendation device 500 according to the embodiment of the present invention mainly includes a user vector acquisition module 501, a recommended clothing determination module 502, and a recommended clothing display module 503.

[0051] The user vector acquisition module 501 is configured to obtain a recommended user identifier in response to a clothing recommendation request, and obtain a user vector representation according to the recommended user identifier; The recommended clothing determination module 502 is configured to perform reverse denoising and decoding processing on the noise picture encoding in combination with the user vector representation to obtain a recommended clothing picture, and the noise picture encoding is obtained by forward diffusion of the clothing picture encoding corresponding to each clothing picture in the clothing picture library; The recommended clothing display module 503 is configured to generate a clothing display video according to the recommended clothing picture and perform recommended clothing display based on the clothing display video.

[0052] According to an embodiment of the present invention, the user vector representation is obtained by inputting user attribute features into the feature embedding layer of the recommendation model and performing embedding vector extraction. Among them, the recommendation model further includes a deep learning network layer, and the recommendation model is trained in the following manner: obtaining the user attribute features, user context features, and item attribute features of multiple users within a specified time period; respectively inputting the user attribute features, user context features, and item attribute features of each user into the feature embedding layer to perform embedding vector extraction, obtaining the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user; inputting the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user into the deep learning network layer to perform recommendation model training.

[0053] According to another embodiment of the present invention, the clothing picture code corresponding to each clothing picture in the clothing picture library is generated by the encoder of the vector diffusion model. Wherein, the vector diffusion model further includes a decoder and a denoising layer; the recommended clothing determination module 502 can specifically be used for: performing reverse denoising processing on the noise picture code by combining the user vector representation through the denoising layer; performing decoding processing on the picture code after reverse denoising through the decoder to obtain the recommended clothing picture.

[0054] According to yet another embodiment of the present invention, the vector diffusion model is trained in the following manner: obtaining the clothing pictures and user vector representations purchased by target users with clothing purchase records within a specified time period; inputting the clothing pictures purchased by the target users into the encoder for image coding to obtain the clothing picture codes corresponding to the clothing pictures; inputting the user vector representations and the clothing picture codes corresponding to the clothing pictures into the denoising layer for forward diffusion and reverse denoising processing to obtain the processed clothing picture codes; inputting the processed clothing picture codes into the decoder for image decoding to perform vector diffusion model training.

[0055] According to yet another embodiment of the present invention, the recommended clothing display module 503 can specifically be used for: inputting the recommended clothing picture and a preset virtual fitting model of a fitting model person image to generate a fitting result picture; inputting the fitting result picture and a preset fitting pose into the pose video generation model to generate a clothing display video.

[0056] According to yet another embodiment of the present invention, the recommended clothing picture includes multiple clothing pictures corresponding to a set of matching clothing.

[0057] According to the technical solution of the embodiment of the present invention, by responding to a clothing recommendation request, obtaining a recommended user identifier, and obtaining a user vector representation according to the recommended user identifier; combining the user vector representation to perform reverse denoising and decoding processing on the noise picture code to obtain the recommended clothing picture, and the noise picture code is obtained by performing forward diffusion on the clothing picture code corresponding to each clothing picture in the clothing picture library; generating a clothing display video according to the recommended clothing picture, and performing recommended clothing display based on the clothing display video. Through the technical solution, by integrating the recommendation system and image and video generation technologies, using the diffusion model to combine the user vector representation to generate clothing that the user is interested in, personalized recommendation results can be generated; on this basis, using virtual fitting technology, generating pictures of the model wearing clothing, and using picture-to-video technology, generating detailed videos of the model showing clothing. Compared with the graphic display recommendation of traditional e-commerce, the dynamic video display can greatly meet the user's personalized clothing selection and matching needs visually, can bring higher reference value to the user, significantly improve the recommendation success rate, and improve the user experience.

[0058] Figure 6An exemplary system architecture 600 is shown to which the method for clothing recommendation or the apparatus for clothing recommendation according to the embodiments of the present invention may be applied.

[0059] like Figure 6 As shown, system architecture 600 may include terminal devices 601, 602, 603, network 604 and server 605. Network 604 is used to provide a medium for communication links between terminal devices 601, 602, 603 and server 605. Network 604 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0060] The user can use the terminal devices 601, 602, 603 to interact with the server 605 through the network 604 to receive or send messages, etc. Various communication client applications can be installed on the terminal devices 601, 602, 603, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for example).

[0061] The terminal devices 601 , 602 , and 603 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers, etc.

[0062] The server 605 may be a server that provides various services, such as a background management server (only an example) that provides support for shopping websites browsed by users using the terminal devices 601, 602, and 603. The background management server may perform processing such as obtaining the identification of the recommended user, obtaining the user vector representation, decoding the noise image, generating the clothing display video, etc. on the received clothing recommendation request and other data, and feed back the processing result (such as the generated clothing display video - only an example) to the terminal device.

[0063] It should be noted that the clothing recommendation method provided in the embodiment of the present invention is generally executed by the server 605 , and accordingly, the clothing recommendation device is generally set in the server 605 .

[0064] It should be understood that Figure 6 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0065] Reference below Figure 7 , which shows a schematic diagram of the structure of a computer system 700 of a terminal device or a server suitable for implementing an embodiment of the present invention. Figure 7 The terminal device or server shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0066] likeFigure 7 As shown in Figure 7 , the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded into a random access memory (RAM) 703 from a storage section 708. In the RAM 703, various programs and data required for the operation of the system 700 are also stored. The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0067] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is mounted on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.

[0068] Specifically, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit (CPU) 701, the above functions defined in the system of the present invention are executed.

[0069] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0070] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0071] The units or modules involved in the embodiments of the present invention can be implemented in software or in hardware. The described units or modules can also be provided in a processor. For example, it can be described as: a processor includes a user vector acquisition module, a recommended clothing determination module, and a recommended clothing display module. Among them, the names of these units or modules do not constitute a limitation to the units or modules themselves in some cases. For example, the user vector acquisition module can also be described as "a module for responding to a clothing recommendation request, acquiring a recommended user identifier, and acquiring a user vector representation according to the recommended user identifier".

[0072] As another aspect, the present invention also provides a computer-readable medium. The computer-readable medium can be included in the device described in the above embodiments; or it can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the device, the device includes: responding to a clothing recommendation request, acquiring a recommended user identifier, and acquiring a user vector representation according to the recommended user identifier; performing reverse denoising and decoding processing on the noise picture encoding in combination with the user vector representation to obtain a recommended clothing picture, where the noise picture encoding is obtained by forward diffusion of the clothing picture encoding corresponding to each clothing picture in the clothing picture library; generating a clothing display video according to the recommended clothing picture, and performing recommended clothing display based on the clothing display video.

[0073] According to the technical solution of the embodiments of the present invention, by responding to a clothing recommendation request, acquiring a recommended user identifier, and acquiring a user vector representation according to the recommended user identifier; performing reverse denoising and decoding processing on the noise picture encoding in combination with the user vector representation to obtain a recommended clothing picture, where the noise picture encoding is obtained by forward diffusion of the clothing picture encoding corresponding to each clothing picture in the clothing picture library; generating a clothing display video according to the recommended clothing picture, and performing recommended clothing display based on the clothing display video, through the technical solution that combines a recommendation system and image / video generation technology, uses a diffusion model in combination with a user vector representation to generate clothing that the user is interested in, personalized recommendation results can be generated; on this basis, using virtual fitting technology, pictures of a model wearing clothing are generated, and using picture-to-video technology, detailed videos of the model showing the clothing are generated. Compared with the graphic display recommendations of traditional e-commerce, the dynamic video display can greatly meet the user's personalized clothing selection and matching needs visually, can bring higher reference value to the user, significantly improve the recommendation success rate, and improve the user experience.

[0074] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for clothing recommendation, characterized in that, Including: In response to a clothing recommendation request, obtain a recommended user identifier, and obtain a user vector representation according to the recommended user identifier; Perform reverse denoising and decoding processing on the noise image encoding in combination with the user vector representation to obtain a recommended clothing image, where the noise image encoding is obtained by forward diffusion of the clothing image encoding corresponding to each clothing image in the clothing image library; Generate a clothing display video based on the recommended clothing image, and perform recommended clothing display based on the clothing display video.

2. The method according to claim 1, characterized in that, The user vector representation is obtained by inputting user attribute features into the feature embedding layer of the recommendation model and extracting embedding vectors; Wherein, the recommendation model further includes a deep learning network layer, and the recommendation model is trained in the following manner: Obtain user attribute features, user context features, and item attribute features of multiple users within a specified time period; Input the user attribute features, user context features, and item attribute features of each user into the feature embedding layer to extract embedding vectors, and obtain the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user; Input the user attribute embedding vector, user context embedding vector, and item attribute embedding vector of each user into the deep learning network layer to train the recommendation model.

3. The method according to claim 1, wherein The clothing image encoding corresponding to each clothing image in the clothing image library is generated by the encoder of the vector diffusion model; Wherein, the vector diffusion model further includes a decoder and a denoising layer; Performing reverse denoising and decoding processing on the noise image encoding in combination with the user vector representation to obtain a recommended clothing image, including: Through the denoising layer, perform reverse denoising processing on the noise image encoding in combination with the user vector representation; Decode the image encoding after reverse denoising through the decoder to obtain a recommended clothing image.

4. The method according to claim 3, wherein The vector diffusion model is trained in the following manner: Obtain the clothing images purchased by target users with clothing purchase records within a specified time period and user vector representations; Input the clothing images purchased by the target users into the encoder for image encoding to obtain the clothing image encoding corresponding to the clothing images; Input the user vector representation and the clothing image encoding corresponding to the clothing images into the denoising layer for forward diffusion and reverse denoising processing to obtain the processed clothing image encoding; Input the processed clothing image encoding into the decoder for image decoding to train the vector diffusion model.

5. The method according to claim 1, wherein Generating a clothing display video based on the recommended clothing image, including: Input the recommended clothing image and a preset virtual fitting model of a fitting model character image to generate a fitting result image; Input the fitting result image and a preset fitting pose into the pose video generation model to generate a clothing display video.

6. The method according to any one of claims 1-5, characterized in that, The recommended clothing image includes multiple clothing images corresponding to a set of dressed clothing.

7. A device for clothing recommendation, characterized in that, Including: A user vector acquisition module, configured to, in response to a clothing recommendation request, obtain a recommended user identifier, and obtain a user vector representation according to the recommended user identifier; A recommended clothing determination module, configured to perform reverse denoising and decoding processing on the noise picture encoding in combination with the user vector representation to obtain a recommended clothing picture, where the noise picture encoding is obtained by forward diffusion of the clothing picture encoding corresponding to each clothing picture in the clothing picture library; A recommended clothing display module, configured to generate a clothing display video based on the recommended clothing picture and perform recommended clothing display based on the clothing display video.

8. An electronic device, characterized in that, Comprising: One or more processors; A storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.