Device and method

JPWO2025248775A1Pending Publication Date: 2025-12-04
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Filing Date
2024-05-31
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing devices fail to accurately depict how eyeglasses affect the appearance of a user's eyes, making it difficult to select suitable eyeglasses that match the user's actual appearance when worn.

Method used

An apparatus and method that includes an acquisition unit for facial and lens information, an estimation unit to determine lens characteristics, and a prompt generation unit to instruct an image generation model to adjust eye size and distortion based on lens information, generating an image of the user wearing eyeglasses as seen by others.

Benefits of technology

Enables users to visualize how they will appear in eyeglasses, accounting for lens effects, facilitating accurate selection without physical try-on.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

A device according to the present invention generates an appearance image of when glasses are being worn, said image approximating how a user would appear to others when actually wearing the glasses. The device comprises a prompt generation unit that generates a prompt that is for providing instruction for generating an appearance image of a state in which the user is wearing the glasses in question and is for providing instruction for generating the appearance image in which the eye region of the user has been adjusted on the basis of lens information.
Need to check novelty before this filing date? Find Prior Art

Description

Apparatus and method

[0001] The present disclosure relates to an apparatus and method for generating images of a user wearing eyeglasses.

[0002] Patent Literature 1 describes a device that generates a composite image of a user wearing eyeglasses. This device generates the composite image by superimposing an image of the eyeglasses on an image of the user's face.

[0003] Japanese Patent Application Laid-Open No. 2003-30494

[0004] When others look at the eyes of a user wearing eyeglasses, for example, the size of the eyes may appear smaller than they actually are due to the influence of the eyeglass lenses. The above-mentioned device simply superimposes an image of the eyeglasses on the user's face image, and does not generate an image that reflects the appearance of the eyes caused by the influence of the lenses. As a result, the user cannot understand how others will see them when actually wearing the eyeglasses, making it difficult to find eyeglasses that suit them.

[0005] Therefore, this disclosure describes an apparatus and method that can generate an image of the user wearing glasses that is similar to how the user would look to others when actually wearing the glasses.

[0006] The device according to the present disclosure includes an acquisition unit that acquires a user's face image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which an image is to be generated, and a prompt generation unit that instructs the generation of an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information, and generates a prompt that instructs the generation of the image image adjusted for the user's eyes based on the lens information.

[0007] According to the present disclosure, it is possible to generate an image of the user wearing glasses that is similar to how the user would look to others when actually wearing the glasses.

[0008] Fig. 1 is a diagram showing the overall configuration of a system including a prompt generation device of the present disclosure. Fig. 2 is a block diagram showing the functional configuration of the prompt generation device. Fig. 3(a) is a diagram showing a facial image of a user when not wearing glasses. Fig. 3(b) is a diagram showing an image image of a user wearing target glasses. Fig. 4 is a diagram showing an example of a prompt generated by a prompt generation unit. Fig. 5 is a flowchart showing the flow of a process for generating an image image in the prompt generation device. Fig. 6 is a diagram showing an example of the hardware configuration of an information processing device.

[0009] The present disclosure will be described with reference to the accompanying drawings. Whenever possible, the same parts are designated by the same reference numerals and redundant description will be omitted.

[0010] FIG. 1 is a diagram showing the overall configuration of a system including a prompt generation device according to the present disclosure. This system generates an image of a face representing a user wearing eyeglasses. As shown in FIG. 1, this system includes a prompt generation device (device) 100 and an image generation model 200. The prompt generation device 100 and the image generation model 200 are communicatively connected via a network (NM). A user terminal 300 is communicatively connected to the prompt generation device 100 via the network.

[0011] The user terminal 300 transmits a request to the prompt generation device 100 to generate an image of what the user will look like when wearing the glasses. The prompt generation device 100 generates a prompt in response to the request for image generation from the user terminal 300, transmits it to the image generation model 200, and receives a response (the generated image). The prompt generation device 100 transfers the image to the user terminal 300. For example, the user terminal 300 displays the image transferred from the prompt generation device 100 on a display screen or the like. This allows the user of the user terminal 300 to check the image of what the user will look like when wearing the glasses.

[0012] The image generation model 200 generates an answer based on the prompt sent from the prompt generation device 100. The image generation model 200 sends the generated answer to the prompt generation device 100. The image generation model 200 may be located within the prompt generation device 100. The prompt generation device 100 or the image generation model 200 may also be located within the user terminal 300.

[0013] The image generation model 200 is, for example, a deep convolutional generative adversarial network (DCGAN), which is a generative artificial intelligence (AI) model that can generate an image in response to a request. Note that DCGAN is just one example, and other image generation models may be used. For example, partial autoencoders (VAE), flow-based models, and diffusion models are available.

[0014] A generative AI model is a model that can generate content in response to a prompt containing input information, according to the instructions, context, question, and output format indicated by the prompt, and return the content as response information. The prompt can also include input information, in which case the generative AI model generates response information targeted at the input information. In this disclosure, a prompt refers to information indicating an instruction or question entered by a user in an interactive system, such as an interaction with a generative AI model or a command line interface (CLI).

[0015] 2 is a block diagram showing the functional configuration of the prompt generation device 100. As shown in FIG. 2, the prompt generation device 100 includes an acquisition unit 110, an estimation unit 120, a selection unit 130, and a prompt generation unit 140.

[0016] The acquisition unit 110 acquires a user's facial image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which images are to be generated. The user's facial image acquired by the acquisition unit 110 is image data of the user's face. For example, this image data of the user's face is a photograph of the user's face. The image generation model 200 generates an image of the user wearing eyeglasses based on the facial image acquired by the acquisition unit 110. Therefore, the facial image acquired by the acquisition unit 110 may be, for example, an image of the user's face when not wearing eyeglasses. For example, the user inputs a facial image to the user terminal 300. In this case, the acquisition unit 110 acquires the facial image input by the user from the user terminal 300. However, the source of the facial image acquired by the acquisition unit 110 is not particularly limited. The user's facial image may be stored in advance in a server or the like. In this case, the acquisition unit 110 acquires the user's facial image from the server.

[0017] The acquisition unit 110 acquires eyeglasses-wearing images in which the user wears eyeglasses and non-glasses-wearing images in which the user does not wear eyeglasses. The eyeglasses-wearing images and non-glasses-wearing images are images that include the user's eyes. In the present disclosure, the eyes include the eyes and the area around the eyes of the face. When the acquisition unit 110 acquires the eyeglasses-wearing images and non-glasses-wearing images, the acquisition unit 110 can use (acquire) the non-glasses-wearing images as the user's facial image described above. For example, the user inputs the eyeglasses-wearing images and non-glasses-wearing images into the user terminal 300. For example, the acquisition unit 110 acquires the eyeglasses-wearing images and non-glasses-wearing images input by the user terminal 300. However, the source of the eyeglasses-wearing images and non-glasses-wearing images acquired by the acquisition unit 110 is not particularly limited. The eyeglasses-wearing images and non-glasses-wearing images may be stored in advance on a server, similar to the user's facial image described above.

[0018] The eyeglasses lens information acquired by the acquisition unit 110 is information that indicates the characteristics of the lenses, etc. The lens information includes information related to the eyeglasses power. Information related to the eyeglasses power is information that indicates the degree of vision correction by the lenses. Information related to the eyeglasses power can also be said to be information related to the refractive index of the lenses. Information related to the eyeglasses power may be the eyeglasses power itself. Information related to the eyeglasses power may be any information that can identify the eyeglasses power. In this embodiment, an example will be described in which the eyeglasses power itself is used as information related to the eyeglasses power.

[0019] The lens information also includes a lens type indicating the type of eyeglass lens. Examples of lens types include a spherical lens, an aspherical lens, a double-sided aspherical lens, a single-focus lens, and a multifocal lens. The lens information may be input by a user. In this case, the user inputs the lens information into the user terminal 300. The acquisition unit 110 may acquire the lens information from the user terminal 300. The lens information may also be estimated by the estimation unit 120. In this case, the acquisition unit 110 acquires the lens information estimated by the estimation unit 120.

[0020] The target glasses in the target glasses information acquired by the acquisition unit 110 are glasses that are the subject of image generation when generating an image of the user's face when wearing the glasses. The target glasses information of the target glasses is information for identifying the target glasses. The target glasses information may be any information that can identify the target glasses, such as the product name, identification number, or model number.

[0021] The target eyeglasses information may be input by a user of the system for generating an image. In this case, the user inputs the target eyeglasses information to, for example, the user terminal 300. The acquisition unit 110 may acquire the target eyeglasses information input by the user from, for example, the user terminal 300. In this way, the acquisition unit 110 can acquire the target eyeglasses information input by the user.

[0022] Furthermore, the selection unit 130 may select target glasses for which an image is to be generated. In this case, the acquisition unit 110 acquires target glasses information of the target glasses selected by the selection unit 130.

[0023] The estimation unit 120 estimates lens information for eyeglasses. Based on the image of the user wearing eyeglasses and the image of the user not wearing eyeglasses, the estimation unit 120 estimates lens information for the eyeglasses worn by the user in the image of the user wearing eyeglasses. Here, the estimation unit 120 estimates the eyeglasses' power as lens information. Here, there is a correlation between the change in eye size between when eyeglasses are worn and when not wearing eyeglasses and the eyeglasses' power in the image of the user wearing eyeglasses and the image of the user not wearing eyeglasses. The estimation unit 120 pre-stores the correlation between the change in eye size and the eyeglasses' power. Therefore, the estimation unit 120 can, for example, calculate the change in eye size using the image of the user wearing eyeglasses and the image of the user not wearing eyeglasses, and estimate the eyeglasses' power based on the stored correlation. The estimation unit 120 can recognize the eyes in the image and calculate the change in size using well-known image processing techniques. Furthermore, the estimation unit 120 may estimate the lens type, which is lens information, based on the image of the user wearing eyeglasses and the image of the user not wearing eyeglasses, based on distortion of the image of the eyes in the image of the user wearing eyeglasses and the image of the user not wearing eyeglasses.

[0024] The selection unit 130 selects target glasses for generating an image. The selection unit 130 selects target glasses based on at least one of the user's attribute information and the shape of each part of the user's face obtained from a facial image. Here, the selection unit 130 can select recommended glasses to be worn by the user as target glasses based on this information. The selection unit 130 may select multiple target glasses.

[0025] For example, when selecting target eyeglasses based on the user's attribute information, the selection unit 130 acquires the attribute information from the user terminal 300 or an external server (not shown). Examples of the user's attribute information include information that comprehensively includes personal information such as "age," "gender," and "past eyeglass purchase history," as well as behavioral tendencies. The selection unit 130 can select target eyeglasses based on the acquired attribute information, for example, by referring to pre-stored eyeglass purchase data for each attribute. Here, the selection unit 130 can select, for example, eyeglasses that are popular (best-selling) among people with the attribute to which the user belongs as the target eyeglasses.

[0026] For example, when selecting target eyeglasses based on the shape of each part of the user's face obtained from the user's face image, the selection unit 130 recognizes the shape of each part from the user's face image acquired by the acquisition unit 110. The selection unit 130 can recognize the shape of the face from the face image by using well-known image processing techniques. For example, the selection unit 130 recognizes the shape of each part of the user's face, such as the face shape (round face, square face, etc.), nose height, and eye shape and size. The selection unit 130 can select target eyeglasses based on the recognized shape of each part of the face, for example, by referring to recommended eyeglasses data stored in advance for each face shape.

[0027] The selection unit 130 can also select target glasses based on review information about the glasses. The review information about the glasses may be impressions from users of the glasses, positive / negative ratings voted on the glasses, or descriptions of the glasses. This review information may be, for example, information available on the Internet (including social networking sites) or information stored in a storage unit that stores multiple pieces of review information.

[0028] The selection unit 130 can select, for example, highly rated glasses as the target glasses based on the review information. In this case, the selection unit 130 can select, as the target glasses, glasses that are highly rated when worn by a person with the same attributes as the user. Furthermore, the selection unit 130 can select, as the target glasses, glasses that are highly rated when worn by a person with the user's face shape (a person with a similar face shape) based on the recognized shapes of each part of the user's face.

[0029] The selection unit 130 may select a plurality of target glasses and present them to the user. The selection unit 130 may present the selected plurality of target glasses as recommendations to the user via the user terminal 300, allowing the user to select a target pair of glasses. In this case, the acquisition unit 110 may acquire target glasses information of the target glasses selected by the user from the target glasses presented as recommendations, as target glasses information of the target glasses selected by the user for generating an image when wearing the glasses. In addition, the acquisition unit 110 may acquire target glasses information of the remaining target glasses selected by the selection unit 130 as target glasses information of the target glasses selected by the selection unit 130 for generating an image when wearing the glasses.

[0030] The prompt generation unit 140 generates a prompt for causing the image generation model 200 to generate an image. This prompt includes a task, which is an instruction content, and input information. The prompt generation unit 140 generates a prompt that instructs the image generation model 200 to generate an image of the user wearing the target eyeglasses, based on the user's face image and target eyeglasses information acquired by the acquisition unit 110. The prompt generation unit 140 also generates a prompt that includes an instruction to generate an image of the user wearing the target eyeglasses, based on the lens information acquired by the acquisition unit 110.

[0031] In this embodiment, the prompt generation unit 140 generates a prompt including an instruction to generate an image in which the size and distortion of the user's eyes have been adjusted based on the lens information, as an image in which the user's eyes have been adjusted based on the lens information. In this way, the prompt includes the target glasses information and the lens information.

[0032] As described above, when a user wearing eyeglasses looks at their eyes, the size of the eyes appears different than they actually are due to the influence of the eyeglass lenses. This change in eye size correlates with the eyeglass prescription. Therefore, the prompt generation unit 140 includes in the prompt an instruction to generate an image taking into account how the eyes appear, which changes depending on the eyeglass prescription. Furthermore, when a user looks at their eyes wearing eyeglasses, the distortion of the eyes varies depending on the type of eyeglass lens. For example, the distortion of the eyes varies depending on the lens type, such as spherical lenses and aspherical lenses. Therefore, the prompt generation unit 140 includes in the prompt an instruction to generate an image taking into account how the eyes are distorted, which differs depending on the lens type.

[0033] The prompt generation unit 140 generates prompts including instructions to generate images of the face of a user wearing the target eyeglasses from different angles. The prompt generation unit 140 may also generate prompts including instructions to generate moving images of the face of a user wearing the target eyeglasses.

[0034] The prompt generation unit 140 transmits the user's facial image and the generated prompt to the image generation model 200, and acquires an image (answer) corresponding to the prompt from the image generation model 200. The prompt generation unit 140 transmits the image acquired from the image generation model 200 to the user terminal 300.

[0035] The image generation model 200 generates an image of what the user will look like when wearing the target eyeglasses, based on the facial image and prompt acquired from the prompt generation device 100. The image generation model 200 transmits the generated image to the prompt generation device 100.

[0036] The image generation model 200 can identify the target glasses for which an image is to be generated based on the target glasses information included in the prompt. For example, the image generation model 200 can acquire images of the target glasses from an external server, etc. In addition, for example, the prompt generation unit 140 may also transmit images of the target glasses to the image generation model 200 when transmitting a prompt, etc. to the image generation model 200.

[0037] The image generation model 200 generates an image of what the user will look like when wearing the target eyeglasses, based on a facial image of the user and an image of the target eyeglasses. For example, as shown in Fig. 3(a), the image generation model 200 acquires a facial image A of the user P along with a prompt. As shown in Fig. 3(b), the image generation model 200 generates an image B based on the acquired facial image A, so as to represent how the user P will look when wearing the target eyeglasses M.

[0038] Furthermore, the image generation model 200 generates an image B in which the size and distortion of the user's eyes have been adjusted based on the lens information included in the prompt. For example, when a third party looks at the eyes of a user wearing eyeglasses, the eyes may appear smaller than they actually are depending on the prescription of the eyeglasses. Furthermore, when the eyes appear smaller, the facial contours around the eyes as seen through the lenses may become discontinuous with the facial contours outside the lenses. Furthermore, when a third party looks at the eyes of a user wearing eyeglasses, the distortion of the user's eyes differs depending on the lens type of the eyeglasses. Therefore, the image generation model 200 generates an image B in which the size of the eyes, the facial contours, and distortion of the eyes have been adjusted based on the prescription and lens type of the eyeglasses so that the user P will look as they do when wearing the target eyeglasses M.

[0039] In the example shown in FIG. 3( b), the lens for the right eye of the target glasses M is a lens with a power (glasses power) for correcting vision. The lens for the left eye of the target glasses M is simply a transparent plate with no power. Therefore, the image generation model 200 adjusts, for example, the size of the area around the right eye R of the user P to a size corresponding to the power of the glasses. In addition, in conjunction with adjusting the size of the area around the right eye R, the image generation model 200 adjusts the position of the facial contour line L near the area around the right eye R as seen through the right eye lens of the target glasses M. In the example shown in FIG. 3( b), the size of the area around the right eye R of the user P is adjusted to be smaller, and accordingly, the position of the facial contour line L near the area around the right eye R as seen through the right eye lens is also shifted toward the center of the face.

[0040] The prompt may also specify multiple pairs of glasses as target glasses for which an image is to be generated. For example, the prompt may include target glasses input by the user and target glasses selected by the selection unit 130 as target glasses for image generation. In this case, the image generation model 200 generates an image of the user wearing each of the multiple target glasses.

[0041] Furthermore, the image generation model 200 generates respective images of the face of the user wearing the target eyeglasses viewed from different angles based on the instructions of the prompt. For example, the image generation model 200 generates respective images of the user viewed from multiple directions, such as an image of the user wearing the target eyeglasses viewed from the front, an image of the user viewed from diagonally forward to the right, an image of the user viewed from above diagonally forward to the left, and an image of the user viewed from diagonally rear to the right.

[0042] The image generation model 200 can generate an image of the user wearing the target glasses by using, for example, the refractive index for each eyeglass power, the degree of distortion depending on the eyeglass lens type, and the results of a skeletal analysis of the user's face based on the user's facial image, as well as a trained model.

[0043] Next, a specific example of a prompt generated by the prompt generating unit 140 will be described. Here, it is assumed that the user inputs information about target glasses for which an image image is to be generated, and that the selection unit 130 selects target glasses recommended for the user.

[0044] As shown in FIG. 4, the prompt generation unit 140 generates a prompt including the following instruction content as an example of a task: "By utilizing the pre-learned refractive index calculation information for each eyeglass prescription and lens type information, please generate an image of the user wearing the target eyeglasses specified below with lenses having the prescription. Also, please pay attention to the following when generating the image: - To check how the image will look when viewed from various angles, generate five or more image images viewed from different angles. - Use the information on "eyeglass prescription" and "lens type" to adjust the size of the eyes and distortion around the eyes in the image. - Also generate an image of the user wearing the recommended eyeglasses without changing the eyeglass prescription and lens type."

[0045] The prompt generating unit 140 generates a prompt including the following information as input information: "Facial image: face.png, Eyeglasses power: -1.5, Lens type: Type A, Target eyeglasses specified by user: BBB, Recommended eyeglasses: CCC, DDD."

[0046] In this example prompt, the content of the task column corresponds to an instruction to generate an image image of the user wearing the target glasses. The task column states, "To check how the image looks from various angles, generate five or more image images viewed from different angles." This corresponds to an instruction to generate image images of the user's face viewed from different angles. The task column states, "Using information about the 'eyeglasses power' and 'lens type', adjust the eye size and eye distortion in the image image." This corresponds to an instruction to generate an image image in which the size and distortion of the user's eyes have been adjusted based on the lens information. The task column states, "Also generate an image image of the user wearing the recommended glasses without changing the glasses power and lens type." This corresponds to an instruction to generate an image image of the user wearing the target glasses selected by the selection unit 130.

[0047] "Facial image: face.png" in the input information field is information specifying the facial image of the user. Note that the facial image data may be transmitted to the image generation model 200 by the prompt generation unit 140 together with a prompt, for example. "Eyeglasses power: -1.5" and "Lens type: Type A" in the input information field correspond to the lens information of the eyeglasses. "Target eyeglasses specified by user: BBB" in the input information field corresponds to the target eyeglasses entered by the user. "Recommended eyeglasses: CCC, DDD" in the input information field corresponds to the target eyeglasses information of the target eyeglasses selected by the selection unit 130.

[0048] Based on this prompt and the user's face image (face.png), the image generation model 200 generates an image of the user wearing the target eyeglasses BBB. The image adjusts the size and distortion of the eyes so that the image matches the lens information (eyeglasses power, lens type). Furthermore, the image generation model 200 also generates an image of the user wearing the target eyeglasses CCC and an image of the user wearing the target eyeglasses DDD.

[0049] The following describes the flow of a method for generating an image of a user wearing target eyeglasses, performed by the prompt generation device 100. FIG. 5 is a flowchart showing the flow of a process for generating an image in the prompt generation device. As shown in FIG. 5, the acquisition unit 110 of the prompt generation device 100 acquires a face image of the user and lens information of the eyeglasses (S101: acquisition step). If the user knows the lens information, the lens information may be input by the user. In this case, the acquisition unit 110 acquires the lens information input by the user. Alternatively, the estimation unit 120 may estimate the lens information based on an image of the user wearing eyeglasses and an image of the user not wearing eyeglasses. In this case, the acquisition unit 110 acquires the lens information estimated by the estimation unit 120. In this way, the lens information may be information input by the user, or information estimated by the estimation unit 120 based on an image of the user wearing eyeglasses and an image of the user not wearing eyeglasses.

[0050] Next, the acquisition unit 110 acquires target eyeglasses information for the target eyeglasses to be used when generating an image of the user wearing the target eyeglasses (S102: acquisition step). Note that the target eyeglasses information may be input (selected) by the user. For example, when shopping online for eyeglasses, the user may input (select) the target eyeglasses information when there are certain eyeglasses for which the user wants to generate an image of the eyeglasses being worn. In this case, the acquisition unit 110 acquires the target eyeglasses information input (selected) by the user. In addition, the target eyeglasses information may be selected by the selection unit 130. For example, the selection unit 130 may select eyeglasses recommended for the user as the target eyeglasses. In this case, the acquisition unit 110 acquires the target eyeglasses information for the target eyeglasses selected by the selection unit 130.

[0051] In this way, the acquisition unit 110 acquires target eyeglasses information input by the user if such information is available, and acquires target eyeglasses information of the target eyeglasses selected by the selection unit 130 if such information is available. In other words, the acquisition unit 110 can acquire at least one of the target eyeglasses information input by the user and the target eyeglasses information of the target eyeglasses selected by the selection unit 130.

[0052] The prompt generation unit 140 generates a prompt to generate an image of the user wearing the target eyeglasses based on the facial image and the target eyeglasses information, and also generates a prompt to generate the image with the size of the user's eyes adjusted based on the lens information (S103: prompt generation step). The prompt generation unit 140 transmits the generated prompt and the user's facial image to the image generation model 200 (S104: prompt transmission step). The image generation model 200 generates an image of the user wearing the target eyeglasses based on the prompt and the user's facial image, and transmits it to the prompt generation device 100. The prompt generation unit 140 acquires the image (answer) generated by the image generation model 200 (S105: image acquisition step). The image generation model 200 transmits the acquired image to the user terminal 300 (S106: transmission step). The user terminal 300 displays the acquired image on a display screen or the like. The user can thereby recognize what the target eyeglasses will look like when worn by checking the displayed image.

[0053] As described above, in the prompt generation device 100 of the present disclosure, the prompt generation unit 140 generates a prompt that instructs the user to generate an image of the user wearing the target eyeglasses based on a facial image and target eyeglass information. The prompt generation unit 140 also includes, in the generated prompt, an instruction to generate the image with the user's eyes adjusted based on the lens information. By generating an image using the image generation model 200 based on the generated prompt, the prompt generation device 100 can generate an image of the user wearing the eyeglasses that closely resembles how the user would look to others if they were actually wearing the eyeglasses.

[0054] This allows the user to check an image of how the target eyeglasses will look to others when worn, without actually wearing them. For example, when purchasing eyeglasses at an online shop, the user can check an image of how the user will look when wearing their favorite eyeglasses in advance.

[0055] The prompt generation unit 140 generates a prompt including an instruction to generate an image in which at least one of the user's eye size and eye distortion has been adjusted based on the lens information. If the prompt includes an instruction to adjust the eye size, the prompt generation device 100 can generate an image that takes into account the change in eye size. If the prompt includes an instruction to adjust the eye distortion, the prompt generation device 100 can generate an image that takes into account the eye distortion. This allows the user to view an image when wearing the glasses that more closely resembles how the user would look to others if they were actually wearing the glasses.

[0056] The lens information includes at least one of information regarding the eyeglasses power and information regarding the lens type. If the lens information includes the eyeglasses power, the prompt generation device 100 can generate an image in which the size of the eyes is adjusted according to the eyeglasses power. If the lens information includes the lens type, the prompt generation device 100 can generate an image in which the distortion of the eyes is adjusted according to the lens type.

[0057] The prompt generation unit 140 generates a prompt including instructions to generate images of the user's face when viewed from different angles. In this case, the prompt generation device 100 can obtain images of the user's face when wearing the target eyeglasses when viewed from various angles. This allows the user to more accurately understand how the user will look when wearing the target eyeglasses.

[0058] The prompt generation device 100 includes an estimation unit 120 that estimates lens information based on an image of a person wearing eyeglasses and an image of a person not wearing eyeglasses. This allows a user to obtain an image that reflects the lens information by inputting an image of a person wearing eyeglasses and an image of a person not wearing eyeglasses, even if the user does not know the lens information.

[0059] The acquisition unit 110 acquires information input by the user as target eyeglasses information for target eyeglasses for which an image is to be generated. This allows the prompt generation device 100 to generate an image in which the user is wearing target eyeglasses selected by the user, for example, their favorite target eyeglasses. The acquisition unit 110 also acquires target eyeglasses information for target eyeglasses selected by the selection unit 130 as target eyeglasses information for target eyeglasses for which an image is to be generated. This allows the prompt generation device 100 to generate an image in which the user is wearing target eyeglasses selected by the selection unit 130, for example, recommended target eyeglasses.

[0060] The prompt generator 140 may generate a prompt including an instruction to generate a video as an image of the user wearing the target eyeglasses. In this case, the user can more accurately grasp the image of wearing the target eyeglasses by watching the video.

[0061] The prompt generation unit 140 generates a prompt that instructs the user to generate an image of the user wearing the target eyeglasses. The prompt generation unit 140 transmits the generated prompt to the image generation model 200 and obtains an answer (image) generated by the image generation model 200. This allows the prompt generation device 100 to generate an image using the image generation model 200.

[0062] The method for generating an image performed by the prompt generation device 100 includes an acquisition step of acquiring a user's face image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which images are to be generated, and a prompt generation step of generating a prompt that instructs the user to generate an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information. The prompt generation step also generates a prompt that includes an instruction to generate an image with the user's eyes adjusted based on the lens information. In this way, this method can generate an image of the user wearing eyeglasses that closely resembles how the user would look to others when actually wearing the eyeglasses.

[0063] The above description is an example of the present disclosure and is not limited to the description. For example, the above-described prompts described in the present disclosure are provided for illustrative purposes and are not limited to the above-described examples. The instructions included in the prompts may be changed as appropriate. For example, eyeglass lens information may include information regarding the lens power and information other than the lens type. The prompt generation device 100 may be configured to generate an image in which either the size or distortion of the user's eyes is adjusted, rather than both.

[0064] The device and method of the present disclosure have the following configuration.

[0065] [1] A device comprising: an acquisition unit that acquires a user's face image, eyeglass lens information, and target eyeglass information for target eyeglasses for which images are to be generated; and a prompt generation unit that generates a prompt that instructs the device to generate an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information, and to generate the image with the size of the user's eyes adjusted based on the lens information. [2] The device described in [1] above, wherein the prompt includes an instruction to generate the image with the size and distortion of the user's eyes adjusted based on the lens information. [3] The device described in [1] or [2] above, wherein the lens information includes eyeglass prescription, or eyeglass prescription and lens type. [4] The device described in any of [1] to [3] above, wherein the prompt includes an instruction to generate the image when the user's face is viewed from different angles. [5] The device according to any one of [1] to [4] above, further comprising an estimation unit that estimates the lens information, wherein the acquisition unit acquires eyeglasses-wearing images in which the user wears eyeglasses and eyeglasses-not-wearing images in which the user does not wear eyeglasses, wherein the estimation unit estimates the lens information based on the eyeglasses-wearing images and the eyeglasses-not-wearing images, and wherein the acquisition unit acquires the lens information estimated by the estimation unit. [6] The device according to any one of [1] to [5] above, wherein the acquisition unit acquires the target eyeglasses information input by the user. [7] The device according to any one of [1] to [6] above, further comprising a selection unit that selects the target eyeglasses based on at least one of attribute information of the user and the shape of each part of the user's face obtained from the face image, wherein the acquisition unit acquires the target eyeglasses information of the target eyeglasses selected by the selection unit. [8] The device according to any one of [1] to [7] above, wherein the prompt includes an instruction to generate a video as the image. [9] The device described in any of [1] to [8] above, wherein the prompt generation unit transmits the facial image and the prompt to an image generation model and obtains the image corresponding to the prompt from the image generation model.

[10] A method including: an acquisition step of acquiring a user's face image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which an image is to be generated; and a prompt generation step of instructing the user to generate an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information, and generating a prompt instructing the user to generate the image with the size of the user's eyes adjusted based on the lens information.

[0066] The block diagrams used to explain the above embodiments show functional blocks. These functional blocks (components) are realized by any combination of hardware and / or software. Furthermore, the method for realizing each functional block is not particularly limited. That is, each functional block may be realized using a single device that is physically or logically coupled, or may be realized using two or more physically or logically separated devices that are connected directly or indirectly (e.g., via wire, wirelessly, etc.) and these multiple devices. The functional block may also be realized by combining the single device or multiple devices with software.

[0067] Functions include, but are not limited to, judgment, determination, judgment, calculation, computation, processing, derivation, investigation, search, confirmation, reception, transmission, output, access, resolution, selection, selection, establishment, comparison, assumption, expectation, consideration, broadcasting, notifying, communicating, forwarding, configuring, reconfiguring, allocating, mapping, and assignment. For example, a functional block (component) that performs transmission is called a transmitting unit or transmitter. As mentioned above, there are no particular limitations on how these functions are implemented.

[0068] For example, the prompt generation device 100 according to an embodiment of the present disclosure may function as a computer that performs processing for the method of generating an image of a user wearing target eyeglasses according to the present disclosure. Fig. 6 is a diagram illustrating an example of the hardware configuration of the prompt generation device 100 according to an embodiment of the present disclosure. The prompt generation device 100 described above may be physically configured as a computer device including a processor 1001, a memory 1002, a storage device 1003, a communication device 1004, an input device 1005, an output device 1006, a bus 1007, and the like.

[0069] In the following description, the term "device" may be interpreted as a circuit, a device, a unit, etc. The hardware configuration of the prompt generation device 100 may be configured to include one or more of the devices shown in the figures, or may be configured to exclude some of the devices.

[0070] Each function in the prompt generating device 100 is realized by loading specified software (programs) onto hardware such as the processor 1001 and memory 1002, causing the processor 1001 to perform calculations, control communication via the communication device 1004, and control at least one of reading and writing data in the memory 1002 and storage 1003.

[0071] The processor 1001, for example, runs an operating system to control the entire computer. The processor 1001 may be configured as a central processing unit (CPU) including an interface with peripheral devices, a control unit, an arithmetic unit, a register, etc. For example, the acquisition unit 110, the estimation unit 120, the selection unit 130, and the prompt generation unit 140 described above may be realized by the processor 1001.

[0072] The processor 1001 also loads programs (program code), software modules, data, etc. from at least one of the storage 1003 and the communication device 1004 into the memory 1002 and executes various processes in accordance with these programs. The programs used are those that cause a computer to execute at least some of the operations described in the above-described embodiments. For example, the acquisition unit 110, the estimation unit 120, the selection unit 130, and the prompt generation unit 140 may be implemented by a control program stored in the memory 1002 and running on the processor 1001, and similar implementations may be used for other functional blocks. While the above-described various processes have been described as being executed by a single processor 1001, they may also be executed simultaneously or sequentially by two or more processors 1001. The processor 1001 may be implemented by one or more chips. The programs may also be transmitted from a network via a telecommunications line.

[0073] The memory 1002 is a computer-readable recording medium and may be configured, for example, by at least one of a read-only memory (ROM), an erasable programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), a random access memory (RAM), etc. The memory 1002 may also be referred to as a register, a cache, a main memory (primary storage device), etc. The memory 1002 may store executable programs (program codes), software modules, etc. for implementing a prompt generation method according to one embodiment of the present disclosure.

[0074] Storage 1003 is a computer-readable recording medium, and may be composed of at least one of, for example, an optical disk such as a CD-ROM (Compact Disc ROM), a hard disk drive, a flexible disk, a magneto-optical disk (e.g., a compact disk, a digital versatile disk, a Blu-ray (registered trademark) disk), a smart card, a flash memory (e.g., a card, a stick, a key drive), a floppy (registered trademark) disk, a magnetic strip, etc. Storage 1003 may also be referred to as an auxiliary storage device. The above-mentioned storage medium may be, for example, a database, a server, or other appropriate medium including at least one of memory 1002 and storage 1003.

[0075] The communication device 1004 is hardware (transmission / reception device) for communicating between computers via at least one of a wired network and a wireless network, and is also referred to as, for example, a network device, a network controller, a network card, or a communication module. The communication device 1004 may be configured to include a high-frequency switch, a duplexer, a filter, a frequency synthesizer, etc. to realize at least one of frequency division duplex (FDD) and time division duplex (TDD). The communication device 1004 may be implemented with a transmitter and a receiver that are physically or logically separated.

[0076] The input device 1005 is an input device (e.g., a keyboard, a mouse, a microphone, a switch, a button, a sensor, etc.) that receives input from the outside. The output device 1006 is an output device (e.g., a display, a speaker, an LED lamp, etc.) that outputs to the outside. Note that the input device 1005 and the output device 1006 may be integrated into one device (e.g., a touch panel).

[0077] Furthermore, each device, such as the processor 1001 and the memory 1002, is connected by a bus 1007 for communicating information. The bus 1007 may be configured using a single bus, or may be configured using different buses between each device.

[0078] Furthermore, the prompt generation device 100 may be configured to include hardware such as a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a field-programmable gate array (FPGA), and some or all of the functional blocks may be realized by the hardware. For example, the processor 1001 may be implemented using at least one of these pieces of hardware. The notification of information is not limited to the aspects / embodiments described in the present disclosure and may be performed using other methods. For example, the notification of information may be performed using physical layer signaling (e.g., downlink control information (DCI) and uplink control information (UCI)), higher layer signaling (e.g., radio resource control (RRC) signaling, medium access control (MAC) signaling, broadcast information (master information block (MIB) and system information block (SIB))), other signals, or a combination thereof. Furthermore, the RRC signaling may be referred to as an RRC message, such as an RRC Connection Setup message, an RRC Connection Reconfiguration message, or the like.

[0079] The order of the procedures, sequences, flowcharts, etc. of each aspect / embodiment described in this disclosure may be changed unless it is consistent. For example, the methods described in this disclosure present elements of various steps using an example order, and are not limited to the particular order presented.

[0080] Input and output information may be stored in a specific location (for example, memory) or may be managed using a management table. Input and output information may be overwritten, updated, or added to. Output information may be deleted. Input information may be sent to another device.

[0081] The determination may be made based on a value represented by one bit (0 or 1), a Boolean value (true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0082] The aspects / embodiments described in this disclosure may be used alone, in combination, or switched depending on the implementation. Notification of predetermined information (e.g., notification that "X is true") is not limited to explicit notification, but may be implicit (e.g., not notifying the predetermined information).

[0083] Although the present disclosure has been described in detail above, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described herein. The present disclosure can be implemented in modified and altered forms without departing from the spirit and scope of the present disclosure as defined by the claims. Therefore, the description of the present disclosure is intended to be illustrative and does not have any limiting meaning on the present disclosure.

[0084] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, threads of execution, procedures, functions, etc., whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise.

[0085] Software, instructions, information, etc. may also be transmitted or received over a transmission medium. For example, if software is transmitted from a website, server, or other remote source using wired technologies (such as coaxial cable, fiber optic cable, twisted pair, Digital Subscriber Line (DSL)), and / or wireless technologies (such as infrared, microwave), then these wired and / or wireless technologies are included within the definition of transmission media.

[0086] The information, signals, etc. described in this disclosure may be represented using any of a variety of different technologies. For example, data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltages, currents, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0087] Note that terms described in this disclosure and terms necessary for understanding this disclosure may be replaced with terms having the same or similar meanings. For example, at least one of a channel and a symbol may be a signal (signaling). Furthermore, a signal may be a message. Furthermore, a component carrier (CC) may be called a carrier frequency, a cell, a frequency carrier, etc.

[0088] Furthermore, the information, parameters, etc. described in the present disclosure may be expressed using absolute values, relative values ​​from a predetermined value, or other corresponding information. For example, a radio resource may be indicated by an index.

[0089] The names used for the above-described parameters are not intended to be limiting in any way. Furthermore, the mathematical expressions using these parameters may differ from those explicitly disclosed in this disclosure. The various channels (e.g., PUCCH, PDCCH, etc.) and information elements may be identified by any suitable names, and therefore the various names assigned to these various channels and information elements are not intended to be limiting in any way.

[0090] In this disclosure, the terms "Mobile Station (MS)," "user terminal," "User Equipment (UE)," "terminal," and the like may be used interchangeably.

[0091] A mobile station may also be referred to by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other suitable terminology.

[0092] As used in this disclosure, the terms "determining" and "determining" may encompass a wide variety of actions. "Determining" and "determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up, searching, inquiring (e.g., searching in a table, database, or other data structure), ascertaining, and the like. "Determining" and "determining" may also include receiving (e.g., receiving information), transmitting (e.g., sending information), input, output, accessing (e.g., accessing data in memory), and the like. Furthermore, "judgment" and "decision" can include regarding resolving, selecting, choosing, establishing, comparing, etc. as having been "judged" or "decided." In other words, "judgment" and "decision" can include regarding some action as having been "judged" or "decided." Furthermore, "judgment (decision)" can be interpreted as "assuming," "expecting," "considering," etc.

[0093] The terms "connected," "coupled," or any variation thereof, refer to any direct or indirect connection or coupling between two or more elements, and may include the presence of one or more intermediate elements between two elements that are "connected" or "coupled" to each other. The coupling or connection between elements may be physical, logical, or a combination thereof. For example, "connected" may be read as "access." As used in this disclosure, two elements may be considered to be "connected" or "coupled" to each other using one or more wires, cables, and / or printed electrical connections, as well as electromagnetic energy having wavelengths in the radio frequency range, microwave range, and optical (both visible and invisible) range, as some non-limiting and non-exhaustive examples.

[0094] As used in this disclosure, the phrase "based on" does not mean "based only on," unless expressly stated otherwise. In other words, the phrase "based on" means both "based only on" and "based at least on."

[0095] As used in this disclosure, any reference to an element using a designation such as "first," "second," etc. does not generally limit the quantity or order of those elements. These designations may be used in this disclosure as a convenient method of distinguishing between two or more elements. Thus, a reference to a first and a second element does not imply that only two elements may be employed or that the first element must in some way precede the second element.

[0096] When the terms "include," "including," and variations thereof are used in this disclosure, these terms are intended to be inclusive, similar to the term "comprising." Furthermore, when the term "or" is used in this disclosure, it is not intended to be an exclusive or.

[0097] In this disclosure, where articles are added by translation, such as a, an, and the in English, the disclosure may include that the nouns following these articles are in the plural form.

[0098] In the present disclosure, the term "A and B are different" may mean "A and B are different from each other." The term may also mean "A and B are each different from C." Terms such as "separate" and "coupled" may also be interpreted in the same way as "different."

[0099] 100...prompt generation device (device), 110...acquisition unit, 120...estimation unit, 130...selection unit, 140...prompt generation unit, 200...image generation model, A...face image, B...image image, M...target glasses, P...user.

Claims

1. A device comprising: an acquisition unit that acquires a user's face image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which an image is to be generated; and a prompt generation unit that instructs the generation of an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information, and generates a prompt that instructs the generation of the image with the user's eyes adjusted based on the lens information.

2. The device according to claim 1, wherein the prompt includes an instruction to generate the image in which at least one of the size and distortion of the user's eyes has been adjusted based on the lens information.

3. The device of claim 1, wherein the lens information includes at least one of information regarding eyeglass power and information regarding lens type.

4. The device of claim 1, wherein the prompt includes instructions to generate images of the user's face viewed from different angles.

5. The device of claim 1, further comprising an estimation unit that estimates the lens information, wherein the acquisition unit acquires eyeglasses-wearing images in which the user is wearing eyeglasses and eyeglasses-free images in which the user is not wearing eyeglasses, the estimation unit estimates the lens information based on the eyeglasses-wearing images and the eyeglasses-free images, and the acquisition unit acquires the lens information estimated by the estimation unit.

6. The device according to claim 1, wherein the acquisition unit acquires the target glasses information input by the user.

7. The device according to claim 1, further comprising a selection unit that selects the target eyeglasses based on at least one of the user's attribute information and the shape of each part of the user's face obtained from the facial image, and the acquisition unit acquires the target eyeglasses information of the target eyeglasses selected by the selection unit.

8. The device of claim 1, wherein the prompt includes instructions to generate a moving image as the image.

9. The device according to claim 1, wherein the prompt generation unit transmits the facial image and the prompt to an image generation model, and obtains the image corresponding to the prompt from the image generation model.

10. A method including: an acquisition step of acquiring a user's face image, eyeglass lens information, and target eyeglass information for the target eyeglasses for which an image is to be generated; and a prompt generation step of instructing the user to generate an image of the user wearing the target eyeglasses based on the face image and the target eyeglass information, and generating a prompt instructing the user to generate the image adjusted for the user's eyes based on the lens information.