Writing prompt for augmented reality beauty looks
A computer-implemented method using generative AI to generate personalized augmented reality visuals based on user input, addressing the limitations of existing AR applications by enabling almost unlimited creativity and aesthetic alignment.
Patent Information
- Application Number
- FR2024004982
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-21
AI Technical Summary
Existing augmented reality applications in beauty lack the ability to provide users with almost unlimited creativity and subjective inspiration, as their content is often limited and does not align with individual desires and preferences.
A computer-implemented method that generates augmented reality visuals through a generative artificial intelligence tool controlled by a descriptive text, allowing users to input their appearance desires, with constraints to ensure aesthetic alignment, and superimposes the generated visuals onto a region of interest in real-time.
Enables users to create personalized and aesthetically pleasing augmented reality looks with almost unlimited creativity, providing immersive and memorable experiences that align with their creative desires.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Augmented reality beauty appearance writing prompt
[0001] Some embodiments relate to a computer-implemented process capable of generating an augmented reality visual in a manner controlled by a request expressed in a descriptive text.
[0002] There are statistics which show that 74% of people from Generation Z (usually defined as people born approximately between 1995 and 2005) use an augmented reality application daily, and which show that 64% of augmented reality application consumers believe that it allows them to be more creative.
[0003] In the world of beauty (i.e., for example, the world of fashion, cosmetics, and appearance), augmented reality can have several uses and different use cases, including:
[0004] Augmented reality applied to beauty can indeed generate ideation, that is to say a creative production process, by allowing the exploration and prototyping of new codes in virtual styles and appearances.
[0005] Augmented reality applied to beauty can be used to reveal concepts of future styles and fashions, before moving on to large-scale production, by offering an audience inspiring experiences of virtual trials.
[0006] Augmented reality applied to beauty can serve as a means of communication about a brand, offering an immersive and memorable experience, for example in the context of a product launch, a seasonal theme, etc...
[0007] Classically, augmented reality filters are accessible in online database systems, usually via library-type applications whose contents are designed and imposed by the service provider and in limited quantity.
[0008] However, in the uses of augmented reality, particularly in the context of cosmetics and beauty, creative inspiration is great and desires and preferences are relatively subjective, so users may suffer frustration when they do not find the offer that corresponds to their creative desire.
[0009] There is therefore a need to remedy this problem of frustration and to offer almost unlimited freedom to the creativity of a user in their augmented reality experience applied to beauty.
[0010] In this regard, a computer-implemented method is proposed, comprising: - an acquisition of an appearance-descriptive text; - the generation of at least one decorative image by a generative artificial intelligence tool controlled by a query determined by said descriptor text; - an acquisition of an image stream; - detection of a region of interest in the image stream; - a generation of an augmented reality visual comprising a mask capable of being superimposed in continuous tracking on the region of interest in the image stream, and incorporating said at least one decorative image; - a supply of a transformed image stream comprising the image stream and the augmented reality visual superimposed on the image stream.
[0011] For example, the region of interest is a face.
[0012] According to one implementation method, said request is determined by said descriptor text, and by at least one constraint automatically introduced into the request to impose at least one predetermined aesthetic characteristic on said decorative image.
[0013] According to one implementation method, said at least one constraint introduced in the query includes at least one of the following constraints: abstract model; no photorealism; no detailed contours; no sharp angles.
[0014] According to one implementation method, said at least one constraint introduced in the query includes at least one of the following constraints: color references; references to a known theme such as a brand, a product, a cultural event.
[0015] According to one embodiment, the generation of the augmented reality visual further includes a realistic simulation of makeup having the shade of a color selected in said decorative image.
[0016] According to one embodiment, the method further comprises providing a digital object including the augmented reality visual suitable for being superimposed in continuous tracking on a region of interest of a third-party image stream.
[0017] A computer program is also proposed comprising instructions which, when the program is executed by a computer, lead the computer to implement the process defined above.
[0018] A computer-readable medium is also proposed comprising instructions which, when executed by a computer, lead the computer to implement the process defined above.
[0019] According to another aspect, a digital object is proposed comprising digital data representative of the augmented reality visual generated by the process defined above, extracted from said transformed image stream, or suitable for being superimposed in a continuous tracking on a region of interest of a third image stream.
[0020] Other advantages and features of the invention will become apparent upon examination of the detailed description of, but not limited to, embodiments and the accompanying drawings, in which:
[0021] [Fig.1A]
[0022] [Fig.1B]
[0023] [Fig.lC]
[0024] [Fig.1D]
[0025] [Fig.1E]
[0026] [Fig.2]
[0027] [Fig.3A]
[0028] [Fig.3B]
[0029] [Fig.3C]
[0030] [Fig.4A]
[0031] [Fig.4B]
[0032] [Fig.5] illustrate ways of implementing the invention.
[0033] Figures IA to 1E illustrate the path of a user of an application, i.e. a computer program, capable of generating an augmented reality visual in a manner controlled by a request expressed by the user, for example in a descriptor text.
[0034] The application is implemented by a computer such as a multifunction phone, a touch tablet, a desktop computer, comprising at least one display screen, possibly touch-sensitive, such as a light-emitting diode "LED" screen, possibly organic "OLED"; a text input means such as a physical keyboard, a keyboard emulated on a touch screen or a microphone associated with a voice-to-text transcription means; at least one image sensor such as a camera sensor, preferably a front camera adapted for self-portrait display (or "selfie" according to the usual contraction of the Anglo-Saxon terms "self" and "photography"); processing means capable of controlling the operations of the computer comprising at least one microprocessor.
[0035] Figure 1A illustrates a hook phase, in which a clear and precise example of the type of EXMP request that the user can submit is presented. This creates a hook effect that piques the user's curiosity and encourages them to continue using the application.
[0036] Figure 1B illustrates a guidance phase, in which simple recommendations are given to help the user make the correct computer settings and grant the application access permissions to the required peripherals, such as the front camera. Once access is granted, the application can, for example, display a continuous VIS video stream from the front camera's field of view. Under normal conditions of use, which may be specified in the recommendations, the user's face is present in the field of vision of the front camera.
[0037] Figure 1C illustrates an input phase in which the user is given a PRMPT writing prompt, allowing them to express their inspiration regarding look and appearance through a short sentence entered using the text input method. The PRMPT writing prompt may, for example, be preceded by a guideline, which establishes the context for using the application. For example, the guideline may be expressed as the beginning of a sentence such as "I want makeup inspired by..." or "I want a style inspired by..." so that the entered text completes the sentence thus begun.
[0038] It should be noted that during the input phase, the user enjoys almost unlimited freedom to express their creativity. For example, the user could write "a field of small pink exotic flowers".
[0039] Figure 1D illustrates a digital creation phase in which the application uses generative artificial intelligence resources to design an appearance (a look) based on the entered text in a few seconds. A PAT patience prompt, for example "loading...", can be displayed during the creation phase, possibly accompanied by an animation to give the impression that a complex process is taking place for the user.
[0040] Fig. 1E illustrates a display phase in which users can view the chosen VIS_AR look in augmented reality, continuously in the image stream of the self-portrait camera.
[0041] Augmented reality is the superimposition of virtual digital elements into an image stream representing a reality, for example, a scene seen in the field of view of an image sensor, in real time. The virtual digital elements (FLR) are embedded in the image stream so as to represent them from a perspective corresponding to that which they would have if they were present in the real-world scene.
[0042] In the above example "a field of small pink exotic flowers" a plurality of small exotic flowers, of one or more varieties, having different shades of pink, can adorn the user's face, displayed in such a way as to give the illusion that the flowers are delicately placed on the face.
[0043] In particular, areas of the face suitable for aesthetically pleasing ornamentation, for example around the eyes, can be identified and favored.
[0044] The user can try another look based on another inspiration, and restart the process of the input, digital creation, and display phases, described above in relation to figures IC to 1E.
[0045] A digital object can also be extracted from the display phase, at the user's command, for example in order to be shared in a digital environment such as a social network, a video game, or even a virtual universe (sometimes called "metaverse").
[0046] The digital object is, for example, a photographic snapshot captured in the continuous display of the video stream incorporating the augmented content, or a short animated clip.
[0047] Alternatively, the digital object can also be 3D (three-dimensional) content, for example, a file containing augmented reality visuals suitable for superimposing into a third-party image stream. The third-party image stream can, for example, be obtained by video capture of a real-world scene at a different time, within the framework of another augmented reality implementation. Alternatively, the third-party image stream can also be virtual, such as, for example, a representation of an avatar chosen to personify the user in a digital environment such as a social network, a video game, or any other virtual world.
[0048] Fig. 2 illustrates the method 200 for generating the augmented reality visual in a manner controlled by the request expressed in a descriptor text, implemented by the application described above in relation to figures IA to 1E.
[0049] Thus, the computer-implemented process 200 begins with a start-up phase 201 comprising, for example, the hooking and guiding phases described previously in relation to Figures IA and IB.
[0050] Thus, the image sensor 2010 (the front camera) is controlled so as to acquire an image stream 202 and transmit the image data from the stream to the software processing means 203 of the process 200.
[0051] The software processing means include in particular access to an augmented reality generation engine (“AR engine”) 203, capable of detecting a region of interest in the images of the stream, and creating a 3D model of this region from the image stream.
[0052] In particular, the augmented reality engine 203 is adapted to identify and model a human face (as an area of interest). 3D modeling can, for example, be done by identifying reference points on the face, also called "trackers," which together form a 3D mesh. The 3D mesh can be used to define anchor points for the addition of virtual content 204.
[0053] The virtual content is supported by a 3D digital object attached to the 3D mesh of the model, the digital object being able to be called mask 204 in the case of a face model.
[0054] Different types of masks can be generated, for example masks located on the contour of each eye, masks covering both eyes transversely, masks occupying a larger part of the face, for example a portion of the face or the entire face, finally the masks can be symmetrical (for example in the same orientation as the symmetry of the face) or asymmetrical.
[0055] For example, the choice of mask model can be made randomly at each implementation of the method. Alternatively, the mask model can be chosen according to the generated image 207 (see below), for example according to an identification of an extent of individual elements made in the image 207, or according to an identification of a particular symmetry or a particular asymmetry in the image, corresponding to a mask model.
[0056] The augmented reality engine may, for example, be based on existing means providing services appropriate to the functions expressed above, such as the tools with trade names "ModiFace" or "8thWall".
[0057] In summary, the ModiFace tool is an augmented reality engine specializing in the field of beauty, and is notably able to build extremely precise and lightweight tracers in terms of data processing, and offers renderings very faithful to reality by automatically adapting to different lighting and skin tones in particular.
[0058] In summary, the 8thWall tool is an augmented reality engine, designed to offer third-party developers the tools needed to create immersive and realistic augmented reality experiences in a web (online) service, with great freedom of development and compatibility.
[0059] In parallel or prior to the augmented reality mask generation steps 203 and 204, the method 200 includes the input phase 205 described previously in relation to [Fig.1C].
[0060] Thus, the user's text input in the writing prompt allows the process to acquire a text intended to express the description of a desired appearance, referred to in this respect as appearance descriptor text.
[0061] From the appearance descriptor text, the process includes a generation of a query 2052, advantageously including an addition of constraints 2051 (see below in relation to Figures 3A-3C) to the appearance descriptor text, in order to control a generative artificial intelligence tool 206 capable of generating an image (called decorative image) from the query 2052.
[0062] In other words, the process 200 comprises a generation of at least one decorative image 207 by a generative artificial intelligence tool 206 controlled by said query determined by said descriptor text.
[0063] The image-generating artificial intelligence tool could, for example, be based on existing means providing services appropriate to this function, such as the tools with the trade names "Dall-E 3", "Midjoumey", "Stable Diffusion", "gettyimages", or "Bria AI". These examples implement artificial intelligence algorithms of the Large Language Model (LLM) type that generate images corresponding to text expressed in natural language.
[0064] Image 207, thus generated by the query based on appearance descriptor text, includes decorative visual content representative of the inspiration communicated by the user.
[0065] The decorative visual content of the image is incorporated 208 into the augmented reality mask. Advantageously, it will be possible, for example, to identify and isolate certain elements of the generated image in order to incorporate them into the augmented reality mask 204.
[0066] Thus, the augmented reality mask 204 generated to virtually fit the user's face in the image stream, is displayed in the image stream so as to show the decorative visual of the image generated according to said request 2052.
[0067] In other words, an augmented reality visual 210 was formed, adapted to enhance the user's face in an aesthetic that follows the appearance inspiration that the user expressed in the input phase 205.
[0068] In this regard, a transformed image stream comprising the initial image stream 202 and the augmented reality visual 210 superimposed on the image stream 202 is provided.
[0069] Optionally and advantageously, the augmented reality visual superimposed on the image stream 202 may further include a realistic simulation of a makeup 2071 having the tint of a selected color 2070 in said decorative image 207.
[0070] Reference is now made to figures 3A to 3C.
[0071] Figures 3A to 3C illustrate examples of adding constraints 2041 in the generation of query 2042, based on appearance descriptor text 204.
[0072] At least one constraint 2041 is automatically introduced into query 2042 to impose at least one predetermined aesthetic characteristic on said decorative image; for example, constraints 2041 may include at least one of the following constraints: abstract model; no photorealism; no detailed contours; no strong angles.
[0073] The constraints are designed to maintain a consistent result that conveys the meaning of the appearance descriptor text. The constraints designed allow to avoid photorealistic images and overly detailed elements, and to keep something abstract enough but still aesthetically appealing to be suitable for a facial ornament.
[0074] On the one hand, a "pre-scriptum" approach is proposed, with a constraint placed before the descriptive text of appearance, to impose an orientation in style and aesthetics, for example "abstract model of".
[0075] On the other hand, a "postscript" approach is proposed in combination or as an alternative to exclude or avoid undesired image characteristics, for example: "non-photorealistic", "no detailed contours", "no strong angles".
[0076] Fig. 3A illustrates a greyscale example of a colour image 207 generated by a query freely using the appearance descriptor text entered by the user, in the case where this text is "pink and purple wildflowers".
[0077] Figure 3B illustrates a greyscale example of a colour image 207 generated by a query using the appearance descriptor text entered by the user "pink and purple wildflowers", adding the postscript constraints "non-photorealistic, without detailed outlines, without strong angles".
[0078] Figure 3C illustrates a greyscale example of a colour image 207 generated by a query using the user-entered appearance descriptor text "pink and purple wildflowers", adding the postscript constraints "non-photorealistic, without detailed outlines, without strong angles", and the prescript constraint "abstract model of".
[0079] It may be noted that the image in [Fig.3C] is the most aesthetically suitable for use in the application implemented by the process described in relation to [Fig.2], described in relation to Figs IA to 1E.
[0080] This type of pre-scriptum and post-scriptum constraint can be implemented by low-level adaptation methods (usually "LoRA" for "Low-Rank Adaptation" in English), that is to say, learning techniques for the fine-tuning of generative artificial intelligence models, which make slight modifications (called checkpoints) to the initial models.
[0081] Figures 4A and 4B illustrate examples of augmented reality visual results 210 generated with other constraints 2041 in the query 2042 controlling the generation of the decorative image 207.
[0082] The decorative image 207 can be guided to be more illustrative or more abstract, more saturated or pastel, for example according to a universe or a stylistic code associated with a brand or a product.
[0083] But in addition, other textual elements, more explicit and rich in meaning, can constrain the query, such as color references or references to a known theme, to correspond to the aesthetics of a brand specific, of a specific product or according to a seasonal cultural event with an identifiable universe such as Halloween or Christmas.
[0084] Fig. 4A illustrates a greyscale example of an augmented reality visual 210 generated with constraints 2041 in the query 2042 commanding the generation of the decorative image 207 in a gothic theme with dark and bright colours.
[0085] The initially entered descriptive text could be in this example "moon in the clouds".
[0086] Fig. 4B illustrates a greyscale example of an augmented reality visual 210 generated with constraints 2041 in the query 2042 commanding the generation of the decorative image 207 following a floral theme with light and matte colours, i.e. “pastel”.
[0087] The initially entered descriptive text could be in this example "wildflowers".
[0088] It may be noted that the augmented reality visuals 210 of figures 4A and 4B further include a realistic simulation of makeup, including blush and lipstick, having shades chosen from the colors of said decorative image.
[0089] Fig. 5 illustrates examples of eyeliner (eye pencil) drawings that can be offered as a realistic simulation of makeup with shades chosen from the colors of the decorative image.
[0090] In particular, it is advantageous to select the main dark color from the generated decorative image and apply it to the lips and eyeliner. The shapes of the pencil strokes are predefined and randomly selected during the makeup simulation 2071.
[0091] It is also proposed to select eyeliners, colours and textures faithful to a brand, corresponding to actual colour references for actual products of the brand.
Claims
Demands
1. A computer-implemented method (200) comprising: - acquiring an appearance descriptor text (205); - generating at least one decorative image (207) by a generative artificial intelligence tool (206) controlled by a query determined (2052) by said appearance descriptor text (205); - acquiring (2010) an image stream (202); - detecting a region of interest in the image stream (203); - generating an augmented reality visual comprising a mask (204) capable of being superimposed in a continuous tracking on the region of interest in the image stream (202), and incorporating (208) said at least one decorative image (207); - providing a transformed image stream (210) comprising the image stream and the augmented reality visual superimposed on the image stream.
2. Method according to claim 1, wherein the region of interest is a face (VIS).
3. A method according to any one of claims 1 or 2, wherein said request (2052) is determined by said descriptor text (205), and by at least one constraint (2051) automatically introduced into the request (2052) to impose at least one predetermined aesthetic characteristic on said decorative image (207).
4. A method according to claim 3, wherein said at least one constraint (2051) introduced in request (2052) includes at least one of the following constraints: abstract model; no photorealism; no detailed contours; no sharp angles.
5. A method according to any one of claims 3 or 4, wherein said at least one constraint (2051) introduced in the request (2052) includes at least one of the following constraints: color references; references to a known theme such as a brand, a product, a cultural event.
6. A method according to any one of claims 1 to 5, wherein the generation of the augmented reality visual (210) further comprises a realistic simulation of a makeup (2071) having the tint of a selected color (2070) in said decorative image (2070).
7. A method according to any one of claims 1 to 6, further comprising the supply of a digital object comprising the reality visual
8.
9.
10. augmented (210) suitable for being superimposed in a continuous tracking over a region of interest of a third image stream. Computer program comprising instructions which, when the program is executed by a computer, cause the computer to implement the method (200) according to any one of claims 1 to 7. Computer-readable support comprising instructions which, when executed by a computer, cause the computer to carry out the method (200) according to any one of claims 1 to 7. Digital object comprising digital data representative of the augmented reality visual generated by the method (200) according to any one of claims 1 to 7, extracted from said transformed image stream, or capable of being superimposed in a continuous tracking on a region of interest of a third image stream.
Citation Information
Patent Citations
Digital makeup palette
CN116830073A
Method and system for using machine-learning for object instance segmentation
US10713794B1
System, device, and method of augmented reality based mapping of a venue and navigation within a venue
US11354728B2
Detecting Augmented-Reality Targets
US20200143238A1
Computer vision systems
US20210279475A1