A device, computer program and method
The method streamlines face merging by displaying matching media options and enabling efficient selection and payment, addressing user challenges in selecting suitable media for face merging.
Patent Information
- Application Number
- PCT/GB2025/050224
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-12
- Filing Date
- 2025-02-06
- Publication Date
- 2025-08-21
AI Technical Summary
Users face challenges in selecting suitable media for face merging due to excessive screen real estate requirements and potential unsuitable media choices, especially when multiple face options are available.
A method and system that identifies and displays only media with matching or greater face counts to the captured image, allowing users to select and merge faces efficiently, with preview and payment mechanisms for final output.
Facilitates efficient face merging by reducing screen clutter and ensuring suitable media selection, providing a streamlined user experience with quality control and payment options.
Smart Images

Figure GB2025050224_21082025_PF_FP_ABST
Abstract
Description
[0001] A DEVICE, COMPUTER PROGRAM AND METHOD
[0002] BACKGROUND
[0003] Field of the Disclosure
[0004] The present technique relates to a device, computer program and method.
[0005] Description of the Related Art
[0006] The “background” description provided herein is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present technique.
[0007] With the advent of Artificial Intelligence, it is now possible to merge faces to create a new face using a diffusion model. However, in many instances, given the amount of media available, it is very hard for a user to decide which faces should be merged. If too many options are provided to a user, the amount of screen real estate required is too high. Further in instances, it is possible that the user may select a piece of media which is not suited for merging.
[0008] It is an aim of the disclosure to address at least one of these two issues.
[0009] SUMMARY
[0010] According to embodiments of the disclosure, there is provided a method, comprising: receiving a captured image containing a number of faces; identifying, from a plurality of media, at least one piece of media containing the same number or greater number of faces as the captured image; displaying only the identified at least one piece media to a user; receiving a user input to select a displayed piece of media; and merging a face from the captured image with a face in the selected piece of media to produce a merged face piece of media.
[0011] The foregoing paragraphs have been provided by way of general introduction, and are not intended to limit the scope of the following claims. The described embodiments, together with further advantages, will be best understood by reference to the following detailed description taken in conjunction with the accompanying drawings.
[0012] BRIEF DESCRIPTION OF THE DRAWINGS
[0013] A more complete appreciation of the disclosure and many of the attendant advantages thereof will be readily obtained as the same becomes better understood by reference to the following detailed description when considered in connection with the accompanying drawings, wherein:
[0014] Figure 1 shows a system 1000 according to embodiments of the disclosure;
[0015] Figures 2A to 2E show a GUI according to embodiments;
[0016] Figure 3 shows the device 300 according to embodiments of the disclosure; and Figure 4 shows a process 400 carried out by the device 300.
[0017] DESCRIPTION OF THE EMBODIMENTS
[0018] Referring now to the drawings, wherein like reference numerals designate identical or corresponding parts throughout the several views.
[0019] Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that within the scope of the appended claims, the disclosure may be practiced otherwise than as specifically described herein.
[0020] Figure 1 shows a system 1000 according to embodiments of the disclosure. The system 1000 generally includes a display 100, a camera 200 and a device 300. In embodiments, the system 1000 is a photobooth, although the disclosure is not so limited. For example, the system 1000 may be embodied as a mobile telephone or any handheld portable electronic apparatus such as a tablet computer or laptop computer. In embodiments, the system 1000 may be a laptop computer or a desktop computer arrangement.
[0021] The display 100 is, in embodiments, a touch screen display which allows a user to interact with and control the device 300 using the touch screen. Of course, the disclosure is not so limited and any kind of user control is envisaged for example, the device 300 may be controlled using gesture recognition. A Graphical User Interface (GUI) according to embodiments which is displayed to the user and with which the user interacts will be described later with reference to Figures 2A to 2D.
[0022] In embodiments, the position of the display 100 is such that the user can see the display when having their image captured. Further, in the embodiments where the display 100 is a touch screen display, the position of the display 100 is such that a user can touch the display 100.
[0023] As noted above, the system 1000 also includes a camera 200. In embodiments, the camera 200 may be any kind of camera such as a still image camera or may be a video camera capable of capturing moving images or may be a camera capable of capturing both still and moving images. In embodiments, the camera is from the Alpha range made by Sony Corporation ®, although the disclosure is not so limited. As noted above, in the embodiments where the system 1000 is a mobile phone, the camera 200 may be the front facing camera. The camera 200 is arranged, in embodiments, to capture an image of a user’s face. This means that the height of the camera, the aperture settings and lighting is configured to capture a user’s face. The image captured by the camera 200 is provided to the device 300. In addition to capturing the image, in embodiments, the camera 200 performs facial recognition on the image(s) it captures. In this case, the position of the recognised face or faces in the captured image is also provided to the device 300.
[0024] The device 300, according to embodiments, will be described later with reference to Figure 3.
[0025] Figures 2 A to 2E show a GUI according to embodiments. The explanation of the GUI will be used to explain the operation of the system according to embodiments.
[0026] The GUI is provided on the display 100. Referring to Figure 2A a capture area 110 is provided. The capture area 110 shows the image that will be captured by the camera 200. In other words, the capture area 110 shows a preview of the image that will be captured by the camera 200. As the display 100 is arranged so that a user can view the display 100, the user can position themselves correctly in front of the camera 200.
[0027] A single participant button 120 and a multiple participant button 130 is provided. The single participant button 120 indicates to the device 300 that only a single participant is to be captured in the image and the multiple participant button 130 indicates that a plurality of participants are to be selected from the image. In embodiments, the number of participant(s) may be automatic and selected by the device 300. In this instance, the device 300 may perform facial recognition on the image to identify the location of the people in the image or may receive the facial recognition information from the camera 200 if provided by the camera 200. The device 300 may identify the participant(s) from the detected faces by their position in the image, the relative size of their face in the image, or by a facial gesture (for example identifying participant(s) that is / are smiling or performing a predefined facial gesture such as blinking at a particular moment). Of course, in embodiments, if the single participant button 120 is pressed and only a single participant is in the image, the device assumes that identified participant is the selected participant.
[0028] In embodiments, the identity of the participant(s) may be made by the user. For example, the user may touch the image of participant(s) on the display 100.
[0029] In embodiments, the device 300 may visually identify the participant(s) by applying a graphic to the selected participant(s). For example, a bounding box may be applied to the face of the selected participant(s). The user may then deselect or select a participant as required.
[0030] In embodiments, the number of selected participants is indicated on the GUI so that the user is notified of the number of selected participants.
[0031] Facial tracking may be provided so that as each participant moves, the bounding box follows the participant to make identification of the selected participant easier. A capture button 140 is also provided. The capture button 140 is, in embodiments, pressed by the user to instruct the camera 200 to capture the image. Of course, the disclosure is not so limited and any appropriate mechanism may be provided which allows a user to capture the image. In embodiments, the user may say a specific word such as “cheese” or “capture” or the like.
[0032] After the image is captured, in embodiments, a preview of the captured image is shown to the user. A specific area in the displayed GUI may be provided. In embodiments, however, the capture area 110 shows the preview of the captured image. The user can approve the preview image or if they are not satisfied, the user can re-take the image. Appropriate mechanisms are envisaged for this such as providing an ‘approve’ and ‘disapprove’ button or allowing the user to issue an appropriate gesture or audible command.
[0033] In embodiments, prior to displaying the preview image, the device 300 analyses the captured image and ensures that all selected participants are located in the image and that no selected participants have their eyes closed or that “red-eye” does not exist in the captured image. In the instance that an undesirable image is captured, the device 300 may re-capture the image automatically, or may apply a filter to remove the undesirable aspect (such as “red-eye ”).In embodiments, the device 300 indicates that a new captured image is required for one or more of the participants. Moreover, in embodiments, each participant in the captured image is cut out from the image and processed individually.
[0034] After the image of the user has been captured and the user has approved the captured image or images, the captured image is then sent to the device 300. In other words, the device receives the captured image that contains a number of faces (i.e. one or more faces). The GUI displays the screen according to Figure 2B. Figure 2B shows the selection of the media image.
[0035] In embodiments, a plurality of pieces of media are displayed to the user. The approved captured image will be merged with a selected one of the displayed pieces of media as will be explained later. In the embodiments of Figure 2B, various iconic still images from movies are shown as the pieces of media. For example, in the embodiments of Figure 2B, a first through eighth piece of media 150A to 15 OH is shown. Although many of the pieces of media are still images taken from movies, in instances, the advertising poster associated with a movie may be shown (see for example the sixth media image 15 OF and seventh media image 150G in Figure 2B).
[0036] Of course, any kind of piece of media is envisaged such as an image of an iconic music performer, politician, or any kind of famous person such as sports people like the Australian cricketer Merv Hughes. Indeed any type of media is envisaged such as video into which a captured still image or images of one or more participant may be merged.
[0037] Figure 2C shows a table of pieces of media available to the user. The table contains information that is associated with each piece of media. This information includes a name which uniquely identifies the piece of media. The name may be a file name for the piece of media or may be a Unique Resource Locator (URL) or the like. Additionally, the information also includes the number of faces within each piece of media. In the embodiments of Figure 2C, the number of faces indicate the number of people or the number of characters in the piece of media whose faces may be merged. Of course, the disclosure is not so limited and the number of faces may indicate the number of faces into which the participant may be merged. In other words, a piece of media may include 3 people, but if only 2 of those people have given permission to allow their face to be merged, the number of faces will be 2.
[0038] In addition, in embodiments, the type of media is provided. The type of media indicates the format of the media, such as a still image or a movie poster as shown, although the disclosure is not so limited. In embodiments, the format of the media may include video clip or the like.
[0039] Finally, a fee associated with each piece of media is provided. This indicates the cost for the user to merge the user’s face with the icons in the piece of media. In embodiments, the fee varies depending upon the type of media. Specifically, a still image is a particular fee and a movie poster, given the size of the media upon which the movie poster will ultimately be printed, is given a second, higher, fee. Of course, the fee may be selected based on any number of criteria including the royalty fee payable to the creator of the piece of media, or the number of faces within the piece of media or the like.
[0040] Of course, other information may be associated with each piece of media. For example, in a non-limiting manner, gender of the people whose faces in the piece of media may be stored as a user may wish to select icons of a particular gender to be displayed.
[0041] As will be appreciated, there will be many pieces of media that can be displayed to the user. In embodiments, appropriate filters are provided such as a particular type of media. For example, a user may be really interested in the movie Skyfall ® and would like to merge his or her face with one of the characters in a particular scene within that movie or may wish to merge his or her face with one of the characters in any scene. Indeed, in embodiments, the user may wish to merge his or her face with one of the characters performing a particular scene in the movie such as an iconic action sequence. This choice is for the user to make.
[0042] In embodiments, the selected users within the image or video captured by the device 300 are used to determine the displayed pieces of media. Specifically, one or more characteristics of the selected users in the captured image or video are used to determine the pieces of media that will be displayed to the user in Figure 2B for selection. This is to ensure that the piece(s) of media which are displayed are suitable for merging. In embodiments, the number of users within the captured image or video will be used to determine the pieces of media that will be displayed to the user. Specifically, the device 300 determines the number of users within the captured image or video (the number of users being determined either automatically or by user input as explained above) and only pieces of media having a corresponding number of icons will be displayed to the user. This makes it easier and quicker for users to select a suitable piece of media as only pieces of media that can be used are displayed to the user. Additionally, this ensures that only pieces of media that are capable of merging the faces are displayed.
[0043] In embodiments, only media having a greater number of icons will be displayed to the user. This allows a user to select from the icons who they wish to merge with. In this instance, the other icons may remain untouched. For example, a first user may wish to merge with a player from a first football team with a second user merging with a player from a second, different football team or one user may wish to merge with a icon who is a hero whereas a second user may wish to merge with an icon who is a villain.
[0044] Moreover, the user may wish to merge with more than one icon displayed on the display or may wish to change from merging with a first icon to merging with a second icon in any image.
[0045] As will be apparent to the skilled person, the pieces of media displayed in Figure 2B include a number of instances where the number of icons does not equal the number of users in image 110. The pieces of media in Figure 2B have been shown for explanation and in reality, only pieces of media shown in 150B, 150C, 150E, and 15 OF would be displayed as the number of faces is the same as the number of selected users.
[0046] Figure 2D shows the preview output of the merged images if the user captured in capture area 110 is merged with Spiderman ® in piece of media 150C. Specifically, the merged image will appear in merged image area 160 on the display 100. The mechanism used for generating the merged image will be explained later.
[0047] The merged image is a preview of the image to ensure that the user is happy with the merged image. If the user is happy with the merged image the user can approve the purchase of the merged image by pressing a purchase button 180A. The purchase button 180A may include the fee to indicate to the user the amount of fee payable.
[0048] If the user is not happy to purchase the merged image, the user can press a reject button 180B and they will be returned to a previous screen.
[0049] In embodiments, a watermark or some other embedded mark or obfuscation of the image may be applied to the merged image to reduce the likelihood of the user illicitly capturing the preview image to avoid paying for the merged image. In Figure 2D the word “sample” 181 is embedded on the merged image. In embodiments, the resolution of the merged image may be low and a higher resolution image may be provided when the purchase button 180A is pressed. In other words, the preview image is a preview version of the merged face piece of media to the user, wherein the preview version of the merged face piece of media is a lower quality version of the merged face piece of media. The term lower quality encapsulates a lower resolution and / or mark embedded in the merged piece of media. Figure 2E shows the final output of the preview merged image of Figure 2D. In response to the purchase button 180A being pressed and the predetermined amount of money derived from the fee in the information of Figure 2C being paid by the user, the device 300 receives an indication that the user has paid the money. The user is then provided with the merged image in the appropriate resolution and / or with any embedded mark removed. In embodiments, this may be as a hard copy print, such as a postcard or a photograph printed on media such as paper or as a printed poster. This may be printed by the system 1000 or may be printed outside of the system and picked up or delivered to the user.
[0050] In embodiments, the merged image may be stored at a central repository or in the device 300 and may be retrieved by the user electronically. This may be achieved by providing a QR code 170. In embodiments, the user scans the QR code using their mobile phone and the image is retrieved. Of course, although a QR code is shown, any kind of machine readable code is envisaged. In other words, the device 300 displays a machine readable code to the user providing access for the user to the merged face piece of media. This is done via the display 200.
[0051] The mechanism used to merge the captured image with the piece of media will now be described. In embodiments, the faces are merged using a diffusion model. In embodiments, the diffusion model may be trained using specific training data suitable for the application described above. For example, in embodiments where the system 1000 is located in a photobooth, the model may be trained on faces captured from a photobooth.
[0052] Further, the captured image may be subjected to pre-processing before being provided to the stable diffusion model. In embodiments, the face or faces of the selected participant(s) may be cropped to a certain size or resolution to meet the input requirements of the diffusion model. In other words, the face or faces may be adapted to the diffusion model before merging with the face in the selected piece of media so that the resultant merged image is more consistent. Figure 3 shows the device 300 according to embodiments of the disclosure. As noted above, the device 300 forms part of a system 1000 that includes other components such as a display 100 and a camera 200. The system 1000 therefore may be embodied as a photobooth or a mobile phone or the like. The device 300 comprises processing circuitry 310 that is configured to perform methods according to embodiments of the disclosure. The processing circuitry 310 is, in embodiments, semiconductor circuitry that may be controlled by a computer program or may be an Application Specific Integrated Circuit or the like. The processing circuitry 310 is connected, via various interfaces, to the display 100 and the camera 200. In addition, the processing circuitry 310 is connected to storage 320 which is configured to store the computer program that controls the processing circuitry. The storage 320 may be solid state storage or may be optically or magnetically readable storage. The storage 320, in embodiments, stores the information shown in Figure 2C. Of course, the disclosure is not so limited and in embodiments, this information is stored on a network such as the internet or on a combination of the storage 320 and the network. Moreover, the pieces of media are stored, in embodiments, in the storage 320 but again the disclosure is not so limited and in embodiments, the pieces of media are stored on a network such as the internet or on a combination of the storage 320 and the network.
[0053] Figure 4 shows a process 400 carried out by the device 300. The process 400 starts and moves to step 402. In step 402, the device 300 receives a captured image containing a number of faces. The process moves to step 404. In step 404, the device identifies from a plurality of media, at least one piece of media containing the same number of faces or a greater number of faces as the captured image. The process moves to step 406. In step 406, the device displays only the identified at least one piece media to a user. The process then moves to step 408. In step 408, the device 300 receives a user input to select a displayed piece of media. The process then moves to step 410. In step 410, the device 300 merges a face from the captured image with a face in the selected piece of media to produce a merged face piece of media. The process then ends.
[0054] In so far as embodiments of the disclosure have been described as being implemented, at least in part, by software -controlled data processing apparatus, it will be appreciated that a non-transitory machine- readable medium carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure.
[0055] It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and / or processors may be used without detracting from the embodiments.
[0056] Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more data processors and / or digital signal processors. The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and / or processors.
[0057] Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to the specific form set forth herein. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the technique.
[0058] Embodiments of the present technique can generally described by the following numbered clauses:
[0059] 1. A method, comprising: receiving a captured image containing a number of faces; identifying, from a plurality of media, at least one piece of media containing the same number or greater number of faces as the captured image; displaying only the identified at least one piece media to a user; receiving a user input to select a displayed piece of media; and merging a face from the captured image with a face in the selected piece of media to produce a merged face piece of media.
[0060] 2. A method according to clause 1, comprising: displaying a preview version of the merged face piece of media to the user, wherein the preview version of the merged face piece of media is a lower quality version of the merged face piece of media.
[0061] 3. A method according to clause 2, comprising: receiving an indication the user had paid a predetermined amount of money; and in response to receiving the indication, displaying a machine readable code to the user providing access for the user to the merged face piece of media.
[0062] 4. A method according to clause 3, wherein the predetermined amount of money depends on the type of piece of media or the number of faces in the piece of media.
[0063] 5. A method according to clause 1 comprising: adapting the face in the captured image before merging with the face in the selected piece of media.
[0064] 6. A device comprising circuitry configured to: receive a captured image containing a number of faces; identify, from a plurality of media, at least one piece of media containing the same number or greater number of faces as the captured image; display only the identified at least one piece media to a user; receive a user input to select a displayed piece of media; and merge a face from the captured image with a face in the selected piece of media to produce a merged face piece of media.
[0065] 7. A device according to clause 6, wherein the circuitry is configured to: display a preview version of the merged face piece of media to the user, wherein the preview version of the merged face piece of media is a lower quality version of the merged face piece of media. 8. A device according to clause 7, wherein the circuitry is configured to: receive an indication the user had paid a predetermined amount of money; and in response to receiving the indication, display a machine readable code to the user providing access for the user to the merged face piece of media.
[0066] 9. A device according to clause 8, wherein the predetermined amount of money depends on the type of piece of media or the number of faces in the piece of media.
[0067] 10. A device according to clause 6 wherein the circuitry is configured to: adapt the face in the captured image before merging with the face in the selected piece of media.
[0068] 11. A computer program comprising computer readable instructions which, when loaded onto a computer, configure the computer to perform a method according to any one of clauses 1 to 5.
Claims
CLAIMS1. A method, comprising: receiving a captured image containing a number of faces; identifying, from a plurality of media, at least one piece of media containing the same number or greater number of faces as the captured image; displaying only the identified at least one piece media to a user; receiving a user input to select a displayed piece of media; and merging a face from the captured image with a face in the selected piece of media to produce a merged face piece of media.
2. A method according to claim 1, comprising: displaying a preview version of the merged face piece of media to the user, wherein the preview version of the merged face piece of media is a lower quality version of the merged face piece of media.
3. A method according to claim 2, comprising: receiving an indication the user had paid a predetermined amount of money; and in response to receiving the indication, displaying a machine readable code to the user providing access for the user to the merged face piece of media.
4. A method according to claim 3, wherein the predetermined amount of money depends on the type of piece of media or the number of faces in the piece of media.
5. A method according to claim 1 comprising: adapting the face in the captured image before merging with the face in the selected piece of media.
6. A device comprising circuitry configured to: receive a captured image containing a number of faces; identify, from a plurality of media, at least one piece of media containing the same number or greater number of faces as the captured image; display only the identified at least one piece media to a user; receive a user input to select a displayed piece of media; and merge a face from the captured image with a face in the selected piece of media to produce a merged face piece of media.
7. A device according to claim 6, wherein the circuitry is configured to: display a preview version of the merged face piece of media to the user, wherein the preview version of the merged face piece of media is a lower quality version of the merged face piece of media.
8. A device according to claim 7, wherein the circuitry is configured to: receive an indication the user had paid a predetermined amount of money; and in response to receiving the indication, display a machine readable code to the user providing access for the user to the merged face piece of media.
9. A device according to claim 8, wherein the predetermined amount of money depends on the type of piece of media or the number of faces in the piece of media.
10. A device according to claim 6 wherein the circuitry is configured to: adapt the face in the captured image before merging with the face in the selected piece of media.
11. A computer program comprising computer readable instructions which, when loaded onto a computer, configure the computer to perform a method according to claim 1.
Citation Information
Patent Citations
Sensing block, battery module assembly comprising the same
KR1020240048328A
Image processing method and related apparatus
WO2023036084A1