Camera device and method
Patent Information
- Application Number
- PCT/EP2026/057598
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-18
- Publication Date
- 2026-09-24
Smart Images

Figure EP2026057598_24092026_PF_FP_ABST
Abstract
Description
[0001] Our ref. : 250072EPWOP 1
[0002] Sony Group Corporation
[0003] CAMERA DEVICE AND METHOD
[0004] TECHNICAL FIELD
[0005] The present disclosure generally pertains to a camera device and a method for guiding a user to recreate an original image.
[0006] TECHNICAL BACKGROUND
[0007] Generally, recreating photos from the past is known, for example, one or more persons may try to recreate a scene from their childhood or historical events to capture an image which represents the original scene depicted in the original image.
[0008] Although there exist techniques for recreating an original image, it is generally desirable to improve the existing techniques.
[0009] SUMMARY
[0010] According to a first aspect, the disclosure provides a camera device for guiding a user to recreate an original image, comprising circuitry configured to:
[0011] receive the original image;
[0012] capture an image of a current scene; and
[0013] input the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0014] According to a second aspect, the disclosure provides a method for guiding a user to recreate an original image, comprising:
[0015] receiving the original image;
[0016] capturing an image of a current scene; and
[0017] inputting the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0018] Further aspects are set forth in the dependent claims, the drawings and the following description.
[0019] BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Embodiments are explained by way of example with respect to the accompanying drawings, in which:
[0021] Fig. 1 schematically illustrates in a block diagram an embodiment of a camera device;Our ref. : 250072EPWOP 2
[0022] Sony Group Corporation
[0023] Fig. 2 schematically illustrates in a flow diagram an embodiment of a method; and
[0024] Fig. 3 schematically illustrates in a block diagram an embodiment of a multi-purpose computer.
[0025] DETAILED DESCRIPTION OF EMBODIMENTS
[0026] Before a detailed description of the embodiments under reference of Fig. 1 is given, general explanations are made.
[0027] As mentioned in the outset, generally, recreating photos from the past is known, for example, one or more persons may try to recreate a scene from their childhood or historical events to capture an image which represents the original scene depicted in the original image.
[0028] However, it has been recognized that such attempts may, in some cases, result in a captured image which has few similarities with the original image regarding a pose of objects and persons or facial expressions of persons.
[0029] Hence, some embodiments pertain to a camera device for guiding a user to recreate an original image, wherein the camera device includes circuitry configured to:
[0030] receive the original image;
[0031] capture an image of a current scene; and
[0032] input the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0033] The camera device may be any device capable of capturing photos (images) such as a smartphone, a tablet, a laptop, smart glasses, an augmented or virtual reality device or the like. The camera device uses Al (“Artificial Intelligence”)-based algorithms to guide the users for the correct posing in order to recreate before and after, and similar style photos. The camera device may be utilized to recreate any type of photos, such as pose in identical manner as actors in the location of a movie or video clip shoot or recreate historical events.
[0034] The artificial intelligence guides the user(s) for the posing and changes using one or more of the following modalities: text, voice and sound, and visual overlays.
[0035] Thus, in some embodiments, the artificial intelligence is configured to guide the user by using at least one of text output, voice output, sound output and visual output.
[0036] The circuitry may include one or more image sensors to capture an image of a scene.
[0037] The circuitry may include one or more processors. A processor may be or may include an application processor, a central processing unit (“CPU”), a graphical processing unit (“GPU”), aOur ref.: 250072EPWOP 3
[0038] Sony Group Corporation
[0039] digital signal processor (“DSP”), a field-programmable gate array (“FPGA”), an application specific integrated circuit (“ASIC”) etc.
[0040] The circuitry may include one or more memory components. A memory component may be or may include volatile and non-volatile memory such as static random-access memory (“SRAM”), dynamic RAM (“DRAM”), non-volatile RAM (“NVRAM”), read-only memory (“ROM”), programmable ROM (“PROM”), electrically PROM (“EPROM”), electrically erasable PROM (“EEPROM”), flash memory (e.g., NOR flash or NAND flash) etc. A memory component may be or may include one or more registers, caches, main memories, hard disk drives, solid-state drives etc.
[0041] The circuitry may include one or more data bus interfaces.
[0042] The circuitry may include one or more communication interfaces, wherein each communication interface may be configured to communicate, for example, via a local area network (LAN), a wireless local area network (WLAN), a mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, etc.
[0043] The functionality of the circuitry may be implemented by typical electronic components configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented in parts by typical electronic components and in parts by software configured to achieve the functionality as described herein. The functionality of the circuitry may be implemented by software configured to achieve the functionality as described herein. As mentioned above, the circuitry is configured to receive the original image.
[0044] The original image or photo could be uploaded to the device manually, an online URL (“Uniform Resource Locator”) may be provided, an image or photo may be taken by the camera device itself as the original image, or the camera device may search and suggest photo recreations depending on the people present and a physical location.
[0045] As mentioned above, the circuitry is configured to capture an image of a current scene.
[0046] For example, the user may operate the camera device to capture an image of a current scene which may include one or more persons and one or more (static) objects. The user may be part of the one or more persons of the current scene.
[0047] As mentioned above, the circuitry is configured to input the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.Our ref.: 250072EPWOP 4
[0048] Sony Group Corporation
[0049] The artificial intelligence may include one or more neural networks (e.g., a Convolutional Neural Network (“CNN”), a recurrent neural network, a linear neural network or a combination thereof), one or more transformer architectures to implement a Large Language Model (“LLM”), a Vision Transformer (“ViT”) or the like.
[0050] The artificial intelligence analyzes the original image or photo that the user(s) wishes to recreate and matches it to the people, objects and background of where the camera device is facing. Then, the artificial intelligence finds good matches and still remaining differences in the scene and guide the user(s) to match the original image or photo as close as possible. This may include to suggest or explain the user where to move, how to rotate the head, how to change a facial expression or what item of clothing or accessory may be changed. This may be done for each “actor” in the scene, for example, each detected person and static foreground object. Moreover, objects in the scene that can be moved, may be guided for more precise placement.
[0051] Hence, in some embodiments, the artificial intelligence is configured to suggest a change in the current scene to guide the user.
[0052] In some embodiments, the suggested change includes at least one of a change of a pose of a detected object and a change of the detected object.
[0053] In some embodiments, the suggested change includes at least one of a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
[0054] Additionally, in some embodiments, the artificial intelligence analyzes the static background to identify the same objects, and the pose of the camera in the original photo may be estimated. If the current placement of the camera device or phone is not matching, a suggestion may be given to adjust the position of the camera.
[0055] Moreover, if automatic zoom is enabled, the camera device may do small adjustments automatically. Given that the exact field of view, lens distortion and other camera device parameters might not match between the camera device that has captured the original image or photo and the current camera device, adjustments may be suggested to make the match as close as possible, and image capture parameters like ISO, exposure time, aperture may be set automatically, for example, to enable algorithms to match the style of the taken photo against the original photo in a postprocessing. The postprocessing may also be done automatically or in response to a user instruction. The postprocessing may include utilizing Al-based style transfer methods.Our ref. : 250072EPWOP 5
[0056] Sony Group Corporation
[0057] Hence, in some embodiments, the artificial intelligence is configured to suggest a change of a pose of the camera device to guide the user.
[0058] In some embodiments, the artificial intelligence is configured to suggest a change of image capture parameters to guide the user.
[0059] In some embodiments, the artificial intelligence is configured to control adjustment of image capture parameters to match the original image.
[0060] In some embodiments, the image capture parameters include an ISO value, a shutter speed, an aperture opening and an exposure time.
[0061] In some embodiments, the artificial intelligence is configured to transfer a style from the original image to the captured image.
[0062] In some embodiments, the circuitry is configured to receive a user input representing an instruction for the artificial intelligence to perform the image style transfer from the original image to the captured image.
[0063] In some embodiments, the artificial intelligence is configured to automatically perform the image style transfer from the original image to the captured image.
[0064] Some embodiments pertain to a (corresponding) method for guiding a user to recreate an original image, wherein the method includes:
[0065] receiving the original image;
[0066] capturing an image of a current scene; and
[0067] inputting the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0068] The method may be performed by the camera device as described herein.
[0069] It has been recognized that also videos may be handled similarly, the videos may range from short videos (e.g., for (social) media networks) to long videos, which may be more frequently used in professional setting, such as a movie and video clip shoot or virtual production.
[0070] Some embodiments pertain to a video camera device for guiding a user to recreate an original video, wherein the camera device includes circuitry configured to:
[0071] receive the original video;
[0072] capture a video of a current scene; andOur ref. : 250072EPWOP 6
[0073] Sony Group Corporation
[0074] input the original video and the captured video into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original video and the captured video, and to guide the user to match the original video.
[0075] The video camera device may be any device capable of capturing videos such as a smartphone, a tablet, a laptop, smart glasses, an augmented or virtual reality device or the like.
[0076] Some embodiments pertain to a (corresponding) method for guiding a user to recreate an original video, wherein the method includes:
[0077] receiving the original video;
[0078] capturing a video of a current scene; and
[0079] inputting the original video and the captured video into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original video and the captured video, and to guide the user to match the original video.
[0080] The methods as described herein are also implemented in some embodiments as a computer program causing a computer and / or a processor to perform the method, when being carried out on the computer and / or processor. In some embodiments, also a non-transitory computer-readable recording medium is provided that stores therein a computer program product, which, when executed by a processor, such as the processor described above, causes the methods described herein to be performed.
[0081] Returning to Fig. 1, there is schematically illustrated in a block diagram an embodiment of a camera device 1, which is discussed in the following.
[0082] The camera device 1 includes a storage 2, a processor 3, an image sensor 4, a display 5 and an audio device 6.
[0083] The camera device 1 stores various data in the storage 2 such as computer program instructions to be executed by the processor 3 and media data such as image data representing an image. For example, the image may be a particular image of a past scene or a historical event. The image may be received from a server in a computer network or may have been captured by the camera device 1 previously. The image may be referred to as original image in the following. A user of the camera device 1 may now wish to recreate the original image.
[0084] The user then operates the camera device 1 to capture an image of a current scene with the image sensor 4, the current scene is assumed to represent a similar scene as depicted in the original image.Our ref. : 250072EPWOP 7
[0085] Sony Group Corporation
[0086] In order to assist or guide the user in recreating the original image, the original image and the captured image of the current scene are processed by the processor 3.
[0087] In particular, the processor 3 executes an artificial intelligence 7 which includes a CNN and a LLM.
[0088] The processor 3 inputs the original image and the captured image of the current scene into the artificial intelligence 7.
[0089] First, the artificial intelligence 7 uses the CNN to determine matches and differences in the original image and the image of the current scene.
[0090] The CNN is configured to perform image analysis of each of the original image and the image of the current scene. The CNN performs, as part of the image analysis, object detection and classification to determine the positions and types of the objects present in each of the original image and the image of the current scene. Moreover, for persons, the CNN performs, as part of the image analysis, skeleton estimation to determine a pose of each detected person in each of the original image and the image of the current scene. Additionally, for persons, the CNN performs, as part of the image analysis, facial expression estimation to determine a facial expression of each detected person in each of the original image and the image of the current scene. Furthermore, the CNN performs, as part of the image analysis, static background analysis to determine a pose of the camera device that has captured the respective image.
[0091] Then, the CNN determines matches and differences in the original image and the image of the current scene, based on the performed image analysis of each of the original image and the image of the current scene.
[0092] The CNN then inputs the determined matches and differences in the original image and the image of the current scene into the LLM.
[0093] Alternatively or additionally to the CNN, the artificial intelligence may also use a ViT for visual analysis. Other known architectures used for providing an artificial intelligence with the ability to perform a visual analysis may be used as well in addition or in alternative to the CNN and / or the ViT.
[0094] Thus, second, the artificial intelligence 7 uses the LLM to explain the matches and / or differences to the user using at least one of text output, voice output, sound output and visual overlays on the image of the current scene and / or the original image to guide the user to match the original image (e.g., in subsequent image captures). For example, the artificial intelligence 7 may use the display 5 and the audio device 6 to guide the user to match the original image.Our ref.: 250072EPWOP 8
[0095] Sony Group Corporation
[0096] The LLM then suggests a change in the current scene to guide the user, for example, a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
[0097] Additionally, the LLM suggests a change of a pose of the camera device 1 to guide the user to match the original image.
[0098] Thereby, the user is guided to match the original image, for example, in subsequent image captures such that the user is iteratively guided to the matching image.
[0099] Fig. 2 schematically illustrates in a flow diagram an embodiment of a method 100, which is discussed in the following.
[0100] The method 100 may be performed by the camera device as described herein, for example, the camera device 1 of Fig. 1 may perform the method 100.
[0101] At 101, an original image is received, as discussed herein.
[0102] At 102, an image of a current scene is captured, as discussed herein.
[0103] At 103, the original image and the captured image are input into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image, as discussed herein.
[0104] At 104, a change in the current scene is suggested to guide the user, as discussed herein.
[0105] At 105, a change of a pose of a camera device that has captured the image of the current scene is suggested to guide the user, as discussed herein.
[0106] At 106, a change of image capture parameters is suggested to guide the user, as discussed herein. At 107, a style is transferred from the original image to the captured image, as discussed herein. It should be recognized that the embodiments describe methods with an exemplary ordering of method steps. The specific ordering of method steps is however given for illustrative purposes only and should not be construed as binding.
[0107] Fig. 3 schematically illustrates in a block diagram an embodiment of a multi-purpose computer 130 which can be used for implementing a camera device.
[0108] The computer 130 can be implemented such that it can basically function as any type of camera device as described herein. The computer has components 131 to 142, which can form a circuitry, such as any one of the circuitries of the camera device as described herein.Our ref. : 250072EPWOP 9
[0109] Sony Group Corporation
[0110] Embodiments which use software, firmware, programs or the like for performing the methods as described herein can be installed on computer 130, which is then configured to be suitable for the concrete embodiment.
[0111] The computer 130 has a CPU 131 (Central Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0112] The computer 130 has a GPU 141 (Graphical Processing Unit), which can execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0113] The CPU 131 and the GPU 141 can commonly execute various types of procedures and methods as described herein, for example, in accordance with programs stored in a read-only memory (ROM) 132, stored in a storage 137 and loaded into a random-access memory (RAM) 133, stored on a medium 140 which can be inserted in a respective drive 139, etc.
[0114] The CPU 131, the ROM 132, the RAM 133 and the GPU 141 are connected with a bus 142, which in turn is connected to an input / output interface 134. The number of CPUs, GPUSs, memories and storages is only exemplary, and the skilled person will appreciate that the computer 130 can be adapted and configured accordingly for meeting specific requirements which arise, when it functions as a camera device.
[0115] At the input / output interface 134, several components are connected: an input 135, an output 136, the storage 137, a communication interface 138 and the drive 139, into which a medium 140 (compact disc, digital video disc, compact flash memory, or the like) can be inserted.
[0116] The input 135 can be a pointer device (mouse, graphic table, or the like), a keyboard, a microphone, a camera, a touchscreen, etc.
[0117] The output 136 can have a display (liquid crystal display, cathode ray tube display, light emittance diode display, etc.), loudspeakers, etc.
[0118] The storage 137 can have a hard disk, a solid-state drive and the like.Our ref. : 250072EPWOP 10
[0119] Sony Group Corporation
[0120] The communication interface 138 can be adapted to communicate, for example, via a local area network (LAN), wireless local area network (WLAN), mobile telecommunications system (GSM, UMTS, LTE, NR etc.), Bluetooth, infrared, etc.
[0121] It should be noted that the description above only pertains to an example configuration of computer 130.
[0122] All units and entities described in this specification and claimed in the appended claims can, if not stated otherwise, be implemented as integrated circuit logic, for example on a chip, and functionality provided by such units and entities can, if not stated otherwise, be implemented by software.
[0123] In so far as the embodiments of the disclosure described above are implemented, at least in part, using software-controlled data processing apparatus, it will be appreciated that a computer program providing such software control and a transmission, storage or other medium by which such a computer program is provided are envisaged as aspects of the present disclosure.
[0124] Note that the present technology can also be configured as described below.
[0125] (1) A camera device for guiding a user to recreate an original image, wherein the camera device includes circuitry configured to:
[0126] receive the original image;
[0127] capture an image of a current scene; and
[0128] input the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0129] (2) The camera device of (1), wherein the artificial intelligence is configured to guide the user by using at least one of text output, voice output, sound output and visual output.
[0130] (3) The camera device of (1) or (2), wherein the artificial intelligence is configured to suggest a change in the current scene to guide the user.
[0131] (4) The camera device of (3), wherein the suggested change includes at least one of a change of a pose of a detected object and a change of the detected object.
[0132] (5) The camera device of (3) or (4), wherein the suggested change includes at least one of a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
[0133] (6) The camera device of any one of (1) to (5), wherein the artificial intelligence is configured to suggest a change of a pose of the camera device to guide the user.Our ref. : 250072EPWOP 11
[0134] Sony Group Corporation
[0135] (7) The camera device of any one of (1) to (6), wherein the artificial intelligence is configured to suggest a change of image capture parameters to guide the user.
[0136] (8) The camera device of any one of (1) to (7), wherein the artificial intelligence is configured to control adjustment of image capture parameters to match the original image.
[0137] (9) The camera device of (7) or (8), wherein the image capture parameters include an ISO value, a shutter speed, an aperture opening and an exposure time.
[0138] (10) The camera device of any one of (1) to (9), wherein the artificial intelligence is configured to transfer a style from the original image to the captured image.
[0139] (11) A method for guiding a user to recreate an original image, wherein the method includes:
[0140] receiving the original image;
[0141] capturing an image of a current scene; and
[0142] inputting the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
[0143] (12) The method of (11), including guiding the user by using at least one of text output, voice output, sound output and visual output.
[0144] (13) The method of (11) or (12), including suggesting a change in the current scene to guide the user.
[0145] (14) The method of (13), wherein the suggested change includes at least one of a change of a pose of a detected object and a change of the detected object.
[0146] (15) The method of (13) or (14), wherein the suggested change includes at least one of a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
[0147] (16) The method of any one of (11) to (15), including suggesting a change of a pose of a camera device that has captured the image of the current scene to guide the user.
[0148] (17) The method of any one of (11) to (16), including suggesting a change of image capture parameters to guide the user.
[0149] (18) The method of any one of (11) to (17), including controlling adjustment of image capture parameters to match the original image.
[0150] (19) The method of (17) or (18), wherein the image capture parameters include an ISO value, a shutter speed, an aperture opening and an exposure time.Our ref.: 250072EPWOP 12
[0151] Sony Group Corporation
[0152] (20) The method of any one of (11) to (19), including transferring a style from the original image to the captured image.
[0153] (21) A computer program comprising program code causing a computer to perform the method according to anyone of (11) to (20), when being carried out on a computer.
[0154] (22) A non-transitory computer-readable recording medium that stores therein a computer program product, which, when executed by a processor, causes the method according to anyone of (11) to (20) to be performed.
Claims
Our ref.: 250072EPWOP 1Sony Group CorporationCLAIMS1. A camera device for guiding a user to recreate an original image, comprising circuitry configured to:receive the original image;capture an image of a current scene; andinput the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
2. The camera device of claim 1, wherein the artificial intelligence is configured to guide the user by using at least one of text output, voice output, sound output and visual output.
3. The camera device of claim 1, wherein the artificial intelligence is configured to suggest a change in the current scene to guide the user.
4. The camera device of claim 3, wherein the suggested change includes at least one of a change of a pose of a detected object and a change of the detected object.
5. The camera device of claim 3, wherein the suggested change includes at least one of a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
6. The camera device of claim 1, wherein the artificial intelligence is configured to suggest a change of a pose of the camera device to guide the user.
7. The camera device of claim 1, wherein the artificial intelligence is configured to suggest a change of image capture parameters to guide the user.
8. The camera device of claim 1, wherein the artificial intelligence is configured to control adjustment of image capture parameters to match the original image.
9. The camera device of claim 7 or 8, wherein the image capture parameters include an ISO value, a shutter speed, an aperture opening and an exposure time.
10. The camera device of claim 1, wherein the artificial intelligence is configured to transfer a style from the original image to the captured image.
11. A method for guiding a user to recreate an original image, comprising:receiving the original image;capturing an image of a current scene; andOur ref.: 250072EPWOP 2Sony Group Corporationinputting the original image and the captured image into an artificial intelligence, wherein the artificial intelligence is configured to determine matches and differences between the original image and the captured image, and to guide the user to match the original image.
12. The method of claim 11, comprising guiding the user by using at least one of text output, voice output, sound output and visual output.
13. The method of claim 11, comprising suggesting a change in the current scene to guide the user.
14. The method of claim 13, wherein the suggested change includes at least one of a change of a pose of a detected object and a change of the detected object.
15. The method of claim 13, wherein the suggested change includes at least one of a change of a pose, a facial expression, and a clothing or an accessory of a detected person.
16. The method of claim 11, comprising suggesting a change of a pose of a camera device that has captured the image of the current scene to guide the user.
17. The method of claim 11, comprising suggesting a change of image capture parameters to guide the user.
18. The method of claim 11, comprising controlling adjustment of image capture parameters to match the original image.
19. The method of claim 17 or 18, wherein the image capture parameters include an ISO value, a shutter speed, an aperture opening and an exposure time.
20. The method of claim 11, comprising transferring a style from the original image to the captured image.