Enhanced creative photography with generative artificial intelligence (AI) assistance
The electronic device uses generative AI to suggest creative compositions and provide real-time recommendations, addressing the limitations of post-production reliance in conventional photography systems by ensuring high-quality image capture and consistency.
Patent Information
- Application Number
- PCT/IB2025/056639
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-03
- Filing Date
- 2025-06-30
- Publication Date
- 2026-01-08
AI Technical Summary
Conventional photography systems rely heavily on post-production techniques, leading to complacency during actual photo shoots, resulting in lower-quality raw images and inconsistent results across multiple images requiring different edits and adjustments.
An electronic device utilizing generative artificial intelligence (AI) assists photographers by suggesting creative compositions based on user inputs, capturing images with different styles through a generative neural network, and providing real-time recommendations for optimal image capture.
Ensures high-quality image capture by intelligently suggesting compositions, allowing photographers to maintain consistent image quality across varying creative compositions without complex algorithms, enhancing the image capture process.
Smart Images

Figure IB2025056639_08012026_PF_FP_ABST
Abstract
Description
ENHANCED CREATIVE PHOTOGRAPHY WITH GENERATIVE ARTIFICIALINTELLIGENCE (Al) ASSISTANCECROSS-REFERENCE TO RELATED APPLICATIONS / INCORPORATION BY REFERENCE
[0001] This application claims priority to Indian Provisional Application No. IN202411051083, filed July 03, 2024, which is hereby incorporated by reference in its entirety.FIELD
[0002] Various embodiments of the disclosure relate to creative photography suggestion systems. More specifically, various embodiments of the disclosure relate to an electronic device and a method to enhance creative photography with generative artificial intelligence (Al) assistance.BACKGROUND
[0003] Advancements in image recognition and processing have led to systems that refine and enhance captured images, with improved accuracy and reliability. In digital photography and computer vision, image quality and clarity may be crucial for effective recognition systems. Enhanced operational capabilities may allow for sophisticated applications across various fields and improve image quality and system performance. Post-production techniques like image enhancement, noise reduction, filtering, segmentation, color space conversion, normalization, resizing, and registration may optimize images for analysis but require complex algorithms. Such complexity can limit photographers’ creative compositions and increase costs and time, especially with large volumes of images. Over-reliance on post-production can compromise the quality of original images. An improvement of the image capture process based on the context of each image may be essential to produce high-quality images.
[0004] Limitations and disadvantages of conventional and traditional approaches will become apparent to one of skill in the art, through comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0005] An electronic device and method to perform enhanced creative photography with generative Al assistance is provided substantially as shown in, and / or described in connection with, at least one of the figures, as set forth more completely in the claims.
[0006] These and other features and advantages of the present disclosure may be appreciated from a review of the following detailed description of the present disclosure, along with the accompanying figures in which like reference numerals refer to like parts throughout.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] FIG. 1 is a block diagram that illustrates an exemplary network environment to enhance creative photography with generative artificial intelligence (GenAI) assistance, in accordance with an embodiment of the disclosure.
[0008] FIG. 2 is a block diagram that illustrates an exemplary electronic device of FIG.1 , in accordance with an embodiment of the disclosure.
[0009] FIGs. 3A, 3B, 3C, and 3D are block diagrams that collectively illustrate an exemplary scenario for enhancement of photography with the generative Al assistance, in accordance with one embodiment of the disclosure.
[0010] FIGs. 4A, 4B, 4C, 4D, 4E, and 4F are block diagrams that collectively illustrate an exemplary scenario for enhancement of photography with the generative Al assistance, in accordance with another embodiment of the disclosure.
[0011] FIG. 5 is a flowchart that illustrates operations of an exemplary method to enhance photography with the generative Al assistance, in accordance with anembodiment of the disclosure.DETAILED DESCRIPTION
[0012] The present disclosure relates to an intelligent composition suggestion system that may utilize generative artificial intelligence (GenAI) to enable visualization of different ways to capture enhanced images by a photographer (or user) such that the images match a certain context. The following described implementation may be found in an electronic device and method that may be configured to enhance creative photography based on generative Al assistance. Exemplary aspects of the present disclosure may provide an electronic device that may be configured to receive a user input associated with a scene. Next, the electronic device may prepare a plurality of prompts based on the user input. The electronic device may generate a plurality of images corresponding to a plurality of composition styles based on application of a generative neural network on the plurality of prompts. The electronic device may then control a display device to display the plurality of images on a user interface (III). The electronic device may then acquire, from a cognitive sensor system, attention information associated with a user of the display device. The electronic device may then select an image from the plurality of images based on the attention information. The electronic device may then determine, from the plurality of composition styles, a first composition style that corresponds to the selected image. The electronic device may then generate recommendation information comprising guidelines for a capture of image data with a second composition style that is same as or substantially similar to the first composition style. The electronic device may then control the display device to display the recommendation information.
[0013] Traditional systems may often rely on post-production image capture techniques, which can lead to complacency during an actual photo shoot, that may potentially result in lower-quality raw images. This may pose a major challenge for achievement of consistent results across multiple images, especially when the multiple images require different levelsof edits and adjustments. In contrast, the disclosed electronic device may intelligently suggest creative compositions associated with a captured image and assist the photographer in the composition process associated with the captured image, which may ensure effective production of the high-quality images based on the suggested creative compositions.
[0014] The electronic device may have collaborative capabilities that allow collaboration of different photographic styles of the composition associated with the captured image. For instance, with respect to each of the photographic styles, the electronic device may assist the photographer to obtain particular photographic styles of the composition. Thus, unlike traditional systems, the disclosed electronic device may not require complex and sophisticated recognition algorithms to capture the images with different creative compositions. Furthermore, the disclosed electronic device may be designed to capture each image with different creative compositions based on the image context and different visualizations set by the photographer to capture each image. Furthermore, the disclosed electronic device may utilize the generative Al to visualize an intent of the photographer based on a set of prompts, to render high-quality images associated with the set of prompts. The visualized intent of the photographer may allow the electronic device to cater to a wider range of use cases and scenarios. Thus, the visualization of the intent may make the electronic device more versatile than traditional systems that utilize the captured images or user-designed sketches to provide recommendations and insights related to a high-quality image capture process.
[0015] The disclosed electronic device may provide features for suggestions related to production of creative compositions of a captured final image as per user requirements, which can help to ensure a robust and improved image capture process. Such features, combined with location identification capabilities, may work together to provide a superior image capture experience for the photographer, based on the image or textual descriptionprovided as the context by the photographer. The disclosed electronic device may provide an enhanced image capture experience based on a suggestion of different creative compositions associated with the context of the image capture process. This may allow the photographer to rapidly switch between different high-quality images without compromise of captured image quality. In contrast, by use of traditional systems, photographers may struggle to maintain consistent quality of captured images, especially when the multiple images require different creative compositions.
[0016] FIG. 1 is a block diagram that illustrates an exemplary network environment to enhance creative photography with generative artificial intelligence (Al) assistance, in accordance with an embodiment of the disclosure. With reference to FIG. 1 , there is shown a network environment 100. The network environment 100 may include an electronic device 102, a server 104, a database 106, and a communication network 138. The electronic device 102 may include (or may be associated with) an image-capture device 108 configured to capture images of a scene 114. There is further shown a user input 112 associated with the scene 114. The user input 112 may include images or a sketch associated with a target object 116 of the scene 114. The electronic device 102 may further include a generative neural network 122. A cognitive sensor system 124 may be associated with the electronic device 102. In FIG. 1 , there is further shown a plurality of images 118 and a plurality of composition styles 120 that may be stored in the database 106. Further, there is shown a user 110 who may be associated with and / or operate the electronic device 102. FIG. 1 further shows attention information 126, which may be associated with the user 110 of a display device (for example, a display device 206A, shown in FIG. 2) associated with the electronic device 102.
[0017] The electronic device 102 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive the user input 112 associated with the scene 114. The electronic device 102 may prepare a plurality of prompts based on the user input 112.The electronic device 102 may generate the plurality of images 118 corresponding to the plurality of composition styles 120 based on application of the generative neural network 122 on the plurality of prompts. The electronic device 102 may control the display device to display the plurality of images 118 on a user interface. The electronic device 102 may acquire, from the cognitive sensor system 124, the attention information 126 associated with the user 110 of the display device. Thereafter, the electronic device 102 may select an image 128 from the plurality of images 118 based on the attention information 126. The electronic device 102 may determine, from the plurality of composition styles 120, a first composition style 130 that corresponds to the selected image 128. The electronic device 102 may generate recommendation information 132 comprising guidelines for a capture of image data 134 with a second composition style 136 that is same as or substantially similar to the first composition style 130. The electronic device 102 may control the display device to display the recommendation information 132. Examples of the electronic device 102 may include, but are not limited to, a computing device, a smartphone, a cellular phone, a mobile phone, a gaming device, a mainframe machine, a server, a computer workstation, a wearable device, an image-capture device, and / or a consumer electronic (CE) device.
[0018] The server 104 may include suitable logic, circuitry, and interfaces, and / or code that may be configured to receive the user input 112 associated with the scene 114. The server 104 may prepare the plurality of prompts based on the user input 112. The server 104 may generate the plurality of images 118 corresponding to the plurality of composition styles 120 based on the application of the generative neural network 122 on the plurality of prompts. The server 104 may control the display device to display the plurality of images 118 on the user interface. The server 104 may acquire, from the cognitive sensor system 124, the attention information 126 associated with the user 110 of the display device. The server 104 may select the image 128 from the plurality of images 118 based on theattention information 126. The server 104 may determine, from the plurality of composition styles 120, the first composition style 130 that corresponds to the selected image 128. The server 104 may generate the recommendation information 132 comprising the guidelines for the capture of the image data 134 with the second composition style 136 that is same as or substantially similar to the first composition style 130. The server 104 may control the display device to display the recommendation information 132.
[0019] The server 104 may be implemented as a cloud server and may execute operations through web applications, cloud applications, HTTP requests, repository operations, file transfer, and the like. Other example implementations of the server 104 may include, but are not limited to, a database server, a file server, a web server, a media server, an application server, a mainframe server, a machine learning server (enabled with or hosting, for example, a computing resource, a memory resource, and a networking resource), or a cloud computing server.
[0020] In at least one embodiment, the server 104 may be implemented as a plurality of distributed cloud-based resources by use of several technologies that are well known to those ordinarily skilled in the art. A person with ordinary skill in the art will understand that the scope of the disclosure may not be limited to the implementation of the server 104 and the electronic device 102, as two separate entities. In certain embodiments, the functionalities of the server 104 can be incorporated in its entirety or at least partially in the electronic device 102 without a departure from the scope of the disclosure.
[0021] In certain embodiments, the server 104 may host the database 106. In such a scenario, the server 104 may receive a request for retrieval of information (such as, the plurality of images 118 and / or the plurality of composition styles 120) from a device, such as, the electronic device 102. Based on the received request, the server may retrieve the requested information and transmit the retrieved information to the electronic device 102. Alternatively, the server 104 may be separate from the database 106 and may becommunicatively coupled to the database 106.
[0022] The database 106 may include suitable logic, interfaces, and / or code that may be configured to store the plurality of images 118 corresponding to the plurality of composition styles 120. The database 106 may be configured to store the plurality of composition styles 120. The database 106 may be derived from data off a relational or non-relational database, or a set of comma-separated values (csv) files in conventional or big-data storage. The database 106 may be stored or cached on a device, such as a server (e.g., the server 104) or the electronic device 102. The device which stores the database 106 may be configured to receive an information retrieval query from the electronic device 102 or the server 104. In response, the device of the database 106 may be configured to retrieve and provide the queried information to the electronic device 102 or the server 104 for generation of creative compositions, based on the received query.
[0023] In some embodiments, the database 106 may be hosted on a plurality of servers stored at the same or different locations. The operations of the database 106 may be executed using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). In some other instances, the database 106 may be implemented using software.
[0024] As used herein, the term “image-capture device” may refer to a device for detection of capture of the image data 134 after the display of the recommendation information 132. Examples of the image-capture device 108 may include, but are not limited to, an image sensor, a wide-angle camera, an action camera, a closed-circuit television (CCTV) camera, a camcorder, a camera with an integrated depth sensor, a cinematic camera, Digital Single-Lens Reflex (DSLR) camera, a Digital Single-Lens Mirrorless (DSLM) camera, a digital camera, a camera phone, a time-of-flight camera (ToF camera), a night-vision camera, and / or other image capture devices.
[0025] As used herein, the term “user” may refer to a photographer who operates and controls the image capture device 108. This may include the person responsible for capture of the images of the target object 116 of the scene 114.
[0026] The generative neural network 122 may be a computational network or a system of artificial neurons, arranged in a plurality of layers, as nodes that may be configured to receive the user input 112 associated with the scene 114. The plurality of layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each layer of the plurality of layers may include one or more nodes (or artificial neurons, represented by circles, for example). Outputs of all nodes in the input layer may be coupled to at least one node of hidden layer(s). Similarly, inputs of each hidden layer may be coupled to outputs of at least one node in other layers of the neural network. Outputs of each hidden layer may be coupled to inputs of at least one node in other layers of the neural network. Node(s) in the final layer may receive inputs from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined from hyper-parameters of the neural network. Such hyper-parameters may be set before, while training, or after training the neural network on a training dataset.
[0027] Each node of the generative neural network 122 may correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters, tunable during training of the network. The set of parameters may include, for example, a weight parameter, a regularization parameter, and the like. Each node may use the mathematical function to compute an output based on one or more inputs from nodes in other layer(s) (e.g., previous layer(s)) of the neural network. All or some of the nodes of the neural network may correspond to same or a different mathematical function.
[0028] In an embodiment, the generative neural network 122 may be a type of an artificial intelligence system (also referred to as an artificial deep neural network) configured to process and understand multiple types of data modalities, such as text,images, audio, 3D data, and video, simultaneously. The generative neural network 122 may extend the capabilities of traditional large language models (LLMs) by integrating various forms of data, enabling a more comprehensive understanding and generation of information across different media types. In some embodiments, the generative neural network 122 may be a large language model, such as a transformer-based decoder-only model, an encoder-decoder model (that uses transformers), or a model that uses neural networks other than transformers. In these or other embodiments, the generative neural network 122 may include multiple encoders specialized for processing different modalities of data (such as image and text). The generative neural network 122 may include a Fusion mechanism to integrate the outputs from various encoders
[0029] In training of the generative neural network 122, one or more parameters of each node of the neural network may be updated based on whether an output of the final layer for a given input (from the training dataset) matches a correct result based on a loss function for the neural network. The above process may be repeated for same or a different input until a minima of loss function may be achieved and a training error may be minimized. Several methods for training are known in art, for example, gradient descent, stochastic gradient descent, batch gradient descent, gradient boost, meta-heuristics, and the like.
[0030] The generative neural network 122 may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The generative neural network 122 may rely on libraries, external scripts, or other logic / instructions for execution by a processing device. The generative neural network 122 may include code and routines configured to enable a computing device to perform one or more operations. Additionally, or alternatively, the generative neural network 122 may be implemented using hardware including a processor, a microprocessor (e.g., to perform or control performance of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the generative neural network 122 may be implemented using a combination of hardware and software.
[0031] In an embodiment, the generative neural network 122 may be a scalable deeplearning model comprising an encoder model, a decoder model, and a set of convolution neural network layers. The scalable deep-learning model may take un-processed data such as, an unprocessed user input, to generate the plurality of images 118 corresponding to the plurality of composition styles 120. The encoder model of the present disclosure may receive the user input 112 associated with the scene 114. Based on the received user input, the encoder model may determine a compressed feature vector associated with the scene 114. An encoded version (i.e. , the compressed feature vector) of the received user input associated with the scene 114 may be transmitted to the decoder model. The decoder model may reconstruct the input dataset such as, the plurality of images 118, back from the encoded version. Thus, the decoder model may decompress the compressed feature vector associated with the scene 114. Each of the set of convolution neural network layers may perform a dot product between two matrices. Herein, a first matrix also known as a kernel, may include a set of learnable parameters and a second matrix may be a portion of a receptive field associated with the corresponding convolution neural network layer. In an embodiment, a kernel size associated with each of the set of convolution neural network layers may be even. That is, the kernel size may be “2”, “4”, “6”, “8”, and so on.
[0032] In an embodiment, the scalable deep-learning model may be a machine learning (ML) model. The ML model may be trained to identify a relationship between inputs, such as, features in a training dataset, and output labels, such as, the user input 112 associated with the scene 114 and the plurality of images 118. The ML model may be defined by its hyper-parameters, for example, number of weights, cost function, input size, number oflayers, and the like. The parameters of the ML model may be tuned and weights may be updated so as to move towards a global minima of a cost function for the ML model. After several epochs of the training on the feature information in the training dataset, the ML model may be trained to output labels associated with identified relationships between the inputs.
[0033] The ML model may include electronic data, which may be implemented as, for example, a software component of an application executable on the electronic device 102. The ML model may rely on libraries, external scripts, or other logic / instructions for execution by a processing device. The ML model may include code and routines configured to enable a computing device, such as the electronic device 102 to perform one or more operations such as, preparation of the plurality of prompts, generation of the plurality of images 118, acquisition of the attention information 126, selection of the image 128, determination of the first composition style 130, and / or generation of the recommendation information 132.
[0034] Examples of the generative neural network 122 may include, but are not limited to, a Generative Adversarial Network (GAN) model, a variational autoencoder (VAE) model, an auto-regressive model, a Generative Pre-trained Transformers (GPT) model, a Bi-directional Encoder Representations from Transformers (BERT) model, a large language model (LLM), a multi-modal generative model, or a neural language model.
[0035] The cognitive sensor system 124 may be configured to acquire the attention information 126 associated with the user 110 of the display device. The cognitive sensor system 124 may further comprise a brain-control interface (BCI) and the attention information 126 may comprise electroencephalogram (EEG) signal data of the user 110 of the electronic device 102. The cognitive sensor system 124 may further comprise an eye gaze tracking sensor and the attention information 126 may comprise gaze information of the user 110 of the electronic device 102. The cognitive sensor system 124 may furthercomprise an imaging sensor and the attention information 126 may comprise images of the user 110 of the electronic device 102. Examples of the cognitive sensor system 124 may include, but are not limited to, eye (or gaze) tracking sensors, electroencephalography (EEG) sensors, galvanic skin response (GSR) sensors, heart rate variability (HRV) sensors, electrocardiography (ECG) sensors, pupil dilation sensors, facial expression recognition sensors, voice analysis sensors, motion and gesture sensors, skin temperature sensors, respiration rate sensors, electromyography (EMG) sensors, blood oxygen level sensors (pulse oximeters), and wearable activity trackers.
[0036] As used herein, the term “attention information” may refer to the information which is selectively emphasized as important features of the captured image with deemphasized less relevant features of the image. The attention information 126 may be acquired from the cognitive sensor system 124 that utilizes dynamic weighting mechanism to select specific parts of the captured image that carry more significance, such as, in case of object detection, image segmentation, image captioning, and other image processing applications. In an embodiment, the attention information 126 may be processed to determine a user preference level for each image of the plurality of images 118. The image 128 may be selected based on the user preference level for the image 128 that is above a threshold level. The first composition style 130 may be determined from the plurality of composition styles 120 that correspond to the selected image 128. The image data 134 may be captured with the second composition style 136 that is same as or substantially similar to the first composition style 130.
[0037] As used herein, the term “recommendation information” may comprise the guidelines for the capture of the image data 134 with the second composition style 136 that Is same or substantially similar to the first composition style 130. The guidelines may include at least one of lighting conditions, an opportunity window, weather condition information, a lens configuration to be used with the image-capture device 108, a suitablebody posture, an imaging angle, a set of imaging parameters, or a movement plan for the capture of the image data 134. In an embodiment, the recommendation information 132 may include at least one of a location recommendation or a geospatial intelligence information. The location recommendation may include scene features specified in the user input 112. The geospatial intelligence information may comprise at least one of locations of vantage points suitable for the capture of the image data 134, and information about the scene features in a surrounding area of the location recommendation. The scene features may include a state of the target object 116 to be included in the image data 134.
[0038] The communication network 138 may include a communication medium through which the electronic device 102 and the server 104 may communicate with one another. The communication network 138 may be one of a wired connection or a wireless connection. Examples of the communication network 138 may include, but are not limited to, the Internet, a cloud network, Cellular or Wireless Mobile Network (such as Long-Term Evolution and 5thGeneration (5G) New Radio (NR)), satellite communication system (using, for example, low earth orbit satellites), a Wireless Fidelity (Wi-Fi) network, a Personal Area Network (PAN), a Local Area Network (LAN), or a Metropolitan Area Network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 138 in accordance with various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of a Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), Zig Bee, EDGE, IEEE 802.11 , light fidelity (Li-Fi), 802.16, IEEE 802.11 s, IEEE 802.11 g, multi-hop communication, wireless access point (AP), device to device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.
[0039] In operation, the electronic device 102 may receive the user input 112 associatedwith the scene 114. The user input 112 may include images of the target object 116 of the scene 114 captured by the user 110 from the image-capture device 108. In an example, the electronic device 102 may retrieve the captured images of the target object 116 as the user input 112 from the database 106. In another example, the electronic device 102 may receive the images of the target object 116 of the scene 114 from the image-capture device 108. Details related to reception of the user input are described further, for example, in FIG. 5 (at 504).
[0040] The electronic device 102 may prepare the plurality of prompts based on the user input 112. In an example, the electronic device 102 may retrieve a prompt template including a base instruction and context associated with the user input 112. The scene information may be extracted from the user input 112, that comprises at least one of a digital image of the scene 114, a textual description of the scene 114, a 3D representation of the scene 114, or a sketch of the scene 114. Thereafter, the electronic device 102 may update the prompt template to include the scene information as the context. The electronic device 102 may then modify the base instruction for each of the plurality of composition styles 120 in the updated prompt template to prepare the plurality of prompts. Details related to the preparation of the plurality of prompts are described further, for example, in FIG. 5 (at 506).
[0041] The electronic device 102 may be configured to generate the plurality of images 118 corresponding to the plurality of composition styles 120 based on application of the generative neural network 122 on the plurality of prompts. In an instance, in case the user input is an image, the electronic device 102 may apply the generative neural network 122 to generate the plurality of images 118 corresponding to the plurality of composition styles 120. In an embodiment, the electronic device 102 may adjust attributes of the image, for instance, geometric shapes, forms, lighting, and shadows associated with the image to generate the plurality of images 118. Alternatively, the electronic device 102 may capturesurface quality of an object in the image and may further use side lighting to emphasize textures of the object. Alternatively, or in addition, the electronic device 102 may alter colors of the object of the image to signify emotions and set the mood. In yet another scenario, the electronic device 102 may adjust a placement of the object, which may create a balanced composition. Alternatively, the electronic device 102 may generate striking and captivating (emotion-filled) images as the plurality of images 118. The electronic device 102 may further be configured to facilitate continuous experimentation with different techniques, styles, and objects for generation of the plurality of images 118. Details related to the generation of the plurality of images are described further, for example, in FIG. 5 (at 508).
[0042] The electronic device 102 may be configured to control the display device to display the plurality of images 118 on the user interface. Details related to the control of the display device are disclosed further, for example, in FIG. 5 (at 510).
[0043] The electronic device 102 may be configured to acquire, from the cognitive sensor system 124, the attention information 126 associated with the user 110 of the display device. For example, the electronic device 102 may selectively emphasize the attention information 126 as important features of the captured image while the less relevant features of the image may be downplayed. The electronic device 102 may acquire the attention information 126 from the cognitive sensor system 124 by use of dynamic weighting mechanism to select specific parts of the captured image that carry more significance. Details related to the acquisition of the attention information are described further, for example, in FIG. 5 (at 512).
[0044] The electronic device 102 may be configured to select the image 128 from the plurality of images 118 based on the attention information 126. In an example, the image 128 may be selected based on the user preference level determined for the image 128 as above a threshold level. The user preference level may be determined for each image ofthe plurality of images 118 based on processing of the attention information 126. Details related to the selection of the image from the plurality of images are described further, for example, in FIG. 5 (at 514).
[0045] The electronic device 102 may be configured to determine, from the plurality of composition styles 120, the first composition style 130 that corresponds to the selected image 128. In an example, the first composition style 130 may refer to the composition of the selected image 128 as per user requirements. The user requirements may be in form of a video, an image, a rough sketch, or a textual description provided as the user input 112. For example, the user input 112 may be provided in the form of the textual description as: “Golden Sunset”. Based on this textual description, a best image setting may be applied to effectively arrange visual elements within a frame of the environment around the sunset, which may render the first composition style 130 corresponding to the selected image.
[0046] In another example, the user input 112 may include words like “pulpy” or “crusty”, which may be draw a mental picture, for instance, the words “pulpy” or “crusty” may activate reflexes. Based on this textual description, a best image setting may be applied to effectively arrange visual elements within a frame of the environment based on the words “pulpy” or “crusty”, which may render the first composition style 130 corresponding to the selected image. Details related to the determination of the first composition style corresponding to the selected image are disclosed further, for example, in FIG. 5 (at 516).
[0047] The electronic device 102 may be configured to generate the recommendation information 132 comprising guidelines for the capture of the image data 134 with the second composition style 136 that is same as or substantially similar to the first composition style 130. In an example, the electronic device 102 may be configured to select, from the plurality of prompts, a first prompt associated with the selected image. The electronic device 102 may further be configured to extract, from the first prompt, aninstruction text associated with the first composition style 130 and the scene 114. The recommendation information 132 comprising the guidelines may be generated further based on the analysis of the extracted instruction text. For example, the user input 112 may be provided in form of the textual description as: “Golden Sunset”. Based on the textual description, recommendations may be provided to the user 110 with respect to locations or exact points from where background of the image associated with the scene 114 may be captured effectively. The electronic device 102 may apply the generative neural network 122 on the recommendation information 132 to perform aesthetic evaluation and recommend composition guidelines associated with the recommendation information 132. Details related to the generation of the recommendation information are disclosed further, for example, in FIG. 5 (at 518).
[0048] The electronic device 102 may be configured to control the display device to display the recommendation information 132. In an example, the recommendation information may refer to guidelines including at least one of lighting conditions, an opportunity window, weather condition information, a lens configuration to be used with the image-capture device 108, a suitable body posture, an imaging angle, a set of imaging parameters, or a movement plan for the capture of the image data 134. Details related to the control of the display device are further disclosed, for example, in FIG. 5 (at 520).
[0049] Traditional systems may often rely on post-production image capture techniques, which can lead to complacency during an actual photo shoot, that may potentially result in lower-quality raw images. This may pose a major challenge for achievement of consistent results across multiple images, especially when the multiple images require different levels of edits and adjustments. In contrast, the disclosed electronic device 102 may intelligently suggest creative compositions associated with a captured image and assist the photographer (e.g., the user 110) in the composition process associated with the captured image, which may ensure effective production of the high-quality images based on thesuggested creative compositions.
[0050] The electronic device 102 may have collaborative capabilities that allow collaboration of different photographic styles of the composition associated with the captured image. For instance, with respect to each of the photographic styles, the electronic device 102 may assist the photographer (e.g., the user 110) to obtain particular photographic styles of the composition. Thus, unlike traditional systems, the disclosed electronic device 102 may not require complex and sophisticated recognition algorithms to capture the images with different creative compositions. Furthermore, the disclosed electronic device 102 may be designed to capture each image with different creative compositions based on the image context and different visualizations set by the photographer (e.g., the user 110) to capture each image. Furthermore, the disclosed electronic device 102 may utilize the generative Al (e.g., the generative neural network 122) to visualize an intent of the photographer based on a set of prompts, to render high- quality images associated with the set of prompts.
[0051] In an embodiment, the electronic device 102 may visualize an intent of the photographer based on factors such as, subject genre, theme, characterization, and styles (e.g. noir, vintage, modem, urban, etc.). Alternatively, or in addition, the electronic device 102 may visualize the intent of the photographer based on creative composition guidance such as, single shot / narrative sequence / movie scene, composition genre, style (for instance, black white or colored, portrait or landscape), visual angles, and shot (long or short). Alternatively, or in addition, the electronic device 102 may use creative elements such as shadows, shadow patterns, texture, silhouette, and other such elements other than the main object or background of the image to visualize the intent of the photographer. The visualized intent of the photographer (e.g., the user 110) may allow the electronic device 102 to cater to a wider range of use cases and scenarios. Thus, the visualization of the intent may make the electronic device 102 more versatile than traditional systemsthat utilize the captured images or user-designed sketches to provide recommendations and insights related to a high-quality image capture process.
[0052] The disclosed electronic device 102 may provide features for suggestions related to production of creative compositions of a captured final image as per user requirements, which can help to ensure a robust and improved image capture process. Such features, combined with location identification capabilities, may work together to provide a superior image capture experience for the photographer (e.g., the user 110), based on the image or textual description provided as the context by the photographer (e.g., the user 110). The disclosed electronic device 102 may provide an enhanced image capture experience based on a suggestion of different creative compositions associated with the context of the image capture process. The suggestions may include factors like choice of lenses, lighting (natural light / artificial light), exposure (underexposure is suggested when there are more shadows and vice versa), depth of field (DoF), framing, composition, and color grading. The suggestions may also include suggestions related to focus aspects such as single-point focus, zone focus, auto-area focus, and the like. The focus aspects and setting may vary for distinct objects, and various techniques, such as portrait focus technique and landscape focus technique, may be applied for updating the focus aspects and settings. This may allow the photographer (e.g., the user 110) to rapidly switch between different high-quality images without compromise of captured image quality. In contrast, by use of traditional systems, photographers may struggle to maintain consistent quality of captured images, especially when the multiple images require different creative compositions.
[0053] FIG. 2 is a block diagram that illustrates an exemplary electronic device of FIG.1 , in accordance with an embodiment of the disclosure. FIG. 2 is explained in conjunction with elements from FIG. 1 . With reference to FIG. 2, there is shown a block diagram 200 of the exemplary electronic device 102. The electronic device 102 may include the generative neural network 122, a circuitry 202, a memory 204, an input / output (I / O) device206, and a network interface 208. The memory 204 may store different types of images. The memory 204 may store the plurality of images 118 corresponding to the plurality of composition styles 120. The input / output (I / O) device 206 may include a display device 206A.
[0054] The circuitry 202 may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be executed by the electronic device 102. The operations may include user input reception, plurality of prompts preparation, an operation for generation of the plurality of images 118, first control of the display device, attention information acquisition, image selection, the first composition style determination, the recommendation information generation, and second control of the display device. The circuitry 202 may include one or more processing units, which may be implemented as a separate processor. In an embodiment, the one or more processing units may be implemented as an integrated processor or a cluster of processors that perform the functions of the one or more specialized processing units, collectively. The circuitry 202 may be implemented based on a number of processor technologies known in the art. Examples of implementations of the circuitry 202 may be an X86-based processor, a Graphics Processing Unit (GPU), a Reduced Instruction Set Computing (RISC) processor, an Application-Specific Integrated Circuit (ASIC) processor, a Complex Instruction Set Computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other control circuits. Various operations of the circuitry 202 for implementation to enhance the creative photography with the generative Al assistance are described further in detail, for example, in FIG. 3A, FIG. 3B, FIG. 3C, FIG. 3D, FIG. 4, FIG. 4A, FIG. 4B, FIG. 4C, FIG. 4D, FIG. 4E, and FIG. 4F.
[0055] The memory 204 may include suitable logic, circuitry, interfaces, and / or code that may be configured to store one or more instructions to be executed by the circuitry 202. The one or more instructions stored in the memory 204 may be configured to execute thedifferent operations of the circuitry 202 (and / or the electronic device 102). The memory 204 may be further configured to store the plurality of images 118 and the plurality of composition styles 120. The memory 204 may further store the user input 112, the attention information 126, and the recommendation information 132. Examples of implementation of the memory 204 may include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Electrically Erasable Programmable Read- Only Memory (EEPROM), Hard Disk Drive (HDD), a Solid-State Drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0056] The I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. For example, the I / O device 206 may receive the user input 112 indicative of a captured image, a sketch, or a textual description associated with the scene 114. The I / O device 206 may also receive the user input 112 that includes the plurality of images 118 associated with the scene 114. The I / O device 206 may render the recommendation information 132 as guidelines for the user 110 to capture images associated with a given scene and context. The I / O device 206 may include the display device 206A. Examples of the I / O device 206 may include, but are not limited to, a display (e.g., a touch screen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of the I / O device 206 may further include braille I / O devices, such as, braille keyboards and braille readers.
[0057] The display device 206A may include suitable logic, circuitry, and interfaces that may be configured to display or render the plurality of images 118 associated with the scene 114. The display device 206A may also render the recommendation information 132. The display device 206A may be a touch screen which may enable a user or users to provide a user-input via the display device 206A. The touch screen may be at least one of a resistive touch screen, a capacitive touch screen, or a thermal touch screen. The display device 206A may be realized through several known technologies such as, but notlimited to, at least one of a Liquid Crystal Display (LCD) display, a Light Emitting Diode (LED) display, a plasma display, or an Organic LED (OLED) display technology, or other display devices. In accordance with an embodiment, the display device 206A may refer to a display screen of a head mounted device (HMD), a smart-glass device, a see-through display, a projection-based display, an electro-chromic display, or a transparent display.
[0058] The network interface 208 may include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication between the electronic device 102 and the server 104, via the communication network 138. The network interface 208 may be implemented by use of various known technologies to support wired or wireless communication of the electronic device 102 with the communication network 138. The network interface 208 may include, but is not limited to, an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, a coder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, or a local buffer circuitry.
[0059] The network interface 208 may be configured to communicate via wireless communication with networks, such as the Internet, an Intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a plurality of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), wideband code division multiple access (W-CDMA), Long Term Evolution (LTE), 5thGeneration (5G) New Radio (NR), code division multiple access (CDMA), time division multiple access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (such as IEEE 802.11 a, IEEE 802.11 b, IEEE 802.11 g or IEEE 802.11 n), voice over Internet Protocol (VoIP), light fidelity (Li-Fi), Worldwide Interoperability for Microwave Access (Wi-MAX), a protocol for email, instant messaging, and a Short Message Service (SMS).
[0060] FIGs. 3A, 3B, 3C, and 3D are diagrams that collectively illustrate an exemplary scenario for enhancement of photography with the generative Al assistance, in accordance with one embodiment of the disclosure. FIGs. 3A, 3B, 3C, and 3D are explained in conjunction with elements from FIG. 1 and FIG. 2. With reference to FIGs. 3A, 3B, 30, and 3D, there is shown an exemplary scenario 300. The exemplary scenario 300 depicts a user input 302 associated with a scene 304. A prompt template 306 is shown including a placeholder 308 for insertion of an instruction text and a context associated with the user input 302. Scene information 312 is further shown which may be extracted from the user input 302 comprising a sketch of the scene 304. In an example, the scene information 312 may comprise the sketch of a cow in the scene 304.
[0061] For example, the sketch of the cow in a dense forest area of the scene 304 may be received as the user input 302. In this scenario, the electronic device 102 may be configured to receive the sketch of the cow as the user input 302. The prompt template 306 may be retrieved that includes the placeholder 308 for the base instruction and the context required for extraction of the scene information 312 from the user input 302. In some instances, the scene information 312 may comprise at least one of a digital image of the scene 304, a textual description of the scene 304, or a 3D representation of the scene 304.
[0062] The electronic device 102 may prepare the plurality of prompts based on the user input 302. The plurality of prompts may include a first prompt 310A and a second prompt 310B associated with the sketch of the cow. Instruction text may be extracted from each of the plurality of prompts. The instruction text may be a base instruction associated with the first composition style 130 and the scene 304. For example, the first prompt 310A may be prepared to include the base instruction as: “Image of a cow with horns in a misty area with a person standing close. Use a 24 millimeter (mm) wide-angle lens to emphasize the vastness of the misty landscape. Position the cow and person slightly off-center to createa balanced composition, with the mist adding an ethereal quality of the scene”. The second prompt 31 OB may be prepared to include the base instruction as: “Image of a cow with horns in a misty area with a person standing close. Use an 85 mm lens with a wide aperture (f / 1.8) to achieve a shallow depth of field, focusing sharply on the cow’s horns and the person’s face. The misty background will blur softly, enhancing the ethereal atmosphere of the photograph”. Thereafter, the electronic device 102 may update the prompt template 306 to include the scene information 312 as the context along with the first prompt 310A and the second prompt 31 OB.
[0063] The electronic device 102 may apply the generative neural network 122 on the plurality of prompts to generate the plurality of images 118 corresponding to the plurality of composition styles 120. For example, the generative neural network 122 may be applied on the first prompt 310A and the second prompt 310B and the scene information 312 to generate the plurality of images 118. The plurality of images 118 may comprise a first image 314A and a second image 314B corresponding to the plurality of composition styles 120. The first image 314A and the second image 314B may be displayed through a user interface 316 of the display device 206A.
[0064] The display device 206A may include a screen to present an option for selection of either the first image 314A or the second image 314B to a user 318. The electronic device 102 may then acquire, from the cognitive sensor system 124, attention information 320 associated with the user 318 of the display device 206A. For example, the acquired attention information 320 may be displayed by the user interface 316 of the display device 206A as either a high EEG response 322A or a low EEG response 322B of the user 318 of the electronic device 102. A gaze duration corresponding to the high EEG response 322A and the low EEG response 322B of the user 318 may be generated. In an instance, the gaze duration corresponding to the high EEG response 322A and the low EEG response 322B may be displayed as 3 second (3s) and 1 second (1 s), respectively.
[0065] The electronic device 102 may be configured to select an image 324 from the plurality of images 118 based on the attention information 320. For example, the first image 314A may be selected as the image 324 from the plurality of images 118 based on the attention information 320. The image 324 may be selected based on the user preference level for the first image 314A of the plurality of images 118 that is above a threshold level. The electronic device 102 may then determine, from the plurality of composition styles 120, the first composition style 130 that corresponds to the selected image 324. In an example, the electronic device 102 may select, from the plurality of prompts, the first prompt 310A associated with the selected image 324. The instruction text (for example, the base instruction of the prompt template 306) associated with the first composition style 130 and the scene 304 may be then extracted from the first prompt 310A. Thereafter, the instruction text may be analyzed for generation of the recommendation information 132. In some instances, the first prompt 310A may be prepared to include the base instruction as: “Based on the image generated by GenAI using the prompt (for example, the first prompt 310A), can you create a detailed photography composition guideline as a recommendation along with location recommendations for a photographer?”.
[0066] In one scenario, the electronic device 102 may generate the recommendation information 132 associated with capture of the image data 134. For example, the generative neural network 122 may be applied to the selected image 324 (for example, the first image 314A) to generate the recommendation information 132. The recommendation information 132 may comprise guidelines 328 for a capture of the image data 134 with the second composition style 136 that is same or substantially similar to the first composition style 130. The display device 206A may then be controlled to display the recommendation information 132 on a first III element 326 of the user interface 316. In some instances, the generated recommendation information 132 comprising the guidelines 328 may be displayed on the first III element 326 as: “24mm wide-angle lensrecommended for perfectly capturing vastness of misty landscape. Capture dust using cow and person as foreground elements to add depth to image and apply lighting conditions during daytime. Use natural elements like paths, fences, or lines in landscape to lead viewer’s eye towards the cow and the person to guide viewer through the image. Balance composition by including other elements on opposite side of image frame. For example, tress, rocks, or other landscape features”.
[0067] In another scenario, the generated recommendation information 132 may further include at least one of a location recommendation 330 and a geospatial intelligence information. The location recommendation 330 may include scene features specified in the user input 302, and the geospatial intelligence information may comprise locations of vantage points suitable for the capture of the image data 134 and information about the scene features in a surrounding area of the location recommendation 330. The display device 206A may display the location recommendation 330 on the first Ul element 326 of the user interface 316. In some instances, the location recommendation 330 may be displayed on the first Ul element 326 as: “Known for its stunning misty landscapes, “Loch Katrine” offers a perfect backdrop for capturing ethereal scenes. The combination of water, hills, and mist creates a magical atmosphere”.
[0068] It should be noted that the exemplary scenario 300 of FIGs. 3A, 3B, 3C, and 3D is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0069] FIGs. 4A, 4B, 4C, 4D, 4E, and 4F are diagrams that collectively illustrate an exemplary scenario for enhancement of the photography with the generative Al assistance, in accordance with another embodiment of the disclosure. FIGs. 4A, 4B, 4C, 4D, 4E, and 4F are explained in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3A, FIG. 3B, FIG. 3C, and FIG. 3D. With reference to FIGs. 4A, 4B, 4C, 4D, 4E, and 4F, there is shown an exemplary scenario 400. The exemplary scenario 400 depicts a user input402 associated with a scene 406. Further, a text prompt 404 associated with the scene406 and a user 408 is shown to capture the image of a landscape scenery as the scene 406 using the image-capture device 108. The text prompt 404 may be a textual description of the scene 406. Scene information 416 is further shown which may be extracted from the user input 402 comprising a digital image of the scene 406. In an example, the scene information 416 may comprise the digital image or a 3D representation of a landscape scenery in the scene 406.
[0070] For example, the digital image or the 3D representation of the landscape scenery in a hilly area of the scene 406 may be received as the user input 402. In this scenario, the electronic device 102 may be configured to receive the digital image of the landscape scenery as the user input 402 along with the text prompt 404. The text prompt 404 may be prepared to include the base instruction as: “Capture a clear and detailed view of a landscape scenery containing all scenic elements”. A prompt template 410 may be retrieved that includes a placeholder 412 for the base instruction and the context required for extraction of the scene information 416 from the user input 402. In some instances, the scene information 416 may comprise the textual description of the scene 406.
[0071] The electronic device 102 may prepare the plurality of prompts based on the user input 402. The plurality of prompts may include a first prompt 414A and a second prompt 414B associated with the digital image of the landscape scenery. The instruction text may be extracted from each of the plurality of prompts to include the base instruction associated with the first composition style 130 and the scene 406. For example, the first prompt 414A may be the instruction text prepared to include the base instruction as: “Image of a landscape scenery in a misty area with a person standing close to tress and other landscape elements to capture a clear and full view. Use a 44mm wide-angle lens to emphasize the vastness of the misty landscape. Imagine dividing image frame into a 3*3 grid and position trees and mountains at their intersections to create a balanced composition”. The second prompt 414B may be prepared to include the base instructionas: “Image of a landscape scenery with a close and clear view of landscape scene elements. Use a 90 mm lens with a wide aperture (f / 1.8) to achieve a shallow depth of field, focusing sharply on the landscape scenery and associated background elements. The misty landscape should be slightly visible through the mist to maintain a sense of mystery and ethereal quality”. Thereafter, the electronic device 102 may update the prompt template 410 to include the scene information 416 as the context along with the first prompt 414A and the second prompt 414B.
[0072] The electronic device 102 may apply the generative neural network 122 on the plurality of prompts to generate the plurality of images 118 corresponding to the plurality of composition styles 120. For example, the generative neural network 122 may be applied to the first prompt 414A and the second prompt 416B and the scene information 416 to generate the plurality of images 118. The plurality of images 118 may comprise a first image 418A and a second image 418B corresponding to the plurality of composition styles 120. The first image 418A and the second image 418B may be displayed through a user interface 420 of the display device 206A.
[0073] The display device 206A may include the screen to present an option for selection of either the first image 418A or the second image 418B by the user 408. The electronic device 102 may then acquire, from the cognitive sensor system 124, attention information 422 associated with the user 408 of the display device 206A. For example, the acquired attention information 422 may be displayed by the user interface 420 of the display device 206A as either a high EEG response 424A or a low EEG response 424B of the user 408 of the electronic device 102. A gaze duration corresponding to the high EEG response 424A and the low EEG response 424B of the user 408 may be generated. In an instance, the gaze duration corresponding to the high EEG response 424A and the low EEG response 424B may be displayed as 7s and 3s, respectively.
[0074] The electronic device 102 may be configured to select an image 426 from theplurality of images 118 based on the attention information 422. For example, the first image 418A may be selected as the image 426 from the plurality of images 118 based on the attention information 422. The image 426 may be selected based on the user preference level for the first image 418A of the plurality of images 118 that is above a threshold level. The electronic device 102 may then determine, from the plurality of composition styles 120, the first composition style 130 that corresponds to the selected image 426. In an example, the electronic device 102 may select, from the plurality of prompts, the first prompt 414A associated with the selected image 426. The instruction text (for example, the base instruction of the prompt template 410) associated with the first composition style 130 and the scene 406 may be then extracted from the first prompt 414A. Thereafter, the instruction text may be analyzed for generation of the recommendation information 132. In some instances, the first prompt 414A may be prepared to include the base instruction as: “Based on the image generated by GenAI using the prompt (for example, the first prompt 310A), can you create a detailed photography composition guideline as a recommendation along with location recommendations and a vantage point for a photographer?”.
[0075] In one scenario, the electronic device 102 may generate the recommendation information 132 associated with the capture of the image data 134. For example, the generative neural network 122 may be applied to the selected image 426 (for example, the first image 418A) to generate the recommendation information 132. The recommendation information 132 may comprise guidelines 430 for a capture of the image data 134 with the second composition style 136 that is same or substantially similar to the first composition style 130. The display device 206A may then be controlled to display the recommendation information 132 on a second III element 428 of the user interface 420. In some instances, the generated recommendation information 132 comprising the guidelines 430 may be displayed on the second III element 428 as: “40mm wide-angle lens recommended for perfectly capturing vastness of landscape scenery. Capture thelandscape scenery with tress, hills, and valleys in close proximity and lighting conditions visible based on time of day. Use empty spaces around the landscape scenery to emphasize the vastness of the landscape scenery to highlight isolation and tranquility of the scene. Misty conditions often provide soft, diffused light which is ideal for creating a serene and calm atmosphere. Avoid harsh lighting to maintain the ethereal quality of the scene. Balance composition by including other elements on opposite side of image frame. For example, tress, rocks, or other landscape scenic features.”
[0076] In another scenario, the generated recommendation information 132 may further include at least one of a location recommendation 432 and the geospatial intelligence information. The location recommendation 432 may include scene features specified in the user input 302, and the geospatial intelligence information may comprise locations of vantage points suitable for the capture of the image data 134 and information about the scene features in the surrounding area of the location recommendation 432. The display device 206A may display the location recommendation 432 on the second Ul element 428 of the user interface 420. In some instances, the location recommendation 432 may be displayed on the second Ul element 428 as: “The Peak District at top of scene provides a variety of landscapes, including misty valleys and rolling hills. It's a great location for capturing interplay of light and mist”.
[0077] Based on the display of the location recommendation 432, a vantage point 436 may be detected for capture of the image data 134 with the second composition style 136. For example, the vantage point 436 may be a location identified as top of the hilly area to capture the image data 134 which is represented by a downward arrow. The vantage point 436 identified may be displayed on a third Ul element 434 of the user interface 420. Based on the display of the vantage point 436, the user 408 may accordingly plan movement activity to reach the vantage point 436 from the current location. The movement activity may be a particular time stamp at which the user 408 needs to capture the image forobtaining the image data 134 with the second composition style 136. A fourth III element 438 may be presented on the screen of the user interface 420 which depicts the guidelines 430 for capture of the image data 134. The guidelines 430 may comprise at least one of the lighting conditions for the capture of the image data 134, the opportunity window for the capture of the image data 134, the weather condition information for the capture of the image data 134, the lens configuration to be used with an image-capture device 108 for the capture of the image data 134, the body posture suitable for the capture of the image data 134, the imaging angle suitable for the capture of the image data 134, or the set of imaging parameters for the image-capture device 108. For example, with reference to FIG. 4F as shown, the user 408 may be in a seated body posture for capture of the landscape scenery as the image data 134 with the second composition style 136, and the user 408 may need to tilt the camera horizontally to get a clear view for capture of the landscape scenery with the landscape scenic elements.
[0078] The electronic device 102 may detect, via a location sensor, a location of the user 408 as proximal to the location recommendation 432. The electronic device 102 may acquire image information based on the detected location of the user 408. Thereafter, the electronic device 102 may detect the scene features in the surrounding area based on the image information. The electronic device 102 may then detect an opportunity window for the capture of the image data 134 based on the detection of the scene features. The scene features may include a state of the target object 116 to be included in the image data 134. The recommendation information 132 may be displayed further based on the detection of the opportunity window. An alert notification may be transmitted to the display device 206A based on the detected opportunity window.
[0079] The landscape scenery as the image data 134 may be captured with the second composition style 136 by the user 408, and the captured image data 134 may then be compared with the guidelines 430 to verify whether the image data 134 have beencaptured as per the user requirements set in the guidelines 430. For example, an image verification window may appear on a fifth III element 440 that represents the verification of the captured image data as per the guidelines 430. The captured image data 134 may be verified based on the guidelines 430 like proper coverage of the image specifications, type of camera, precise image coordinates, ideal lighting conditions, correct time stamp, an emotional state of the user.
[0080] In an example, the electronic device 102 may detect the captured image data via the image-capture device 108 after the display of the recommendation information 132. The electronic device 102 may acquire, from the cognitive sensor system 124, a cognitive signal indicative of an emotional state of the user 408 at a time of the capture of the image data 134. Thereafter, the electronic device 102 may identify emotional aspects of the scene in the image data 134. The electronic device 102 may then determine a match between the emotional aspects of the scene in the image data 134 and the emotional state of the user 408.
[0081] It should be noted that the exemplary scenario 400 of FIGs. 4A, 4B, 4C, 4D, 4E, and 4F is for exemplary purposes and should not be construed to limit the scope of the disclosure.
[0082] FIG. 5 is a flowchart that illustrates operations of an exemplary method to enhance the photography with the generative Al assistance, in accordance with an embodiment of the disclosure. FIG. 5 is described in conjunction with elements from FIG. 1 , FIG. 2, FIG. 3, FIG. 3A, FIG. 3B, FIG. 3C, FIG. 4A, FIG. 4B. FIG. 4C, FIG. 4D, FIG. 4E, and FIG. 4F. With reference to FIG. 5, there is shown an exemplary flowchart 500. An exemplary method depicted in the flowchart 500 may include operations from 502 to 520 that may be implemented by the electronic device 102 of FIG. 1 or by the circuitry 202 of FIG. 2. The exemplary method 500 may start at 502 and proceed to 504.
[0083] At block 504, a user input associated with the composition of a scene may bereceived. The circuitry 202 may be configured to receive the user input 112 associated with the composition of the scene 114. Details related to the reception of the user input are described further, for example, in FIG. 3A (at 302) and FIG. 4A (at 402).
[0084] At block 506, a plurality of prompts may be prepared based on the user input. The circuitry 202 may be configured to prepare the plurality of prompts based on the user input 112. Details related to the preparation of the plurality of prompts are described further, for example, in FIG. 3B (at 310A and 310B) and FIG. 4B (at 414A and 414B).
[0085] At block 508, a plurality of images may be generated corresponding to a plurality of composition styles based on application of the generative neural network on the plurality of prompts. The circuitry 202 may be configured to generate the plurality of images 118 corresponding to the plurality of composition styles 120 based on the application of the generative neural network 122 on the plurality of prompts. Details related to the generation of the plurality of images are described further, for example, in FIG. 3B (at 314A and 314B) and FIG. 4B (at 418A and 418B).
[0086] At block 510, a display device may be controlled to display the plurality of images on a user interface. The circuitry 202 may be configured to control the display device 206A to display the plurality of images 118 on the user interface. Details related to the first control of the display device are further described, for example, in FIG. 3C (at 316) and FIG. 4C (at 420).
[0087] At block 512, attention information associated with a user of the display device may be acquired from a cognitive sensor system. The circuitry 202 may be configured to acquire, from the cognitive sensor system 124, the attention information 126 associated with the user 110 of the display device 206A. Details related to the acquisition of the attention information are described further, for example, in FIG. 3C (at 320) and FIG. 4C (at 422).
[0088] At block 514, a first image may be selected from the plurality of images based onthe attention information. The circuitry 202 may be configured to select the first image 128 from the plurality of images 118 based on the attention information 126. Details related to the selection of the first image are described further, for example, in FIG. 3C (at 322A) and FIG. 4B (at 424A).
[0089] At block 516, a first composition style that corresponds to the selected first image may be determined from the plurality of composition styles. The circuitry 202 may be configured to determine, from the plurality of composition styles 120, the first composition style 130 that corresponds to the selected first image 128. Details related to the determination of the first composition style are described further, for example, in FIG. 3C (at 324) and FIG. 4C (at 426).
[0090] At block 518, recommendation information comprising guidelines for capture of image data with the second composition style may be generated. The second composition style 136 may be same or substantially similar to the first composition style 130. The circuitry 202 may be configured to generate the recommendation information 132 comprising the guidelines 328 for the capture of the image data 134 with the second composition style that is same or substantially similar to the first composition style 130. Details related to the generation of the recommendation information are described further, for example, in FIG. 3D (at 326, 328, and 330) and FIG. 4D (at 428, 430, and 432).
[0091] At block 520, the display device may be controlled to display the recommendation information. The circuitry 202 may be configured to control the display device 206A to display the recommendation information 132. Details related to the second control of the display device are described further, for example, in FIG. 3D (at 326) and FIG. 4D (at 428). Control may pass to end.
[0092] Although the exemplary method 500 is illustrated as discrete operations, such as 502, 504, 506, 508, 510, 512, 514, 516, 518 and 520, the disclosure is not so limited.Accordingly, in certain embodiments, such discrete operations may be further divided intoadditional operations, combined into fewer operations, or eliminated, depending on the implementation without detracting from the essence of the disclosed embodiments.
[0093] Various embodiments of the disclosure may provide a non-transitory computer- readable medium and / or storage medium having stored thereon, computer-executable instructions by a machine and / or a computer to operate an electronic device (for example, the electronic device 102 of FIG. 1 ). Such instructions may cause the electronic device 102 to perform operations that may include receipt of user input (e.g., the user input 112) associated with a scene 114. The operations may further include preparation of a plurality of prompts (e.g., the first prompt 310A and the second prompt 31 OB) based on the user input 112. The operations may further include generation of a plurality of images (e.g., the plurality of images 118) corresponding to a plurality of composition styles (e.g., the composition styles 120), based on application of generative neural network 122 on the plurality of prompts. The operations may further include control of a display device (e.g., the display device 206A) to display the plurality of images 118 on a user interface. The operations may further include acquisition, from a cognitive sensor system 124, an attention information 126 associated with a user 110 of the display device 206A. The operations may further include selection of an image 128 from the plurality of images 118 based on the attention information 126. The operations may further include determination, from the plurality of composition styles 120, a first composition style 130 that corresponds to the selected image 128. The operations may further include generation of recommendation information 132 comprising guidelines (for e.g., the guidelines 328) for the capture of the image data 134 with a second composition style 136 that is same as or substantially similar to the first composition style 130. The operations may further include control of the display device 206A to display the recommendation information 132.
[0094] Exemplary aspects of the disclosure may provide an electronic device (such as, the electronic device 102 of FIG. 1 ) that includes circuitry (such as, the circuitry 202). Thecircuitry 202 may be configured to receive a user input (e.g., the user input 112) associated with a scene 114. The circuitry 202 may be configured to prepare a plurality of prompts (e.g., the first prompt 310A and the second prompt 31 OB) based on the user input 112. The circuitry 202 may be configured to generate a plurality of images (e.g., the plurality of images 118) corresponding to a plurality of composition styles (e.g., the composition styles 120), based on application of generative neural network 122 on the plurality of prompts. The circuitry 202 may be configured to control a display device (e.g., the display device 206A) to display the plurality of images 118 on a user interface. The circuitry 202 may be configured to include acquisition, from a cognitive sensor system 124, an attention information 126 associated with a user 110 of the display device 206A. The circuitry 202 may be configured to select an image 128 from the plurality of images 118 based on the attention information 126. The circuitry 202 may be configured to determine, from the plurality of composition styles 120, a first composition style 130 that corresponds to the selected image 128. The circuitry 202 may be configured to generate recommendation information 132 comprising guidelines (for e.g., the guidelines 328) for the capture of the image data 134 with a second composition style 136 that is same as or substantially similar to the first composition style 130. The circuitry 202 may be configured to control the display device 206A to display the recommendation information 132.
[0095] The circuitry 202 may be configured to retrieve a prompt template 306 that includes a base instruction and context associated with the user input 112. The circuitry 202 may be configured to extract a scene information from the user input 112. The scene information may comprise at least one of a digital image of the scene 114, a textual description of the scene 114, a 3D representation of the scene 114, or a sketch of the scene 114. The circuitry 202 may be configured to update the prompt template 306 to include the scene information as the context. The circuitry 202 may be configured to modify, for each composition style of the plurality of composition styles 120, the baseinstruction in the updated prompt template to prepare the plurality of prompts.
[0096] The circuitry 202 may be further configured to process the attention information 126 to determine a user preference level for each image of the plurality of images 118. The image 128 may be selected based on the user preference level for the image 128 being above a threshold level. The circuitry 202 may be further configured to select, from the plurality of prompts, the first prompt 310A associated with the selected image 128. The circuitry 202 may be further configured to extract, from the first prompt 310A, an instruction text associated with the first composition style 130 and the scene 114. The recommendation information 132 comprising the guidelines 328 may be generated further based on the analysis of the extracted instruction text.
[0097] The circuitry 202 may be further configured to detect, via a location sensor, a location of the user 110 as proximal to the location recommendation 432. The circuitry 202 may be further configured to acquire image information based on the detected location of the user 110. The circuitry 202 may be further configured to detect the scene features in the surrounding area based on the image information. The circuitry 202 may be further configured to detect an opportunity window for the capture of the image data 134 based on the detection of the scene features. The scene features include a state of a target object 116 to be included in the image data 134. The circuitry 202 may be further configured to display the recommendation information 132 based on the detection of the opportunity window. An alert notification may be transmitted by the circuitry 202 to the display device 206A based on the detected opportunity window.
[0098] The circuitry 202 may be further configured to detect the captured image data via an image-capture device 108 after the display of the recommendation information 132. The circuitry 202 may be further configured to acquire, from the cognitive sensor system 124, a cognitive signal indicative of an emotional state of the user 110 at a time of the capture of the image data 134. The circuitry 202 may be further configured to identifyemotional aspects of the scene in the image data 134. The circuitry 202 may be further configured to determine a match between the emotional aspects of the scene in the image data 134 and the emotional state of the user 110.
[0099] The present disclosure may also be positioned in a computer program product, which comprises all the features that enable the implementation of the methods described herein, and which when loaded in a computer system is able to carry out these methods. Computer program, in the present context, means any expression, in any language, code or notation, of a set of instructions intended to cause a system with information processing capability to perform a particular function either directly, or after either or both of the following: a) conversion to another language, code or notation; b) reproduction in a different material form.
[0100] While the present disclosure is described with reference to certain embodiments, it will be understood by those skilled in the art that various changes may be made, and equivalents may be substituted without departure from the scope of the present disclosure. In addition, many modifications may be made to adapt a particular situation or material to the teachings of the present disclosure without departure from its scope. Therefore, it is intended that the present disclosure is not limited to the embodiment disclosed, but that the present disclosure will include all embodiments that fall within the scope of the appended claims.
Claims
CLAIMSWhat is claimed is:1 . An electronic device, comprising: circuitry configured to: receive a user input associated with a scene; prepare a plurality of prompts based on the user input; generate a plurality of images corresponding to a plurality of composition styles based on application of a generative neural network on the plurality of prompts; control a display device to display the plurality of images on a user interface; acquire, from a cognitive sensor system, attention information associated with a user of the display device; select an image from the plurality of images based on the attention information; determine, from the plurality of composition styles, a first composition style that corresponds to the selected image; generate recommendation information comprising guidelines for a capture of image data with a second composition style that is same as or substantially similar to the first composition style; and control the display device to display the recommendation information.2 The electronic device according to claim 1 , wherein the circuitry is further configured to: retrieve a prompt template that includes a base instruction and context associated with the user input;extract, from the user input, scene information comprising at least one of a digital image of the scene, a textual description of the scene, a 3D representation of the scene, or a sketch of the scene; update the prompt template to include the scene information as the context; and modify, for each composition style of the plurality of composition styles, the base instruction in the updated prompt template to prepare the plurality of prompts. The electronic device according to claim 1 , wherein the cognitive sensor system comprises a Brain-Control Interface and the attention information comprises electroencephalogram (EEG) signal data of a user of the electronic device. The electronic device according to claim 1 , wherein the cognitive sensor system comprises an eye gaze tracking sensor and the attention information comprises gaze information of a user of the electronic device. The electronic device according to claim 1 , wherein the cognitive sensor system comprises an imaging sensor and the attention information comprises images of a user of the electronic device. The electronic device according to claim 1 , wherein the circuitry is further configured to process the attention information to determine a user preference level for each image of the plurality of images, wherein the image is selected based on the user preference level for the image that is above a threshold level.The electronic device according to claim 1 , wherein the circuitry is further configured to: select, from the plurality of prompts, a first prompt associated with the selected image; and extract, from the first prompt, an instruction text associated with the first composition style and the scene, wherein recommendation information comprising the guidelines is generated further based on analysis of the extracted instruction text. The electronic device according to claim 1 , wherein the guidelines include at least one of: lighting conditions for the capture of the image data, an opportunity window for the capture of the image data, weather condition information for the capture of the image data, a lens configuration to be used with an image-capture device for the capture of the image data, a body posture suitable for the capture of the image data, an imaging angle suitable for the capture of the image data, a set of imaging parameters for the image-capture device, or a movement plan required for the capture of the image data. The electronic device according to claim 1 , wherein the recommendation information further includes at least one of: a location recommendation that includes scene features specified in the user input, and geospatial intelligence information comprising at least one of:locations of vantage points suitable for the capture of the image data, and information about the scene features in a surrounding area of the location recommendation.
10. The electronic device according to claim 9, wherein the circuitry is further configured to: detect, via a location sensor, a location of the user as proximal to the location recommendation; acquire image information based on the detected location of the user; detect the scene features in the surrounding area based on the image information; and detect an opportunity window for the capture of the image data based on the detection of the scene features.11 . The electronic device according to claim 10, wherein the scene features include a state of a target object to be included in the image data.
12. The electronic device according to claim 10, wherein the recommendation information is displayed further based on the detection of the opportunity window.
13. The electronic device according to claim 10, wherein the circuitry is further configured to transmit an alert notification to the display device based on the detected opportunity window.
14. The electronic device according to claim 1 , wherein the circuitry is further configured to:detect the capture of the image data via an image-capture device after the display of the recommendation information; acquire, from the cognitive sensor system, a cognitive signal indicative of an emotional state of the user at a time of the capture of the image data; identify emotional aspects of the scene in the image data; and determine a match between the emotional aspects of the scene in the image data and the emotional state of the user.
15. A method, comprising: in an electronic device: receiving a user input associated with a scene; preparing a plurality of prompts based on the user input; generating a plurality of images corresponding to a plurality of composition styles based on application of a generative neural network on the plurality of prompts; controlling a display device to display the plurality of images on a user interface; acquiring, from a cognitive sensor system, attention information associated with a user of the display device; selecting an image from the plurality of images based on the attention information; determining, from the plurality of composition styles, a first composition style that corresponds to the selected image; generating recommendation information comprising guidelines for a capture of image data with a second composition style that is same as or substantially similar to the first composition style; andcontrolling the display device to display the recommendation information.
16. The method according to claim 15, further comprising: retrieving a prompt template that includes a base instruction and context associated with the user input; extracting, from the user input, scene information comprising at least one of a digital image of the scene, a textual description of the scene, a 3D representation of the scene, or a sketch of the scene; updating the prompt template to include the scene information as the context; and modifying, for each composition style of the plurality of composition styles, the base instruction in the updated prompt template to prepare the plurality of prompts.
17. The method according to claim 15, further comprising processing the attention information to determine a user preference level for each image of the plurality of images, wherein the image is selected based on the user preference level for the image that is above a threshold level.
18. The method according to claim 15, wherein selecting, from the plurality of prompts, a first prompt associated with the selected image; and extracting, from the first prompt, an instruction text associated with the first composition style and the scene,wherein recommendation information comprising the guidelines is generated further based on analysis of the extracted instruction text.
19. The method according to claim 15, further comprising: detecting the capture of the image data via an image-capture device after the display of the recommendation information; acquiring, from the cognitive sensor system, a cognitive signal indicative of an emotional state of the user at a time of the capture of the image data; identifying emotional aspects of the scene in the image data; and determining a match between the emotional aspects of the scene in the image data and the emotional state of the user.
20. A non-transitory computer-readable medium having stored thereon, computerexecutable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising: receiving a user input associated with a scene; preparing a plurality of prompts based on the user input; generating a plurality of images corresponding to a plurality of composition styles based on application of a generative neural network on the plurality of prompts; controlling a display device to display the plurality of images on a user interface; acquiring, from a cognitive sensor system, attention information associated with a user of the display device; selecting an image from the plurality of images based on the attention information;determining, from the plurality of composition styles, a first composition style that corresponds to the selected image; generating recommendation information comprising guidelines for a capture of image data with a second composition style that is same as or substantially similar to the first composition style; and controlling the display device to display the recommendation information.
Citation Information
Patent Citations
Photography assistant and method for assisting a user in photographing landmarks and scenes
US20110314049A1
Glass-type terminal and method for controlling the same
US20150381885A1
Client terminal, display control method, program, and system
US20160080643A1
Imaging device, information terminal, control method for imaging device, and control method for information terminal
US20190253615A1
Image Composition Instruction Based On Reference Image Perspective
US20190320113A1