Apparatus and method for generating photorealistic synthetic images

The apparatus and method iteratively refine image searches using deep learning to generate photorealistic synthetic images, addressing complexity and inefficiencies in existing technologies, enabling quick and intuitive image generation aligned with user ideas.

WO2025175340A1PCT designated stage Publication Date: 2025-08-28BERSERQ PTE LTD +3

Patent Information

Application Number
PCT/AU2025/050133
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2025-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

The generation of photorealistic synthetic images is hindered by complex user interfaces, logistical challenges, and the need for intricate prompt engineering, which is time-consuming and unintuitive for non-technical users, leading to inefficient and suboptimal outputs that fail to capture the intended message or aesthetic.

Method used

An apparatus and method that iteratively refines image searches based on user inputs, using a deep learning model to generate synthetic images by progressively aligning with the user's visual idea, simplifying the process through semantic analysis and user-friendly interaction, eliminating the need for complex prompts.

Benefits of technology

This approach drastically reduces the time required to produce photorealistic images, enhances user engagement, and ensures outputs closely match the user's visual concept, improving productivity and creativity without requiring advanced technical knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050133_28082025_PF_FP_ABST
    Figure AU2025050133_28082025_PF_FP_ABST
Patent Text Reader

Abstract

A method and apparatus for generating photorealistic synthetic images by receiving multiple forms and multiple instances of user input corresponding to a user's visual idea, executing an iterative image search to identify pre-existing images semantically aligned with the user's visual idea, and using an image synthesis deep learning model to generate at least one synthetic image based on the multiple forms and instances of user inputs.
Need to check novelty before this filing date? Find Prior Art

Description

APPARATUS AND METHOD FOR GENERATING PHOTOREALISTIC SYNTHETIC IMAGESFIELD OF THE INVENTION

[0001] The present invention relates generally to digital image processing and, more specifically, to an apparatus and method for quickly and automatically generating photorealistic synthetic images that align closely with a user's visual idea.BACKGROUND

[0002] The generation of custom photo-quality digital images has traditionally been complex, requiring users to navigate cumbersome interfaces and master intricate software, posing a barrier for non-technical users. Conventional methods are time-consuming and inefficient, particularly for refining ideas or experimenting with visual concepts.

[0003] The traditional production of photo-quality digital images is constrained by factors such as weather conditions, subject availability, and the coordination of photographers, lighting technicians, and other support staff, leading to logistical challenges, increased costs, and delays. These limitations hinder rapid visual communication with clients, impacting efficiency, customer engagement, and responsiveness in fast-paced markets.

[0004] Ensuring photoshoot images meet the required standards for social media and marketing campaigns depends on clear client-brief communication with photographers. Additionally, effective micro-targeting requires diverse visuals tailored to specific customer segments and the ability to rapidly adapt to market trends and feedback, enhancing campaign effectiveness and responsiveness.

[0005] A “prompt” is an explicit text-based instruction typed by a user and provided to an Al model as input, guiding the Al model to produce a desired output. For example, this is a typical prompt: “Extreme close-up of a 24-year-old Nordic woman, standing inMarrakech during magic hour, cinematic film shot in 70mm.” Such a prompt exemplifies complexity that is not intuitive for most users unfamiliar with visual terminology such as “extreme close-up,” “magic hour,” “cinematic film shot,” and “shot in 70mm,” as well as their significance to a photo’s composition and style. Currently, prompt engineering is necessary for generative Al models to produce meaningful outputs, but the complexity and nuances of prompt design significantly impact effectiveness. Al models are highly sensitive to input prompts, where minor changes in wording, structure, or phrasing can lead to vastly different results, making prompt crafting more intuitive than systematic.Designing effective prompts requires understanding the model’s capabilities, structuring input strategically, and experimenting iteratively, which can be time-consuming and frustrating, especially for users without Al or programming expertise. The challenge is exacerbated when Al generates pixel-based visual outputs, as users must articulate their vision with sufficient detail, requiring knowledge of composition, lighting, and style, which most users lack. Furthermore, crafting precise prompts demands linguistic proficiency, which can disadvantage those from diverse linguistic or educational backgrounds. Unlike professionals accustomed to precise instructions, such as lawyers, writers, and radiologists, the average user may struggle to articulate requests in a way that maximises Al performance, potentially widening the digital divide and limiting accessibility to Al technologies.

[0006] The cognitive load of dictating or typing intricate prompts can deter user engagement with Al, especially when accuracy and detail are critical. Dictation and detailed typing require users to formulate thoughts, structure them for Al interpretation, and communicate them effectively, which can be challenging for those unaccustomed to providing detailed instructions. This often leads to frustration, errors, and suboptimal outputs. The time investment required for crafting precise prompts reduces efficiency, particularly in fast-paced environments where quick results are needed. As a result, Al tools intended to enhance productivity and creativity may instead become a barrier to adoption due to inefficiencies in prompt formulation.

[0007] Additionally, many users often struggle to translate abstract concepts or moods into the detailed, concrete terms required for accurate Al-generated images, resulting inoutputs that fail to capture the intended message or aesthetic. Refining prompts through iteration can be cumbersome and unintuitive, as users may struggle to identify which prompt elements to adjust, leading to an inefficient and frustrating trial-and-error process.

[0008] While some generative Al models can produce photo-quality images, achieving realism indistinguishable from real photographs remains challenging. Uncanny valley effects may cause discomfort, as Al-generated images often fall just short of lifelike accuracy. Realistic lighting and shadows are difficult to replicate, as Al struggles to authentically model light direction, intensity, and colour, which are essential for photorealism but difficult to articulate in prompts. Material and texture fidelity also presents challenges, as Al may misinterpret descriptions of surfaces like fabric, metal, or organic matter, leading to unrealistic renderings. Capturing dynamic motion and interactions further complicates generation, often requiring overly detailed prompts that exceed current Al capabilities. Additionally, Al models may struggle to accurately reflect specific cultural contexts or historical periods, leading to anachronisms or inaccuracies due to a lack of nuanced contextual understanding.

[0009] Stock photos typically lack originality, and there is a persistent difficulty in finding ones that precisely match a user’s intent, wasting countless hours for a compromised visual representation that often fails to fully convey the specific message, emotion, or atmosphere intended, leading to diluted brand identity and diminished engagement with the target audience. Any uniqueness of a stock photo is immediately undermined when multiple entities use the same stock images, which can dilute brand identity and reduce the impact of marketing efforts. Many stock photos are perceived as cheesy or inauthentic, featuring posed models and unrealistic scenarios that fail to resonate with audiences. This can weaken the message's credibility and detract from the user's intended narrative. Popular stock images can become overused, leading to audience fatigue. When consumers repeatedly see the same images across different contexts, it can diminish engagement and effectiveness of the visual content.

[0010] It is, accordingly, an object of the present invention to ameliorate one or more of the technical problems mentioned.SUMMARY OF THE INVENTIONIn a first aspect, the invention provides an apparatus comprising: a processor configured to: receive first user input comprising words that describe a user’s visual idea; execute an image search to retrieve pre-existing images semantically close to the described visual idea based on the received words; iteratively receive second user input indicating which of the retrieved images aligns with the user's visual idea, and in response to each iteration, progressively refine and execute the image search to retrieve increasingly focused pre-existing images that are semantically closer to the images indicated by the user as aligning with the user's visual idea; and generating at least one synthetic image based on the first user input and the second user input(s) using an image synthesis deep learning model.The processor may be further configured to: receive third user input comprising plain text instructions for further refinement of the at least one synthetic image; and generate at least one additional synthetic image based on the first user input, the second user input(s), and the third user input.The second user input may further comprise indications of which retrieved images do not align with the user's visual idea.The iteratively receiving of second user inputs may continue until reaching a saturation point where the user determines that it is unlikely that additional pre-existing images will more closely align with the user’s visual idea, and upon reaching the saturation point, the user initiates the generation of the synthetic images.The image synthesis deep learning model may be any one from the group consisting of: a latent diffusion model comprising a U-Net, autoencoder and text-encoder; an autoregressive text-to-image generation model; or a text-to-image Transformer model.The visual idea may comprise a mental image.The semantic closeness may be a predetermined threshold or user-configurable threshold in high-dimensional space, determined by any one from the group consisting of: cosine similarity, squared Euclidian distance, Manhattan distance, dot product and Hamming distance.The second user inputs may be stored in a memory in the form of a positive indicators list corresponding to images aligning with the user's visual idea and negative indicators list corresponding to images that do not align with the user's visual idea.The positive and negative indicators list may be used to weight a query for the image search by adjusting the importance of specific features or attributes present in the images in the positive indicators list while decreasing the weight of features present in the images in the negative indicators list.The processor may be further configured to:disambiguate the received words based on the sentence structure or surrounding words using a language model.The language model may comprise any one from the group consisting of: attention mechanism, Structured State Space sequence model, Monarch matrices, and combined set of linear projections using long convolutions and element- wise multiplication.The apparatus may further comprise: a recommendation engine to recommend the pre-existing images that are in the form of emojis based on semantic and visual relationships between user-selected emojis; and a visualisation component to display the impact of each emoji addition or subtraction on the generation of emojis through incremental thumbnails or an animated sequence; and a versioning control component to enable branching and exploration of different creative directions by the user without losing the original sequence of user-selected emojis.The apparatus may further comprise a merging tool for algorithmically blending image elements from different branches to suggest a coherent mix of user-selected emojis.The apparatus may further comprise a rollback feature to allow users to revert to an earlier point in the generation process or undo changes.In a second aspect, there is provided a method for generating synthetic images, the method comprising:receiving, by a processor, first user input comprising words that describe a user's visual idea; executing, by the processor, an image search to retrieve pre-existing images semantically close to the described visual idea based on the received words; iteratively receiving, by the processor, second user input indicating which of the retrieved images aligns with the user's visual idea, and in response to each iteration, progressively refining and executing the image search to retrieve increasingly focused pre-existing images that are semantically closer to the images indicated by the user as aligning with the user's visual idea; and generating, by the processor, at least one synthetic image based on the first user input and the second user input(s) using an image synthesis deep learning model.

[0011] The system may further comprise a plurality of interconnected modules, wherein an image database stores reference images and associated metadata, a preference capture module records user selections and behavioural data, a selection and reordering interface allows users to arrange reference images via a drag-and-drop mechanism, and a weighting engine assigns numerical weight values to selected reference images based on selection order, behavioural signals, historical preferences, image similarity metrics, and global trends, wherein the assigned weight values are processed by an Al image generation module to generate Al-produced images by blending attributes of the reference images in accordance with the computed weight distribution, and wherein a refinement module presents the Al-produced images to the user, allowing iterative selection adjustments, reordering of reference images, modification of weightings, or expression of new preferences, with updates feeding back into the weighting engine and Al image generation module to progressively refine the generation of successive Al-produced images.

[0012] The system may further comprise an iterative image refinement workflow, wherein user preferences are dynamically captured and applied to generate Al-synthesised outputs, wherein at initialisation, the system retrieves and displays an initial set ofreference images from an image database based on a machine learning model trained to predict reference images aligning with demographic attributes or prior user interactions, wherein a preference capture module records user actions, including "like" and "dislike" selections, and behavioural metrics such as the time interval between the presentation of reference images and user response, wherein a selection and reordering interface allows users to select a subset of preferred images from those marked as "liked" and provides a ranking mechanism enabling users to reposition the preferred images based on preference, and wherein a weighting engine assigns weight values to the preferred images such that the image placed in the first position receives the highest weight, with subsequent positions receiving progressively lower weights.

[0013] The system may further comprise an adaptive evolution agent module, wherein an Al model tracks, analyses, and evolves with individual user preferences over extended periods, wherein the adaptive evolution agent module dynamically refines default suggestions, weighting biases, and refinement pathways by continuously learning from user interactions within the preference capture module, the weighting engine, and the refinement module.

[0014] The present invention aims to improve the productivity of creatives in the art and design industry by drastically cutting down the time they spend searching for the perfect stock photos for their projects, such as blog posts, content marketing, and PowerPoint presentations. It also aims to eliminate the need for learning complex prompt syntax or techniques through abstraction, simplifying the user interaction by removing the necessity for long, detailed word prompts. This offers simplicity and ease of use by being a user- friendly alternative, steering away from complex products currently available that require highly complex prompts to be conceived and painstakingly typed in by the user. These complex products sometimes require an expert-level understanding of visual elements (e.g., composition, lighting, perspective, style). Users must correctly select from a range of pre-defined, static options, parameters, and settings presented in nested drop-down lists. This process, which lacks flexibility, often generates synthetic images of marginal aesthetic quality and limited utility.

[0015] The present invention provides functionality to iteratively refine image searches based on progressively received single-action user inputs, ensuring that each iteration brings the search results closer to the user’s visual idea. Semantic analysis is performed to interpret the user's inputs, and is synergistically integrated with an image synthesis deep learning model. This combination generates photorealistic images that progressively align more closely with the user’s visual idea with each iteration.

[0016] Further aspects, advantages, and features of embodiments of the invention will be apparent to persons skilled in the relevant arts from the following description of various embodiments. It will be appreciated, however, that the invention is not limited to the embodiments described, which are provided in order to illustrate the principles of the invention as defined in the foregoing statements and in the appended claims, and to assist skilled persons in putting these principles into practical effect.BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Embodiments of the invention will now be described with reference to the accompanying drawings, in which like reference numerals indicate like features, and wherein:Figure l is a process flow diagram of a method for generating synthetic images in accordance with an embodiment of the present invention.Figure 2 is a search user interface provided by a search query component of the system of Figure 11.Figure 3 is a search results gallery provided by an image discovery and preference gathering (IDPG) component of the system of Figure 2.Figure 4 is the search results gallery where two image search result thumbnails are user- selected as “liked” via the heart icon indication.Figure 5 is a similar image gallery indicating that a similar image locator function of an orchestrator component of the system is searching for visually similar images to a selected image (left side).Figure 6 is the similar image gallery showing visually similar images to the selected image (left side) found by the similar image locator function.Figure 7 is the similar image gallery showing several images that are user-selected as “liked” via the heart icon indication.Figure 8 is the similar image gallery showing the user has clicked the generate button to create photorealistic synthetic images based on the liked images and user’s search query.Figure 9 is a synthetic image gallery provided by an image finalisation component of the system displaying the photorealistic synthetic images.Figure 10 is the refined image gallery provided by an interactive refinement component of the system displaying refined photorealistic synthetic images.Figure 11 is a system diagram of a system for generating synthetic images in accordance with an embodiment of the present invention.Figure 12 is a model diagram of a diffusion model used in the system of Figure 2.Figure 13 is a model diagram of a language model used in the system of Figure 2.Figure 14 is a model diagram of a multi-modal vision and language model used in the system of Figure 2.DETAILED DESCRIPTION OF EMBODIMENTS

[0018] Fig. 11 is a block diagram illustrating a system embodying the present invention. A public communications network, such as the Internet 120, is used for messaging between a secure system 200, one or more user endpoint devices 130. Generally speaking, the endpoint devices 130 may be any suitable computing, communications and / or processing appliances having the ability to communicate via the Internet 120, for example using web browser software and / or other connected applications.

[0019] In this specification, terms such as ‘processor’, ‘computer’, and so forth, unless otherwise required by the context, should be understood as referring to a range of possible implementations of devices, apparatus and systems comprising a combination of hardware and software. This includes single-processor and multi-processor devices and apparatus, including portable devices, desktop computers, and various types of server systems, including cooperating hardware and software platforms that may be co-located or distributed. Hardware may include conventional personal computer architectures, or other general-purpose hardware platforms. Software may include commercially available operating system software in combination with various application and service programs. Alternatively, computing or processing platforms may comprise custom hardware and / or software architectures. For enhanced scalability, computing and processing systems may comprise cloud computing platforms, enabling physical hardware resources to be allocated dynamically in response to service demands. While all of these variations fall within the scope of the present invention, for ease of explanation and understanding the exemplary embodiments described herein are based upon single-processor general- purpose computing platforms, commonly available operating system platforms, and / or widely available consumer products, such as desktop PCs, laptop PCs, smartphones, and so forth.

[0020] In particular, the term ‘processing unit’ is used in this specification (including the claims) to refer to any suitable combination of hardware and software configured to perform a particular defined task, such as generating and transmitting authentication data, receiving and processing authentication data, or receiving and validating authenticationdata. Such a processing unit may comprise an executable code module executing at a single location on a single processing device, or may comprise cooperating executable code modules executing in multiple locations and / or on multiple processing devices. For example, in some embodiments of the invention, processing may be performed entirely by code executing on a server 110, while in other embodiments corresponding processing may be performed cooperatively by code modules executing on the secure system and server 110. For example, embodiments of the invention may use application programming interface (API) code modules 140, installed at the system 200, or at another third-party system, configured to operate cooperatively with code modules executing on the server 110 in order to provide the secure system with authentication services.

[0021] Software components implementing features of the invention may be developed in any suitable programming language or environment familiar to those skilled in software engineering. Examples include C, TypeScript, and Python, though other languages may be used based on system.

[0022] In the exemplary system, each endpoint device 130 includes a processor, a communications interface, one or more user VO interfaces, and local storage comprising volatile (RAM) and non-volatile memory (ROM, flash memory). The local storage contains program instructions and transient data for device operation. The endpoint device 130 storage contains program instructions and data necessary for normal operation, including operating system files (e.g., Windows, Android, iOS, MacOS) and other application software. It also stores instructions executed by the processor to perform operations related to an embodiment of the invention. As shown in Fig. 11, the server 110 comprises a processor 111 interfaced with non-volatile storage 113 (e.g., hard disk drive, NVMe M.2 drive) and volatile memory (RAM), which contains program instructions and transient data for server operation. The storage 113 holds operating system files, application software, and program instructions that, when executed by processor 111, enable the server to perform operations related to an embodiment of the invention. During operation, instructions and data are transferred from storage to volatile memory for execution. The processor I l l is operably associated with a communications interface, facilitating network access for data transmission. During use, the volatile memory holdsprogram instructions transferred from storage 113, executing processing tasks that embody features of the invention.

[0023] Referring to Figs. 1 and 11, the system 200 comprises a series of interconnected frontend components 201, each serving a distinct function within an AIP generation workflow that aligns with a user journey starting with receiving user intent to final image refinement (see Fig. 10). The frontend components 201 include the HomePage, orchestrator 202, search query 203, image discovery and preference gathering (IDPG) 204, interactive refinement 205, image finalisation 206, and image display card (IDC) 207.

[0024] The HomePage component (not shown) renders the initial page for users entering the system 200 (in the form of a web application), which may be in the form of a Single- Page Application (SPA). Navigation from the HomePage component to the orchestrator component 202 marks the beginning of user interaction with the application's core features. It is built using a functional programming approach, a modem web framework for enhanced client-side navigation and image handling, and a styling library for dynamic layout design. A navigation hook programmatically manages user navigation within the application. The HomePage component is structured to provide a user-friendly navigation interface, featuring a layout composed of container elements arranged hierarchically.

[0025] The orchestrator component 202 provides the web application's operational logic and manages state transitions across various stages and coordinates the AIP generation workflow. The orchestrator component 202 includes image interaction modules allowing users to like or dislike images, thus influencing the selection state of image search results (Sis). It orchestrates a multi-step AIP generation process within a web application framework. It is responsible for initiating the image search process via the search query component 203, displaying the Sis using the IDPG component 204, and managing user engagement with Al-generated images / synthetic photos (AIPs) using an interactive refinement component 205 (see Fig. 9). The orchestrator component 201 facilitates user navigation through the sequential processes of image search, gathering user feedback via their image results selection, refinement based on user preferences and user feedback, andinteraction with an iterative probabilistic generation model, (for example, a diffusion model) 252 for generating AIPs. A request management mechanism 208 handles requests and a state management mechanism 209 handles state management. The process flow is supported by various components designated for distinct phases of the user intent gathering and analysis and AIP generation cycle, including the search query 203, IDPG 204, interactive refinement 205, and image finalisation 206 components.

[0026] State management within the orchestrator component 202 is multi-faceted, tracking the user’s current state and storing user preferences for liked and disliked Sis. States are managed related to image search activities, the discovery of similar images, and initiating the generation of AIPs.

[0027] The image search initiator function 210 initiates image search requests. A similar image locator function 211, synthetic image generation trigger function 212 and the image refinement handler function 213 identify image similarities and facilitate the generation and refinement of AIPs.

[0028] The rendering logic of the orchestrator component 202 is dynamic and uses conditional rendering to display various pages aligned with the process's current stage, and is responsive to user actions and backend API responses.

[0029] Referring to Fig. 2, the search query component 203 is the initial step enabling users to input search parameters to find images corresponding to their preferences. The search query component 203 has form management over form state, input validation, and submission. Inputs from users are forwarded to the orchestrator component 202, which then orchestrates the transition to the IDPG component 204 for communicating user intent. In one example, the search query component 203 is a functional React component to provide image search within a web application context. A form management library handles form operations, including submission, input validation, and validation procedures, enabling users to input keywords or phrases (USPs) in the search input field 221 for retrieving pre-existing images. In the depicted example, the USP is “roses for valentines day”.

[0030] An image search is activated upon form submission, sequentially invokes the workflow advancement function to transition through the workflow stages, subsequently executing the image search initiator function 210 to relay search data for subsequent processing. The user interface of the search query component 203 is structured around a centrally positioned form 220. The search input field 221 is visually prominent. Below this, a "search" button 222 is pressed to submit search queries.

[0031] Referring to Figs. 3 and 4, the IDPG component 204 is configured to display Sis that align with the user's search query, integrating functionalities for users to express their preferences by liking 230, disliking 231, finding 232 similar images 235 to the Sis, and initiating generation 233 of AIPs based on user selected image search results (Pls) (see Fig. 8). The orchestrator 202 receives, as an input, user interactions to proceed to the interactive refinement component 205 for generating AIPs (see Fig. 9). The IDPG component 204, in one example, is a React functional component, to display Sis within the web app 200 and incorporates a responsive masonry layout to accommodate the dynamic display of images relative to screen size and provides the framework for a flexible and responsive image layout that responsively adjusts the number of columns according to screen width, optimising image result display. Each SI is displayed in thumbnail form that is encapsulated in an IDC component 207. The IDPG component 204 also incorporates form handling, responsive design, and user feedback mechanisms, in order to provide an intuitive interface for image discovery, and learning, understanding and interpreting user intent and their preferences based on their selection of Sis.

[0032] The system 200 includes an input management hook for search input validation, along with a state control hook and an effect handling hook to manage state transitions and ensure UI responsiveness to loading and error signals. The IDPG component 204 comprises instructional text for guiding user interactions with Sis, conditional rendering to display Pls alongside similar images 235 (Fig. 6), or all Sis when no PI is available. Overlay components communicate loading states or errors to the user. Event handlers include a search initiation function to process user-inputted search criteria and a similar image locator function to designate a PI as a positive indication, retrieving similar images 235 based on the Pls.

[0033] Referring to Fig. 10, upon displaying the AIPs that were influenced by user preferences identified in the IDPG component 204, the interactive refinement component 205 invites users to refine these outcomes further through textual inputs entered into a text input field. This allows the AIPs to be progressively modified and triggers further AIP generation cycles. In one example, the interactive refinement component 205 is a React functional component and displays the AIPs (see Fig. 9). AIPs can be edited or refined through the textual inputs 240 in the form of brief editing instructions from the user (see Fig. 10), expressing user preferences via like or dislike buttons against certain AIPs, and / or instigating subsequent rounds of AIP generation based on certain AIPs. In the depicted example, the editing instruction is “beach background”.

[0034] Form management and validation ensure efficient handling of form inputs and validations. Synthetic image interaction parameters 214 manage user actions such as search submissions, image refinement, and preferences (liking / disliking images). These parameters enable users to specify refinements or edits to AIPs.

[0035] The interactive refinement component 205 incrementally adjusts and enhances AIPs to correspond with the user’s conceptual visualisation more closely. This iterative process involves a series of feedback loops wherein the user can specify adjustments or modifications, either through direct manipulation tools or descriptive inputs, which the system 200 interprets and applies to the AIP. The system 200 finely adjusts successive AIPs through a series of iterations, based on user input. This approach is dynamic and interactive ensuring that the AIPs more accurately reflect the user's intended visual concept. It gives users a high level of control for generating AIPs and effectively "steers" the generation towards producing AIPs that closely match the one envisioned in their mind. Integrating user feedback directly into the AIP generation process provides a personalised and flexible means for creative expression, allowing users to shape AIPs until they achieve a desired visual outcome that aligns with their aesthetic vision. Users are not merely passive recipients of AIPs but active participants in shaping the final outcome. This iterative approach uses the combination of earlier described components for quick and convenient user input (to express the user intent clearly without requiring a high level of visual vocabulary) at each stage, ensures that the AIPs progressively alignwith the user's specific visual ideas, thus enhancing the overall customisation and relevance of the synthetic images.

[0036] The UI structure of the interactive refinement component 205 includes forms for submitting briefly written, refinement text-based instructions, for users to further personalise the AIPs in a progressive step-wise manner. There is also a responsive layout that displays AIPs, each encapsulated within the IDC component 207, enabling individual user actions such as liking or disliking on each image. Event handlers within the component 205 manage search submissions and capturing refinement inputs from users, to activate the corresponding processes.

[0037] Referring to Fig. 10, the image finalisation component 206 displays the AIPs, with an input field 240 for receiving user input on iterative refinement, final touches or additional edits to the AIPs. From the user’s perspective, this fosters a sense of continuous improvement and personalisation. Compared to available image generators, the system 200 generates AIPs with high compositional coherence, strong alignment with the user’s intent and aesthetic quality without requiring the user to think of and write complex textual prompts and avoid excessive trial and error rounds of generating AIPs, which are the primary sources of delay and user frustration.

[0038] The IDC component 207 is designed for displaying AIPs and is used by the IDPG component 204, the interactive refinement component 205, and the image finalisation component 206. The IDC component 207 displays AIPs within a card format, and embeds interactive elements for user engagement to like, dislike, and potentially download or generate more AIPs (based on additional like / dislike of Sis). These interactive elements enable users to express and communicate their preferences corresponding to their mental image of the visual concept, without requiring detailed textual input. The IDC component 207 enables user feedback to directly influence curation of Sis and generation of AIPs.

[0039] Key properties under image interaction properties include an image identifier for the SI intended for display, callback functions activated upon the liking or disliking of a SI by a user and a list object cataloguing arrays of liked and disliked image identifiers.There are event handling mechanisms include direct linkage of the like and dislike buttons to their respective callback functions.

[0040] An image identifier renders a SI and construct the source URL based on configuration settings. Also, user interaction occurs by pressing like and dislike buttons, whose appearances are contingent on the current preference state of the SI. These interactions are orchestrated via callback functions. There is also conditional button visibility to ensure action buttons are prominently visible only upon mouse hover, maintaining a clean look while preserving functionality access. Also, the user selection state for each SI is dynamically mirrored to influence the visibility and design of the like and dislike buttons.

[0041] Regarding data flow and user interactions, the web app 200 commences at the HomePage component, directing users towards the orchestrator component 201 through interactive engagement. User inputs, encompassing search terms, likes / dislikes, and userbased refinement instructions, are present in the frontend components 201, influencing the procedural logic and flow of the AIP generation process. The orchestrator component 201 harmonises state management and user selection dissemination, ensuring a cohesive experience across the workflow. Callback functionalities embedded within parent components are transmitted downstream, empowering user actions to modulate the application's state and trigger necessary API interactions or navigational adjustments. Integrating external APIs for image searching, generation, and similarity assessment allows for easy extensibility of discrete functionality of the web app 200.

[0042] The user journey involves several phases: the exploration phase on the search page (see Fig. 2), where users input their USPs leading to the execution of image search APIs that fetch Sis matching the criteria; the guidance phase (see Figs. 3 to 8), where users express their preferences through interactive icons under each thumbnail, guiding the diffusion model 252 for generating AIPs, with the option to explore further similar images 235 or input new USPs for additional Sis; and the creation phase (see Figs. 9 and 10), where users, upon clicking "Generate," generate and view the AIPs, with their actionspost-viewing indicating satisfaction levels and potential desires for refinement through additional keywords or commands for more finely nuanced AIP variations.

[0043] Referring to Figs. 3 to 10, users may indicate their like or dislike of Sis and AIPs through the heart or X icons. This causes their classification to change to Pls. This feedback refines search results and / or guides AIP generation accuracy. A backend endpoint tracks these user interactions, storing URLs in "like" or "dislike" lists, which then inform the construction of input for a language model 253 to blend image captions generated for Sis, aiming for accuracy by incorporating semantic meaning in the form of words about Pls while excluding those from disliked ones.

[0044] A history of interactions for progression and accuracy in Sis, Pls and AIPs may be maintained. Managing data in session-specific lists for positively and negatively marked images to improve search contextuality and AIP accuracy reflecting real user intent.

[0045] Referring to Fig. 13, the language model 253 is trained on a text corpus to learn the statistical structure of visual vocabulary and image descriptions, enabling it to predict the next word in a sequence or generate coherent text. The model 253 comprises an embedding layer 531 converting words or tokens of an input sequence 530 into vectors. Positional encodings 532 are added to these vectors to maintain word order, to enable the model 253 to understand the sequence, enhanced by an attention mechanism 533 evaluating word importance based on their context within the input sequence 530. A feedforward network 535 introduces non-linear properties. Add & norm layers 534 incorporate residual connections that are batch normalised to stabilise the learning process. The output layer 536 of the model 253 translates the processed features into a sequence of tokens as a prediction, using softmax for probability distribution over all possible next tokens. Training adjusts neuron weights via backpropagation based on the difference between the model's predictions and the actual outcomes. Inference involves taking a sequence 530 of tokens as input, processing it through its layers, and outputting a sequence of tokens as a prediction. This process involves calculating probabilities for each possible next token and selecting the most likely one, using beam search, greedy search, sampling, top-k sampling, top-p sampling, temperature scaling, Gumbel-Softmaxsampling, penalising repeated N-grams, to maintain a set of the most probable sequences at each step.

[0046] Merging multiple image elements

[0047] The system 200 allows users to manually edit AIPs by placing, resizing, and rotating a user-provided or cropped image over a section of an AIP or another image, enabling visual-based composition without the complexity of traditional graphic editing software. A screen-snipping tool provides a user interface for selecting image segments using free-form, rectangular, or lasso selection, with operating system APIs (e.g., Win32 API) capturing the selected region. Image segmentation and cropping algorithms extract the segment, and users can adjust size, orientation, and apply basic edits before integrating it into an AIP. Blending algorithms ensure seamless integration by matching colour, texture, and lighting between segments and the target image. Techniques include gradient blending to smooth transitions, Poisson image editing for adjusting illumination and colour tones, and feature matching with warping to maintain consistent geometry and perspective. GANs and deep learning models synthesise transitional areas, while Fourier Transform and frequency domain analysis align textures and details. Colour correction algorithms match segment histograms to ensure consistent saturation, brightness, and contrast, and edge detection with smoothing filters mitigates harsh segment edges using feathering techniques. A layer management machine learning model predicts and organises segments in a layered structure, allowing control over composition depth, opacity, and blending modes. A real-time preview function provides instant feedback for refinement. Users can upload images or use the web application’s search function to select from a photo library. An image cropping component allows cropping via bounding box or freehand drawing, with the cropped region available for pasting onto an AIP. Users can resize, move, and rotate the pasted region through mouse interactions, while an "Al merge" button activates automatic blending, ensuring a natural and cohesive composite image. This visual-based editing approach eliminates reliance on text-based prompts, streamlining Al-generated image manipulation.

[0048] Semantic image search

[0049] The system 200 uses a semantic image search coupled with user-driven selection to ascertain user intent and the contextual meaning of their chosen images. Instead of textual queries, users select images (Pls) that align with their desired outcome. Through multiple selection rounds or iterations, the system 200 refines its understanding of user preferences by analysing positive (heart icon clicks) and negative (X icon clicks) feedback displayed when the user hovers over a thumbnail version of an SI. This process continually narrows the searching phase to better align with user preferences, thereby enhancing the relevance of Sis and influencing the generation of AIPs. It incrementally narrows the exploration area within the high-dimensional vector space. Each iteration of user feedback effectively guides the focus on more relevant regions of this space, optimising the search and generation functions. This refinement process improves identifying or generating images that closely match the user's desired criteria.

[0050] In one example, the semantic image search processes the USP, potentially with exclusions, targets an image corpus. A function processes the user query to identify and separate positive and negative search criteria, computes the semantic similarity between the query and the Pls, and returns a list of images that best match the query. The query is divided into positive and negative segments using a delimiter, enabling users to articulate their preferences and dislikes regarding their desired visual content. In processing the positive segment of the query, the function checks for references to specific images within the image corpus, identified by their unique IDs and, optionally, supplemented with additional textual descriptions. For such referenced images, their embeddings are retrieved (step 180) from a vector database, and should there be any supplementary textual descriptions, embeddings for these are computed (step 181). These embeddings are then aggregated through a concatenation process to generate an overall holistic vector representation that encapsulates the query's positive attributes.

[0051] To ascertain the semantic relevance between the query and the images within the image corpus, the method performs a dot product calculation of the positive query's embeddings against those of all images in the corpus. This computation serves as a measure of semantic similarity, providing a basis for the subsequent normalisation of similarity scores. Normalisation ensures consistency in comparison across varied queries.In instances where the query includes negative segments, the method proceeds to compute embeddings for these criteria as well. Image similarity may be computing using a TensorFlow image embedder. The similarity scores are then adjusted to penalise images aligning with these negative aspects, further refining the search results.

[0052] The Sis are sorted for display based on their adjusted similarity scores and semantic relevance, selecting the top matches as defined by the user-specified number of results. The method retrieves its path from the dataset for each top-matching image, complementing these details with the source information from the image corpus. The output is the return of a list of tuples. Each tuple encompasses the image path, and index, collectively representing the images that most closely align with the user's query.

[0053] Referring to Fig. 1, the system 200 integrates a series of functionalities to guide users from initial input to the generation and refinement of AIPs. Users begin by entering their USP (step 101), for example, "dog with a bone," which articulates their desired photo outcomes. This search is initiated by clicking (step 102) the "Search" button or pressing enter, prompting calls to image search APIs such as Google Images, Yandex, and Pinterest (step 103), which then return URLs of reference images (Sis) (step 104). The Sis are displayed (step 105), where users can express their preferences directly on the thumbnails through like (heart icon) or dislike (X icon) buttons (step 106).

[0054] Hovering over a thumbnail of an SI displays some visual cues: an X or a heart icon. Clicking the heart icon classifies the SI as a PI, where Pls are stored in a list (step 107). It also causes the remaining Sis to be reordered (step 182) within the display to prioritise the display order of Sis based on visual similarity or semantic similarity to the PI, using a multi-modal vision and language model 254 such as Contrastive Language- Image Pre-training (CLIP) model to assess similarity and relevance to Pls (step 99), improving the relevance of future Sis retrieved and order they are shown to the user where the top left position in the masonry layout is the most relevant and the bottom right position is the least relevant. Further exploration is encouraged as clicking on a SI or related images button (step 108) triggers the display of images similar 235 to the SI (see Figs. 5 and 6).

[0055] Referring to Fig. 8, upon clicking the "Generate" button (step 109), the CLIP model 254 or caption APIs generate detailed descriptions for Pls (step 170) and disliked Sis for exclusion, which, along with any user comments provided (step 171), form an enriched search prompt (ESP) (step 172). The ESP is then summarised into a Summarised Search Prompt (SSP) (step 173) and guides the generation (step 174) of AIPs that are closely aligned with the user's intentions. Users may further refine these AIPs by submitting additional instructions (step 171) (see Fig. 10) and repeating the summarisation and generation process (steps 172 to 174) for progressively more personalised results. Throughout this process, user interactions are carefully tracked, ensuring a progression towards generating AIPs that accurately reflect user preferences.

[0056] Referring to Fig. 14, the CLIP model 254 has a vision encoder 551 and a language encoder 552, which are trained simultaneously on a dataset of text and image pairs. The encoders 551, 552 transform the input text / image data 553, 554 into a vector space. Image pre-processing 555 standardises the input images 554. In the vision encoder 551, the images are divided into image patches 556. Patch embedding 557 converts image into a format suitable for processing by the model 254, effectively translating raw visual information into a high-dimensional vector space. The embedded patches are passed through ViT B / 32 Transformer layers 558 to identify and encode the complex relationships and patterns within the visual data. The output of these Transformer layers is a visual feature vector 559 which is a representation of the image's finer details and overall structure. In the language encoder 552, the input text is tokenised 561 and word embeddings 562 are generated for each token. The embeddings are passed through Transformer layers 563 that outputs a textual feature vector 564 which is a representation of the input text that encapsulates its semantic content. Next, two feature projections 560, 565 are used: one maps the visual feature vector 559 into the multimodal embedding space 570, and another separately maps the textual feature vector 564 into the same multimodal embedding space 570. This dual mapping process ensures that both feature vectors 559, 564 are coherently aligned within the multimodal embedding space 570. Next, the model 254 performs a similarity calculation 571 within this space 570 which quantifies the degree of alignment or similarity between the input image 554 and text 553, where during training, the model 254 learns to associate images with text by maximising the similarity between correctly matched pairs and minimising it for mispatched pairs.Based on the calculated similarities, a classification head or output layer 572 computes similarities between images and text.

[0057] In some embodiments, the system 200 can retrieve Sis within two seconds and generate AIPs in lower resolution within ten seconds. In some instances, AIPs can be generated in less than one second.

[0058] Backend server

[0059] At the backend 250, the architecture of the backend server 251 may be written in Python code and supports the web app 200. This server 251 uses a high-performance web framework for building RESTful APIs and enables efficient handling of HTTP requests emanating from associated frontend components 201. In one example, the backend server's operations may interface with various external API services 140 to conduct image searches, generate image captions, and facilitate generation of AIPs, subsequently relaying processed results back to the frontend 201 for user interaction. External API services 140 may include image search engines, similar image retrieval, image caption generation and language models. However, internal, and locally hosted microservices are also envisaged in addition to or as a replacement for external API services interfaces. This may offer the benefit of increased data security and reduced latency.

[0060] The initial setup of the backend server 251 incorporates cross-origin resource sharing (CORS) middleware which allows the web app 200, which is a web server gateway interface (WSGI) app, to serve CORS headers for multiple configured domains and permits cross-origin requests. This allows communication between the frontend 201 to communicate with the backend server 251, particularly when they are hosted at remote sites.

[0061] The backend server 251 also has functions for image caption generation and / or generating detailed image descriptions to derive textual descriptions from visual content, condensing image descriptions into succinct phrases, are used and iteratively refine image descriptions based on user feedback, for example, through their selections or additionaltext input. This approach uses natural language processing to synthesise and interpret image characteristics in line with user preferences.

[0062] The backend server 251 links the web app's frontend 201 with the capabilities of microservices, external image search and Al models. The backend server 251 handles requests for generating AIPs and refining Sis with a custom client designed for interaction with a diffusion model (for example, a latent diffusion model) 252. The server 251 has a modular design and can integrate with multiple iterative probabilistic generation models (one at a time, or in parallel) for generating AIPs or to replace as necessary models with inferior performance over time.

[0063] The backend server 251 defines various endpoints to manage operations such as health checks, image search to retrieve Sis, processing similar images, generating AIPs and refining AIPs, checking the generation status, and retrieving or downloading AIPs. The structured interaction between the frontend 201 and these backend endpoints involves the transmission of requests accompanied by parameters like: search queries and user preferences, followed by the backend's processing of these requests through external API calls or internal logic.

[0064] The system 200 eliminates the need for skills in illustration or advanced photo editing software. A hybrid approach of combining image search and AIP generation is embodied in a workflow that starts with selecting pre-existing images (Pls) as inspiration to narrow the exploration area within the high-dimensional vector space. The diffusion model 252 then generates synthetic images (AIPs) based on these Pls, optionally incorporating user-specified modifications.

[0065] Image synthesis deep learning model

[0066] Referring to Fig. 12, in a first embodiment, the image synthesis deep learning model may be a latent diffusion model 252. An input image 580 is encoded 581 into a latent space representation 582. This encoding step 581 reduces the high-dimensional input image 580 into a compact, lower-dimensional latent vector. In the latent space, themodel 252 applies a diffusion process 583, gradually adding Gaussian noise to the latent representation over a series of steps, influenced by a predetermined noise schedule. This noise schedule outlines the increment of noise addition at each step, simulating a Markov chain process that incrementally moves the latent representation towards a state resembling pure noise 584.

[0067] The model 252 uses external conditioning data 587, including semantic maps or textual descriptions 585, along with representations and images 586, to direct the AIP generation process. This conditioning data 587 is integrated into the model 252 to interact with the noisy latent 584 ensuring that the generated outputs remain aligned with the semantic or textual cues. During the conditioning step 588, the model 252 infuses the noisy latent representation with the conditioning data 587, steering the generation towards outputs that faithfully reflect the intended context or description.

[0068] Next, a denoising step 589 learns a denoising function to estimate the original latent representation from the noisy latent version 584. The model 252 iteratively refines this function, aiming to reverse the effects of the added noise and recover a denoised or reconstructed latent representation 590 that closely approximates the original latent space 582 before the diffusion process began. Once a satisfactory reconstructed latent representation 590 is achieved, the model 252 proceeds to the decoding step 591. A decoder translates the reconstructed latent representation 590 back into the highdimensional space of images. This decoding step 591 transforms the denoised latent representation into a visually coherent and contextually aligned output image 592.

[0069] In a second embodiment, the image synthesis deep learning model may be a latent diffusion model guided by a CLIP model. The AIP generation process starts with an initial state of stochastic visual noise and evaluates and scores the progressive iterations of the image as it is incrementally enhanced in terms of clarity and coherence. These evaluations are based on the likelihood of the image conforming to a predefined textual prompt. In the initial phase, the visual noise bears minimal resemblance to the targeted image concept. As the process advances, the diffusion model methodically modifies the image, guided by CLIP'S feedback, to increasingly resemble the desired outcomeinfluenced by the prompt. This iterative enhancement is analogous to the meticulous sculpting of a figure from a block of marble, where the initial noise represents the raw material, and CLIP'S scoring directs the diffusion model on where to "carve" to achieve an image that aligns closely with the textual prompt at each successive step. The initial noise is a fundamental component of the AIP generation workflow as there is an iterative interaction between the CLIP model and the diffusion model in steering the image from an indistinct state towards a detailed representation that accurately reflects the given prompt.

[0070] In a third embodiment, the image synthesis deep learning model may have an architecture with a two-stage pipeline. Initially, initial latents are generated using a foundational diffusion model. These latents are then further refined using a high- resolution refinement model. The architecture incorporates a convolutional UNet framework (with about 2.6B parameters), which is augmented by self-attention mechanisms, enhanced upscaling layers, and cross-attention features tailored for text-to- image synthesis. A significant adaptation within this architecture is the strategic reallocation of computational resources towards lower-level features in the UNet, using a transformer-based approach. There is a heterogeneous distribution of transformer blocks across different levels of the UNet, optimising the model's efficiency and text conditioning through the use of a powerful pre-trained text encoder, OpenCLIP ViT-bigG, combined with CLIP ViT-L. This combination enables detailed text embedding and conditions the model on these embeddings to produce images. Micro-conditioning the UNet model on the original image resolution involves embedding the original dimensions of the input images as additional conditioning information. This technique enables the model to dynamically adjust to different image sizes and resolutions during both the training and inference phases. By incorporating the original resolution directly into the model's conditioning, it gains the ability to adjust its processing to the specific dimensions of each input image, thereby enhancing the flexibility of the model and improving the fidelity and detail of the AIPs. This approach enables generation at multiple input resolutions, optimising the model's performance across a broad spectrum of image sizes.

[0071] Further conditioning techniques, such as cropping parameters, mitigate common synthesis issues like cropped objects. By conditioning on crop coordinates during training, the model learns to generate images with intended cropping effects, giving users control over the synthesis process. Combined with size-conditioning, this technique offers a refined approach to AIP generation, catering to specific user requirements.

[0072] Multi-aspect training enables the model to handle images of various sizes and aspect ratios by partitioning data into different buckets based on aspect ratio. This approach, coupled with conditioning on target size, allows the model to generate images that better match real-world distribution and user preferences. The improved autoencoder within the system 200 enhances the reconstruction performance, providing better local, high-frequency details in AIPs. The diffusion model 252 is trained by pretraining on an internal dataset, followed by optimisation steps at different resolutions and the application of multi-aspect training. This multi-stage procedure enables the generation of high-quality images tailored to specific aspect ratios and resolutions.

[0073] GANs are unsuitable for generating the AIPs of the present invention compared to a diffusion model because a diffusion model uses a controlled, iterative process of adding and then removing noise, which enables more stable training and results in higher-quality images with greater detail and less susceptibility to mode collapse compared to the adversarial training approach of GANs that results in a limited variety of outputs when the generator finds a particular set of outputs that consistently fools the discriminator into believing they are real, leading the generator to converge prematurely to those outputs. The diffusion model 252 of the present invention has an adjustable noise reduction, incorporates conditional information at various phases of the generation phase, and an iterative refinement process by gradually refining noise.

[0074] Example use cases

[0075] The system 200 supports diverse applications across industries by generating high- fidelity AIPs based on textual descriptions, enhancing creativity, efficiency, and personalisation. In web and app development, the system can rapidly generate designmockups, streamlining early-stage visualisation and stakeholder approval without requiring extensive graphic design expertise. In the film and video game industries, the system assists with concept art generation, transforming descriptive inputs into character, setting, or scene visualisations, accelerating pre-production workflows. By enabling quick iterations and refinements, the system enhances customisation and personalisation, reducing production time while expanding creative possibilities.

[0076] The system 200 is designed for modular, scalable, and seamlessly integrated image synthesis model updates, ensuring adaptability to Al advancements without disrupting operations. A modular architecture assigns distinct components to specific tasks within the AIP generation workflow, allowing individual image synthesis models to be updated or replaced with minimal system-wide impact. An API layer abstracts model integration complexity, isolating changes and preserving application logic integrity. A configuration-driven approach enables dynamic interaction adjustments through a centralised configuration repository. Continuous Integration and Deployment (CI / CD) pipelines automate testing and deployment, ensuring compatibility and reducing downtime, while version control mechanisms allow for managing different model iterations and rollback capabilities to maintain stability. A user feedback loop collects responses to model outputs, informing further refinements.

[0077] Interactive Search Tool

[0078] The system 200 primarily relies on brief textual input for image searching and refinement but may incorporate voice commands through natural language processing (NLP) and speech recognition, allowing users to verbally articulate visual ideas and feedback. Additional input methods, including gesture recognition and augmented reality interfaces, may further enhance accessibility, particularly for users with disabilities, enabling a more dynamic and immersive creative process.

[0079] An Interactive Search Tool (1ST) provides gesture-based interactivity, allowing users to engage with images and videos using circling, highlighting, scribbling, or tapping. The 1ST integrates multiple search modalities, including image recognition, text,and voice inputs, for a fluid multi-modal search experience. It features focused search capabilities within photos or videos, identifying untagged items of interest, such as products, people, or objects, and enabling search for similar items within visually complex scenes.

[0080] Users can customise queries by directly drawing or highlighting within the 1ST, offering precise, user-driven search refinement. This interactive approach enhances search accuracy and personalisation, improving how users navigate and interact with digital content.

[0081] The system 200 uses Adversarial Diffusion Distillation (ADD) as a training algorithm, enabling fast sampling from large-scale pre-trained diffusion models with only 1 to 4 steps while maintaining high image quality. The one-step diffusion model is trained to generate images that fool a discriminator (similar to a GAN classifier) and match the multi-step diffusion model's output. Score Distillation Loss quantifies differences between the ADD-student model’s samples and the DM-teacher model’s predictions, using a squared L2 norm as the distance metric. The ADD teacher model refines diffused outputs from the ADD student model instead of raw outputs to mitigate out-of-distribution issues. The method applies exponential weighting, which reduces higher noise contributions, and score distillation sampling weighting, aligning distillation loss with the SDS (Score Distillation Sampling) objective. This approach enhances reconstruction target visualisation, supports successive denoising steps, and introduces a noise-free score distillation objective, extending the SDS framework.

[0082] The system 200 provides a responsive user interface supporting adaptable touchscreen interactions, including mouse clicks, taps, swipes, and pinches. The interface highlights hovered images by darkening them with a border, where a green border indicates selection. The "Enter" key functions as a "Search" button, while the "Esc" key acts as a "Clear" button, resetting the search input field.

[0083] The system 200 processes the Enriched Search Prompt (ESP) through a language model 253, generating a Summarised Search Prompt (SSP), which is then forwarded tothe diffusion model 252 for AIP generation. Alongside the SSP, additional parameters such as model specifics, image dimensions, sampler settings, Low-Rank Adaptations (LoRAs), and textual inversions are specified to enhance output quality. A GPU 112 enables parallel processing, generating high-quality photorealistic images with at least 512x512 resolution, preferably 1024x 1024 resolution.

[0084] The system 200 stores user interaction data in a PostgreSQL database, including search queries, clicks, image URLs, and timestamps, creating a record of user engagements. This data is used to fine-tune diffusion models 252 by mapping initial search queries to accepted AIPs, identified through downloads or other positive actions.

[0085] In another embodiment, the process may be simplified by removing the need for any text prompt altogether. Instead, the system 200 takes the Pls and uses them as direct inputs into the diffusion model 252. One advantage of this approach is that it reduces the cognitive load on the user as they don't need to come up with the perfect text prompt or learn or understand visual vocabulary. All they need to do is select the Sis they like. This method might be more intuitive for users who "know what's good when they can actually see it".

[0086] Context-aware buttons

[0087] In another embodiment, the system 200 includes a promptless interaction mode that integrates machine learning algorithms and UI design principles for an intuitive user experience. Image analysis algorithms dynamically adjust UI settings by detecting key elements in uploaded images. For example, when a landscape photo is identified using computer vision techniques, the system provides landscape enhancement tools such as colour correction sliders. An editable word string component, leveraging NLP and computer vision, generates a detailed textual description of the image, allowing users to refine Al interpretations without traditional prompts. The system includes interactive buttons and sliders for adjusting settings, themes, and filters, ensuring accessibility for users without Al expertise. A one-click generation mechanism powered by a backend Al model synthesises input parameters into a refined image. At the same time, reinforcementlearning algorithms adapt Al output based on user feedback, including likes, downloads, and interaction patterns. The promptless interaction mode incorporates Reinforcement Learning from Human Feedback (RLHF), dynamically learning from user responses, mouse movements, and eye tracking to improve predictive accuracy. Expert users can access predefined controls for technical parameters (e.g., ISO, exposure, focal length) and compositional elements (e.g., depth of field, aspect ratio, colour tone), with an "Advanced Options" section for fine-tuned adjustments while maintaining ease of use.

[0088] Smart modes

[0089] In one embodiment, the system 200 comprises a smart modes component 55 that automatically adjusts image settings based on keywords in the user’s query or captions from selected images, generating an augmented SSP as input to the diffusion model. A language model 56 analyses user-provided text and suggests image adjustments, such as aspect ratio, colour tone, and depth of field, to match the intended theme. For example, "portrait" prompts a shallow depth of field, while "landscape" adjusts for a wider aspect ratio and deep field depth.

[0090] The system 200 includes smart modes such as Macro Mode for close-up details and Night Mode for low-light conditions, each configuring exposure, focal length, and colour temperature to match the selected theme. The system 200 also supports specialised modes, including Architectural and Wildlife Photography, optimising settings for distinct photographic styles. A machine learning model 57 refines these adjustments by analysing historical data and user feedback, ensuring alignment with evolving trends and user preferences.

[0091] The system 200 applies algorithmic adjustments to generate a preliminary visual scene, configuring a virtual camera type and lens, adjusting lighting conditions, and setting the camera shot angle for the intended perspective. The diffusion model then processes the refined augmented SSP to generate AIPs that adhere to photographic best practices while capturing the user's vision. Smart modes streamline AIP generation, enabling users to express intent through natural language while the system 200 handlescomplex technical optimisations, making Al-generated photography accessible to users of all expertise levels.

[0092] Video generation

[0093] In another embodiment, the system 200 may include a video diffusion model capable of generating short-form photorealistic videos of up to one minute in length. The model initialises a noise-filled video sequence during inferencing and progressively refines it by iteratively denoising frames. It supports full video generation from scratch and video length augmentation for previously Al-generated sequences. To ensure subject continuity, the model processes multiple frames concurrently and applies a forwardlooking perspective to maintain consistency even when elements exit and re-enter the frame. Architecturally, the video diffusion model is built on a Transformer framework, leveraging patch-based representations of videos and images, similar to tokens in language models, to support scalability and training across varied durations, resolutions, and aspect ratios.

[0094] The model undergoes recaptioning-based training, where highly descriptive captions are generated for training videos, refining the model’s visual -textual comprehension and improving context-aware image / video synthesis. This involves analysing video datasets and autonomously generating detailed textual descriptions, capturing nuanced visual elements that might otherwise be overlooked. These recaptured descriptions are fed back into the training cycle, enhancing the model’s ability to interpret complex visual scenes and generate accurate representations during inference.

[0095] The video diffusion model enables the generation of complex video scenes featuring multiple characters, dynamic motions, and detailed backgrounds while ensuring spatial-temporal consistency. It processes textual prompts to position elements within realistic physical contexts, facilitating the creation of emotionally expressive characters and ensuring continuity in shot composition, character consistency, and visual style across multiple frames, thereby improving physical realism and spatial-temporal dynamics modelling.

[0096] Hierarchical Image Synthesis Pipeline

[0097] The Hierarchical Image Synthesis Pipeline (HISP) uses a tripartite model architecture (Stages A, B, and C) for hierarchical image compression and reconstruction, enabling high-quality image generation from a highly compressed latent space. Stage C functions as the latent generator, transforming user inputs into compact 24x24 latent representations, forming the foundation for decoding and reconstruction. Stages A and B serve as latent decoders, decompressing latents into high-resolution images with a high compression rate. Decoupling text-conditional generation (Stage C) from the decoding process (Stages A and B) enables isolated training or fine-tuning of Stage C, integrating ControlNet and LoRAs, reducing computational costs by approximately 16 times compared to similar models. While Stages A and B allow optional fine-tuning, training Stage C is prioritised for efficiency and cost-effectiveness, as fine-tuning Stages A and B yields marginal benefits. Stage C models include IB and 3.6B parameters, with the 3.6B model preferred for higher output quality and the IB model optimised for lower hardware requirements. Stage B models include 700M and 1.5B parameters, with the 1.5B model preferred for high-fidelity detail reconstruction. HISP’s modular architecture enables VRAM-efficient inference (-20GB), with smaller model variants available at the cost of output quality. HISP extends text-to-image generation to support image variations and image-to-image transformations by extracting CLIP embeddings from an input image and reintegrating them into the model. Image-to-image generation introduces noise to an existing image and generates variations from the noisy input. HISP also supports inpainting, outpainting, canny edge detection, and 2x super-resolution, enabling image upscaling and latent refinement in Stage C for enhanced detail and fidelity.

[0098] Al-assisted image editing

[0099] The Multi-Modal Language Model (MLLM) 279 is trained to edit an input image into a desired output based on brief textual instructions by generating clear directives that translate vague or imprecise user input into structured guidance for the diffusion model 252. A pre-trained summariser 280 condenses detailed editing commands into concise instructions to enhance processing efficiency. The MLLM 279 bridges linguistic input andvisual objectives by interpreting special visual tokens, which are transformed into visual guidance through an editing mechanism 281 that directs the diffusion model 252 to modify the image as per user intent. The image-editing diffusion model 252 operates in a compressed image space, applying denoising and refinement processes to achieve the desired outcome based on the provided visual guidance. The MLLM 279 is trained using input-goal-instruction triplets, generating summarised explanations that refine images toward the specified goal, with instruction loss measuring deviation from ideal guidance to improve the model’s ability to produce concise, comprehensive directives. These visual tokens function as latent visual imagination cues, guiding image refinement, while editing loss calculations further enhance accuracy through iterative training, ensuring the system precisely executes user instructions and aligns outputs with the envisioned outcome.

[0100] Mood-based image adjustment may be implemented via interface 282, where users select a predefined mood, triggering a mood Al model 283 to modify the image’s colour scheme, brightness, and contrast based on mapped image processing parameters. The mood Al model 283, trained on mood-labelled datasets, analyses the original image’s saturation, contrast, and brightness as a baseline and applies targeted modifications while preserving image integrity, preventing over-saturation or detail loss. Adjustments are applied adaptively across different image regions based on their initial visual characteristics, ensuring AIPs remain cohesive. The system may incorporate an optional feedback loop for post-adjustment analysis and fine-tuning, allowing manual user refinements for personalised transformations. The mood Al model improves recognition and replication of nuanced mood variations, enhancing the accuracy of mood-based adjustments and ensuring a more authentic and contextually aligned visual output.

[0101] Emotion-responsive filters 284 may integrate facial recognition technology to analyse facial expressions detected by a face detector algorithm 285 in images. Based on the recognised emotion, the system 200 performs a lookup operation and applies corresponding image filters and adjustments. For example, happy expressions may trigger brighter and more vibrant filters, while contemplative moods may result in subdued and softer tones. This dynamic adjustment process uses machine learning models trained onfacial expressions and associated emotional states, ensuring adaptive and contextually relevant modifications to the image.

[0102] A pan and zoom component 286 enhances image quality and detail retention at varying zoom levels using predictive algorithms, enabling users to closely examine or edit specific areas without loss of clarity. Super-resolution models 287, such as ESRGAN or SRGAN, upscale images beyond their original resolution by reconstructing fine details, while edge detection algorithms like Canny or Sobel enhance image edges through sharpening filters such as Unsharp Mask or Laplacian filters. The system 200 uses adaptive interpolation techniques, including bicubic interpolation and Lanczos resampling, to estimate and reconstruct missing pixel data while differentiating between textured and smooth regions to prevent artifacts. When users pan or zoom beyond image borders, a diffusion model generates additional content that blends seamlessly with the existing image, assisted by content-aware fill techniques that extend patterns, textures, and colours. Reinforcement learning models personalise image enhancements by adapting zoom-level adjustments based on user interaction history. The system 200 may further use depth estimation algorithms to generate depth maps for realistic 3D perspective adjustments, while generative Al models extrapolate existing content to predictively expand images, ensuring logically consistent scene extensions.

[0103] A content-aware fill component 287 uses pixel prediction algorithms to seamlessly fill blank or newly created spaces by matching the surrounding content’s colour, texture, and pattern. The system 200 uses convolutional neural networks (CNNs) to analyse adjacent pixels, extracting feature vectors that encode colour distributions, textural patterns, and structural elements. A content-fill diffusion model 288 processes these extracted features to generate matching pixels, ensuring high-fidelity synthesis. The model, trained on diverse image datasets, predicts and reconstructs complex textures and patterns. To ensure seamless integration, refinement algorithms apply edge-smoothing and gradient adjustments, eliminating visible transitions between the original and generated content. A validation module 289 assesses the fill’s consistency against the surrounding context using similarity metrics, iteratively refining the output until it meets predefined quality standards.

[0104] Interactive 3D perspective adjustments allow users to modify an image’s viewing angle by simulating a 3D perspective shift in real-time. Al models analyse the 2D image to infer depth, segment objects into distinct layers, and estimate each segment's depth relative to the original viewpoint. Upon user interaction, the system recalculates rotation, translation, and scaling transformations to adjust the perspective accordingly. CNNs identify and segment image content, while GANs generate visually coherent transformations. A dynamically optimised rendering engine processes Al outputs in real time, ensuring smooth, natural viewpoint transitions for an immersive user experience.

[0105] A smart focus tracking component 290 has a deep learning model to identify and enhance the main subjects in an image by dynamically adjusting focus and depth of field (DoF) during panning or zooming. A CNN 291, trained on annotated datasets, discerns focal points and computes optimal focus parameters, including focal length, aperture size, and focal distance, to maximise subject clarity. Upon user interaction, the system 200 analyses the image, applies a focus adjustment algorithm, and simulates optical blurring to create a bokeh effect, ensuring the main subject remains prominent. At the same time, the background or foreground is softly blurred. The Al model maintains natural aesthetics, preventing artificial distortions.

[0106] An auto-resize component 292 for social platform display adjusts images to meet platform-specific requirements, including resolution, aspect ratio, file size, and format. An adaptive resizing algorithm retrieves platform specifications from a dynamically updated database and analyses the image’s original properties before applying scaling, cropping, and padding to fit display guidelines. Bicubic interpolation preserves sharpness during scaling, while content-aware cropping ensures key elements remain intact. Padding is applied to meet size constraints without distorting content, optimising image appearance across different social media platforms.

[0107] An automatic story generation component 293 uses natural language processing (NLP) models to generate narratives from a text caption, enhancing the storytelling dimension of images. The system 200 preprocesses user inputs, including search queries, refinement instructions, and image captions, to extract key themes,characters, settings, and emotions. A Transformer-based NLP model 294 of component 293, trained on a large corpus of literature and storytelling data, generates contextually relevant stories by constructing a context vector that encodes the semantic meaning of the input text. The model 294 then iteratively predicts subsequent words, leveraging attention layers to maintain coherence by prioritising key input elements, ensuring that the generated narrative aligns with the user’s text and associated image selections.

[0108] Real-time feedback, augmentation and generation

[0109] The system 200 may include predictive text and auto-complete functionality powered by machine learning models trained on extensive visual language datasets, enabling accurate search query predictions based on initial inputs, user history, and contextual trends to enhance accuracy.[001 10] As users initiate searches, the system 200 may fetch and display search results in real-time, providing an immediate feedback loop that allows for instant refinement without requiring full query completion, reducing cognitive load, while a recommender Al model 295 analyses past interactions and user history to pre-select likely relevant images, minimising manual selection effort by presenting a refined subset of Sis for user adjustment.[001 1 1 ] The integration of the diffusion model 252 with the search and selection process enables generation of AIPs as soon as a user preference is indicated. By pre- loading and partially processing images in the background, the system 200 can significantly reduce wait times for the user. To continuously improve the system's accuracy in predicting user preferences, an incremental learning approach from user interactions may be implemented. This feedback loop would refine the recommender Al model's understanding of user preferences over time. High-performance compute, such as NVIDIA H100 GPUs, alongside an optimised backend architecture for parallel processing and load balancing, enable simultaneous user request handling with minimal latency, while efficient data caching for frequently searched terms and popular images furtherreduces response times, collectively delivering an almost real-time experience that enhances usability and efficiency.

[0112] Voice analysis and hand gestures[001 13] The system 200 may include voice input functionality for synthetic photo generation, enabling users to provide extended verbal descriptions without requiring complex visual terminology. A language model interprets spoken words received via a microphone and translates them into structured instructions for guiding the diffusion model 252 in producing images aligned with the user’s vision. The system 200 analyses voice intonation, pitch, pacing, and pauses using a speech emotion recognition model to extract storytelling nuances and emotional cues, generating diffusion model instructions that enhance the synthetic photo’s emotional depth and narrative significance. Additionally, the system 200 may capture voice input via video recording such as a camera or web cam, allowing users to supplement verbal descriptions with hand gestures or sketches, which are analysed and converted into image elements, integrating natural, expressive movements into the Al-driven synthesis process for greater clarity and creative alignment.

[0114] Enhanced non-linguistic user input[001 15] The system 200 may include a non-linguistic user input component 260 that allows users to provide visual cues beyond text, voice, or sketching by capturing photographs of themselves or others to guide the Al’s creative process. A "capture moment" button initiates a countdown timer, enabling users to pose before the system captures and displays the image for approval, with an option to retake it to ensure the expression and pose align with their intent. The captured image is processed in pixel data form, eliminating the need for textual or sketched articulation. The system 200 uses a facial expression recognition model 261 to analyse the captured expression and predict emotional states such as happiness, surprise, or contemplation, with this data influencing aspects of the AIPs, including mood, lighting, and subject expressions. For example, a detected smile may prompt the diffusion model 252 to generate vibrant, joyful imagesreflecting the user's emotional state. Additionally, the system 200 may interpret body language and gestures using an image classifier, analysing poses to infer intended actions in the AIPs. If a user extends an arm, the Al model may generate images where subjects replicate the gesture, such as reaching for an object or pointing, ensuring alignment between user intent and the generated visual output.[001 16] Emoji generation[001 17] In another embodiment, an emoji generation workflow is provided. Users select a set of initial pre-existing emojis, serving as seed inputs for the image synthesis deep learning model, by querying a database or library of emojis for user-selection. Next, the algorithm surfaces a range of suggested emojis through a recommendation engine that analyses semantic and visual relationships between emojis. Users can refine their selection by adding or removing emojis, facilitated through a drag-and-drop interface or simple selection tools within the application.[001 18] A visualisation mechanism displays the progression and impact of each emoji addition or subtraction on the final emoji generation, through a series of thumbnails showing incremental changes or an animated sequence. The system 200 incorporates a versioning control mechanism similar to those used in software development, allowing users to create branches at any point in the emoji selection timeline. This enables the exploration of different creative directions without losing the original sequence of selections.[001 19] For blending two branches, a merging tool within the system 200 allows users to combine elements from different branches, involving algorithmic blending where the system 200 suggests a coherent mix based on the properties of the selected emojis from both forks. If a user wishes to undo changes or revert to an earlier point in the generation process, the system 200 provides a rollback feature, implemented through an undo function or by allowing users to click on a previous point in the timeline to make it the current state.

[0120] After the iterative process of refinement and branching, the final emoji composition is processed by the model, which synthesises a coherent visual output based on the arranged emojis using a diffusion model to generate high-quality AIPs that represent the combined attributes of the selected emojis. The AIP is then presented to the user, with options for further refinement or approval before saving or sharing the AIP.

[0121] Constraint-driven visual / non-text refinement for image generation by a diffusion model

[0122] In another embodiment, the web application 200 is configured to enable users to access a reference image library via an interactive user interface. In some embodiments, the system 200 displays the Sis in a sequential browsing format, where users can indicate their preferences by selecting or rejecting individual Sis through a binary input mechanism. The user interface may enable users to navigate through Sis one at a time, progressing forward or backward based on their interactions. In alternative embodiments, the system 200 may present Sis in a horizontally or vertically scrollable carousel, wherein users can rapidly browse multiple Sis and express preferences through direct selection.

[0123] The system 200 comprises a plurality of interconnected modules 1001, 1002, 1003, 1004, 1005, 1006, each configured to facilitate the selection, refinement, and synthesis of AIPs based on user interactions. The image database 1001 stores Sis and associated metadata, including style classification, similarity scores, and quality ratings, providing input to the preference capture module 1002, which records user selections, behavioural data, and preference indications such as "like" or "dislike" inputs, along with selection speed and interaction patterns. The selection and reordering interface 1003 retrieves user-preferred Sis from the database, allowing users to arrange them via a drag- and-drop mechanism, where positioning influences Al-generated outputs. The updated ranking is transmitted to the weighting engine 1004, which assigns numerical weight values to selected Sis based on selection order, behavioural signals, historical preferences, image similarity metrics, and global trends. These weighted preferences are processed by the Al image generation module 1005, which creates AIPs using generative architecturessuch as latent diffusion models, GANs, or Transformer architecture with a self-attention mechanism, blending attributes of the Sis in accordance with the computed weight distribution. The AIPs are then presented in the refinement module 1006, where users can iteratively adjust selections, reorder Sis, modify weightings, or express new preferences, with updates feeding back into the weighting engine 1004 and Al image generation module 1005 to progressively refine the creation of successive AIPs.

[0124] The system 200 is configured to facilitate an iterative image refinement workflow, wherein user preferences are dynamically captured and applied to generate AI- synthesised outputs. At initialisation, the system 200 retrieves and displays an initial set of Sis from the image database 1001. The selection of the initial Sis may be based on a machine learning model trained to predict Sis aligning with demographic attributes or prior user interactions. The preference capture module 1002 records user actions, including "like" and "dislike" selections, and behavioural metrics, such as the time interval between presentation of Sis and user response. The system 200 allows the user to select a subset of Pls from those marked as "liked" (i.e. Pls) the selection and reordering interface 103 provides a ranking mechanism where Pls can be repositioned based on user preferences. The system 200 assigns weight values to the Pls, wherein the image placed in the first position is assigned the highest weight, with subsequent positions receiving progressively lower weights.

[0125] The weighting engine 1004 computes the final weight distribution for each selected SI based on a weighted function that may include the order of selection, the speed of the "like" action, and additional quality factors. The weighting calculation may be represented as:FlncdWelght(i) = BaseWeight(ot) x SpeedFactor(st) x QualityFactor^qi) where:• FinalWeight^i) represents the final computed weight assigned to image i.• BaseWeight(c>i) is the base weight derived from the order position o(of image i.• SpeedF actor (st) is a multiplier that adjusts weight based on the speed of the "like" action Sj.• QualityF actor (cq) accounts for the quality rating qLof image i, influencing its overall contribution.

[0126] The computed weights are transmitted to the Al image generation module 1005, which synthesises an AIP by blending the selected reference images (Pls) according to their respective weight distributions. The AIP reflects the dominant attributes of higher- weighted images while incorporating secondary influences from lower-weighted references.

[0127] The iterative refinement process allows the user to progressively align the AIP with their desired aesthetic or functional requirements without reliance on textual prompts. The system 200 continuously adapts to user interactions, ensuring that refinements are performed within the constraints of previously established preferences and selections.

[0128] This embodiment of the present invention for constraint-driven visual / non- text refinement for image generation eliminates the need for substantial text-based input, enabling a fully visual and behaviour-driven approach to Al image generation, making it highly accessible to users who are non-technical or use non-alphabetic languages. It enhances personalisation by incorporating implicit behavioural data such as selection speed, reorder actions, and historical preferences, allowing for adaptive weighting based on user behaviour, image quality, and trending styles. The iterative fine-tuning mechanism enables users to refine outputs intuitively through simple reorder adjustments, while the system’s scalable and extensible architecture allows seamless integration of additional preference signals or advanced Al models without complicating the user interface.

[0129] In yet another embodiment, the system 200 may incorporate custom weight adjustments via the selection and reordering interface 1003, allowing users to assign specific influence levels to individual reference images using a slider or intensity control rather than relying solely on order-based weighting. For example, the system 200 mayenable users to set Image A to 50% influence and Image B to 10%, dynamically modifying the weighting engine 1004's calculations. Additionally, a "Lock Influence" option may be provided, allowing users to freeze an image’s contribution to ensure its stylistic impact remains unchanged throughout the refinement process.

[0130] The system 200 may further include interactive image refinement mechanisms 1060 within the refinement module 1006, allowing users to specify areas of interest through pinch-to-zoom or crop functions. Users may select image regions that should be emphasised (e.g., focusing on a subject’s pose while ignoring the background), with the weighting engine 1004 adjusting the influence accordingly. Furthermore, the system 200 may enable sketch or annotate refinements, allowing users to draw rough edits, such as "increase size" or "adjust colour," which the Al image generation module 1005 interprets to modify the generated output.

[0131] The system 200 may further support multi-touch gesture controls 1070, allowing users to swipe left or right to increase or decrease an image’s influence instead of using drag-and-drop functionality. This interaction feeds directly into the selection and reordering interface 1003, which transmits updates to the weighting engine 1004 for realtime recalibration.

[0132] The system 200 may implement Al-based image retrieval through the image database 1001, automatically suggesting visually similar images based on user- selected Pls. If a user selects multiple dark, moody portraits, the preference capture module 1002 may analyse their characteristics and retrieve additional Sis with similar lighting, contrast, and composition, reducing manual search effort.

[0133] To enhance real-time user feedback, the system 200 may incorporate a preview heatmap 1080 within the refinement module 1006, visually representing how much influence each selected image contributes to the final Al-generated output. Additionally, a before-and-after toggle may allow users to compare the impact of their adjustments, ensuring transparency in the refinement process.

[0134] The system 200 may also create individual user preference profiles via the preference capture module 1002, tracking historical selections and refinement behaviours. Users may name and save preferred styles (e.g., “Dark Cinematic” or “Vintage Editorial”), enabling the weighting engine 1004 to auto-tune future generations based on past behaviours, optimising iteration speed and reducing manual adjustments.

[0135] In another embodiment, the system 200 may include a live Al influence mixer 1007 within the selection and reordering interface 1003, allowing real-time blended previews that update dynamically as users reorder images. A multi-layered latent diffusion preview engine 1090 may generate low-resolution visualisations, such as heatmaps or soft morphing effects, to illustrate how each selected image contributes to the final synthesis before committing changes. The system 200 may also allow users to hover over an image, triggering an overlay highlight that visually represents its influence on elements like colour, texture, and subject pose, enhancing user control and eliminating trial-and-error refinements.

[0136] The weighting engine 1004 may implement a three-tier weighting system 1030, segmenting reference image influence into Tier 1 : Dominant Element, determining core subjects (e.g., main figure or object structure); Tier 2: Stylistic Influence, affecting lighting, colour, and texture; and Tier 3 : Background Context, influencing environment, depth, and scene details. Users may assign tier labels to images, and the Al processes these separately, ensuring precise composition structuring and preventing blending inconsistencies.

[0137] A temporal adaptation model 1009 may be incorporated into the preference capture module 1002, allowing the system 200 to detect recurring user refinements over time. The system 200 may preconfigure smart weight suggestions at the start of each session based on past behaviour patterns, such as users consistently adjusting brightness or preferring softer textures. The Al model automatically aligns initial weightings to these learned preferences, reducing manual adjustments while progressively refining image generation predictability. By integrating automated weighting assignment, real-time adaptation, and visual weight projections, this module 1009 eliminates the need formanual slider adjustments, allowing Al model to learn, balance, and optimise influence distributions based entirely on user interaction patterns.

[0138] The system 200 may introduce an Al-driven morphing tool 1010 within the refinement module 1006, enabling users to morph between multiple styles and compositions using a progressive interpolation slider instead of generating discrete image versions. The system 200 may use latent-space interpolation to compute smooth transitions between reference styles, allowing users to gradually adjust Al weighting across stylistic influences (e.g., shifting between “Cyberpunk” and “Classic Oil Painting”) for continuous aesthetic refinement.

[0139] A gesture-based refinement system 1011 may be integrated into the selection and reordering interface 1003, leveraging Al heuristics to interpret micromovements such as tap-hold, swipe-tilt, or hover linger as implicit refinement cues. For example, a faster swipe may trigger the system 200 to increase style intensity, a slower drag may indicate subtle weighting adjustments, and a tap-hold may signal the system 200 to prioritise specific elements in synthesis. These real-time heuristics eliminate explicit UI interactions, creating a fluid, organic user experience.

[0140] The system 200 may incorporate a predictive completion engine 1012 within the Al image generation module 1005, allowing a deep learning model 1092 to suggest and autofill missing image components. If a selected reference lacks critical details (e.g., missing background or inconsistent lighting), the system 200 may use pretrained multimodal embeddings to compare user-selected references against large-scale image datasets and predict missing elements. For example, if a portrait-style reference lacks a background, the deep learning model 1092 may auto-suggest contextual environments based on common artistic conventions. This reduces manual workload, allowing users to focus on refining rather than rebuilding missing elements.

[0141] These improvements aim to eliminate the need for manual weighting, enabling the system 200 to become fully responsive to micro-interactions while facilitating continuous real-time image transformation. By incorporating predictivemodelling, gesture-based refinements, dynamic previews, and structured weighting hierarchies, the system 200 aligns Al-driven image generation with user intent and iterative design workflows.

[0142] In a further embodiment, the system 200 may include a smart intent- triggered preview module 1027, enabling dynamic, resource-efficient previews based on user intent. This module 1027 may leverage mouse tracking with predictive timing, progressively loading previews only when user behaviour indicates engagement, thereby minimising unnecessary compute usage while maintaining a responsive interface.

[0143] The selection and reordering interface 1003 may incorporate a mouse tracking with predictive timing module 1028, wherein the system 200 monitors user interaction patterns. If a user hovers over an image for a threshold duration, this module 1028 may generate a low-fidelity preview overlay. The threshold may be determined based on historical browsing behaviour stored in the preference capture module 1002. Conversely, if a user moves quickly past an image, the system 200 suppresses preview generation, conserving computational resources.

[0144] To enhance responsiveness, the system 200 may use a progressive detail loading component 1021, wherein previews initially appear as fuzzy heatmap overlays and progressively sharpen if the user maintains engagement. The Al image generation module 1005 may control the progressive refinement of preview detail, ensuring previews are computationally lightweight while providing relevant visual feedback.

[0145] The system 200 may further include an action-based refinement suggestion mechanism 1022, wherein the preference capture module 1002 detects user hesitation between two images. If the system 200 determines the user is indecisive, this mechanism 1022 may auto-generate a blended preview, allowing the user to assess an intermediate representation without committing to a selection.

[0146] By dynamically adjusting preview generation based on user behaviour, the smart intent-triggered preview module 1027 provides instantaneous feedback whileoptimising system efficiency, ensuring computational resources are only used when necessary.

[0147] The system 200 may include an adaptive morphing module 1008, enabling Al-driven progressive morphing between stylistic variations without requiring manual user adjustments. This module 1008 may infer optimal morphing points based on user history, selected image styles, and implicit behavioural tracking, dynamically generating refined outputs that reduce the need for iterative manual tweaks.

[0148] The preference capture module 1002 may store user history, tracking past refinements, such as preferred contrast levels, colour tones, or stylistic influences. The Al image generation module 1005 may use this data to generate Al-guided morphing paths, automatically suggesting two or three optimal transition paths between styles. These morph paths may be derived from historical user preferences, such as recurring selections of darker tones or high contrast, selected image characteristics, including dominant styles like cinematic or natural, and implicit behavioural tracking, which considers selection speed and refinement frequency.

[0149] The refinement module 1006 may use intent-based activation, ensuring morphing sliders appear only when a user lingers over two or more different outputs, signalling an intent to refine. This prevents unnecessary UI clutter and optimises workflow efficiency.

[0150] The system 200 may further include guided suggestions, wherein the adaptive morphing module 1008 places a "Best Fit" indicator on the morph variant most aligned with inferred user preferences. This ensures the Al surfaces the most contextually relevant refinement option before the user commits to a selection.

[0151] By integrating Al-driven morphing paths, intent-aware activation, and guided refinement suggestions, the adaptive morphing module 1008 provides a smooth, predictive evolution of images, reducing manual intervention while enhancing precision in stylistic transitions.

[0152] The system 200 may include an auto-weighting module 1019, enabling AI- driven dynamic weight distribution for selected reference images (i.e. Pls) without requiring manual adjustments. This module 1019 may analyse user selection patterns, reordering behaviour, and interaction speed to automatically assign proportional weights, ensuring seamless, behaviour-driven refinement.

[0153] To enable live weight adjustment via micro-interactions, the system 200 may detect prolonged user focus on a specific image within the selection and reordering interface 1003 and automatically increase its weight in the weighting engine 1004. Conversely, if a user rapidly moves an image downward in ranking, the system 200 may reduce its weighting influence, adapting to user intent in real time.

[0154] The system 200 may further include an Al-projected weights display within the refinement module 1006, providing a soft visualisation overlay that dynamically indicates each reference image’s contribution to the final synthesised output. This ensures transparency while maintaining an intuitive, non-intrusive user experience.

[0155] The system 200 may include a constrained vector space navigation module 1020, which refines Al-generated outputs by restricting generative exploration within a user-defined aesthetic space, ensuring refinements remain aligned with user preferences while eliminating off-mark variations. This module 1020 may dynamically adjust exploration boundaries based on historical user interactions, stylistic preferences, and implicit behavioural data. This module 1020 may also integrate trend prediction, wherein the system 200 monitors aesthetic preferences over multiple sessions and dynamically auto-adjusts the constrained space to align with evolving user behaviours. If a user consistently favours a specific colour scheme, composition style, or subject emphasis, system 200 may refine future generations accordingly, ensuring outputs remain progressively tailored to the user’s artistic trajectory. By restricting generative randomness, enabling controlled stylistic expansion, and adapting to evolving user preferences, the constrained vector space navigation module 1020 ensures that refinements remain laser-focused, allowing the Al model to self-correct search paths and generate outputs that consistently align with the user’s evolving creative style.

[0156] The Al image generation module 1005 may implement Best-Fit Generations, wherein the system 200 limits synthesis to variations within the constrained vector space, preventing the generation of irrelevant or stylistically inconsistent outputs. The preference capture module 1002 may track user-selected reference images, weighting adjustments, and refinement history, ensuring that new generations only explore stylistic deviations within user-established constraints. The preference capture module 1002 may store session notes and contextual hints, auto-tagging past refinements (e.g., “User last adjusted lighting intensity”) to provide intelligent resumption prompts, guiding users upon return. The preference capture module 1002 may store data on long-term aesthetic preferences, allowing the system 200 to automatically adjust default weightings and refinement biases without requiring explicit user input.

[0157] The system 200 may further include an "Explore Further" function within the refinement module 1006, allowing users to expand their constrained vector space selectively. When activated, the system 200 retrieves adjacent stylistic variations using a confidence scoring mechanism, wherein the preference capture module 1002 assigns similarity scores (e.g., 85% similarity to past liked images) to determine which alternative styles are most likely to align with user intent.

[0158] The system 200 may include an Al-driven image completion module 1021, enabling predictive enhancements to Al-generated images while maintaining user control over modifications. This module 1021 may analyse generated outputs for missing or incomplete elements, providing subtle, non-intrusive auto-fill suggestions where necessary, ensuring refinements remain user-directed rather than automatic.

[0159] The Al image generation module 1005 may implement subtle auto-fill suggestions, wherein the system 200 detects contextual gaps in an image, such as missing backgrounds or incomplete subject details, and generates soft-fill recommendations without immediately applying them. The system 200 may reference pre-trained multimodal embeddings to infer visually coherent fill options, ensuring that generated completions maintain stylistic consistency with the existing image.

[0160] The system 200 may further include a "Would You Like to Add?" query component within the refinement module 1006, wherein Al-generated suggestions for missing elements are surfaced only when the system 200 detects an incomplete visual context. Instead of applying auto-fills directly, the system 200 may allow users to accept, reject, or modify the Al-generated suggestions, maintaining user autonomy over refinements.

[0161] To enhance flexibility, the system 200 may provide a controlled expansion mechanism within the selection and reordering interface 1003, allowing users to expand or contract Al-generated fill areas using drag-and-expand gestures. The system 200 dynamically adjusts the extent of content synthesis based on user interaction, ensuring that modifications remain intuitive and precise. By incorporating subtle auto-fill suggestions, user-guided completion prompts, and interactive expansion controls, the AI- driven image completion module 1011 ensures Al assists without overwhelming the user, creating a seamless and intuitive refinement experience that preserves artistic intent while optimising composition integrity.

[0162] The system 200 may include an ALassisted gesture recognition module 1012, enabling natural, expressive, and intuitive user interactions for refining Al- generated images without relying on rigid UI controls. This module 1012 may detect micro-gestures and hand movements, translating them into real-time adjustments within the selection and reordering interface 1003 and the refinement module 1006. The gesture recognition module 1012 may implement micro-gesture recognition, wherein Al interprets user interactions to refine Al-generated outputs. Light finger taps signal a preference for specific image detail, swipes over an area indicate elements of increased importance, and fast taps adjust the weighting engine 1004 to increase the influence of a particular element, dynamically modifying image synthesis based on real-time user gestures. By eliminating rigid parameter controls and enabling fluid, artistic user interactions, the gesture recognition module 1012 transforms Al image refinement into a natural, intuitive process, enhancing usability while maintaining precision in Al-driven composition adjustments.

[0163] The system 200 may further include an interactive brush Al refinement mechanism, allowing users to "paint" influence adjustments directly onto an image. The system 200 may analyse tap frequency, duration, and pressure sensitivity, dynamically adjusting composition weighting, stylistic emphasis, or local refinements. These refinements may be processed by the Al image generation module 1005, ensuring seamless adjustments that align with user intent.

[0164] The system 200 may include an adaptive evolution agent module 1013, enabling an Al model to track, analyse, and evolve with individual user preferences over extended periods. This module 1013 may dynamically refine default suggestions, weighting biases, and refinement pathways by continuously learning from user interactions within the preference capture module 1002, weighting engine 1004, and refinement module 1006.

[0165] The system 200 may implement long-term style tracking, wherein the adaptive evolution agent module 1013 monitors changes in user preferences over weeks or months, adjusting default generation parameters based on recurring selection patterns. If a user consistently prefers certain lighting conditions, composition styles, or subject emphasis, the system 200 adapts its preconfigured weighting distributions and refinement recommendations to align with evolving user intent.

[0166] To enhance personalisation, the system 200 may incorporate a dynamic creativity profile 1024, wherein Al categorises users into adaptive creative profiles (e.g., “Cinematic,” “Minimalist,” “Surrealist”) based on their historical refinements. The Al image generation module 1005 may use these profiles to suggest optimised generation paths, ensuring Al-generated outputs progressively align with user-specific aesthetics.

[0167] The system 200 may further include predictive exploration nudges 1015, wherein the adaptive evolution agent module 1013 identifies user stagnation patterns, such as repeated generation of visually similar outputs, and provides subtle variation suggestions within the refinement module 1006. These nudges encourage controlledcreative exploration, allowing users to expand stylistic preferences while maintaining consistency with their established aesthetic inclinations.

[0168] By integrating long-term preference tracking, adaptive creative profiling, and predictive exploration guidance, the adaptive evolution agent module 1013 transforms Al from a static tool into an evolving creative collaborator, ensuring a progressively personalised and dynamic user experience.

[0169] With these additional modules and configuration of the system 200, it enables a transition from manual Al control to an adaptive creative companion for users, dynamically refining outputs within constrained stylistic boundaries based on engagement patterns. The adaptive evolution agent module 1013 analyses real-time user interactions to automate refinements and optimise style adaptation, weight distribution, and generation recommendations. By continuously tracking engagement signals and micro-interactions, the system 200 evolves without requiring manual intervention, ensuring Al-driven image generation remains fluid, intuitive, and aligned with user intent.

[0170] The system 200 may include an Al memory and session management module 1014, enabling structured storage of user progress, version tracking, and adaptive refinement pathways. This module 1014 may maintain persistent, session-aware Al memory, allowing for rollback functionality, contextual resumption, and non-intrusive trend discovery while optimising user control over creative evolution. This module 1014 may include a stateful Al mechanism, persisting user progress across multiple sessions. The system 200 may store session-specific adjustments, allowing users to resume from the exact state of a prior workspace, including image selections, influence weightings, and refinements applied.

[0171] The system 200 may include an inspiration discovery module 1025, surfacing relevant creative trends while ensuring non-intrusive exploration. This module 1025 may operate independently of active workflows, preventing unwanted distractions. This module 1025 may implement tag-based filtering, enabling users to filter AI- generated inspirations by themes (e.g., cinematic, surreal, editorial). Additionally, a"Blend with My Style" feature within the weighting engine 1004 may allow selective integration of new trends into the user’s existing creative profile.

[0172] The Al image generation module 1005 may generate an Al-curated inspiration gallery, categorising emerging trends into an opt-in "Explore" section, allowing users to browse creative styles without affecting ongoing projects.

[0173] The system 200 may integrate UI adaptation mechanisms within the user interaction module 1016, minimising unnecessary controls while ensuring accessibility for advanced refinements. This module 1016 may implement auto-collapsing controls, wherein unused refinement settings are temporarily minimised but dynamically reappear if user behaviour suggests a need for them.

[0174] The system 200 may include hotspot hints that fade over time, dynamically reducing tooltips after repeated interactions to maintain a decluttered interface. Users may toggle between "Minimal Mode" (drag / drop-based controls) and "Pro Mode" (full parameter adjustments).

[0175] The system 200 may include a gesture-driven refinement module 1017, enabling Al-driven image adjustments based on mouse movement and micro-gestures rather than conventional parameter sliders.

[0176] The system 200 may implement tactile weighting adjustments within the selection and reordering interface 1003, allowing users to dynamically "paint" influence strength onto an image by dragging over specific areas. The Al image generation module 1005 may interpret this data to adjust composition balance, subject prominence, and stylistic influence.

[0177] The system 200 may further support interactive Al zones, where user taps on specific image areas are interpreted as priority regions for refinement. Additionally, "Undo via Natural Movement" functionality may allow users to swipe backward to reverse an action instead of requiring explicit button interactions.

[0178] The system 200 may integrate a personalised Al preference learning component within the adaptive evolution agent module 1013, tracking recurring user choices over extended periods.

[0179] System 200 may incorporate preferred styles tagging, and auto-suggesting refinements based on past behaviours. Additionally, the Al image generation module 1005 may use a creative mood prediction model, adjusting mode settings based on user engagement patterns over time.

[0180] The integration of Al-driven session memory, contextual rollbacks, adaptive UI adjustments, and personalised refinement tracking transforms the system 200 into a structured creative memory system. By persisting iterative refinements, intelligently surfacing inspirations, and evolving alongside user preferences, the system 200 enables frictionless, session-aware Al-driven image generation while maintaining full creative autonomy for the user.

[0181] It should be appreciated that while particular embodiments and variations of the invention have been described herein, further modifications and alternatives will be apparent to persons skilled in the relevant arts. In particular, the examples are offered by way of illustrating the principles of the invention, and to provide a number of specific methods and arrangements for putting those principles into effect.

[0182] Accordingly, the described embodiments should be understood as being provided by way of example, for the purpose of teaching the general features and principles of the invention, but should not be understood as limiting the scope of the invention, which is as defined in the appended claims.

Claims

CLAIMS1. An apparatus comprising: a processor configured to: receive first user input comprising words that describe a user’s visual idea; execute an image search to retrieve pre-existing images semantically close to the described visual idea based on the received words; iteratively receive second user input indicating which of the retrieved images aligns with the user's visual idea, and in response to each iteration, progressively refine and execute the image search to retrieve increasingly focused pre-existing images that are semantically closer to the images indicated by the user as aligning with the user's visual idea; and generating at least one photorealistic synthetic image based on the first user input and the second user input(s) using an image synthesis deep learning model.

2. The apparatus of claim 1, wherein the processor is further configured to: receive third user input comprising plain text instructions for further refinement of the at least one synthetic image; and generate at least one additional synthetic image based on the first user input, the second user input(s), and the third user input.

3. The apparatus of claim 1, wherein the second user input further comprises indications of which retrieved images do not align with the user's visual idea.

4. The apparatus of claim 1, wherein the iteratively receiving of second user inputs continues until reaching a saturation point where the user determines that it is unlikely that additional pre-existing images will more closely align with the user’s visual idea, and upon reaching the saturation point, the user initiates the generation of the synthetic images.

5. The apparatus of claim 1, wherein the image synthesis deep learning model is any one from the group consisting of: a latent diffusion model comprising a U-Net, autoencoder and text-encoder; an autoregressive text-to-image generation model; or a text-to-image Transformer model.

6. The apparatus of claim 1, wherein the visual idea comprises a mental image.

7. The apparatus of claim 1, wherein the semantic closeness is a predetermined threshold or user-configurable threshold in high-dimensional space, determined by any one from the group consisting of: cosine similarity, squared Euclidian distance, Manhattan distance, dot product and Hamming distance.

8. The apparatus of claim 3, wherein the second user inputs are stored in a memory in the form of a positive indicators list corresponding to images aligning with the user's visual idea and negative indicators list corresponding to images that do not align with the user's visual idea.

9. The apparatus of claim 8, wherein the positive and negative indicators list weights a query for the image search by adjusting the importance of specific features or attributes present in the images in the positive indicators list while decreasing the weight of features present in the images in the negative indicators list.

10. The apparatus of claim 1, wherein the processor is further configured to: disambiguate the received words based on the sentence structure or surrounding words using a language model.

11. The apparatus of claim 10, wherein the language model comprises any one from the group consisting of: attention mechanism, Structured State Space sequence model, Monarch matrices, and combined set of linear projections using long convolutions and element-wise multiplication.

12. The apparatus of claim 1, further comprising: a recommendation engine to recommend the pre-existing images that are in the form of emojis based on semantic and visual relationships between user-selected emojis; and a visualisation component to display the impact of each emoji addition or subtraction on the generation of emojis through incremental thumbnails or an animated sequence; and a versioning control component to enable branching and exploration of different creative directions by the user without losing the original sequence of user-selected emojis.

13. The apparatus of claim 12, further comprising a merging tool for algorithmically blending image elements from different branches to suggest a coherent mix of user- selected emojis.

14. The apparatus of claim 13, further comprising a rollback feature to allow users to revert to an earlier point in the generation process or undo changes.

15. A method for generating synthetic images, the method comprising: receiving, by a processor, first user input comprising words that describe a user's visual idea; executing, by the processor, an image search to retrieve pre-existing images semantically close to the described visual idea based on the received words; iteratively receiving, by the processor, second user input indicating which of the retrieved images aligns with the user's visual idea, and in response to each iteration, progressively refining and executing the image search to retrieve increasingly focused pre-existing images that are semantically closer to the images indicated by the user as aligning with the user's visual idea; and generating, by the processor, at least one photorealistic synthetic image based on the first user input and the second user input(s) using an image synthesis deep learning model.

Citation Information

Patent Citations

  • Apparatus and method for converting text to image based on learning

    KR102293950B1

  • Response Deriving Method From Multi Input Data In AI Chat-bot and System Thereof

    KR102592859B1

  • Apparatus and method for training a learning system to detect event

    US20180114095A1

  • Electronic device and control method therefor

    US20200058298A1

  • System and method for generating photorealistic synthetic images based on semantic information

    US20230090801A1

Cited By

  • Image retrieval method and device and electronic equipment

    CN121412407A

  • Power prediction method, device and equipment and readable storage medium

    CN121710204A

  • Image detection method based on edge calculation and attention perception compression

    CN121937846A

  • An image detection method based on edge computing and attention-aware compression

    CN121937846B

  • Generative adversarial network architecture search method for text-guided image generation task

    CN122133725A