Multimodal semantic analysis and image search
A multimodal system using a visual language model addresses the limitations of traditional image search by iteratively refining searches based on user input and expert feedback, enhancing context understanding and accuracy in image retrieval.
Patent Information
- Application Number
- JP2025560320
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-18
- Filing Date
- 2024-04-19
- Publication Date
- 2026-04-16
AI Technical Summary
Traditional image search and semantic analysis methods rely on keyword-based approaches and visual similarity techniques that fail to understand the context, meaning, or intent behind user queries, particularly in complex or ambiguous searches, requiring extensive labeled datasets and manual effort for training.
A multimodal system using a trained visual language model performs semantic analysis to identify related concepts, iteratively refines search results based on user input, and incorporates a feedback loop for continuous improvement, enabling accurate and contextually sensitive image retrieval.
The system efficiently and accurately retrieves semantically similar images by understanding broader context and relevance, reducing manual effort and improving search accuracy over time through user interaction and domain expert feedback.
Smart Images

Figure 2026512490000001_ABST
Abstract
Description
Technical Field
[0001] Related Application Information This application claims priority to U.S. Provisional Application No. 63 / 460,952, filed on April 21, 2023, U.S. Provisional Application No. 63 / 468,283, filed on May 23, 2023, and U.S. Patent Application No. 18 / 639,500, filed on April 18, 2024, the contents of each of which are hereby incorporated by reference in their entirety.
[0002] The present invention relates to systems and methods for image search and semantic analysis, and more particularly, to an interactive multi-modal system and method for dynamically composing and semantically classifying images using machine learning and vision language models to process complex semantic queries.
Background Art
[0003] Description of Related Art In the fields of semantic and information retrieval, traditional methods employ keyword-based approaches for text and basic visual similarity techniques for images. These traditional systems rely on simple matching of text or visual patterns without analyzing or understanding the context, meaning, or intent behind the user query. As a result, such methods are insufficient to handle complex, ambiguous, or multi-meaning queries. For example, a search for "Apple" could relate to the technology company, the fruit, or related concepts such as "Apple Music" or "Apple pie recipe," which traditional (e.g., keyword-based) searches cannot effectively distinguish. In image retrieval, traditional methods have relied on visual similarity (e.g., evaluating images based on color, texture, shape, and pattern), limiting their ability to understand deeper semantic content and thematic connections between images. This limitation is particularly pronounced in applications requiring nuance analysis, such as identifying thematic relevance across diverse image datasets or understanding the sentiment conveyed by images in social media analysis.
[0004] Furthermore, traditional systems and methods require extensive and accurately labeled datasets to train models for effective search and retrieval. This need presents significant challenges, including the considerable manual effort required to create and maintain these datasets, and the difficulty in covering a wide range of semantic nuances across different domains. Moreover, traditional approaches require considerable manual effort for training and fine-tuning. These limitations of traditional systems highlight the need for more advanced, intuitive, and contextually sensitive semantic retrieval techniques that can bridge the gap between simple keyword and visual pattern matching and the rich, nuanced understanding required in today's diverse and dynamic data landscape. [Overview of the Initiative]
[0005] According to one aspect of the present invention, a method is provided for identifying and retrieving semantically similar images from a database. This method includes: performing semantic analysis of an input query using a trained visual language (VL) model to identify semantic concepts related to the input query; retrieving a preliminary set of images from a database containing images annotated with semantic information based on the identified semantic concepts; and, for each image in the preliminary set of images, extracting related concepts using a tokenizer that identifies related concepts by comparing each image to a predefined label space. A ranked list of related concepts is generated and presented to the user based on their frequency of occurrence in the preliminary set of images. The preliminary set of images is narrowed down based on the user's selection of specific related concepts from the ranked list of related concepts by combining the input query with the selection of specific related concepts, and further semantic analysis is performed iteratively until a threshold condition is met to obtain an additional set of images that are semantically similar to the combined input query and the selection of specific related concepts.
[0006] According to another aspect of the present invention, a system is provided for identifying and retrieving semantically similar images from a database. The system includes a processor and memory that stores instructions, when executed by the processor, causing the processor to perform semantic analysis on an input query using a trained visual language (VL) model to identify semantic concepts related to the input query, and to retrieve a preliminary set of images from a database containing images annotated with semantic information based on the identified semantic concepts. For each image in the preliminary set, related concepts are extracted using a tokenizer that compares each image to a predefined label space, and a ranked list of related concepts is generated and presented to the user based on their frequency of occurrence in the preliminary set of images. The preliminary set of images is narrowed down based on the user's selection of specific related concepts from the ranked list by combining the input query and selection, and further semantic analysis is iteratively performed until a threshold condition is met, to obtain an additional set of images that are semantically similar to the combined input query and the selection of specific related concepts.
[0007] According to another aspect of the present invention, a computer program product is provided for identifying and retrieving semantically similar images from a database. The computer program product includes a computer-readable storage medium containing program instructions, the program instructions, which are executable by a hardware processor, cause the hardware processor to perform semantic analysis on an input query using a trained visual language (VL) model to identify semantic concepts related to the input query, retrieve a preliminary set of images from a database containing images annotated with semantic information based on the identified semantic concepts, and for each image in the preliminary set, extract related concepts using a tokenizer that compares each image to a predefined label space. A ranked list of related concepts is generated and presented to the user based on their frequency of occurrence in the preliminary set of images. The preliminary set of images is narrowed down based on the user's selection of specific related concepts from the ranked list by combining the input query and selection, and further semantic analysis is performed iteratively until a threshold condition is met to obtain an additional set of images that are semantically similar to the combined input query and the selection of specific related concepts.
[0008] These features and advantages, as well as other features and advantages, will become apparent from the following detailed description of exemplary embodiments of the present invention, which should be read in conjunction with the accompanying drawings. [Brief explanation of the drawing]
[0009] This disclosure provides further details in the following description of preferred embodiments with reference to the following drawings.
[0010] [Figure 1] This block diagram illustrates an exemplary processing system to which the present invention can be applied, according to embodiments of the present invention.
[0011] [Figure 2]This figure illustrates a high-level diagram of an exemplary system and method for artificial intelligence (AI)-based multimodal semantic analysis and image retrieval that responds to diverse input queries, according to embodiments of the present invention.
[0012] [Figure 3] This figure illustrates an exemplary system and method for domain-specific artificial intelligence (AI)-based multimodal semantic analysis and image retrieval that responds to diverse input queries, according to embodiments of the present invention.
[0013] [Figure 4] This figure illustrates an exemplary system and method for iteratively enhancing artificial intelligence (AI)-based multimodal semantic analysis and image retrieval based on feedback from human-in-the-rope experts, according to embodiments of the present invention.
[0014] [Figure 5] This figure illustrates an exemplary system and method for advanced semantic query interpretation and image retrieval using artificial intelligence (AI)-based multimodal semantic analysis and image retrieval according to embodiments of the present invention.
[0015] [Figure 6] This figure illustrates an exemplary method for domain-specific artificial intelligence (AI)-based multimodal semantic analysis and image retrieval that responds to diverse input queries, according to an embodiment of the present invention.
[0016] [Figure 7] This figure illustrates an exemplary method for artificial intelligence (Al)-based multimodal semantic analysis and image retrieval that responds to a variety of input queries for tattoo identification and image retrieval, according to embodiments of the present invention.
[0017] [Figure 8]This figure illustrates an exemplary method for iterative artificial intelligence (AI)-based multimodal semantic analysis and image retrieval in response to diverse input queries, according to an aspect of the present invention.
[0018] [Figure 9] This figure illustrates an exemplary system for artificial intelligence (AI)-based multimodal semantic analysis and image retrieval that responds to diverse input queries, according to an aspect of the present invention. [Modes for carrying out the invention]
[0019] According to embodiments of the present invention, a system and method are provided for a multimodal semantic analysis and image retrieval system that efficiently and accurately searches and semantically analyzes images using an advanced visual language model that responds to user queries (e.g., text, images, videos, etc.). This system and method can take, for example, combinations of visual and text data as input and interpret and identify the subtle meanings behind various image queries. At its core, the present invention integrates an advanced neural network meticulously trained on diverse sets of image-text pairs to extract complex semantic concepts from visual input. The system's capabilities surpass conventional retrieval methods because it not only recognizes visual patterns but also understands the broader context and relevance associated with the images. Through an interactive interface, users can refine search results by selecting relevant concepts from a ranked list, iteratively guiding the system to a more accurate and contextually rich set of images.
[0020] In some embodiments, the present invention can incorporate a feedback loop that domain experts can use to impart knowledge and further fine-tune the ability of an AI model to accurately identify and classify images. This feedback mechanism enables the system to continuously evolve, adapt, and improve its semantic understanding over time. Additionally, a robust computing network supports the demanding tasks of data processing and image retrieval, maintaining the efficiency of the system. The system can include multiple system components, such as, for example, a camera for image acquisition, an image retrieval device, a user interface and a server interface, and an integrated framework that simplifies complex semantic searches using artificial intelligence (AI)-based multimodal semantic analysis and image retrieval techniques. According to an aspect of the present invention, this system is excellent at managing an extensive image database, can utilize a tokenizer to break down text information into semantic units, and can effectively bridge the gap between raw data and meaningful content.
[0021] The embodiments described herein may include elements that are entirely hardware, entirely software, or both hardware and software. In a preferred embodiment, the present invention is implemented in software, which includes, but is not limited to, firmware, resident software, microcode, etc.
[0022] Embodiments can include a computer program product accessible from a computer-usable medium or computer-readable medium that provides program code for use by or in relation to a computer or any instruction execution system. The computer-usable medium or computer-readable medium can include any device that stores, communicates, propagates, or transports a program for use by or in relation to an instruction execution system, apparatus, or device. The medium can be magnetic, optical, electronic, electromagnetic, infrared, or semiconductor system (or apparatus or device) or propagation medium. The medium can include computer-readable recording media such as semiconductor memory, solid state memory, magnetic tape, removable computer diskettes, random access memory (RAM), read-only memory (ROM), rigid magnetic disks, optical disks, etc.
[0023] Each computer program is tangibly stored on a machine-readable storage medium or device (such as a program memory or magnetic disk) readable by a general purpose or special purpose programmable computer, and when the storage medium or device is read by a computer and the procedures described herein are executed, it configures and controls the operation of the computer. Also, the system of the present invention can be considered to be embodied in a computer-readable storage medium composed of a computer program, and the storage medium configured as such causes the computer to operate in a specific predefined manner and execute the functions described herein.
[0024] A data processing system suitable for storing and / or executing program code may include at least one processor directly or indirectly coupled to memory elements via a system bus. The memory elements may include local memory used during the actual execution of the program code, bulk storage, and cache memory for temporarily storing at least some of the program code to reduce the number of times the code is retrieved from bulk storage during execution. Input / output devices or VO devices (including, but not limited to, keyboards, displays, and pointing devices) may be connected to the system directly or via an intermediary VO controller.
[0025] Network adapters can also be integrated into a system to enable it to connect to other data processing systems or remote printers or storage devices via a private or public network. Modems, cable modems, and Ethernet cards are just a few of the types of network adapters currently available.
[0026] Aspects of the present invention will be described below with reference to flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present invention. Each block in the flowcharts and / or block diagrams, as well as combinations of blocks within the flowcharts and / or block diagrams, can be implemented by computer program instructions.
[0027] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. Each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for performing a specified logical function. In some alternative embodiments of the present invention, the functions described within a block may be executed in a different order than that shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or in reverse order, or in any other order depending on the functionality of a particular embodiment.
[0028] Furthermore, it should be noted that each block in a block diagram and / or flowchart, and combinations of blocks within a block diagram and / or flowchart, can be implemented by a specific-purpose hardware system that performs a particular function / operation, or by a combination of dedicated hardware and computer instructions that conform to this principle.
[0029] Here, the same number refers to a drawing in which the same or similar element is represented. First, referring to Figure 1, an exemplary processing system 100 to which the principle of the present invention can be applied is illustrated exemplary according to an embodiment of the principle of the present invention.
[0030] In some embodiments, the processing system 100 may include at least one processor (CPU) 104 operably coupled to other components via a system bus 102. A cache 106, read-only memory (ROM) 108, random access memory (RAM) 110, input / output (VO) adapter 120, sound adapter 130, network adapter 140, user interface adapter 150, and display adapter 160 are operably coupled to the system bus 102.
[0031] The first storage device 122 and the second storage device 124 are operably coupled to the system bus 102 by the VO adapter 120. The storage devices 122 and 124 may be any type, such as disk storage devices (e.g., magnetic disk storage devices or optical disk storage devices), solid-state magnetic devices, etc. The storage devices 122 and 124 may be the same type of storage device or different types of storage devices.
[0032] The speaker 132 is operably coupled to the system bus 102 by the sound adapter 130. The transceiver 142 is operably coupled to the system bus 102 by the network adapter 140. The display device 162 is operably coupled to the system bus 102 by the display adapter 160. According to aspects of the present invention, a visual language (VL) model can be used in combination with a semantic search engine 164 for text and / or image processing tasks, and can further be coupled to the system bus 102 by any suitable connection system or method (e.g., WiFi, wired, network adapter, etc.).
[0033] The first user input device 152 and the second user input device 154 are operably coupled to the system bus 102 by a user interface adapter 150. The user input devices 152 and 154 can be one or more, such as a keyboard, mouse, keypad, image capture device, motion sensing device, microphone, or a device incorporating at least two of the functions of the aforementioned devices. The VL model 156 can be included in a system comprising one or more storage devices, communication / network devices (e.g., WiFi, 4G, 5G, wired connection), hardware processors, etc., according to aspects of the present invention. In various embodiments, other types of input devices can also be used while maintaining the spirit of the principles of the present invention. The user input devices 152 and 154 may be of the same type or different types. The user input devices 152 and 154 are used to input and output information to and from the system 100, according to aspects of the present invention. The VL model 156 can process the received input, and the semantic search engine 164 can be operably connected to the system 100 for semantic retrieval and image retrieval tasks, according to aspects of the present invention.
[0034] Of course, the processing system 100 may include other elements (not shown) that are readily conceivable to those skilled in the art, and certain elements may be omitted. For example, as will be readily understood to those skilled in the art, the processing system 100 may include various other input and / or output devices depending on its specific implementation. For example, various types of wireless and / or wired input and / or output devices may be used. Furthermore, as will be readily understood to those skilled in the art, various configurations of additional processors, controllers, memories, etc., may also be used. Given the teachings of the present principle provided herein, these and other variations of the processing system 100 are readily conceivable to those skilled in the art.
[0035] Furthermore, it should be understood that the systems 200, 300, 400, 500, and 900 described later with respect to Figures 2, 3, 4, 5, and 9 are systems for implementing each embodiment of the present invention. Part or all of the processing system 100 can be implemented in one or more elements of systems 200, 300, 400, 500, and 900 according to aspects of the present invention.
[0036] Furthermore, it should be understood that the processing system 100 may perform at least a portion of the methods described herein, including, for example, at least a portion of methods 200, 300, 400, 500, 600, 700, and 800 as described later with respect to Figures 2, 3, 4, 5, 6, 7, and 8, respectively. Similarly, some or all of systems 200, 300, 400, 500, and 900 may be used to perform at least a portion of methods 200, 300, 400, 500, 600, 700, and 800 of Figures 2, 3, 4, 5, 6, 7, and 8, respectively, according to aspects of the present invention.
[0037] As used herein, the terms “hardware processor subsystem,” “processor,” or “hardware processor” may refer to a processor, memory, software, or combination thereof that works together to perform one or more specific tasks. In useful embodiments, a hardware processor subsystem may include one or more data processing elements (e.g., logic circuits, processing circuits, instruction execution devices, etc.). One or more data processing elements may be included in a central processing unit, a graphics processing unit, and / or a separate processor-based or arithmetic element-based controller (e.g., logic gates, etc.). A hardware processor subsystem may include one or more onboard memories (e.g., caches, dedicated memory arrays, read-only memory, etc.). In some embodiments, a hardware processor subsystem may include one or more memories (e.g., ROM, RAM, basic input / output system (BIOS, etc.)) that may be onboard or offboard, or dedicated to the hardware processor subsystem.
[0038] In some embodiments, the hardware processor subsystem may include and execute one or more software elements. These one or more software elements may include an operating system and / or one or more applications and / or specific code to achieve a particular result.
[0039] In other embodiments, the hardware processor subsystem may include dedicated special circuits that perform one or more electronic processing functions to achieve a particular result. Such circuits may include one or more application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), and / or programmable logic arrays (PLAs). These and other variations of the hardware processor subsystem are also considered according to embodiments of the present invention.
[0040] Referring here to Figure 2, a high-level diagram is illustrative of an exemplary system and method 200 for artificial intelligence (AI)-based multimodal semantic analysis and image retrieval in response to diverse input queries, according to an embodiment of the present invention.
[0041] In various embodiments, block 202 may utilize a domain-specific image dataset that can encompass a wide range of image types, such as tattoos, faces, or automobiles. This dataset is carefully curated, and the images are annotated with semantic tags and descriptors. Such a dataset is useful for a search system, providing a reference repository from which semantic comparisons can be derived. The images in this dataset are preprocessed to ensure that features are extractable and that each image is associated with metadata that represents its semantic content. Thus, this repository provides a rich database from which input queries can be semantically matched.
[0042] Block 204 represents the reception of input queries. This is the stage where the system interfaces with the user and accepts the submission of images for semantic analysis. Preprocessing of these images includes normalization, resolution adjustment, and, if necessary, feature enhancement for optimization for encoding. This preprocessing step is crucial for removing noise and standardizing the input, enabling more effective feature extraction and ensuring consistency throughout the semantic retrieval process. In Block 206, the visual encoder acts as a transformation mechanism, converting visual data from the input query into numerical feature vectors. Using a convolutional neural network or a similar advanced machine learning architecture, the visual encoder transforms the essence of the input visual information into a form that summarizes the most important visual attributes such as shape, texture, and pattern. The output is a concise and comprehensive representation of the image's visual information, ready for semantic comparison.
[0043] In various embodiments, the output from block 206 is fed into block 201, where the vectorized query image resides in a common feature space. This space is an abstract construct that transcends the original modality, regardless of whether the semantic similarity is textual or visual. Here, the feature vectors of the query image are ready to be matched with other vectors representing the semantic content of images and words in the database, facilitating modality-independent comparisons. Block 208, the common feature space, is depicted as a melting pot of multimodal data, where each vector is compared regardless of whether it originates from visual or textual analysis. This comparison is based on semantic proximity, where close vectors within this space are considered semantically similar. This semantic matchmaking is at the heart of the system, enabling the identification of context and subtle meanings.
[0044] In block 210, the tokenizer can break down text information into constituent elements (e.g., words or tokens that the system can process). The word list within the tokenizer contains semantically significant terms that may be relevant to the search domain. These tokens can be used for semantic analysis of the text and act as a bridge between the raw text data and semantically enhanced vectors that can be incorporated into a common feature space. Block 212 shows a text encoder, which converts text tokens into text-based feature vectors. This encoder uses natural language processing to interpret and encode the semantic significance of tokens, aligning them with the system's semantic understanding established within the common feature space.
[0045] In various embodiments, blocks 203, 205, 207, 209, 211, 213, and 215 can be used for the identification and association of concepts. These blocks represent individual concepts identified by the system and are the result of the search process. Each concept is a node in a common feature space linked to the query image, indicating semantic relevance. In this exemplary example, the concept nodes from "tattoo" to "tiger" are not merely independent points, but interconnected nodes that constitute a semantically rich network that the system uses to understand the query in a broader semantic context. The above processing, from the initial reception of the image query to the semantic linking of concepts within a multimodal framework, represents a revolutionary leap in speed, accuracy, and reduction of processor requirements in the search technique, according to aspects of the present invention.
[0046] Referring here to Figure 3, an exemplary diagram is shown illustrating a system and method 300 for domain-specific artificial intelligence (AI)-based multimodal semantic analysis and image retrieval in response to diverse input queries, according to an embodiment of the present invention.
[0047] In various embodiments, System and Method 300 provides a robust solution for analyzing and retrieving tattoo images. The system integrates various components, from image capture to semantic output, enabling domain experts to derive actionable insights from complex visual data. Integration of the system with state-of-the-art visual language models facilitates a new approach to interpreting tattoos that conventional visual similarity algorithms cannot decipher. While tattoo identification will be discussed later, it should be noted that, according to embodiments of the present invention, all types of images can be analyzed and interpreted. In this exemplary embodiment, the object of interest shown in Block 302 represents a tattoo, an image that the user wants to analyze for its semantic content. The object may contain symbols or elements that require sophisticated interpretation, such as determining affiliation or meaning, which are often associated with specific groups or activities. Block 304 includes various cameras or imaging devices, ranging from high-resolution digital SLR cameras to smartphone cameras, to capture high-fidelity digital images of the object of interest. These imaging devices are tuned to capture the subtle details of the tattoo, ensuring that the visual data is accurately represented for subsequent semantic analysis.
[0048] Once acquired, the image can undergo digital processing in block 306. This step may include adjusting the resolution, enhancing contrast, and applying filters to highlight defining features of the tattoo. These preprocessing tasks can be used to ensure that the digital representation of the image retains important visual cues for accurate semantic analysis. Block 308 represents an interface for the user to interact with the system, such as a dedicated application on a tablet or smartphone, or a web portal on a desktop computer. Through this interface, the user submits a digital image and, if necessary, attaches additional contextual information or specific questions they wish to resolve through semantic retrieval processing. This interface may also include drag-and-drop functionality, allowing the user to combine multiple concepts (randomly or in a specified order) to further refine their search.
[0049] The server device 310 can function as the central processing hub of the system. The server device 310 can support the execution of the visual language model and have the necessary computing resources to manage interactions with the database. The server can be configured to handle multiple concurrent user requests and perform real-time processing to meet the demands of time-constrained forensic investigations. In block 312, the core of semantic analysis is the VL model. This advanced model uses a deep neural network to perform cross-modal analysis, enabling the understanding of the contextual and symbolic meaning of visual data. According to aspects of the present invention, this model can be trained on diverse datasets, enabling the generalization and accurate interpretation of a wide range of visual concepts. Block 314 refers to a comprehensive database or dataset containing numerous images annotated with detailed semantic information. The repository is dynamic and continuously updated with new data and annotations, ensuring that the system's output is based on the most up-to-date information available. This resource can be used to maintain the system's speed, accuracy, and semantic relevance.
[0050] In various embodiments, output 316 represents the final product delivered to the user after processing and analysis of the digital image. These outputs may include semantic associations that the system has determined to be relevant to the image query. These may be presented in a clear and user-friendly format, either separately from or within the same interface used to submit the query. Blocks 318, 320, 322, and 324 may determine and provide to the user detailed semantic association outputs. These blocks represent specific types of semantic associations and outputs that the system may identify and present. Block 318 outlines the system's ability to associate objects of interest with crimes in real time, thereby assisting criminal investigations and profiling. Block 320 highlights the system's ability to link images to known gangs and subcultures, providing useful investigative insights to law enforcement agencies. Block 322 relates to high-level semantic derivation, where the system can infer beyond visible elements to understand the broader meaning and context of an image (such as a tattoo image). According to aspects of the present invention, block 324 emphasizes the system's adaptability and user-centric design, encompassing the system's expertise in associating any query type submitted by the user with images.
[0051] Referring here to Figure 4, an exemplary diagram is shown illustrating a system and method 400 for iteratively enhancing artificial intelligence (AI)-based multimodal semantic analysis and image retrieval based on human-in-the-loop expert feedback, according to an embodiment of the present invention.
[0052] In various embodiments, System and Method 400 is a diagram of an advanced adaptive AI system designed to perform semantic analysis and retrieval of images. This system is unique in that it incorporates domain-specific feedback, enabling continuous learning and enhancement of semantic analysis capabilities. The detailed depiction of the feedback loop and the incorporation of human expertise highlight the inventive step, according to aspects of the invention, of fusing machine learning with domain-specific knowledge to bring about advancements that surpass the capabilities of conventional image retrieval systems.
[0053] In various embodiments, in block 402, the core AI system is depicted as the primary processor for image data and semantic interpretation. It is responsible for initial analysis, applying machine learning and deep learning techniques to extract and understand the visual content of images submitted as queries. The AI system incorporates sophisticated algorithms that can identify complex patterns and learn from vast datasets to form initial hypotheses about the semantic content of each image. Block 404 represents the AI system's ability to develop self-learning representations. These representations can be complex multidimensional vectors or feature maps that the system autonomously constructs from the input data without explicit external labeling. These are the result of unsupervised learning by the AI, enabling it to identify and code unique structures and patterns within the data, forming the basis for subsequent semantic interpretation.
[0054] Block 406 demonstrates the application of enhanced semantic analysis by an AI system for target image retrieval. Based on improved high-level concepts and insights provided by domain experts, the system can search for new image sets that better match precise semantic criteria. These search results can be returned to the domain experts for further analysis, continuing the feedback and refinement cycle. In Block 408, the AI system can infer relatively high-level concepts from self-learned representations. These concepts may be abstractions or generalizations that the system identifies as semantically important and potentially meaningful to domain experts. This step can include the AI's ability to infer beyond individual data points by leveraging its learned knowledge to categorize and conceptualize visual data into broader themes and ideas.
[0055] Block 410 describes the role of a domain expert in the search process. In this exemplary embodiment, the expert can scrutinize the concepts generated by the AI and apply their expertise and contextual understanding to evaluate and refine the AI's output. The domain expert's analysis may include identifying nuances and subtle points that the AI's algorithmic approach might overlook, thereby enhancing the insights gained from the system. The human-in-the-loop feedback mechanism shown in Block 412 can represent an iterative component of the system. This can include dynamic interactions between the domain expert and the AI system (e.g., human-in-the-loop), where feedback from the expert's analysis can inform and refine the AI's learning process. This feedback loop allows the system to evolve its understanding, improve semantic representation and high-level conceptual identification over time, and tailor its learning to the expert's requirements, according to aspects of the invention.
[0056] Referring here to Figure 5, an exemplary diagram is shown illustrating a system and method 500 for advanced semantic query interpretation and image retrieval using artificial intelligence (AI)-based multimodal semantic analysis and image retrieval, according to an embodiment of the present invention.
[0057] In various embodiments, in block 502, the search process is initiated when the user enters a text query, such as "Tesla," into the query initiation interface. This interface is designed to accommodate queries with a broad semantic landscape and is sophisticated enough to understand that a term may refer to multiple entities or concepts, such as a prominent historical figure, a modern technology company, or a comprehensive scientific principle. While "Tesla" is described as a query for illustrative purposes, it should be understood that according to aspects of the invention, any query can be entered and analyzed. In block 504, an intelligent search system core can be utilized, which can be an advanced AI mechanism that processes text queries and maps the input to potential visual representations using state-of-the-art semantic analysis. The AI can retrieve images from a vast database of tagged images and leverage deep learning to associate text with precise visual representations across various categories.
[0058] The AI system in Block 506 can aggregate image sets related to the theme of "scientists," particularly those concerning the scientist Nikola Tesla. This demonstrates the AI's nuanced ability to extract historical and academic significance from queries and assemble search results with images that embody the legacy of scientist Tesla. Block 501 is a model of targeted image search, presenting the first variation of Nikola Tesla's face. This image is one of several representations identified and retrieved by the AI system, illustrating the diversity in the semantic context of the historical figure that the query may represent. In Block 503, the system provides a second variation of Nikola Tesla's face. Each version selected by the AI is distinct, reflecting the system's diverse archive access and its ability to provide a comprehensive visual arrangement by distinguishing between and within image categories. Block 505 shows a third variation of Nikola Tesla's face. By presenting multiple instances of related images, the system caters to users searching for a specific depiction or those who desire a wide range of options for comparison and analysis.
[0059] In Block 508, the AI integrates a category specifically for Tesla logos, distinguishing them from other image categories. This separation allows users to quickly find brand-related images that are distinct from personal or conceptual associations with Tesla. Block 507 displays the first variation of the Tesla logo image. The search system uses visual recognition technology to distinguish and search for various designs and iterations of the Tesla brand logo. Block 509 follows with a second variation of the Tesla logo image, further demonstrating the AI system's ability to diversify the search and present the user with a range of logo options, each potentially reflecting different eras or design changes.
[0060] Block 511 shows a third variation of the Tesla logo image, adding to the range of the brand's visual identity accessible by the system. This exemplifies the AI's search capabilities and sensitivity to visual nuances within the brand image. The AI's advanced knowledge is demonstrated in Block 510, where it can select a series of images related to Tesla vehicles, presenting a clear category that appeals to users interested in automotive design, technology, and evolution. This category reflects the system's ability to distinguish product-specific images. In Block 513, the first variation of a Tesla vehicle is found. This may represent a specific model or type, demonstrating the AI system's detailed recognition capabilities to correspond to the nuances of automotive design that symbolize the Tesla brand. Block 515 shows the system searching for a second variation of a Tesla vehicle, highlighting the system's ability to provide diverse images within the same product category, allowing users to explore various models and designs of Tesla vehicles. Block 517 presents a third variation of a Tesla vehicle, completing the set of three vehicle images provided. This diversity allows users to compare, analyze, or simply admire the various vehicle designs that Tesla offers.
[0061] Block 512 illustrates the AI's intellectual leap in linking queries to the broad concept of electricity, presenting a series of images that summarize the essence of electrical innovation. These may include photographs of inventions, symbols, or conceptual artwork related to electricity, all connected to the theme of electricity embodied in Tesla's work. The search system in Block 519 provides a first variation of images related to the concept of electricity, which may be textual descriptions or abstract illustrations. This demonstrates the AI system's broad semantic understanding and its ability to interpret and retrieve images based on abstract concepts. In Block 521, a second variation of images representing the concept of electricity is retrieved, offering different visual perspectives on the theme. This variation demonstrates the depth and breadth of the system in curating images that, while thematically connected, offer various representations of the same concept. Block 523 completes this exemplary descriptive search set with a third variation of images related to the concept of electricity, solidifying the system's ability to present a rich tapestry of conceptual images. This scope allows users to understand multiple visual interpretations and the deep connection between Tesla and the field of electrical science, according to aspects of the present invention.
[0062] Next, referring to Figure 6, an illustrative diagram shows a domain-specific artificial intelligence (AI)-based multimodal semantic analysis and image retrieval method 600 according to an embodiment of the present invention, which responds to a variety of input queries.
[0063] In various embodiments, in block 602, the multimodal semantic retrieval system begins by receiving an input query that may be any type of media (e.g., text, video, image, etc.). The system can handle complex queries that have ambiguous meanings or require a deep, domain-specific understanding. The input is preprocessed to improve text clarity and image quality for subsequent analysis. This step can involve understanding the format of the query and preprocessing it to prepare it for subsequent semantic or conceptual analysis. The system ensures that queries, encompassing a wide range of subjects from specific objects to abstract concepts, are accurately interpreted for further processing.
[0064] In block 604, an advanced visual language (VL) model can be used to perform semantic analysis and retrieval of the top k semantically similar images, including semantic analysis of input queries to identify and retrieve the top k semantically similar images. Models trained on a wide range of image-text pairs can leverage multimodal capabilities to understand the nuances behind queries, significantly outperforming traditional keyword-based or visual similarity-based searches. This process can be used to bridge the gap between visual concepts and textual meanings, enabling more refined searches that recognize the intent and context of the query. In block 606, once the top k images have been retrieved, the system can apply an advanced tokenization process to extract relevant concepts from each image. This can include a detailed comparison of images against a comprehensive label space, which can be user-defined or set by default in the model's tokenizer. The extraction process can include refinement stages to eliminate redundant or overlapping concepts and / or non-discriminatory labels, ensuring that only clear and relevant concepts are considered. According to aspects of the present invention, a rich set of semantically important concepts can be obtained for each image, providing a basis for understanding and classifying the nuances of the content.
[0065] In block 608, the user-interactive presentation of extracted concepts may include presenting the extracted concepts to the user in an interactive and user-friendly manner. The system can rank these concepts based on their frequency of occurrence across the searched images and display them in a way that facilitates easy and user-friendly review and selection by the user. This interactive presentation not only enhances user engagement but also allows the user to refine their search by gaining a deeper understanding of the semantic landscape surrounding the initial query. Block 610 focuses on refining the search by combining the original query with selected concepts based on the user's selection of relevant concepts. This combination process is configurable to match the format of the initial query and integrates textual and visual information in a way that respects the semantic coherence of the user's intent using advanced algorithms. The formulation of the refined query highlights the flexibility of the invention and its ability to adapt the search strategy based on user input, ensuring search output that closely matches the user's needs.
[0066] In block 612, a refined query, enhanced with a user-selected concept, can undergo additional semantic retrieval processing to provide a more refined semantic search. This may involve complex interactions of the system's visual and language processing components to find something that closely matches the semantic parameters of the refined query. This search step can be performed precisely by leveraging the system's comprehensive understanding of both textual and visual semantics, producing results that are not only visually similar but also conceptually consistent with the user's refined query. In block 614, the latest search results are presented to the user via a user interface, displaying a curated selection of images or text that closely match the refined query parameters. This display is designed to be informative and easy to navigate, allowing the user to quickly find results that satisfy their search criteria. The detailed display may include metadata and conceptual tags for each result, allowing the user to gain a deeper understanding of the semantic basis behind each match.
[0067] In Block 616, search results can be continuously improved by iteratively integrating user feedback. This feedback loop allows users to directly input information about the relevance and accuracy of search results, which the system can then use to further refine semantic analysis and search processing. By incorporating user feedback, the system can learn and evolve to gain a deeper understanding of user intent and preferences, gradually yielding more accurate and relevant search results over time. Block 618 demonstrates an example of the invention's application to forensic analysis, where it can be used for searching, matching, and searching tattoos. According to aspects of the invention, the system's ability to interpret the meaning behind tattoos and associate them with specific entities such as gangs or crimes demonstrates its usefulness in real-world applications and its value in assisting law enforcement agencies and domain experts.
[0068] Next, referring to Figure 7, an exemplary diagram is shown illustrating a method 700 for automated artificial intelligence (Al)-based multimodal semantic analysis and image retrieval, according to an embodiment of the present invention, which responds to a variety of input queries for real-time tattoo identification and image retrieval.
[0069] In various embodiments, an input query can be received in block 702. While the following exemplary embodiments describe images particularly relevant to tattoo images for ease of explanation, according to aspects of the present invention, the invention is applicable to all kinds of queries for any subject type. The input query can be, for example, a text description (e.g., "dragon" or "tribal") and / or a visual input (e.g., an image of a tattoo). This step may include preprocessing to identify the nature of the query, focusing on specific aspects related to the image and symbolism of tattoos, thereby preparing for targeted semantic analysis.
[0070] Block 704 utilizes a visual language model trained to focus on tattoo designs and their associated meanings. This step may include performing a detailed semantic analysis of the input query, recognizing the subtle expressions and cultural significance of tattoos, and capturing the symbolic and often personal meanings embedded in the tattoo art beyond basic visual features. Block 706 can retrieve a curated list of tattoo designs that are semantically similar to the input query. This search can be performed based on a comprehensive database of tattoo images annotated with rich semantic information, including cultural, symbolic, and aesthetic aspects. This step can utilize the model's ability to interpret the multifaceted meanings behind tattoos, ensuring that the search matches the user's intent. In Block 708, for each retrieved tattoo design, the system can extract relevant concepts using advanced tokenization processing tailored specifically for tattoo-related content. This may include, for example, analyzing the design against a specialized label space encompassing a wide range of tattoo styles, symbols, and their meanings. According to aspects of the invention, this processing can be refined to exclude irrelevant concepts and highlight those central to understanding the significance of each tattoo.
[0071] In various embodiments, block 710 allows the extracted concepts to be presented interactively to a forensic expert, who can then consider and select the concepts deemed most relevant to the forensic investigation. This step facilitates expert involvement in the semantic layers of tattoos and helps refine the search based on deeper insights into the meaning and relevance of tattoos. In block 712, leveraging the forensic expert's selection of relevant tattoo concepts, the system can refine the search to focus more closely on tattoos that match forensic criteria. This refinement process can incorporate the expert's expertise and selected concepts to adjust the search parameters, ensuring that the resulting tattoo designs are particularly interesting for a specific forensic investigation. In block 714, the refined search can output precise search results for tattoo designs that are not only semantically similar to the original query but also meet the forensic criteria specified in the refinement process. This step can identify a selection of targeted tattoo designs, each with detailed semantic annotations explaining their relevance and importance in a forensic context. In block 714, refined searches can generate precise search results for tattoo designs that are not only semantically similar to the original query but also meet specified forensic criteria through the refinement process. This step allows for the extraction of targeted tattoo designs with detailed semantic annotations explaining their relevance and importance in a forensic context.
[0072] Block 716 allows for the presentation of refined search results to forensic professionals, highlighting tattoo designs deemed potentially important in the investigation. This presentation is rich in detailed information, providing insights into the cultural, symbolic, and aesthetic meanings of each design and offering a comprehensive overview to support forensic analysis. Block 718 allows for the iterative integration of forensic professional feedback on the relevance and accuracy of the retrieved tattoo designs into the system, contributing to continuous system improvement. According to aspects of the present invention, this feedback can be used to continuously refine the system's semantic understanding and search accuracy, increasing its usefulness in forensic applications over time, strengthening the system's ability to understand and interpret the complex meanings of tattoos and tattoo images (or other subjects of interest), and ensuring that the search process is accurate and appropriate to the needs of forensic professionals.
[0073] Referring next to Figure 8, an exemplary diagram is shown illustrating a method 800 of iterative artificial intelligence (AI)-based multimodal semantic analysis and image retrieval in response to diverse input queries, according to an embodiment of the present invention. This computer implementation method 800 can identify and retrieve semantically similar images from a database and exhibits a complex and systematic approach incorporating advanced AI semantic analysis, a user-interactive system, and iterative refinement techniques. According to aspects of the present invention, the method 700 is highly adaptive and user-centric, ensuring that the retrieved images precisely match the complex semantic intent of the user's query.
[0074] In various embodiments, block 802 enables robust semantic analysis of user input queries using an advanced visual language model. Trained on diverse datasets, this model interprets input queries multidimensionally, understanding deeper intent and broader context. It considers various linguistic nuances, cultural references, and semantic interpretations that may be relevant to the query, forming a comprehensive semantic file that underlies image retrieval. Block 804 details the process of searching a preliminary set of images from a meticulously annotated database. This set can be identified based on the semantic profile established from the input query. The database may contain images tagged with extensive semantic information, facilitating searches that go beyond superficial visual similarity and into the realm of semantic matching.
[0075] In block 806, a detailed concept extraction process is performed on each image in the preliminary set. A tokenizer capable of incorporating advanced natural language processing capabilities examines each image and cross-references it with a comprehensive predefined label space. Because this space contains a wide range of latent concepts, highly relevant semantic information can be extracted from each image. Block 808 generates a ranked list of concepts. The relevance of this ranked list is determined not only by frequency of occurrence but also by semantic weights. Semantic weights can be calculated using a weighting mechanism that takes into account the depth of relevance of each concept to the input query. This hierarchical ranking helps prioritize the concepts in user reviews, ensuring that the most relevant concepts are highlighted.
[0076] Block 810 allows the user to be presented with a ranked list of relevant concepts through an interactive and intuitive user interface. This interface may include visualization tools such as graphs or heatmaps to help the user understand the prevalence and relevance of each concept. The user is encouraged to make informed choices about the concepts best suited to their search objectives while referring to the list. Block 812 demonstrates a process of refining the initial image selection in real time based on user interaction. Input queries can be dynamically combined with the concepts selected by the user, and the visual language model can apply these refined criteria to subsequent searches. Further refinement can be performed in real time in response to user input, allowing the image set to continuously adjust to better match the user's expectations. Block 814 allows for iterative semantic analysis and retrieval of new image sets, leveraging additional feedback and refined search criteria. This analysis and search loop can be used to improve the accuracy and relevance of search results. Furthermore, according to aspects of the present invention, this iterative process may also incorporate machine learning techniques to adapt the parameters of the visual language model to improve future searches based on the user's interaction patterns.
[0077] In block 816, it is possible to evaluate whether the iterative refinement process has met predetermined threshold conditions. These conditions can be defined by the amount of image data retrieved, user satisfaction ratings, determination by calculation of semantic alignment quality, or other user-defined criteria or thresholds. This evaluation mechanism ensures that the search process is efficient and meets user requirements. In block 818, a curated set of final images, semantically refined and validated through iterative user feedback, can be presented to the user. According to aspects of the present invention, this final display may include an advanced display algorithm that organizes the images in a manner best suited to the user's needs, including, for example, grouping by conceptual relevance or other criteria specified by the user during the search process.
[0078] Next, referring to Figure 9, an exemplary diagram is shown illustrating a system 900 for artificial intelligence (AI)-based multimodal semantic analysis and image retrieval that responds to a variety of input queries, according to an embodiment of the present invention.
[0079] In various embodiments, camera 902 can function as the primary image acquisition device that captures visual data to be processed by the system. This may be a high-resolution digital camera (or other image acquisition device) capable of taking detailed photographs that the system uses for semantic analysis. Camera 902 can capture even subtle details for accurate semantic interpretation of visual content. Block 904 refers to a visual language model that forms the core of the system's semantic understanding capabilities. This artificial intelligence-powered model can interpret the visual information acquired by camera 902. By understanding context and extracting semantic data, it can process and analyze images, achieving advanced search capabilities that surpass conventional keyword matching. The database or image dataset 906 functions as the system's repository, storing a vast number of images annotated with semantic information. It is configured to support complex queries, enabling efficient storage and retrieval of images. This dataset can be continuously updated and refined to maintain the accuracy and relevance of search results.
[0080] In various embodiments, block 908 incorporates both a visual encoder and a text encoder, capable of converting raw images and text into encoded features that the system can understand and compare. The visual encoder processes image data, and the text encoder processes text information, enabling the system to comprehensively understand both visual and text content. The visual feature extraction engine 910 is a dedicated subsystem capable of analyzing images and extracting characteristic features. Using techniques such as edge detection, pattern recognition, and color analysis, it can decompose images into feature sets that can be used for matching and searching. Block 912 includes a tokenizer that processes text data from the encoder and database to identify distinct concepts within the content. It works by decomposing text into tokens representing basic units of meaning, which can then be used to map the text to visual content. Bus 901 is the system's communication backbone, enabling data transfer and command flow between various components. The operation of the camera, encoder, extraction engine, and tokenizer are all synchronized, resulting in a seamless workflow.
[0081] Block 914 allows for the continuous learning and improvement of visual language (VL) models using a neural network training device. This advanced device can iteratively train VL models using a corpus of annotated images and associated text data, employing machine learning algorithms, particularly deep learning neural networks. The training device can improve the model's ability to accurately interpret and process semantic information by combining supervised learning, unsupervised learning, and reinforcement learning techniques.
[0082] Training can involve feeding the model a large number of image-text pairs, and the model learns to associate specific visual patterns with corresponding semantic labels. The device can implement a backpropagation algorithm, adjusting the weights within the neural network to minimize the difference between the model's output and known labels. This is a process known as error reduction. This iterative training can be carried out over multiple epochs, enabling the model to recognize a vast range of semantic concepts, from simple object recognition to complex abstract ideas expressed in visual form. Further techniques such as transfer learning can be applied to training in block 914. In transfer learning, a model pre-trained on a general dataset is fine-tuned with domain-specific data to improve its performance on specific tasks. Another aspect of training may include adversarial training, which involves introducing perturbations or "noise" into the training data to improve the model's robustness and ability to generalize from incomplete inputs.
[0083] In some embodiments, in addition to the core training function, block 914 can also perform hyperparameter tuning to optimize the model's performance. This includes tuning the learning rate, the number of layers and neurons in the neural network, the activation function, and other model parameters that affect the training results. According to aspects of the present invention, the training device can validate the model's performance using a separate validation dataset to prevent overfitting and ensure that the model maintains high accuracy even when exposed to new, unknown data. The computing network 916 can provide the processing power (e.g., cloud) necessary to handle the system's intensive computational tasks. It also facilitates connectivity between system components, supports cloud-based operation, and enables the system to operate at scale.
[0084] Block 918 is an artificial intelligence system that forms the basis of semantic analysis and retrieval functions. It integrates machine learning, computer vision, and natural language processing to build a powerful AI core capable of handling complex semantic queries. The image retrieval device 920 can retrieve images semantically linked to a user's query by querying an image database, for example executing a search command, and retrieving relevant images based on semantic analysis performed by the AI system. Block 922 details a user device or system interface, which may be a web portal, mobile app, or desktop app, and serves as the point of interaction between the system and the user. This interface is where the user enters queries, receives image results, and provides feedback to the system. A server device 924, which may include one or more processor devices, can be used to manage the overall operation of the system, hosting AI components and processing image and text data. According to aspects of the present invention, all parts of the system cooperate efficiently to manage resources and optimize the performance of a sophisticated system architecture for semantic image retrieval processing.
[0085] In this specification, any reference to “one embodiment” or “embodiment” of the present invention, and to other modifications, means that certain features, structures, properties, etc., described in relation to the embodiment are included in at least one embodiment of the present invention. Therefore, expressions such as “in one embodiment” or “in an embodiment” appearing elsewhere in this specification, and any other modifications, do not necessarily all refer to the same embodiment. However, it should be understood that, considering the teachings of the present invention provided herein, features of one or more embodiments can be combined.
[0086] For example, in the case of "A / B," the use of any of the following " / ," "and / or," or "at least one," such as "A and / or B" or "at least one of A and B," will be understood as intended to include the selection of only the first listed option (A), only the second listed option (B), or both options (A and B). As further examples, in the case of "A, B, and / or C" and "at least one of A, B, and C," such expressions are intended to include the selection of only the first listed option (A), only the second listed option (B), only the third listed option (C), only the first and second listed options (A and B), only the first and third listed options (A and C), only the second and third listed options (B and C), or all three options (A, B, and C). This can be extended as many times as there are listed items.
[0087] The foregoing is illustrative and representative in all respects, but not restrictive, and the scope of the invention disclosed herein is determined not from the detailed description, but from the claims, which are interpreted in accordance with the maximum extent permitted by patent law. The embodiments shown and described herein are merely illustrative of the invention, and those skilled in the art should understand that various modifications can be implemented without departing from the scope and spirit of the invention. Those skilled in the art can implement various other combinations of features without departing from the scope and spirit of the invention. Thus, while aspects of the invention have been described with the detail and specificity required by patent law, what is claimed and intended to be protected by the patent is as stated in the appended claims.
Claims
1. A computer implementation method for identifying and searching for semantically similar images from a database, (406) Performing semantic analysis of an input query using a trained visual language (VL) model to identify semantic concepts related to the input query, Based on the identified semantic concept, obtain a preliminary set of images from the database containing images annotated with semantic information (804), For each image in the aforementioned set of images, the associated concept is extracted using a tokenizer that identifies the associated concept by comparing each image with a predefined label space (806), (810) To generate and present to the user a ranked list of the related concepts based on their frequency of appearance within a preliminary set of the aforementioned images, Based on the user's selection of a specific related concept from the ranked list of related concepts, the input query is combined with the selection of the specific related concept to narrow down the preliminary set of images (812), A method comprising: iteratively performing further semantic analysis until a threshold condition is met to obtain an additional set of images that are semantically similar to the combined input query and the selection of the particular related concept (814).
2. The method according to claim 1, wherein the preliminary set of images is a set of the top k images from the database, where k is a predetermined number, and the images are ranked according to a determined semantic similarity to the input query.
3. The method according to claim 1, wherein the input query is text, and the semantic analysis includes using the language encoder portion of the VL model to find semantically similar images based on text concepts in the predefined label space.
4. The method according to claim 1, wherein the input query is an image, and the semantic analysis includes mapping the visual concepts of the input image to textual concepts in the predefined label space using the visual encoder portion of the VL model.
5. The method according to claim 1, further comprising narrowing down the predefined label space before extracting the related concepts by removing redundant and non-identifiable labels in order to improve the accuracy of concept extraction.
6. The method according to claim 1, further comprising using a feedback loop to iteratively refine the search for the input query based on the selection of additional relevant concepts by the user.
7. The method according to claim 1, wherein the user's selection of related concepts includes enabling the user to combine multiple concepts in a specified order for a more refined search by utilizing the drag-and-drop functionality of the user interface.
8. A system for identifying and searching for semantically similar images from a database, Processor (104), The system has a memory (110) for storing instructions, and when an instruction is executed by the processor, it is executed by the system. A trained visual language (VL) model is used to perform semantic analysis of the input query and identify semantic concepts related to the input query (406). Based on the identified semantic concepts, a preliminary set of images is obtained from the database containing images annotated with semantic information (804), A tokenizer is used to compare each image with a predefined label space to extract the associated concepts for each image in the preliminary set (806), A ranked list of the related concepts is generated and presented to the user based on their frequency of appearance within the preliminary set of images (810). Based on the user's selection of a specific related concept from the ranked list, the input query is combined with the selection to narrow down the preliminary set of images (812), Further semantic analysis is performed iteratively until the threshold condition is met (814). A system (814) that causes the system to obtain an additional set of images that are semantically similar to the combined input queries and the selection of the specific related concepts.
9. The system according to claim 8, wherein the preliminary set of images is a set of the top k images from the database, where k is a predetermined number, and the images are ranked according to a determined semantic similarity to the input query.
10. The system according to claim 8, wherein the input query is text, and the semantic analysis includes processing by a language encoder component of the VL model to find semantically similar images based on text concepts in the predefined label space.
11. The system according to claim 8, wherein the input query is an image, and the semantic analysis includes processing by a visual encoder component of the VL model to map the visual concepts of the input image to textual concepts in the predefined label space.
12. The system according to claim 8, wherein the instruction further causes the system to narrow down the predefined label space by removing redundant and non-identifiable labels in order to improve the accuracy of concept extraction.
13. The system according to claim 8, wherein the instruction further causes the system to employ a feedback loop to iteratively refine the search for the input query based on the user's selection of the additional relevant concepts.
14. The user interface will be further enhanced with drag-and-drop functionality, allowing the user to combine multiple concepts in a specified order for more refined searches. The system according to claim 8.
15. A computer program product for identifying and searching for semantically similar images from a database, wherein the computer program product comprises a computer-readable storage medium into which program instructions are embedded, and the program instructions, which are executable by a hardware processor, are transmitted to the hardware processor. A trained visual language (VL) model is used to perform semantic analysis of the input query and identify semantic concepts related to the input query (406). Based on the identified semantic concepts, a preliminary set of images is obtained from the database containing images annotated with semantic information (804), A tokenizer is used to compare each image with a predefined label space to extract the associated concepts for each image in the preliminary set (806), A ranked list of the related concepts is generated and presented to the user based on their frequency of appearance within the preliminary set of images (810). Based on the user's selection of a specific related concept from the ranked list, the input query is combined with the selection to narrow down the preliminary set of images (812), Further semantic analysis is performed iteratively until the threshold condition is met (814). A computer program product that causes (814) to obtain an additional set of images that are semantically similar to the combined input query and the selection of the particular related concept.
16. The computer program product according to claim 15, wherein the preliminary set of images is a set of the top k images from the database, where k is a predetermined number, and the images are ranked according to a determined semantic similarity to the input query.
17. The computer program product according to claim 15, wherein the input query is text, and the semantic analysis includes processing by a language encoder component of the VL model to find semantically similar images based on text concepts in the predefined label space.
18. The computer program product according to claim 15, wherein the input query is an image, and the semantic analysis includes processing by a visual encoder component of the VL model to map the visual concepts of the input image to textual concepts in the predefined label space.
19. The computer program product according to claim 15, wherein the hardware processor improves the accuracy of concept extraction by narrowing down the predefined label space by removing redundant and non-identifiable labels.
20. The computer program product according to claim 15, wherein the hardware processor uses a feedback loop to iteratively refine the search for the input query based on the user's selection of the additional relevant concepts.