Electronic device for searching for images and methods thereof
By analyzing and reconstructing user queries to align with AI models, the electronic device enhances the accuracy of image search results, addressing the issue of biased outcomes in existing AI systems.
Patent Information
- Application Number
- PCT/KR2024/021146
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-24
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-31
AI Technical Summary
Existing artificial intelligence models for image search often provide biased or inaccurate results due to data format inconsistencies, failing to accurately reflect all user-entered criteria.
An electronic device performs word and sentence analysis on user queries, reconstructing their format, structure, and content to enhance accuracy by transforming the query before inputting it into an AI model for image search.
This process improves the accuracy of image search results by aligning user intent with the AI model's output, ensuring that search results reflect the user's intended criteria more effectively.
Smart Images

Figure KR2024021146_31072025_PF_FP_ABST
Abstract
Description
Electronic devices for searching images and methods thereof
[0001] The present invention relates to an electronic device and method for searching an image.
[0002] As artificial intelligence technology advances and becomes more widespread, artificial intelligence models are being used in various fields.
[0003] To improve the accuracy of answers provided by AI models, it's crucial to understand the user's intent. Specifically, when using AI models for image search, accurate search results must reflect all user-entered criteria. However, depending on the data format used for training, the model's search performance can be biased, or search results can only reflect one of the user's entered criteria, resulting in reduced accuracy.
[0004] Accordingly, the need for a method to provide more accurate search results has arisen.
[0005] According to at least one embodiment of the present disclosure, an electronic device includes an interface for receiving a user query, a memory storing an artificial intelligence model, and a processor.
[0006] The processor performs at least one of word analysis and sentence analysis on the user query input through the interface, reconstructs the user query by transforming at least one of the format, structure, and content of the user query based on the analysis result, and inputs the reconstructed user query into the artificial intelligence model to search for an image corresponding to the user query.
[0007] In addition, a method for searching for an image using an electronic device includes the steps of receiving a user query, performing at least one of word analysis and sentence analysis on the user query, and reconstructing the user query by converting at least one of a format, a structure, and a content of the user query based on the analysis result, and inputting the reconstructed user query into an artificial intelligence model to search for an image corresponding to the user query among a plurality of images stored in the electronic device.
[0008] FIG. 1 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.
[0009] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.
[0010] FIG. 3 is a diagram illustrating a software module according to at least one embodiment of the present disclosure.
[0011] FIG. 4 is a diagram illustrating an image search method using an artificial intelligence model according to at least one embodiment of the present disclosure.
[0012] FIG. 5 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to determine the importance of a word.
[0013] FIG. 6 is a diagram illustrating that the accuracy of image search is improved after query transformation according to at least one embodiment of the present disclosure.
[0014] FIG. 7 is a diagram illustrating a case where an image search is performed in a server device according to at least one embodiment of the present disclosure.
[0015] FIG. 8 is a flowchart illustrating an image search method according to at least one embodiment of the present disclosure.
[0016] FIG. 9 is a flowchart illustrating a method for determining the importance of words included in a query according to at least one embodiment of the present disclosure.
[0017] The terms used in the various embodiments of this disclosure have been selected from widely used, current terms, taking into account the functions of this disclosure. However, these terms may vary depending on the intentions of those skilled in the art, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the description of the relevant disclosure. Therefore, the terms used in this disclosure should be defined based on the meaning of the terms and the overall content of this disclosure, rather than simply their names.
[0018] In this disclosure, expressions such as “has,” “can have,” “includes,” or “may include” indicate the presence of a corresponding feature (e.g., a component such as a number, function, operation, or part), and do not exclude the presence of additional features.
[0019] The expression "at least one of A and / or B" should be understood to mean either "A" or "B" or "A and B".
[0020] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.
[0021] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that the component may be directly coupled to the other component, or may be connected through another component (e.g., a third component).
[0022] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this disclosure, terms such as "comprise" or "comprises" are intended to indicate the presence of a feature, number, step, operation, component, part, or combination thereof described in the specification, but should be understood not to preclude the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0023] In the present disclosure, a "module" or "part" performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Furthermore, multiple "modules" or multiple "parts" may be integrated into at least one module and implemented as at least one processor (not shown), excluding any "modules" or "parts" that need to be implemented as specific hardware.
[0024] In this disclosure, the term user may refer to a person using an electronic device or a device used by the person.
[0025] An embodiment of the present disclosure will be described in more detail with reference to the attached drawings below.
[0026] FIG. 1 is a drawing for explaining the operation of an electronic device according to at least one embodiment of the present disclosure.
[0027] According to FIG. 1, the electronic device (100) can recognize a user query (10), analyze its contents, search for images that meet search conditions, and provide results.
[0028] Although FIG. 1 exemplifies a situation in which an electronic device (100) utters a user query (10), the user query (10) may be input in various ways. For example, the user query (10) may be input using input methods such as a keyboard, soft keyboard, joystick, or remote control in addition to voice recognition.
[0029] The electronic device (100) may be implemented as various types of electronic devices, such as a server device, a personal computer (PC), a laptop PC, a smartphone, a tablet PC, a set-top box, a TV, a kiosk, and other home appliances. If the electronic device (100) is implemented as a server device or a set-top box, an external display device may also be connected and used. For convenience of explanation, the present specification will describe a case where the electronic device (100) is implemented as a smartphone equipped with a display.
[0030] A user query (10) may be a command that a user inputs into an electronic device (100). The user query (10) may be described by various terms such as prompt, user input, search request, etc., but in the following description, it will be described as 'query' or 'user query'.
[0031] When an electronic device (100) receives a query (10) from a user to search for an image, it performs at least one of word analysis and sentence analysis on the input user query (10).
[0032] The electronic device (100) reconstructs the user query (10) by transforming at least one of the form, structure, and content of the user query (10) based on the analysis result, and inputs the reconstructed user query (10) into an artificial intelligence model to search for an image corresponding to the user query. In the present disclosure, reconstructing may be an operation of editing the user query (10) to increase the accuracy of the image search intended by the user. Reconstructing may also be described as editing, modifying, converting, etc., but is described as reconstructing in the present disclosure.
[0033] The image being searched may be an image already stored in the terminal device, or an image uploaded to the Internet and stored on a server device.
[0034] For example, when a user inputs a query (10) such as 'Find a picture of a man wearing white clothes on a mountain' into an electronic device (100), the electronic device (100) can extract keywords and semantic units such as 'mountain', 'white clothes', and 'man' through word analysis, and can determine through sentence analysis that the query (10) is not a compound sentence and is written in a realistic style.
[0035] Based on the analysis results, the electronic device (100) can reconstruct the initially input query (10) 'Find a picture of a man wearing white on a mountain' into a sentence in the form of 'man wearing white' + 'mountain' or 'man wearing white' - 'mountain'. Through the above-described process, the electronic device (100) can provide more accurate search results because it reconstructs the sentence by removing unnecessary parts for search in the user query (10), changing the order, or adding some words. The method for determining which components among the format, structure, and content of the user query (10) the electronic device (100) should change and other specific examples will be described later based on other drawings.
[0036] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to at least one embodiment of the present disclosure.
[0037] According to FIG. 2, the electronic device (100) includes an interface (110), a processor (120), and a memory (130).
[0038] An interface (110) is a configuration created for interaction between two or more systems, devices, programs, or users. The interface (110) may include at least one of a communication interface, an input / output interface, and a user interface.
[0039] A communication interface is a component that communicates with external devices via wired or wireless communication. The communication interface can receive user queries from external devices such as terminals, wireless speakers, remote controls, and microphones. The communication interface may also be referred to as a communication unit.
[0040] The input / output interface is a configuration for transmitting or receiving various signals, data, etc. from a device connected via a wire. When an external microphone is connected through the input / output interface, the electronic device (100) can receive a user query (10) input to the external microphone through the input / output interface. The input / output interface may include various ports, such as a USB port or HDMI port, for connecting to various external devices, such as a microphone, keyboard, joystick, etc. The input / output interface may also be referred to as a connection port.
[0041] The user interface is a configuration for directly receiving various user commands from the user. The user interface can be implemented using a touch screen, a touch pad, buttons, etc. For example, if implemented using a touch screen, the user can directly input a user query (10) by drawing by touching the touch screen with his or her hand or a touch pen, or can input a user query (10) using a soft keyboard displayed on the touch screen.
[0042] In the present disclosure, a user can input a query into an electronic device (100) through various types of interfaces (110) described above.
[0043] The processor (120) is configured to control the overall operation of the electronic device (100).
[0044] The processor (120) may include one or more of a digital signal processor (DSP), a microprocessor, a central processing unit (CPU), a micro controller unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), an ARM processor, or an artificial intelligence (AI) processor, or may be defined by the terms thereof. In addition, the processor (120) may be implemented as a system on chip (SoC) or large scale integration (LSI) having a built-in processing algorithm, or may be implemented in the form of a field programmable gate array (FPGA). The processor (120) may perform various functions by executing computer executable instructions stored in the memory (130).
[0045] The memory (130) is a configuration for storing various software, commands, control codes, and data required for the operation of the electronic device (100). The memory (130) may be implemented as at least one of various memories, such as DRAM (dynamic RAM), SRAM (static RAM), SDRAM (synchronous dynamic RAM), OTPROM (one time programmable ROM), PROM (programmable ROM), EPROM (erasable and programmable ROM), EEPROM (electrically erasable and programmable ROM), mask ROM, flash ROM, flash memory, hard drive, or solid state drive (SSD).
[0046] According to at least one embodiment of the present disclosure, the memory (130) may store at least one artificial intelligence model. Specifically, the memory (130) may store an artificial intelligence model for image retrieval. Alternatively, the memory (130) may further store various artificial intelligence models for performing analysis operations on user queries, operations for grouping and transforming user queries, and the like. In some embodiments, when these operations are implemented rule-based rather than artificial intelligence models, the memory (130) may store data on various databases, operation algorithms, and the like used for the analysis operations, grouping operations, transformation operations, and the like.
[0047] The processor (120) performs various operations using the memory (130).
[0048] Specifically, the memory (130) may store data on a transformation algorithm that matches multiple types to classify images and queries that are the subject of a search. Through this, the processor (120) may analyze and classify the input user query (10). The user query (10) may be classified into various types. For example, the user query may be classified into the first to fourth types.
[0049] Among these, Type 1 refers to a user query that describes only simple facts, Type 2 refers to a user query that describes complex facts, and Type 3 refers to a user query that includes sentimental expressions. Type 4 refers to a user query that contains at least two of Types 1 through 3.
[0050] When a user query (10) is input, the processor (120) can perform a word analysis operation that identifies at least one of a keyword and a semantic unit included in the user query and a sentence analysis operation that identifies at least one of a sentence length of the user query (10), a sentence form of the user query (10), and a writing style of the user query. The processor (120) identifies the user query as one of the first type and the second type based on the sentence length and sentence form of the user query. Usually, the longer the sentence, the more likely it is to be classified as a compound sentence.
[0051] In addition, if at least one of the keywords, meaning units, and writing styles of the user query (10) corresponds to a sentimental expression, the processor (120) identifies the user query as a third type. Here, a sentimental expression refers to an expression of an abstract and sensory feeling. For example, the expression “There is a flower” is a realistic expression, but the expression “A pretty flower is blooming alone” includes an expression that represents the user’s abstract feeling, so the processor (120) classifies it as a sentimental expression.
[0052] As described above, when the classification task of the user query (10) is completed, the processor (120) converts the query according to the classified type.
[0053] If the user query (10) is of the first type that describes only simple facts, the processor (120) converts the format of the user query based on a conversion algorithm that matches the first type, and if the user query (10) is of the second type that describes complex facts, the processor (120) converts the structure of the user query based on a conversion algorithm that matches the second type. Here, a complex fact means a sentence in which two or more facts are combined.
[0054] In addition, if the user query (10) is of the third type including sentimental expressions, the processor (120) converts the content of the user query based on a conversion algorithm matching the third type. If the user query (10) belongs to at least two types among the first to third types, i.e., the fourth type, the processor (120) converts at least two of the format, structure, and content of the user query based on data stored in the memory (130).
[0055] At this time, converting the format of the user query (10) means dividing the sentences included in the user query into semantic units and reconstructing the sentences so that the colors and objects among the keywords included in the user query are displayed in a state in which they are connected to each other.
[0056] In addition, transforming the structure of a user query (10) means first listing the main semantic units with high importance based on the importance of each semantic unit extracted from the user query, and reconstructing the sentence so as to separate the relative distance between keywords corresponding to colors and keywords corresponding to places among the keywords included in the user query.
[0057] Additionally, changing the content of the user query (10) means reconstructing the sentence to add a higher concept word to identify at least one keyword among the keywords included in the user query.
[0058] For example, the processor (120) may add a preset first word to a keyword indicating a color among keywords included in the user query (10), add a preset second word to a keyword indicating an action, and add a preset third word to a keyword indicating a background. The processor (120) may reconstruct the user query by adding a preset fourth word to a keyword indicating an object excluding the background. The first to fourth words may be set as upper concept words indicating that the previous or subsequent word is a color, an action, a background, or an object, respectively.
[0059] As an example of a query that includes object and color information, assume that a user inputs the query 'A Photo of a man wearing a white shirt on the mountain'. At this time, the processor (120) can extract the keywords and semantic units 'man wearing a white shirt' and 'on the mountain'. A semantic unit can be the smallest unit that has a complete meaning. In the above-described query example, man, white, and shirt are extracted as keywords, and among them, white shirt can be extracted as a single semantic unit.
[0060] A keyword may consist of a single word, but is not necessarily limited thereto, and may also consist of multiple words containing one or more semantic units. For example, "man wearing a white shirt" may be extracted as a single keyword. After extracting each word within a user query (10), the processor (120) can identify keywords or semantic units based on the relationships and meanings between the words.
[0061] When converting the format of a user query, the processor (120) can group natural language sentences by meaning, or convert them into 'man wearing a white shirt' + 'on the mountain' by connecting symbols such as '+' or '-', or convert a color object query into a "color - object" format and reconstruct it into a form such as "man wearing a white shirt - on the mountain".
[0062] When transforming the structure of a user query, the processor (120) can transform the query into "man wearing a white shirt on the mountain" by placing information with high importance, such as keywords or semantic units, at the beginning of the sentence. In addition, the processor (120) can separate the keyword "white shirt" corresponding to color from the keyword "on the mountain" corresponding to location, thereby transforming the query into "on the mountain, man wearing a white shirt."
[0063] When converting the content of a user query, the processor (120) adds a preset first word when color information is specified. The first word may be a word of a superordinate concept encompassing specific colors. For example, 'color' may be used as the first word. The processor (120) may convert an existing sentence into a query that more clearly includes color information of the object to be found by reconstructing it into 'man wearing a white 'color' shirt on the 'general color' mountain. In addition, the processor (120) may add a preset third word for a keyword indicating a background, and a preset fourth word for a keyword indicating an object excluding the background. For example, words such as 'background' may be preset as the third word, and 'foreground' may be preset as the fourth word, and these may be stored in the memory (130). The processor (120) may add these words and reconstruct it into 'man wearing a white shirt 'foreground' on the mountain 'background', thereby converting the query so that the object and the background are more clearly distinguished.
[0064] When changing the format, structure, and content of a user-entered query, the query is reconstructed to include all of the above. In this case, by extracting and concatenating keywords and semantic units and attaching superordinate concepts to colors, objects, and backgrounds, the user query can be reconstructed as "man wearing white-colored shirt foreground" and "general-colored mountain background."
[0065] As another example, assume that a user enters a query such as 'A photo of a child is standing in the spray of water in a park', which includes object and action information. In this case, the processor (120) can extract 'child is standing', 'in the spray of water', and 'in a park' as keywords and semantic units.
[0066] When changing the format of a user query, the above-mentioned semantic units can be grouped and converted into 'a child is standing', 'in the spray of water', 'in a park', or they can be grouped into an action-object format and converted into a standing-child in the spray of water in a park.
[0067] When restructuring a user query, you can move important information earlier in the sentence, transforming it into 'a child is standing in a spray of water in a park', or physically distancing the action object query from the location query, transforming it into 'a standing child in the spray of water in a park'.
[0068] When changing the content of a user query, the processor (120) may add a preset second word to a keyword representing an action. The word “action,” which is a superordinate concept of various actions, may be used as the second word. The processor (120) may add this second word to convert it into “a standing ‘action’ child in the spray of water in a park.” Alternatively, as described above, the word “foreground” may be added to the object, and the word “background” may be added to the background, to convert it into “a child ‘foreground’ is standing in the spray of water in a park ‘background.’”
[0069] When changing the format, structure, and content of a user query, the processor (120) extracts and connects semantic units, and adds words corresponding to superordinate concepts to keywords representing actions, objects, and backgrounds, thereby converting the query into 'standing child foreground', 'in the spray of water background', and 'in a park background'.
[0070] The processor (120) can classify the type of the user query (10) and convert the format as described above using a rule-base.
[0071] Rule-based systems are systems or programs that perform operations based on predefined rules. They are algorithmic techniques that apply predefined rules through statistical analysis. Rule-based systems define both conditions and outcomes. In other words, if the predefined conditions are met, the action is performed.
[0072] For example, in a state where a database of words expressing unclear emotions is pre-stored in the memory (130), the processor (120) can group the user query into the third type if at least some of the keywords or semantic units included in the user query include at least one word from the database. On the other hand, if the words included in the database are not included in the user query, the processor classifies the user query into the first or second type.
[0073] Additionally, the processor (120) may group the user query into the first type if there is only one word expressed as a verb or adjective in the user query. On the other hand, if there are multiple verbs or adjectives, or if the query includes words such as “and, but, or” used as a conjunction between sentences, the processor (120) may group the user query into the second type corresponding to a complex fact description. In addition, if the query includes words for excluding or adding specific conditions during image search, such as “except, and additionally,” the processor (120) may group the query into the second type.
[0074] Alternatively, the processor (120) may classify the user query as a second type if the length of the user query is longer than a certain length, and as a first type if the length is shorter than the certain length. For example, if a user query such as 'A photo of a child is standing in the spray of water in a park' is input, a total of 15 words are extracted. If the classification criteria for classifying the first and second types are set to 10 words and stored in the memory (130), the processor (120) may group the input user query as a second type.
[0075] Alternatively, if the type is classified based on the number of verbs, the processor (120) may group the user query described above into the first type since the two words “is standing” constitute one verb.
[0076] In this way, even the same user query may be classified differently into various types depending on the classification criteria or classification algorithm.
[0077] The manufacturer of the electronic device (100) or the manufacturer of the artificial intelligence model can select the classification criterion or classification algorithm with the highest accuracy by referring to the search results obtained by reconstructing the query according to each classification method while applying various classification methods as described above.
[0078] The method of classifying and transforming queries is not limited to rule-based operations as described above. These operations can also be performed by an artificial intelligence model stored in memory (130). This will be described below with reference to FIG. 3.
[0079] The processor (120) performs a search using a query converted through the above-described process.
[0080] As an example of a method for the processor (120) to search for an image, a method of comparing the similarity of feature vectors of images can be used. Specifically, the processor (120) can extract a feature vector representing a caption of an image or a feature of the image itself. The processor (120) can match an image and a feature vector and store them together in the memory (130). When a user query is input, the processor (120) extracts a feature vector of an image that the user query is to search for, and then compares the extracted feature vector with the feature vectors of images stored in the memory (130) to calculate the similarity. A feature vector is a vector representing features extracted from data, and may include information such as pixel values and color histograms of an image.
[0081] The formula for calculating similarity is as follows:
[0082]
[0083] In the formula for calculating similarity, and means the transformed feature vector, Is and It means the angle formed by .
[0084] At this time, since the feature vector does not have a - value, the similarity has a value between 0 and 1.
[0085] The processor (120) determines that if the similarity is 1, the two vectors are completely identical vectors, and if the similarity is 0, the two vectors are uncorrelated vectors.
[0086] The processor (120) ranks images using the calculation result values so that images with higher similarity are positioned higher, and provides the user with search results with images with higher rankings first. The method for searching images is not limited to the method using the processor (120), and there is also a method using an artificial intelligence model. This will be described later.
[0087] FIG. 3 is a diagram illustrating a software module according to at least one embodiment of the present disclosure. The software module may be stored in memory (130).
[0088] According to FIG. 3, the software module may include a module for analyzing queries (111), a module for grouping queries (112), and a module for transforming queries (113).
[0089] The query analysis module (111) analyzes words in a user query to extract keywords and semantic units, and analyzes sentences to determine whether they are compound sentences and to analyze the writing style.
[0090] The query grouping module (112) classifies queries into Type 1 to Type 4 based on the results of query analysis. The query types and specific classification methods have been described above.
[0091] The query conversion module (113) changes at least one of the query format, structure, and content. The process by which the processor (120) receives and converts a query is described above based on FIG. 2.
[0092] In Fig. 3, a case is illustrated where multiple software modules (111, 112, 113) are configured to sequentially perform operations such as analyzing a user query, grouping the types, and then converting them according to the types. However, according to another embodiment, the type grouping process may be omitted. That is, the processor (120) may analyze a user query and then immediately convert at least one of the format, structure, and content according to the analysis result.
[0093] The processor (120) may perform a query transformation process using an artificial intelligence model in addition to one based on a rule-based method.
[0094] An artificial intelligence model is a computer system that learns and infers based on data, recognizes patterns, and automatically performs various tasks such as prediction, classification, and generation.
[0095] The functions related to artificial intelligence according to the present disclosure can be operated through a processor (120) and a memory (130).
[0096] The processor (120) may be composed of one or more processors. In this case, one or more processors may be implemented as a general-purpose processor such as a CPU, AP, DSP, etc., a graphics-only processor such as a GPU, VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU.
[0097] If the processor (120) is a processor dedicated to artificial intelligence, it may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0098] At least one artificial intelligence model among a first artificial intelligence model that analyzes words and sentences of a query, a second artificial intelligence model that classifies queries by type, a third artificial intelligence model that converts queries, and a fourth artificial intelligence model that handles image search may be stored in the memory (130). The processor (120) may perform at least one of the various operations described above using at least one artificial intelligence model stored in the memory (130).
[0099] Specifically,
[0100] The processor (120) uses an artificial intelligence model to perform word analysis and sentence analysis of the user query and to reconstruct the user query as described above.
[0101] Additionally, the processor (120) inputs the reconstructed user query into the fourth artificial intelligence model to search for an image corresponding to the reconstructed user query among a plurality of images stored in the memory (130), thereby deriving more accurate search results.
[0102] FIG. 4 is a diagram illustrating an image search method using an artificial intelligence model according to at least one embodiment of the present disclosure.
[0103] According to FIG. 4, the artificial intelligence model is stored in the memory (130) in a state in which learning is completed using a large amount of learning data. Here, being created through learning means that a basic artificial intelligence model is learned using a learning algorithm using a large amount of learning data, thereby creating a predefined operation rule or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed in the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above. The manufacturer of the electronic device (100) or the provider of the artificial intelligence model may input a plurality of images and their caption data or feature vectors as learning data into the artificial intelligence model, and then provide feedback on the output result, thereby training the artificial intelligence model.
[0104] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations by calculating the results of previous layers and the multiple weights. The multiple weights of the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated during the learning process to reduce or minimize the loss or cost values obtained by the artificial intelligence model.
[0105] Artificial neural networks may include deep neural networks (DNNs), such as, but not limited to, convolutional neural networks (CNNs), deep neural networks (DNNs), recurrent neural networks (RNNs), restricted boltzmann machines (RBMs), deep belief networks (DBNs), bidirectional recurrent deep neural networks (BRDNNs), generative adversarial networks (GANs), or deep Q-networks.
[0106] As described above, training data for training an artificial intelligence model can be configured in the form of pairs of images and captions. Training of the artificial intelligence model can be performed using the processor (120) of an electronic device (100) equipped with the artificial intelligence model, or by the processor of a learning device separately provided for training the artificial intelligence model. For convenience of explanation, the following description will be based on the case where the processor (120) performs training.
[0107] Captions can be input by the learner directly observing the image, or can be automatically acquired from the image using a separate program. When an image to be learned and a caption for the image are input (410), the processor (120) performs at least one of word analysis and sentence analysis on the input caption, and based on the analysis result, transforms at least one of the format, structure, and content of the caption to reconstruct (transform) the caption (420). The method for reconstructing the caption can be performed according to an algorithm identical or similar to the above-described reconstructing of the user query.
[0108] The processor (120) can train an artificial intelligence model (400) using training data including images and reconstructed captions. This allows for more accurate search results by matching the format of the learned captions with the format of the user-entered and reconstructed query.
[0109] When a user query is input (460) while the artificial intelligence model learned in this way is installed in an electronic device (100), the processor (120) reconstructs the query (450) through the classification and conversion process described above and inputs it into the artificial intelligence model (400). The artificial intelligence model (400) derives an image that matches or has a high degree of similarity to the image the user is looking for among the stored images (440) as a search result value (430).
[0110] One method for an artificial intelligence model to search for images can be through Visual Language Pretraining (VLP). This involves pretraining a model using a large amount of image-caption data. The trained artificial intelligence model receives an image as input, extracts key features of the image, and generates a caption (text) associated with the extracted image features. Captions for image data previously stored in memory (130) can be manually entered by the user, but the artificial intelligence model can also automatically generate captions through the aforementioned process. When a user inputs a query, the processor (120) calculates the similarity between the query and the caption for the image and derives a search result.
[0111] Meanwhile, as described in the above section, if the processor (120) needs to change the structure of a user query, the processor (120) may adjust the order based on the importance of each keyword or semantic unit included in the user query. The importance may be determined in various ways.
[0112] FIG. 5 is a diagram illustrating a method for an electronic device according to at least one embodiment of the present disclosure to determine the importance of a word.
[0113] When implemented in the form of a mobile phone as shown in FIG. 5, the electronic device (100) may further include a display in addition to the configuration of FIG. 2. When a user query is input, the electronic device (100) may control the display to display a UI screen (510) including an area (511) that displays the input user query. By looking at the UI screen (510), the user can intuitively check whether the user query he or she inputted was properly recognized. If the user query displayed in the area (511) is different from the query he or she uttered, the user may re-enter the query or directly modify the user query by touching the area (511). If the input user query is correct, the user may also press a confirmation button. Menus or confirmation buttons for re-entry or query modification may be additionally displayed on the UI screen (510), but are omitted in FIG. 5. FIG. 5 illustrates a state where a user has input the query 'man wearing a white shirt on the mountain'.
[0114] When the user presses the confirmation button or a certain amount of time passes without any selection, the electronic device (100) can control the display to display a new UI screen (520) including an area (521, 522) for displaying various words extracted from the input user query.
[0115] While viewing words (521), the user can sequentially select a selection area (522) corresponding to each word (521), or select at least one word deemed important. The processor (120) may use all words selected by the user as important words, or may assign different levels of importance depending on the order in which the user selected the words. For example, if the word "mountain" is selected first, the processor (120) may identify that word as the most important keyword.
[0116] In this way, when important words are selected from among the words included in the query, the processor (120) controls the display to display a UI screen (530) including an area (531) indicating the selected words and their importance, and menus (532, 533) for confirming or canceling the importance. Fig. 5 shows a state in which the user has selected words in the order of mountain, white shirt, and man.
[0117] Since the user selected mountain first, mountain is ranked first in the importance ranking results. The user can check the importance ranking results, and if satisfied, press the Confirm button (532), or if dissatisfied, press the Cancel button (533). If the Confirm button (532) is pressed, the processor (120) stores the setting value for the importance of the word in the memory (130). The processor (120) reconstructs the query according to the importance and searches for the image. If the Cancel button (533) is pressed, the screen returns to where the user can select important words again. Although FIG. 5 illustrates a case where only some of the words included in the input user query are displayed as important word candidates, according to another embodiment, the processor (120) may display all of the words that make up the user query. For example, various particles such as “on” and “the”, as well as prepositions, indefinite articles, and adverbs, may all be displayed. The user may select important keywords from among these words. When configured in this way, the keyword extraction module or semantic unit extraction module of Fig. 3 may be omitted or simplified.
[0118] In addition, although FIG. 5 describes a case where a user selects an important word from a UI screen (520) that displays important word candidates extracted from a user query, the user may also select an important word from the user query itself. That is, in the example of FIG. 5, the electronic device (100) may identify an important word selected by the user from among the user queries displayed in an area (511) on the first UI screen (510) and then immediately display a third UI screen (530) that displays the importance of each word. In this case, the second UI screen (520) may be omitted.
[0119] The electronic device (100) can provide more accurate search results by receiving feedback on important words (keywords) from the user.
[0120] FIG. 6 is a diagram illustrating that the accuracy of image search is improved after query transformation according to at least one embodiment of the present disclosure.
[0121] According to FIG. 6, the left photo (610) shows the image search result when a user inputs the query 'man wearing a white shirt on the mountain' in a conventional electronic device. According to FIG. 6, a case is shown where a photo of a man (612) standing on a snow-covered mountain (611) is searched. Although the user included not only a man standing on a mountain but also a man wearing a white shirt as search conditions, the accuracy of the search was not high, so it can be seen that results including the conditions desired by the user were provided. In addition, an error occurred in which a photo with a white snow-covered mountain in the background was provided as a search result when a white shirt should be searched for.
[0122] The right side shows a search result photo (620) obtained by reconstructing a query according to the query modification and transformation method described above in an electronic device (100) according to at least one embodiment of the present disclosure. By reconstructing the query as 'man wearing white color shirt foreground', 'general color mountain background', unnecessary parts for the search are removed, and the color of the clothes worn by the object, the man, and the color of the mountain in the background are clearly distinguished, thereby providing the user with accurate search results.
[0123] In the above various embodiments, the electronic device (100) has been described as receiving a user query, reconstructing it, and then performing an image search. However, the user query reconstructing operation or the image search operation may also be performed by an external device connected to the electronic device (100). For example, when a server device communicates with the electronic device (100), the server device may receive a user query provided by the electronic device (100) and perform the above-described operations.
[0124] FIG. 7 is a diagram illustrating a case where an image search is performed in a server device according to at least one embodiment of the present disclosure.
[0125] According to FIG. 7, the electronic device (100) can access the server device (200) through a web search via the Internet. When the electronic device (100) is connected, the server device (200) provides page data for composing a web page to the electronic device (100) via the Internet, and the electronic device (100) can display the web page based on the received page data.
[0126] A user can search for a desired image on a web page. That is, when a user inputs a query, the server device (200) can reconstruct the user query as described above, and then search for images matching the reconstructed user query from the server device (200)'s own database or a plurality of external devices (700-1 to 700-n) connected via the Internet. For example, if the server device (200) is a server device operated by the Patent Office, the user can input a user query to search for a design. The server device (200) can reconstruct the user query as described above, and then search for and provide designs matching the reconstructed user query from among the designs stored in its own database.
[0127] Alternatively, if the server device (200) is a web portal server, the server device (200) may reconstruct a user query and then transmit the reconstructed query to various connected external devices (700-1 to 700-n). Accordingly, when image data matching the reconstructed query is searched for in each of the external devices (700-1 to 700-n), the server device (200) may provide some content and link information of the web pages of the external devices to the electronic device (100). Alternatively, the electronic device (100) may perform the work of reconstructing the user query, and the image search may be performed in the server device (200). In this case, the electronic device (100) may not be equipped with an artificial intelligence model for image search.
[0128] FIG. 8 is a flowchart illustrating an image search method according to at least one embodiment of the present disclosure.
[0129] According to FIG. 8, the electronic device can receive a user query (S810), and as described above, the user query can be input in various ways.
[0130] The electronic device performs at least one of word analysis and sentence analysis on an input user query (S820), and reconstructs the user query (S840) by converting at least one of the format, structure, and content of the user query based on the analysis result (S830).
[0131] The electronic device inputs the reconstructed user query into an artificial intelligence model and searches for an image corresponding to the user query among a plurality of images stored in the electronic device (S850).
[0132] Specific examples of this have been described above, so duplicate descriptions are omitted.
[0133] The control method of FIG. 8 can be performed by an electronic device (100) having the configuration described in FIGS. 1 and 2, but is not necessarily limited thereto, and can also be performed by a device having a different configuration from that of FIG. 2.
[0134] FIG. 9 is a flowchart illustrating a method for determining the importance of words included in a query according to at least one embodiment of the present disclosure.
[0135] According to Fig. 9, the importance of words included in a user query can be directly selected by the user to provide more accurate search results.
[0136] Specifically, when a user query is input (S910), the electronic device can display words extracted from the user query (S920). In this state, when a user selection for important words is input (S930), the importance of each word is determined (S940) and stored (S950) based on the user selection.
[0137] The electronic device can perform actions such as rearranging the order of words within a query based on their importance. A detailed description of this has been provided in the above section, so a duplicate description will be omitted.
[0138] The control method of FIG. 9 can be performed by an electronic device (100) having the configuration described in FIGS. 1 and 2, but is not necessarily limited thereto, and can also be performed by a device having a different configuration from that of FIG. 2.
[0139] The programs or instructions for performing the various information processing methods described above may be provided stored on a non-transitory, readable medium. The non-transitory, readable medium may be loaded and used in a device capable of recalling the instructions stored in the storage medium and performing operations according to the recalled instructions. Accordingly, when the program or instructions stored on the non-transitory, readable medium are executed by a processor, the processor may directly, or under the control of the processor, utilize other components to perform the operations described in the various embodiments described above.
[0140] A non-transitory computer-readable medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specific examples of non-transitory computer-readable media include CDs, DVDs, hard disks, Blu-ray discs, USBs, memory cards, and ROMs.
[0141] Instructions may include code generated or executed by a compiler or interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" means that the storage medium does not contain signals and is tangible, but does not distinguish between whether data is stored semi-permanently or temporarily on the storage medium.
[0142] Furthermore, according to one embodiment of the present disclosure, the method according to the various embodiments described above may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be implemented as a product distributed online through an application store, in addition to the non-transitory readable recording medium described above. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0143] Accordingly, in a non-transitory computer-readable medium storing one or more instructions executed by a processor of an electronic device to cause the electronic device to perform an operation, the operation of the medium includes the steps of receiving a user query, performing at least one of word analysis and sentence analysis on the user query, and reconstructing the user query by converting at least one of a format, a structure, and a content of the user query based on the analysis result, and inputting the reconstructed user query into an artificial intelligence model to search for an image corresponding to the user query among a plurality of images stored in the electronic device.
[0144] While the present invention has been described with reference to the attached drawings, the scope of the present invention is determined by the claims described below and should not be construed as being limited to the aforementioned embodiments and / or drawings. Furthermore, it should be clearly understood that improvements, modifications, and variations apparent to those skilled in the art, as defined in the claims, are also included within the scope of the present invention.
Claims
1. In electronic devices, An interface for receiving user queries; Memory where artificial intelligence models are stored; and Processor; including; The above processor, By performing at least one of word analysis and sentence analysis on the user query entered through the above interface, Reconstructing the user query by transforming at least one of the format, structure, and content of the user query based on the analysis results, An electronic device that inputs the reconstructed user query into the artificial intelligence model to search for an image corresponding to the user query.
2. In paragraph 1, The above memory is, Store data about the conversion algorithm that matches multiple types, The above processor, Identifying the type of the user query based on the above analysis results, If the user query is of the first type that describes a simple fact, the format of the user query is converted based on a conversion algorithm that matches the first type, If the user query is of the second type that describes a complex fact, the structure of the user query is transformed based on a transformation algorithm that matches the second type, If the user query is of the third type that includes a sentimental expression, the content of the user query is converted based on a conversion algorithm that matches the third type, An electronic device that converts at least two of the format, structure, and content of the user query based on data stored in the memory, if the user query belongs to at least two types among the first to third types.
3. In paragraph 1, The above processor, When an image to be learned and a caption for the image are input, at least one of word analysis and sentence analysis is performed on the caption, Reconstructing the caption by transforming at least one of the format, structure, and content of the caption based on the analysis results, An electronic device that trains the artificial intelligence model using training data including the image and the reconstructed caption.
4. In paragraph 2, The above processor, Performing a word analysis operation that identifies at least one of the keywords and semantic units included in the user query and a sentence analysis operation that identifies at least one of the sentence length of the user query, the sentence form of the user query, and the writing style of the user query, Identifying the user query as one of the first type and the second type based on the sentence length and sentence form, An electronic device that identifies the user query as the third type if at least one of the keyword, the semantic unit, and the style of the user query corresponds to the sentimental expression.
5. In paragraph 2, The above processor, If the above user query is of the first type, The sentences included in the user query are divided into semantic units, and the format of the user query is changed so that the keywords included in the user query are displayed in a state where colors and objects are connected to each other. If the user query is of the second type, based on the importance of each semantic unit extracted from the user query, the main semantic unit with high importance is described first, The structure of the user query is changed to separate the relative distance between keywords corresponding to color and keywords corresponding to place among the keywords included in the user query, An electronic device that changes the content of the user query to add a higher concept word for identifying at least one keyword among keywords included in the user query if the user query is of the third type.
6. In paragraph 5, including display; The above processor, Control the display to display words extracted from the user query, An electronic device that determines the importance of the displayed word based on a user selection for the displayed word and stores the determined importance in the memory.
7. In paragraph 1, The above processor, An electronic device that reconstructs the user query by adding a preset first word to a keyword indicating a color among keywords included in the user query, a preset second word to a keyword indicating an action, a preset third word to a keyword indicating a background, and a preset fourth word to a keyword indicating an object excluding the background.
8. In paragraph 1, The above memory further stores at least one artificial intelligence model and a plurality of images for processing the user query, The above processor, Using at least one artificial intelligence model, word analysis and sentence analysis for the user query and an operation of reconstructing the user query are performed, An electronic device that inputs the reconstructed user query into the artificial intelligence model and searches for an image corresponding to the reconstructed user query among a plurality of images stored in the memory.
9. In a method for searching an image using an electronic device, Step of receiving a user query; A step of performing at least one of word analysis and sentence analysis on the user query, and reconstructing the user query by transforming at least one of the format, structure, and content of the user query based on the analysis result; An image search method comprising: a step of inputting the reconstructed user query into an artificial intelligence model and searching for an image corresponding to the user query among a plurality of images stored in the electronic device; 10. In paragraph 9, The steps for reconstructing the above user query are: A step of identifying the type of the user query based on the analysis results; and If the user query is of the first type that describes a simple fact, the format of the user query is converted based on a conversion algorithm that matches the first type, If the user query is of the second type that describes a complex fact, the structure of the user query is transformed based on a transformation algorithm that matches the second type, If the user query is of the third type that includes a sentimental expression, the content of the user query is converted based on a conversion algorithm that matches the third type, An image search method, comprising: a step of converting at least two of the format, structure, and content of the user query based on stored data, if the user query belongs to at least two types among the first to third types.
11. In paragraph 9, A step of performing at least one of word analysis and sentence analysis on an image to be learned and a caption for the image when the image is input; A step of reconstructing the caption by transforming at least one of the format, structure, and content of the caption based on the analysis results; and An image retrieval method further comprising: a step of training the artificial intelligence model using learning data including the image and the reconstructed caption.
12. In paragraph 10, The step of identifying the type of the user query based on the above analysis results is: A step of performing a word analysis operation for identifying at least one of a keyword and a semantic unit included in the user query and a sentence analysis operation for identifying at least one of a sentence length of the user query, a sentence form of the user query, and a writing style of the user query; and Identifying the user query as one of the first type and the second type based on the sentence length and sentence form, An image search method, comprising: a step of identifying the user query as the third type if at least one of the keyword, the semantic unit, and the style of the user query corresponds to the sentimental expression.
13. In paragraph 9, The steps for reconstructing the above user query are: A step of identifying the type of the user query based on the analysis results; and If the above user query is of the first type, describing a simple fact, The sentences included in the user query are divided into semantic units, and the format of the user query is changed so that the keywords included in the user query are displayed in a state where colors and objects are connected to each other. If the above user query is of the second type describing a complex fact, Based on the importance of each semantic unit extracted from the user query, the main semantic unit with high importance is described first, and the structure of the user query is changed to separate the relative distance between the keyword corresponding to color and the keyword corresponding to place among the keywords included in the user query. If the above user query is of the third type containing sentimental expressions, An image search method comprising a step of changing the content of the user query to add a higher concept word for identifying at least one keyword among the keywords included in the user query.
14. In paragraph 13, A step of displaying words extracted from the above user query; An image search method further comprising: a step of determining and storing the importance of the word based on a user selection for the displayed word; 15. In paragraph 9, The steps for reconstructing the above user query are: An image search method comprising the steps of: adding a preset first word to a keyword indicating a color among keywords included in the user query; adding a preset second word to a keyword indicating an action; adding a preset third word to a keyword indicating a background; and adding a preset fourth word to a keyword indicating an object excluding the background.
Citation Information
Patent Citations
Method and device for emphasizing and displaying image in order of adaptability
JP1996287086A
Picture retrieval method, device therefor and record medium recording picture retrieval program
JP2000076287A
Data retrieval system, data retrieval method, and computer program
JP2008146219A
Image retrieval method, image retrieval device and image retrieval program
JP2016006668A
Machine learning program, retrieval program, machine learning apparatus, and method
JP2023056798A