Image retrieval method, image generation method, image retrieval program, image retrieval apparatus, and image retrieval system

The image search method addresses the inefficiency of manual tagging by using machine learning to generate models that accurately identify desired images, enhancing search capabilities in large image collections.

JP2025139470APending Publication Date: 2025-09-26DENSO TEN LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024038427
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-12
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Adding tags to images stored on a device is time-consuming and becomes increasingly difficult as the number of images increases, leading to frequent inability to search for desired images effectively.

Method used

An image search method that uses a computer to search for images matching user-input search words by collecting and performing similarity searches with external server images, generating trained models to accurately determine image content using machine learning.

Benefits of technology

Enables efficient and accurate image searches by generating models that can determine the presence of specific objects in images, improving search efficiency and reducing the need for manual tagging.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025139470000001_ABST
    Figure 2025139470000001_ABST
Patent Text Reader

Abstract

To provide an image retrieval method, an image generation method, an image retrieval program, an image retrieval apparatus, and an image retrieval system which allow a user to suitably execute the image retrieval of a desired image.SOLUTION: An image retrieval method according to the present application is used for retrieving a storage image which matches a retrieval word inputted by a user among storage images, in which images which match a retrieval word are collected from an external server, and a storage image which matches the retrieval word is outputted by similarity retrieval with the collected images.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an image search method, an image generation method, an image search program, an image search device, and an image search system. [Background technology]

[0002] Conventionally, there has been known a technique for searching for a camera image (hereinafter referred to as a desired image) that shows a specific object (such as an object or background, hereinafter referred to as a desired object) that a user wants to view from among a large number of camera images (hereinafter referred to as images stored in the terminal) stored in a terminal device such as a smartphone. This technique proposes a technique that allows a search for an image stored in the terminal that shows the desired object by adding a search tag related to the object to each image stored in the terminal and searching the tag by the name of the desired object, etc. (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Application Publication No. 2018-55730 Summary of the Invention [Problem to be solved by the invention]

[0004] However, adding tags to images stored on a device is time-consuming and becomes increasingly difficult as the number of images stored on a device increases, resulting in frequent inability to search for a desired image from among the images stored on a device.

[0005] The present application has been made in consideration of the above, and aims to provide an image search method, an image generation method, an image search program, an image search device, and an image search system that allow users to preferably perform image searches for desired images. [Means for solving the problem]

[0006] The image search method of the present invention is an image search method executed by a computer that searches among stored images for images that match a search word entered by a user, collects images that match the search word from an external server, and outputs the stored images that match the search word by performing a similarity search with the collected images. [Effects of the Invention]

[0007] In the present invention, an appropriate image search for an input search word can be performed by similarity search for images having characteristics of image data to be extracted using the input search word. [Brief explanation of the drawings]

[0008] [Figure 1A] FIG. 1A is a diagram illustrating an example of the configuration of an image search system according to an embodiment. [Figure 1B] FIG. 1B is a diagram showing a specific example of processing executed in the image search system according to the embodiment. [Figure 2] FIG. 2 is a diagram illustrating an example of the configuration of a terminal device according to the embodiment. [Figure 3] FIG. 3 is a diagram showing an example of the saved image information. [Figure 4] FIG. 4 is a diagram illustrating an example of model information. [Figure 5] FIG. 5 is a flowchart showing the processing performed by the controller. [Figure 6] FIG. 6 is a diagram illustrating a learning process for multiple objects. [Figure 7] FIG. 7 is a diagram illustrating a learning process for multiple objects. [Figure 8] FIG. 8 is a diagram showing another example of an image search system. DETAILED DESCRIPTION OF THE INVENTION

[0009] Hereinafter, embodiments of an image search method, an image generation method, an image search program, an image search device, and an image search system will be described in detail with reference to the accompanying drawings. Note that the present invention is not limited to the following embodiments.

[0010] First, an image search system according to an embodiment will be described with reference to Figures 1A and 1B. Figure 1A is a diagram showing an example of the configuration of the image search system according to an embodiment. Figure 1B is a diagram showing a specific example of processing executed in the image search system according to an embodiment.

[0011] 1A, an image search system S according to an embodiment includes a terminal device 1 that functions as an image search device, and a search server 100. The terminal device 1, which includes a storage unit that stores images stored in the terminal, and the image search device that searches for a desired image that a user wishes to view from among the images stored in the terminal, are not necessarily configured as an integrated unit, and each device may be configured separately. The image search method according to an embodiment is executed by the terminal device 1 (image search device).

[0012] The terminal device 1 is a terminal device carried by a user, and can be any type of terminal device such as a smartphone, a desktop PC, a notebook PC, or a tablet PC.

[0013] The terminal device 1 stores a plurality of terminal-saved images in its own storage unit 4 (storage device). The terminal-saved images include, for example, images captured by a camera function (not shown) of the terminal device 1, images downloaded from a website, and images downloaded from images posted by other users on a social networking service (SNS). The terminal-saved images are images that depict the real world, but are not limited to this and may also be illustration images such as anime images.

[0014] The terminal device 1 includes a controller 3. The controller 3 searches (extracts) a desired image containing a desired object (in which the desired object is shown) from among a plurality of terminal-stored images. Specifically, the controller applies each terminal-stored image to an AI (artificial intelligence) model that determines whether the desired object is included, and obtains a determination result indicating whether the desired object is included in each terminal-stored image (or the probability that a specific object is included). Then, based on the determination result, the controller determines whether each terminal-stored image is a desired image, and searches (outputs) the desired image.

[0015] The search server 100 is a server device that provides a search service. When the search server 100 receives a search word input from an external device, it provides search results that match the search word. In the present disclosure, the search server 100 provides images (images showing objects that match the search word) as search results. Note that instead of directly providing images, the search server 100 may provide information on various websites for acquiring the relevant images, such as information such as a URL indicating the link destination of the relevant image, as search results.

[0016] Next, an overview of an image search method according to an embodiment will be described with reference to Fig. 1A. When the terminal device 1 enters an image search state, for example, when a user performs an image search operation, the terminal device 1 displays a screen (described later in Fig. 1B) for searching for images stored in the terminal on a display (not shown).

[0017] The terminal device 1 (controller 3) receives from the user on the screen a search word (e.g., the name of a desired object) to search for a desired image (hereinafter referred to as a user-input search word).Then, the terminal device 1 transmits a search request to the search server 100, requesting a search process for images within the search server 100 or on the Internet using the user-input search word (step S1).

[0018] In the present disclosure, the terminal device 1 (controller 3) determines whether the models stored in the storage unit 4 include a model suitable for determining the search word entered by the user. If a suitable model is not included, the terminal device 1 (controller 3) sends a search request to the search server 100. If a suitable model is included, the terminal device 1 (controller) does not send a search request to the search server 100, but instead uses the model suitable for the determination to determine whether each terminal-stored image matches the search word, thereby searching for terminal-stored images (desired images) that match the search word. Note that the terminal device 1 (controller 3) may always send a search request to the search server 100 without performing these branching processes.

[0019] Next, the search server 100 performs a search process to search for images that match the search word in accordance with the search request, and outputs search results (images: referred to as server search result images) that match the search word. The terminal device 1 collects these server search result images as learning images (step S2). The terminal device 1 extracts a predetermined number of learning images suitable for learning from the server search result images; the extraction method will be described in detail later.

[0020] Next, the terminal device 1 (controller 3) generates training data using the collected training images as input data and the user-input search words as correct answer data. The terminal device 1 (controller 3) then performs machine learning on a model using the generated training data to generate a model that searches for device-stored images using the user-input search words (step S3) (hereinafter referred to as the trained model). Through this training, the trained model becomes a model that, when an image stored on the device is input, determines whether the image matches the user-input search words. The model here is configured, for example, to include a neural network, and outputs the probability that the input image corresponds to each class corresponding to each output layer (in the trained model described above, the input image corresponds to a class that matches the user-input search words and a class that does not match), and determines whether the input image matches the user-input search words by discriminating the probability using an appropriate threshold. In other words, the trained model outputs the accuracy of the image elements of the object indicated by the user-input search words in the input image. Note that details of the model generation method will be described later. Since this model discriminates the probability of the class of the desired object using an appropriate threshold value to determine whether the desired object is included in the image to be judged, the output layer may be reduced to one class indicating that it is the desired object by model compression processing (e.g., pruning, quantization, distillation, etc.).

[0021] Next, the terminal device 1 (controller 3) applies (inputs) each terminal-stored image to the generated trained model and determines whether each terminal-stored image matches the user-input search word, thereby searching for terminal-stored images that match the search word (step S4).Then, the searched terminal-stored images (that match the user-input search word) are provided to the user.

[0022] Next, a specific example of the image search method according to the embodiment will be described with reference to Fig. 1B. Fig. 1B shows an example in which the terminal device 1 displays a display screen for searching terminal-saved images stored in the storage unit 4 on the display, and the user enters "shiitake mushroom" in a search box for entering a search word.

[0023] In such a case, the terminal device 1 (controller 3) first determines whether or not a model capable of determining the presence or absence of the user-input search word "shiitake" (a model that has been trained with training data of images of "shiitake") exists among the trained models (image search models) stored in the storage unit 4. If a model capable of determining the presence or absence of "shiitake" exists, the terminal device 1 uses the model to search for terminal-stored images that include the image element of "shiitake."

[0024] On the other hand, if there is no model capable of determining the presence or absence of "shiitake mushroom" among the trained models stored in the storage unit 4, the terminal device 1 (controller) transmits a request signal (including a search command and search conditions (search word "shiitake")) to the search server 100 to cause the search server 100 to perform an image search for the user-input search word "shiitake mushroom," thereby causing the search server 100 to perform an image search for the user-input search word "shiitake mushroom." That is, the terminal device 1 makes a search request to the search server 100 (step S1). Next, the terminal device 1 acquires, as learning images, images that match the user-input search word "shiitake" (images that include an image of "shiitake mushroom" as an image element), which are the results of the image search in the search server 100. That is, the terminal device 1 collects learning images from the search server 100 (step S2).

[0025] The terminal device 1 (controller 3) then uses the collected training images as input data and trains the model using training data with the user-input search word "shiitake" as correct answer data, thereby generating a trained model capable of determining the presence or absence of "shiitake" (step S3). FIG. 1B illustrates a neural network as the model. The terminal device 1 trains the neural network using the training images to generate a model having an output layer class indicating the presence or absence of a new target, "shiitake." This output layer class outputs the probability that an object corresponding to the class is included in the terminal-stored image as a determination result. The terminal device 1 then selects the generated model as an image search model, inputs the terminal-stored image into the image search model, and determines the presence or absence of "shiitake" based on the determination result output from the image search model. That is, the terminal device 1 performs an image search using the image search model (step S4). The terminal device 1 then displays the terminal-stored images determined to include "shiitake" on the screen as search results.

[0026] As described above, in the image search method according to the embodiment, when a search word related to an object (for example, shiitake mushroom) is input, learning images that match the search word are collected from the search server 100 to generate learning data, and machine learning of the model is performed. As a result, a model that can accurately determine the presence or absence of the object is generated, and an image search for any object can be performed with high accuracy.

[0027] 1A shows an example in which the source of the learning images is the search server 100, but it is also possible to collect images generated by a generation AI (Artificial Intelligence) as learning images. This point will be described in detail later.

[0028] Next, an example of the configuration of the terminal device 1 will be described with reference to Fig. 2. Fig. 2 is a diagram showing an example of the configuration of the terminal device 1 according to an embodiment. Note that Fig. 2 shows components necessary for explaining the features of this embodiment, and general components are omitted. As shown in Fig. 2, the terminal device 1 includes a communication unit 2, a controller 3, a storage unit 4, and a camera 5. The terminal device 1 may be a so-called computer device. Note that the terminal device 1 may be configured to include an input device such as a keyboard and an output device such as a display.

[0029] The communication unit 2 is an interface for communicating data with other devices via a network, and is, for example, a network interface card (NIC).

[0030] The controller 3 includes a processor that performs arithmetic processing and the like. The processor may be configured to include, for example, a CPU (Central Processing Unit). The controller 3 may be configured with one processor or multiple processors. When configured with multiple processors, the processors may be connected to each other so that they can communicate with each other. Note that when the terminal device 1 (image search device) is configured as a cloud server, the CPU that configures the processor may be a virtual CPU.

[0031] In this embodiment, the functions of the controller 3 are realized by the processor executing arithmetic processing in accordance with a program stored in the storage unit 4.

[0032] The scope of this embodiment may include a computer program that causes a processor (computer) to realize at least some of the functions of the terminal device 1. The scope of this embodiment may also include a computer-readable nonvolatile recording medium that records such a computer program. The nonvolatile recording medium may be, for example, the nonvolatile memory described above, an optical recording medium (e.g., an optical disk), a magneto-optical recording medium (e.g., a magneto-optical disk), a USB memory, an SD card, or the like.

[0033] Furthermore, each function of the controller 3 may be realized by a single program, or may be realized by a separate program for each function. Each function may also be realized as a separate server. As described above, each function may be realized by having a processor execute a program, i.e., by software, but may also be realized by other methods. At least a portion of each function may be realized using, for example, an ASIC (Application Specific Integrated Circuit) or an FPGA (Field Programmable Gate Array).

[0034] The storage unit 4 includes a volatile memory and a nonvolatile memory. The volatile memory may include, for example, a random access memory (RAM). The nonvolatile memory may include, for example, a read-only memory (ROM), a flash memory, or a hard disk drive. The nonvolatile memory stores computer-readable programs and data. Note that at least some of the programs and data stored in the nonvolatile memory may be obtained from another computer device connected via a wired or wireless connection, or from a portable recording medium.

[0035] The camera 5 is an imaging device built into the terminal device 1. The camera 5 is capable of capturing not only still images but also moving images.

[0036] 2, in this embodiment, the storage unit 4 stores saved image information 41 and model information 42. Fig. 3 is a diagram showing an example of the saved image information 41. Fig. 4 is a diagram showing an example of the model information 42.

[0037] The saved image information 41 is information about an image captured by the camera 5. As shown in Fig. 3, the saved image information 41 includes items such as "image ID," "image information," and "tag information."

[0038] The "image ID" is image ID data, which is identification information for identifying the data record of each image in the saved image information 41. The image ID is also the primary key of the data record in the saved image information 41. In other words, in the saved image information 41, a data record is configured for each image ID data, and data for each item associated with the image ID data is stored in that data record.

[0039] "Image information" is information about the image captured by camera 5. "Tag information" is various information related to the image stored on the device, and is used for various processes such as search processing. For example, "tag information" is information indicating the object (type) contained (pictured) in the image stored on the device. For example, the image stored on the device identified by image ID "P1" is associated with tag information of "wedding". This indicates that the image stored on the device is an image related to a wedding. The "tag information" of the object contained in this image stored on the device is then used for conventional image searches (searches by specifying the name of the object contained in the image stored on the device), etc.

[0040] Here, for easier understanding, the image search function of the terminal device 1 in this embodiment will be described. The terminal device 1 stores a trained model suitable for searching for multiple items, etc. that are predicted to be frequently searched for in images stored in the terminal. This becomes the model for existing image search in the terminal device 1. In addition, a pre-trained model is stored for searching for objects, etc. that are not search targets of this existing model. Then, when a search for a desired object that cannot be searched for using the existing model is instructed, the pre-trained model is trained using training data consisting of images containing the desired object, and a trained model suitable for searching for the desired object is generated. After that, images containing the desired object are searched for using the trained model from the images stored in the terminal.

[0041] Note that data characterizing the trained model (such as weight values ​​of each layer in a neural network) may be stored, and when searching for images containing the same desired object again, the data characterizing the model may be read and set in the pre-trained model to generate a trained model suitable for searching for the desired object. Furthermore, if the performance of the trained model deteriorates (for example, if the shape of the object changes over time), it is preferable to update the trained model or the stored data characterizing the model by re-learning (for example, at an appropriate update learning period).

[0042] The model information 42 is information about the model used in the image search model 31. As shown in Fig. 4, the model information 42 includes items such as "model ID," "model parameter information," "type," and "class." The model information 42 is updated every time the controller 3 updates (or generates) a model.

[0043] The "model ID" is model ID data, which is identification information for identifying the data record of each model in the model information 42. The model ID is also the primary key of the data record in the model information 42. In other words, in the model information 42, a data record is configured for each model ID data, and data for each item linked to the model ID data is stored in that data record.

[0044] "Model parameter information" is information about the parameters of a model. Parameters are, for example, weight values ​​of each layer in a neural network. By setting the data of "model parameter information" in the pre-training model, a trained model that searches for images containing the relevant object (the object stored in the "class") is generated. The "model parameter information" is updated (or newly generated (if there is no model of the same class, new model information is newly generated)) when the controller 3 trains the model through machine learning.

[0045] "Type" is information indicating the type of model information, and stores either "Existing" indicating that the model information was originally included in the image search model 31, or "Added" indicating that the model information was newly added through learning using learning images. "Class" is information indicating the type of object to be searched when searching for images stored on the terminal.

[0046] In other words, the model identified by model ID "M1" is a model that determines whether any of "wedding ceremony," "bicycle," "mountain," etc. is included in the image stored on the device. Also, model ID "M2" is a model that determines whether "shiitake mushroom" is included in the image stored on the device. Furthermore, model ID "M5" is a model that determines whether both "banana" and "milk" are included.

[0047] Next, a description will be given of the process 10 performed by the controller 3. Fig. 5 is a flowchart showing the process performed by the controller 3. The flowchart shown in Fig. 5 is executed when an application for viewing terminal-stored images is launched in the terminal device 1 (when the user performs an image search start operation).

[0048] In step S101, the controller 3 displays a search window on the display, accepts (acquires) a search word (name of desired object) for a desired object entered by the user in the search window, and proceeds to step S102. The search word may be a word entered by the user when searching for a specific object among images stored in the device, or may be a word entered to generate a model in advance, as described below. In other words, the controller 3 accepts as search words words entered for search or for pre-learning. Note that the determination of whether the search word is for search or pre-learning can be made, for example, by displaying a search button or pre-learning button in the search window and determining which button is selected.

[0049] In step S102, the controller 3 searches for data records (model IDs) that contain the received search word (name of the desired object) in the "class" of the model information 42, and proceeds to step S103. In step S103, the controller 3 determines whether or not a corresponding data record has been detected. If a data record has been detected (Yes), the controller 3 proceeds to step S104; if not (No), the controller 3 proceeds to step S105. In step S104, the controller 3 selects the model information of the data record (model ID) that corresponds to the search word as the image search model (selects (regenerates) a trained model to be used for the search), and proceeds to step S110. Through this process, the image search model becomes a trained model that determines whether the desired object is included in an input image.

[0050] On the other hand, if there is no data record in the "class" of the model information 42 in which the object indicated by the search word exists, that is, if there is no model information that can be used to search for the desired object in images stored on the terminal, the controller 3 generates (learns) a model for searching for the object indicated by the search word.

[0051] Specifically, in step S105, the controller 3 transmits the received search word and a command to provide images containing the desired object corresponding to the search word to the search server 100, which is an external server, and proceeds to step S106. As a result, the search server 100 performs an image search using the search word and collects images containing the desired object that matches the search word from within the search server 100 or from other servers in cyberspace. The search server 100 then transmits (provides) the collected images to the terminal device 1. In step S106, the controller 3 receives the images containing the desired object transmitted from the search server 100, and proceeds to step S107. Since these received images are images containing the desired object, they can become learning data to which the desired object name has been assigned as tag data (correct answer data). In other words, the controller 3 collects learning images to which the correct answer data for the search word (desired object name) has been assigned from the search server 100.

[0052] In step S107, the controller 3 extracts a predetermined number of training images suitable for predetermined learning from the collected training images, and then proceeds to step S108. If the number of collected training images is less than the predetermined number, the controller 3 may generate extended training images by lightly processing the training images, such as enlarging or reducing them at various magnifications, slightly modifying the brightness or color, or slightly distorting the images, to the extent that the type of object in the image can be distinguished, and add these images to the training images. In step S108, the controller 3 trains a pre-training model using the extracted training images, and then proceeds to step S109. This generates a model that determines whether an image contains a desired object. In step S109, the controller 3 generates a new data record in the storage unit 4 for the model information of the created model, assigns a new model ID to it, and stores it as part of the model information 42. Then, the controller 3 proceeds to step S110. In other words, the controller 3 stores in the storage unit 4 model information of a newly generated model that searches for an image containing a desired object. In this embodiment, the additional model other than the existing model (model ID: M1) has one class (a model that outputs the probability that an image contains the desired object) or two classes (a model that outputs the probability that an image contains the desired object and the probability that an image does not contain the desired object). However, it is also possible to apply a model that outputs the probability that an image contains multiple types of objects (the number of classes is more than one). In this case, the controller 3 may perform model training by adding a "class" of a model to which a model ID has already been assigned in the model information 42. In other words, training may be performed by adding a class to the output layer of an existing model. After completing model training, the controller 3 updates the model information 42 stored in the storage unit 4 according to the newly generated model (or the newly added class).

[0053] In step S110, the controller 3 sets the trained model generated and stored in steps S108 and S109, or the trained model generated (selected) in step S104, as the image retrieval model 31, sequentially inputs terminal-stored images to the image retrieval model 31, determines whether each terminal-stored image is an image containing a desired object (image search), and proceeds to step S111. In step S111, the controller 3 displays the terminal-stored image determined in step S110 to be an image containing a desired object on the display of the terminal device 1 as an image containing the object desired by the user, and proceeds to step S112. Note that the determination of whether the terminal-stored image is an image containing the desired object can be made by discriminating the output of the trained model (the probability that the image is an image containing the desired object) using a preset determination threshold (if the output is equal to or greater than the determination threshold, the image is determined to be an image containing the desired object). In step S112, the controller 3 assigns tag information indicating the name of the object to the terminal-stored image determined to be an image containing the desired object, and ends the process. By assigning tags in this way, when the same search word is input next time or later, image search can be performed by tag search without using a model. Note that, since tag information is not assigned to images newly saved after the image search using a model, image search using a model is preferable when performing a comprehensive search including newly captured images.

[0054] When training the model, the controller 3 may adjust the amount of model training depending on the number of device-saved images stored in the memory unit 4. Specifically, the greater the number of device-saved images, the greater the amount of model training. For example, the controller 3 collects more training images as the number of device-saved images increases. Specifically, when extracting images to be used for training from the training images collected from the search server 100, the controller 3 extracts more images as the number of device-saved images increases, thereby increasing the amount of training. Furthermore, when the number of device-saved images increases to or exceeds a predetermined additional training threshold, the controller 3 may perform additional training using unused training images. When the number of device-saved images is large, the device-saved images will contain many images containing objects in various forms, making the judgment more difficult. However, the judgment accuracy of the model increases depending on the number of device-saved images, thereby enabling appropriate image search. Furthermore, when the number of device-saved images is small, the difficulty of judgment using the model is low, so the required image search accuracy can be maintained, and the search wait time can be shortened by reducing the amount of training.

[0055] The controller 3 may delete models that have not been searched for for a long period of time from among the models of the type "added" in the model information 42. Specifically, the controller 3 deletes from the model information 42 models for which a predetermined deletion threshold period or more has passed since they were last used as image search models 31. This prevents the memory capacity of the storage unit 4 from being overwhelmed. Specifically, as the number of added models increases with the number of searches, the storage capacity for storing model information increases. In this case, models with the oldest IDs from among the models of the type "added" in the model information may be deleted preferentially. Alternatively, past search history may be stored, and models with the oldest most recent search date and time may be deleted preferentially.

[0056] The controller 3 can also search for images stored in the terminal using multiple words (multiple desired objects) as search words. Here, the learning process, search process, and other processes when searching for images containing multiple desired objects as search words will be described with reference to Figs. 6 and 7.

[0057] 6 and 7 are diagrams illustrating the learning process for multiple objects. FIG. 6 shows an example of generating a model with a new model ID indicating the class of each of multiple desired objects, and FIG. 7 shows an example of generating a model with a new model ID indicating a class that includes all of the multiple desired objects. In FIGS. 6 and 7, it is assumed that the search words entered are "shiitake mushroom, person, mandarin orange." In other words, this example shows a case where the user wants to search for images stored on the device that contain all of "shiitake mushroom," "person," and "mandarin orange."

[0058] First, an example of generating classes for each of a plurality of objects will be described with reference to Fig. 6. In the example shown in Fig. 6, the controller 3 detects that a plurality of desired objects have been input in the search words. The detection method may be, for example, detecting spaces or punctuation marks between each word to detect whether a plurality of desired objects have been input.

[0059] Next, if none of the desired objects indicated by the search word are present in the "class" of the model information 42, the controller 3 causes the search server 100 to perform an image search using the input multiple desired objects (object names) as search words. That is, the controller 3 collects training images of "shiitake mushrooms," training images of "people," and training images of "tangerines." The controller 3 then performs machine learning using each training image to train a model individually and generate a model for each class (a model with a new model ID). That is, the controller 3 generates a class model that determines the presence or absence of "shiitake mushrooms" (Yes / No), a class model that determines the presence or absence of "people" (Yes / No), and a class model that determines the presence or absence of "tangerines" (Yes / No). Note that if the target of the search word is partially present in the "class" of the model information 42, training images other than the present targets are collected. Specifically, if "shiitake mushroom" exists as a "class" in the model information 42, learning images of "people" and "tangerines" are collected separately, and learning is performed individually using each learning image.

[0060] When performing an image search using the model, the controller 3 selects the model as the image search model 31 and inputs the device-stored image to the image search model 31. The image search model 31 outputs a determination result indicating whether the stored image contains each of the specific objects (classes) that the model can determine, or the probability that the specific object is included. The controller 3 also extracts stored images that are determined to contain all of the objects indicated by the search words from the determination results for each of the three classes ("shiitake mushroom," "person," and "orange"). If the determination result is a probability, the controller 3 determines and extracts stored images with a probability equal to or greater than a predetermined threshold as containing the object. The extracted stored images are then displayed as search results for the search word. This makes it possible to search for stored images that contain all of the multiple objects. In the example shown in FIG. 6, for example, when a device-stored image (referred to as a determination target image) that shows only a shiitake mushroom among "shiitake mushroom," "person," and "orange" is determined as follows: When the image to be determined is input into the image search model for "shiitake mushrooms," a determination result (Yes) is obtained that the device-stored images contain "shiitake mushrooms." Similarly, when the image to be determined is input into the image search model for "people," a determination result (No) is obtained that the device-stored images do not contain "people." Similarly, when the image to be determined is input into the image search model for "tangerines," a determination result (No) is obtained that the device-stored images do not contain "tangerines." Then, the controller 3 performs an AND logic operation on each of the determination results to obtain "No" as the output answer for "images containing all of 'shiitake mushrooms,' 'people,' and 'tangerines.'" In other words, out of "shiitake mushrooms," "people," and "tangerines," a device-stored image that only shows shiitake mushrooms is not an image that contains all of the desired objects "shiitake mushrooms," "people," and "tangerines," and is therefore not output as a search result. By performing the above processing on all saved device-stored images, the controller 3 extracts device-stored images that contain all of the desired objects "shiitake mushrooms," "people," and "tangerines," and extracts them as search results.

[0061] Furthermore, the controller 3 can also perform a search using the model to search for images stored on the device that contain objects of at least one of the three classes. In this case, the controller 3 performs an OR logical operation on each image search model instead of an AND logical operation, and outputs the search results. Furthermore, if there are zero (or fewer than a predetermined number of) images stored on the device that contain all of the desired objects of the three classes, the controller 3 may output search results that include stored images that contain objects of two classes, or even stored images that contain objects of one class, with an appropriate comment (such as "no matching images, few matching images") attached. In this case, the controller 3 combines AND logical operations and OR logical operations to extract images stored on the device that contain two or one of the desired objects of the three classes, and output these as search results.

[0062] In this way, by performing an image search by combining separate learning models trained on images containing each object, the controller 3 can search not only for stored images that contain all of the objects, but also for stored images that contain any combination of objects or a single object. In other words, it is possible to generate a model that can perform highly versatile searches.

[0063] Next, an example of generating a class that includes all of the multiple objects will be described with reference to Fig. 7. First, as in the above, the controller 3 detects whether multiple desired objects have been entered in the search words. The detection method is, for example, to detect whether multiple objects have been entered by detecting spaces or punctuation marks between each word.

[0064] Next, the controller 3 causes the search server 100 to perform an image search using search terms that contain all of the input desired objects. Specifically, the controller 3 performs an image search using "shiitake mushroom," "person," and "orange" as search terms using an AND logic operation, and collects training images that contain all of "shiitake mushroom," "person," and "orange." The controller 3 then performs machine learning using the collected training images to train the image search model 31 and generate a trained model that detects images that contain all of the desired objects (shiitake mushroom, person, and orange). Specifically, the controller 3 generates a model that determines whether or not "shiitake mushroom, person, orange" is present in an input image. When performing an image search using this model, the controller 3 extracts device-stored images that are determined to contain the desired objects based on a class determination result indicating that all three desired objects are included, or, if the determination result is a probability value, extracts device-stored images with a probability equal to or greater than a predetermined threshold, and displays these as search results. This allows the controller 3 to search for device-stored images that contain all of the desired objects.

[0065] In this way, by generating one model that detects an image that includes all of the multiple targets, the controller 3 can reduce the processing load associated with the learning process compared to generating separate models that detect images that include each of the multiple targets shown in Fig. 6. Furthermore, since the number of models to be stored is reduced, the storage capacity of the storage unit 4 can be saved.

[0066] In the above, an example has been given in which the search server 100 is used as the external server, but the present invention is not limited to this. For example, a server having a generation AI that generates images (generation AI server) may be used as the external server. This point will be explained using FIG. 8. In the following, other examples of the method for collecting learning images will be explained, and since the processing other than the method for collecting learning images (learning processing, etc.) is the same as in the above embodiment, explanations may be omitted.

[0067] Fig. 8 is a diagram showing another example of an image search system S. As shown in Fig. 8, the image search system S includes a terminal device (image search device) 1 and a generation AI server 200 which is an external server.

[0068] The generation AI server 200 is a server that has a generation AI that generates new images based on input prompts. The generation AI server 200 inputs prompts input from outside into the generation AI and provides the images generated by the generation AI to the outside.

[0069] In such an image search system S, the terminal device 1 accepts input of a search word from a user. The process of accepting the input of a search word is the same as that described above, and therefore a description thereof will be omitted.

[0070] Next, the terminal device 1 generates a prompt including the input search word and a command statement requesting image generation, and inputs the generated prompt to the generation AI server 200. The terminal device 1 generates a prompt including a word indicating the target, such as "Please generate an image of XX," and a command statement requesting image generation. As a result, the terminal device 1 collects images generated by the generation AI server 200 in accordance with the input prompt as learning images. Note that the terminal device 1 may include sentences in the prompt that specify the viewpoint, color, size, quantity, etc. of the target image.

[0071] In this way, in the image search system S shown in Figure 8, by collecting training images generated by the generation AI, it is possible to prevent the mixing of training images containing incorrect objects, thereby improving the learning accuracy of the model.

[0072] Further advantages and modifications will readily occur to those skilled in the art. Therefore, the invention in its broader aspects is not limited to the specific details and representative embodiments shown and described above. Accordingly, various modifications may be made without departing from the spirit or scope of the general inventive concept as defined by the appended claims and their equivalents. [Explanation of symbols]

[0073] 1. Terminal device (image search device) 2. Communications Department 3 Controller 4 Storage section 5. Camera 31 Image Search Model 41 Saved image information 42 Model Information 100 search servers 200 Generation AI Server S Image Search System

Claims

1. An image search method executed by a computer that searches stored images for images that match a search word entered by a user, Collecting images that match the search word from an external server; Outputting the saved images that match the search word by similarity search with the collected images. How to search for images.

2. the images collected from the external server are images for model training, The model is trained using the training images as correct answer data, the stored images are applied to the trained model, and the stored images that match the search word are output based on the judgment result of the model. The image retrieval method according to claim 1 .

3. the external server is a search server that performs an image search based on an input word; Sending the search word to the search server; The images searched by the search server are collected as the learning images. The image retrieval method according to claim 2 .

4. The external server is a generation AI server having a generation AI that generates an image based on an input word, sending a prompt to the generation AI server, the prompt including the search word and a command requesting image generation; Collecting images generated by the generation AI based on the prompt as the learning images. The image retrieval method according to claim 2 .

5. When the input of the search word is received, it is determined whether or not the model corresponding to the search word is stored, and if the model is not stored, the learning image is generated and the model corresponding to the search word is generated by learning. The image retrieval method according to claim 2 .

6. If the search word includes a plurality of words indicating a plurality of objects, the learning images are collected for each of the words, and the model is generated for each of the words. The image retrieval method according to claim 2 .

7. If the search word includes a plurality of words indicating a plurality of objects, the learning images including all of the objects indicated by the plurality of words are collected. The image retrieval method according to claim 2 .

8. The more the number of images stored in the storage device, the more the model is trained. The image retrieval method according to claim 2 .

9. The greater the number of images stored in the storage device, the greater the number of images to be collected for learning. The image retrieval method according to claim 8.

10. 1. An image generation method for generating learning images to be used in learning a model that determines whether or not a saved image stored in a storage device matches a search word entered by a user, comprising: A prompt including the search word and a command requesting image generation is input to a generation AI that generates an image, and the generation AI generates the learning image. Image generation method.

11. An image search program for searching stored images that match a search word entered by a user, Collecting images that match the search word from an external server; outputting the stored images that match the search word by performing a similarity search with the collected images; An image search program that runs on a computer.

12. An image search device for searching stored images for images that match a search word entered by a user, the image search device having a controller: The controller Collecting images that match the search word from an external server; Outputting the saved images that match the search word by similarity search with the collected images. Image search device.

13. An external server that provides images; an image search device that searches for images that match a search word entered by a user from among images stored in a storage device; Equipped with The image search device includes: Collecting images that match the search word from an external server; Outputting the saved images that match the search word by similarity search with the collected images. Image search system.

Citation Information

Patent Citations

  • Image search apparatus and image search method

    JP2018055730A