Search model generation system, search device, search model generation method, and program
The search model generation system addresses domain dependence in ReID by using image generation AI to create adaptable search models, enhancing accuracy through iterative evaluation and retraining.
Patent Information
- Application Number
- JP2024001338
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-09
- Publication Date
- 2025-07-22
AI Technical Summary
Existing ReID technologies face domain dependence issues due to variations in clothing and shooting location biases, necessitating re-learning of models based on time and season.
A search model generation system that includes a teacher data generation device using image generation AI to create images with associated identifying information, a search model generation device for model training, and a model evaluation device for accuracy assessment, facilitating model retraining.
Enables efficient retraining of search models to adapt to varying conditions, improving search accuracy and adaptability.
Smart Images

Figure 2025107848000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a search model generation system, a search device, a search model generation method, and a program.
Background Art
[0002] There is a technology called ReID (Re-Identification). ReID is a technology that enables re-identifying a person by vectorizing feature quantities such as the person's clothing, hairstyle, and belongings and quantitatively expressing them. With ReID, for example, it becomes possible to search by sorting people in order of similarity based on images or videos taken at different locations, track the same person, narrow down multiple people by specific attributes, and so on. For example, Patent Document 1 discloses a technology for performing narrowing-down search by attributes and searching for people. In order to perform ReID, learning is performed using teacher data to create a search model.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, domain dependence occurs in the feature quantities of people. For example, clothing varies depending on the season and time. Also, there is a bias in the people photographed depending on the shooting location. Therefore, it is necessary to re-learn the model depending on time, season, shooting location, and so on.
[0005] An object of the present invention is to provide a search model generation system, a search device, a search model generation method, and a program that can facilitate re-learning of a model.
Means for Solving the Problem
[0006] One aspect of the present invention is a search model generation system including a search model generation device that generates a search model for searching for a person based on teacher data in which an image of a person generated by an image generation AI is associated with information for identifying the person, and a model evaluation device that evaluates the search accuracy of the search model and outputs the evaluation result.
[0007] One aspect of the present invention is a search model generation method including a search model generation step of generating a search model for searching for a person based on teacher data in which an image of a person generated by an image generation AI is associated with information for identifying the person, and a model evaluation step of evaluating the search accuracy of the search model and outputting the evaluation result.
Advantages of the Invention
[0008] According to the present invention, retraining of the model can be facilitated.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Modes for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. (First Embodiment) FIG. 1 is a diagram showing the configuration of a search model generation system 1 according to the first embodiment. The search model generation system 1 generates a search model for searching for a person. The search model generation system 1 includes a teacher data generation device 11, a search model generation device 12, and a model evaluation device 13.
[0011] The teacher data generation device 11 generates teacher data by generating an image of a person using an image generation AI based on a prompt input by the user P. The teacher data generation device 11 outputs the generated teacher data to the search model generation device 12. The image generation AI is, for example, a text-to-image model that takes a prompt as an input and outputs an image that matches the input prompt. The text-to-image model is not particularly limited, and examples include Stable Diffusion and DALL-e.
[0012] By inputting a plurality of different prompts to the teacher data generation device 11, images of a plurality of different persons are generated. The prompt includes common text and unique text. The common text indicates conditions common to the persons to be searched by the search model. The common conditions are, for example, time period and location. Depending on whether the time period is summer or winter, there is a tendency for the clothing, hairstyle, and items carried by the person, which are characteristic quantities of the person, to be different. Also, depending on the location, there is a tendency for the clothing, hairstyle, and items carried by the person, which are characteristic quantities of the person, to be different. The individual text indicates individual conditions different from the common conditions indicated by the common text. The individual conditions are, for example, the gender, height, build, and color of the clothes of the person.
[0013] By inputting one prompt to the teacher data generation device 11, a plurality of different images of a person under the same conditions are generated. As described above, the teacher data generation device 11 generates a plurality of images of persons who satisfy the conditions common to the persons to be searched by the search model and have different individual conditions according to each individual condition. That is, the teacher data generation device 11 outputs data in which individual conditions of persons are associated with a plurality of images. The individual condition may be anything as long as it is information for identifying the person in the image. For example, it may be a prompt including individual conditions and common conditions, or simply a symbol. The teacher data generation device 11 outputs data in which individual conditions are associated with a plurality of images as teacher data.
[0014] FIG. 2 is a diagram showing an example of teacher data. In the teacher data shown in FIG. 2, a plurality of images 1A to 1C are generated from the prompt of "common condition + individual condition C1", a plurality of images 2A to 2C are generated from the prompt of "common condition + individual condition C2", and a plurality of images 3A to 3C are generated from the prompt of "common condition + individual condition C3". In the teacher data shown in FIG. 2, the individual condition C1 is associated with the plurality of images 1A to 1C, the prompt of the individual condition C2 is associated with the plurality of images 2A to 2C, and the prompt of the individual condition C3 is associated with the plurality of images 3A to 3C.
[0015] The search model generation device 12 generates a search model for searching for a person from an image based on the teacher data input from the teacher data generation device 11. The search model generation device 12 generates a search model that searches for the persons appearing in the plurality of images generated from the same prompt as the same person. The search model generation device 12 uses, as teacher data, data in which one prompt or individual conditions are associated with a plurality of images, and generates a search model by a machine learning method. The machine learning method is not particularly limited. For example, it is Metric Learning (deep metric learning).
[0016] The search model is, for example, a model that vectorizes the feature amounts of the persons appearing in the image, calculates a quantitative value, and searches for images with close calculated values as images in which the same person appears. At this time, the search model generation device 12 generates an estimation model so that the values calculated by the generated estimation model from the plurality of images generated from the same prompt are as close as possible.
[0017] The search model generation device 12 outputs the generated search model to the model evaluation device 13.
[0018] The model evaluation device 13 evaluates the search model input from the search model generation device 12. The model evaluation device 13 evaluates the search accuracy of the search model using evaluation data in which a person and a plurality of images of the person are linked. The model evaluation device 13 outputs an evaluation result. The evaluation result may include data regarding the search accuracy by the search model and the data of the person for whom the search model succeeded in the search and the person for whom the search failed among the evaluation data. The evaluation data may be teacher data generated by the teacher data generation device 11 and data not used for the generation of the search model.
[0019] Upon receiving the evaluation result, the user P can change the prompt input to the teacher data generation device 11 so that a search model with higher search accuracy is generated.
[0020] FIG. 3 is a flowchart showing the operation of the search model generation system 1 according to the first embodiment. When the user P inputs a prompt, the teacher data generation device 11 generates an image of a person (step S11). At this time, the teacher data generation device 11 outputs a plurality of sets of data in which one prompt or individual conditions and a plurality of images are linked. The search model generation device 12 generates a search model for searching for a person from the image based on the image input from the teacher data generation device 11 (step S12). The model evaluation device 13 evaluates the search model generated by the search model generation device 12 (step S13). The model evaluation device 13 outputs an evaluation result (step S14).
[0021] In the search model generation system 1 according to the first embodiment, the teacher data generation device 11 generates an image of a desired person by an image generation AI, and the search model generation device 12 generates a search model for searching for a person using the teacher data including the generated image. As a result, the search model can be trained to search for a person under desired conditions.
[0022] (Second Embodiment) FIG. 4 is a diagram showing the configuration of the search model generation system 1 according to the second embodiment. The search model generation system 1 according to the second embodiment includes a prompt generation device 14 in addition to the search model generation system 1 according to the first embodiment.
[0023] The prompt generation device 14 generates a prompt and outputs it to the teacher data generation device 11. The teacher data generation device 11 according to the second embodiment generates an image based on the prompt input from the prompt generation device 14 and generates teacher data. The prompt generation device 14 generates a prompt by using a generation AI that generates text. The AI that generates text refers to, for example, a text-to-text model that generates text based on text or an image-to-text model that generates text based on an image. Examples of the generation AI that generates text include GPT and Llama.
[0024] The generated teacher data may be divided into hierarchical category based on the content of the prompt. For example, the teacher data is divided into categories to which the characteristics (clothing, hairstyle, belongings, etc.) of the person in the image described in the prompt belong. For example, when the prompt is "a 30-year-old man wearing black suit and glasses", the teacher data is divided into categories of "black suit", "glasses", "30-year-old", and "man". The category of "black suit" may be further divided into the category of "suit", and the category of "suit" may be further divided into the category of "upper garment".
[0025] A prompt is generated by the prompt generation device 14. Then, teacher data is generated by the teacher data generation device 11, a search model is generated by the search model generation device 12, the search model is evaluated by the model evaluation device 13, and an evaluation result is output. The evaluation result is input to the prompt generation device 14. The prompt generation device 14 generates a prompt based on the evaluation result. The teacher data generation device 11 stores a plurality of pieces of teacher data generated by a plurality of prompts, may generate new teacher data based on the newly generated prompt, and update the stored teacher data.
[0026] The teacher data generation device 11, for example, based on the search accuracy of the search model indicated by the evaluation result, discards the teacher data generated by the prompt when the search accuracy has not improved, and retains the teacher data generated by the prompt when the search accuracy has improved. By repeating this, it is considered possible to retain teacher data with improved search accuracy and improve the search accuracy by the search model. This can be achieved, for example, by the prompt generation device 14 using a text-to-text model to generate a prompt, the teacher data generation device 11 generating teacher data based on the prompt, and discarding and retaining the teacher data. Here, the prompt generation device 14 can prevent the generation of teacher data discarded by the teacher data generation device 11 by generating a prompt that does not include text indicating the lowest-level category to which the discarded teacher data belongs, and can shorten the learning time.
[0027] The prompt generation device 14 may generate a prompt using an image-to-text model. When using the image-to-text model, the prompt generation device 14 identifies the category into which the training data that causes the search accuracy calculated by the model evaluation device 13 to decrease is classified. By inputting the training data that causes the search accuracy to decrease into the image-to-text model, text indicating the category of the training data is generated. The prompt generation device 14 then generates a prompt that does not include text indicating the category into which the training data that causes the search accuracy to decrease is classified. It is considered that this can improve the search accuracy of the search model.
[0028] When the prompt generation device 14 uses an image-to-text model, compared with the case of using a text-to-text model, based on the training data that causes the search accuracy to decrease, the combination of categories of the training data that directly causes the search accuracy to decrease can be identified. Therefore, it is considered that the search accuracy of the search model can be improved more quickly. When the prompt generation device 14 uses an image-to-text model, since the reason for the factor that the search accuracy of the search model decreases can be explained compared with the case of using a text-to-text model, it is easy for the user to check the progress of the learning process and add an evaluation.
[0029] The prompt generation device 14 outputs the newly generated prompt to the training data generation device 11. By repeating this, a search model with high search accuracy can be generated. When the search accuracy indicated by the evaluation result is equal to or higher than a predetermined value, the prompt generation device 14 may terminate the operation without generating a newly generated prompt.
[0030] FIG. 5 is a flowchart showing the operation of the search model generation system 1 according to the second embodiment. First, when an instruction is input to the prompt generation device 14, the prompt generation device 14 generates a prompt (step S20). The operations of steps S21 to S23 are the same as the operations of steps S11 to S13 in the first embodiment. The model evaluation device 13 outputs the evaluation result of the search model to the prompt generation device 14 (step S24). The prompt generation device 14 determines whether the search accuracy indicated by the evaluation result is equal to or greater than a predetermined value (step S25).
[0031] If the search accuracy is equal to or greater than the predetermined value (step S25: YES), the prompt generation device 14 ends the operation. The search model generated immediately before the end of the operation is used by a search device, which will be described later, for searching for a person. If the search accuracy is less than the predetermined value (step S25: NO), the prompt generation device 14 generates a new prompt based on the evaluation result and outputs the generated prompt to the teacher data generation device 11 (step S26). As a result, images are generated based on different prompts, and different teacher data are generated.
[0032] In the search model generation system 1 according to the second embodiment, the prompt generation device 14 generates a prompt, and the teacher data generation device 11 generates an image based on the generated prompt. Therefore, in the second embodiment, compared with the first embodiment, it is not necessary for a human to input a prompt, and the search model can be generated more easily.
[0033] (Search System) FIG. 6 is a diagram showing the configuration of the search system 2. The search system 2 is a system that captures a video and searches for a person appearing in the video. The search system 2 includes a photographing device 21 and a search device 22.
[0034] The photographing device 21 captures a video including a person. The photographing device 21 is installed, for example, in a street and captures a video of a moving person. The search system 2 includes a plurality of photographing devices 21 installed at different locations. The imaging device 21 outputs the captured video to the search device 22.
[0035] The search device 22 searches for the people shown in the video obtained from the imaging device 21. The search device 22 stores the search model generated by the search model generation system 1. The search device 22 uses the search model to search for the same people shown in the video. The search device 22 outputs the search results.
[0036] The search device 22 may update the search model based on the identification result of the person performed manually. FIG. 7 is a diagram showing the search device 22 and the identified person U. The search device 22 outputs an image of the person searched from the video to the identified person U. The identified person U checks the image of the person and identifies the image in which the same person is shown. The identified person U inputs the identification result to the search device 22. As a result, the search device 22 receives the identification result of the image in which the same person who could not be searched is shown. The search device 22 updates the search model based on the identification result.
[0037] By updating the search model based on the identification result by the identified person U, the search model is updated to search for the people shown in the video captured by the imaging device 21. As a result, the search device 22 can update the search model according to the timing and location of the video captured by the imaging device 21.
[0038] <Other Embodiments> As described above, one embodiment of the present invention has been described in detail with reference to the drawings. However, the specific configuration is not limited to the above, and various design changes and the like can be made without departing from the gist of the present invention.
[0039] The processing of the teacher data generation device 11, the search model generation device 12, the model evaluation device 13, the prompt generation device 14, or the search device 22 in the above-described embodiment may be realized by a computer using software. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" shall include hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, a CD-ROM, etc., and a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, something that dynamically holds a program for a short time, and also includes something that holds a program for a certain time, like a volatile memory inside a computer system that becomes a server or a client in that case. Also, the above program may be for realizing a part of the above-described functions, and may further be something that can be realized in combination with a program already recorded in the computer system, or may be realized using a programmable logic device such as an FPGA (Field Programmable Gate Array).
Explanation of Signs
[0040] 1 Search model generation system, 11 Teacher data generation device, 12 Search model generation device, 13 Model evaluation device, 14 Prompt generation device, 2 Search system, 21 Photographing device, 22 Search device
Claims
1. A search model generation device that generates a search model for searching for a person based on teacher data in which an image of a person generated by an image generation AI is associated with information for identifying the person, A model evaluation device that evaluates the search accuracy of the search model and outputs an evaluation result, A search model generation system comprising the above.
2. A prompt generation device that generates a prompt to be input to the image generation AI based on the evaluation result, further comprising, The search model generation device generates an image of a person by inputting the prompt to the image generation AI. The search model generation system according to Claim 1.
3. The prompt generation device changes a prompt generated based on the evaluation result and inputs the changed prompt to the image generation AI. The search model generation system according to Claim 2.
4. The prompt generation device generates the prompt based on the tendency of the teacher data that causes the search accuracy, which is the evaluation result, to decrease. The search model generation system according to Claim 3.
5. Updating a search model for searching for a person based on an image identified as depicting the same person. Search device.
6. A search model generation step of generating a search model for searching for a person based on teacher data in which an image of a person generated by an image generation AI is associated with information for identifying the person, A model evaluation step of evaluating the search accuracy of the search model and outputting an evaluation result. A search model generation method having the above.
7. A program for causing a computer to execute the search model generation method according to Claim 6.
Citation Information
Patent Citations
Person tracking device, person tracking system, person tracking method, and person tracking program
JP2022030846A
Image analysis system, image analysis method, and image analysis program
JP6947899B1