Learning model selection support system, terminal, and learning model selection support method
Patent Information
- Application Number
- JP2025023739
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2026-08-27
AI Technical Summary
【0010】 本開示によれば、学習モデルの選定を容易にすることができる。
Smart Images

Figure 2026137559000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to a learning model selection support system, a terminal, and a learning model selection support method.
Background Art
[0002] Conventionally, an evaluation device that evaluates learning models from multiple viewpoints has been known. This evaluation device inputs evaluation data to a learning model to be evaluated, evaluates the quality of the functional or non-functional aspects, and displays an evaluation result screen on which the evaluation results of multiple learning models are superimposed on a display (see Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] It is difficult to intuitively understand the evaluation results of the learning models by the evaluation device of Patent Document 1, and there is room for improvement.
[0005] The present disclosure has been devised in view of the above-described conventional situation, and aims to facilitate the selection of a learning model.
Means for Solving the Problems
[0006] [[ID=<<MASK_E>>45]] The present disclosure provides a learning model selection support system that supports the selection of a learning model, includes a display device and a processing device, the processing device displays a setting screen for setting a plurality of images and selection conditions for a learning model on the display device, performs a performance evaluation of each learning model based on the plurality of images and the selection conditions set on the setting screen, and displays a selection result screen including at least one result of the performance evaluation of each learning model on the display device.
[0007] Furthermore, this disclosure provides a terminal that assists in the selection of a learning model, comprising a display device and a processor, wherein the processor displays a settings screen on the display device for setting a plurality of images and selection conditions for a learning model, and displays a selection results screen on the display device that includes at least one of the performance evaluation results for each learning model based on the plurality of images and selection conditions set on the settings screen.
[0008] Furthermore, this disclosure provides a learning model selection support method for assisting the selection of a learning model, comprising: displaying a setting screen on a display device for setting a plurality of images and selection conditions for a learning model; performing a performance evaluation of each learning model based on the plurality of images and selection conditions set on the setting screen; and displaying a selection result screen on the display device that includes at least one result of the performance evaluation of each learning model.
[0009] Furthermore, any combination of the above components, as well as any conversion of the expressions of this disclosure between methods, apparatus, systems, storage media, computer programs, etc., are also valid forms of this disclosure. [Effects of the Invention]
[0010] According to this disclosure, the selection of a learning model can be made easier. [Brief explanation of the drawing]
[0011] [Figure 1] Block diagram showing an example configuration of the support system according to Embodiment 1. [Figure 2] Schematic diagram showing an example of the settings screen according to Embodiment 1 [Figure 3] Schematic diagram showing an example of the image evaluation result screen according to Embodiment 1. [Figure 4] Schematic diagram showing the first example of the selection result screen according to Embodiment 1. [Figure 5] Schematic diagram showing a second example of the selection result screen according to Embodiment 1. [Figure 6] Schematic diagram showing a third example of the selection result screen according to Embodiment 1. [Figure 7] Flowchart showing the overall processing of the support system according to Embodiment 1 [Figure 8] Flowchart showing the model performance evaluation process by the support system according to Embodiment 1 [Figure 9] Flowchart showing the selection result output process by the support system according to Embodiment 1 [Modes for carrying out the invention]
[0012] The embodiments will be described in detail below, with reference to the drawings as appropriate. However, unnecessary details may be omitted. For example, detailed explanations of already well-known matters and redundant explanations of substantially identical configurations may be omitted. This is to avoid the following explanation becoming unnecessarily verbose and to facilitate understanding by those skilled in the art. The accompanying drawings and the following explanation are provided to enable those skilled in the art to fully understand this disclosure and are not intended to limit the subject matter described in the claims.
[0013] (The circumstances that led to this form of disclosure) The evaluation device described in Patent Document 1 displays the evaluation of a learning model using numerical information, specifically by displaying the numerical value itself or by using a graph or table based on the numerical value. However, for someone who does not understand the learning model, such as a person who is not an expert or researcher, it is difficult to judge the quality of the learning model based on such numerical information. For example, even if the processing time when using the learning model is displayed as 0.8 seconds, it is difficult for someone who does not understand the learning model to judge how good the evaluation is. Therefore, it is preferable to be able to intuitively evaluate a learning model and easily select a learning model, even for someone who does not have a high level of understanding of learning models.
[0014] In the following embodiments, a learning model selection support system, a terminal, and a learning model selection support method for assisting in easily selecting a learning model will be described.
[0015] (Embodiment 1) [System Configuration] FIG. 1 is a block diagram showing a configuration example of the support system 5 according to Embodiment 1. The support system 5 is a system for assisting in the selection of a learning model. In the description of this embodiment, it is assumed that the process using the learning model is face authentication processing. The support system 5 assists in selecting a learning model suitable for the field. Here, the field refers to the application site of the learning model, and examples include the entrance and exit of a facility. The support system 5 assists in selecting, for example, a learning model used for face authentication when a person passes through at a field such as the entrance and exit of a facility. Note that the process using the learning model may be authentication processing other than face authentication processing, or may be a process other than authentication processing. Also, the application site of the learning model is not limited to the entrance and exit of a facility, and any site where the process using the learning model (for example, face authentication processing) can be applied is acceptable.
[0016] The support system 5 shown in FIG. 1 includes a server device 10 and a terminal 20. The server device 10 performs various processes related to the selection of a learning model. The terminal 20 acquires information necessary for various processes related to the selection of a learning model, or displays information or a screen obtained by various processes related to the selection of a learning model.
[0017] The server device 10 may be configured as an on-premises type server device, or may be configured in a cloud type on a network. The server device 10 may be configured by one computer, or may be configured in a distributed manner by a plurality of computers.
[0018] The server device 10 includes a processor 11, a memory 12, and a communication device 13.
[0019] The processor 11 may be configured using, for example, a Central Processing Unit (CPU), a Digital Signal Processor (DSP), or a Graphical Processing Unit (GPU). The processor 11 may also be configured using various integrated circuits (for example, a Large Scale Integration (LSI) or a Field Programmable Gate Array (FPGA)). The processor 11 implements various functions by executing programs held in memory 12. The processor 11 comprehensively controls each part of the server device 10 and performs various processing.
[0020] For example, processor 11 performs screen control processing to control the screen displayed by terminal 20 or other display device, as described later. Processor 11 also performs access processing to a model performance list showing a list of the performance of the learning models. Processor 11 also performs graph creation processing to create a graph related to the learning models (e.g., evaluation of the learning models). Processor 11 also performs table creation processing to create a table that holds information related to the learning models (e.g., evaluation of the learning models). Processor 11 also performs attribute estimation processing to have the learning model estimate the attributes contained in the image. The attributes contained in the image will be described later. Processor 11 also performs model selection processing to select a learning model.
[0021] Memory 12 includes primary storage devices (e.g., Random Access Memory (RAM) or Read Only Memory (ROM)). Memory 12 may also include secondary storage devices (e.g., Hard Disk Drive (HDD) or Solid State Drive (SSD)) or tertiary storage devices (e.g., optical disc or SD card). Memory 12 may also be an external storage medium and may be detachable from the server device 10. Memory 12 stores various data, information, or programs.
[0022] In this embodiment, the support system 5 inputs images containing a person's face into each of several learning models and selects, for example, a learning model with high authentication accuracy, or in other words, recommends it. Hereinafter, images containing a person's face may be referred to as face images. Memory 12 holds face images of the person to be facially recognized. Memory 12 also acquires and holds captured images for input into the learning models. Hereinafter, captured images for input into the learning models may be referred to as input images. It is preferable that the input images be captured at the application site in order to select a learning model suitable for the application site.
[0023] In this embodiment, it is assumed that a tag is assigned to the input image before it is input to the learning model. The tag may be, for example, one used to identify a person. Specifically, the name of a person may be assigned as a tag to an input image that shows a person. In the following description, it is assumed that the input image is tagged with a tag that can identify the person shown in the input image. It is also assumed that one person is shown in one input image. In other words, it is assumed that one tag is assigned to one input image. The tagging of the input image may be performed, for example, by the operation of the user's terminal 20. Alternatively, the tagging of the input image may be performed by a device that can implement both face recognition and tagging functions. For example, the processor 21 and memory 22 of the terminal 20 may work together to identify the person shown in the input image and assign a tag corresponding to that person to the input image. Alternatively, these processes may be performed by the processor 11 and memory 12 of the server device 10 working together, or by an external device (not shown).
[0024] The learning model calculates a matching score between the input image and pre-set registered images, based on the input image. The matching score indicates the degree of similarity between the images. For example, the higher the matching score calculated by the learning model, the higher the similarity between the registered image and the input image.
[0025] A registered image is an image that is pre-set to calculate a matching score with the input image. Hereafter, the matching score between the input image and the registered image may simply be referred to as the matching score of the input image.
[0026] Here, among multiple input images tagged to represent the same person, the image with the highest quality score may be set as the registered image of that person. The quality score indicates the quality (image quality) of the input image as an image of a person, and may be calculated from, for example, the clarity of the person in the input image, the resolution of the input image, the color balance, noise, etc. The quality score may be calculated using, for example, known technology such as AI technology. The quality score may also be calculated by, for example, the server device 10 or the terminal 20.
[0027] For example, consider a case where each of multiple input images is assigned a tag for a different person, such that the person in the input image corresponds to the person indicated by the tag. In this case, a registered image corresponding to each of the multiple tags, in other words, a registered image corresponding to each of the multiple people, is set. Simply put, a registered image is set for each person. As a result, when an input image of a certain person is input into the learning model, a matching score is calculated between that input image and the registered image corresponding to that person among the multiple registered images.
[0028] Furthermore, the processor 11 can calculate the authentication accuracy of the learning model based on the matching score calculated by the learning model. Authentication accuracy is an indicator of how accurately the learning model can authenticate a person in the input image, or how accurately it cannot. For illustrative purposes, in this embodiment, the processor 11 determines that the person in the input image and the person in the registered image are the same person if the matching score of the input image is equal to or greater than a predetermined first threshold. The processor 11 also determines that the person in the input image and the person in the registered image are different people if the matching score of the input image is less than the first threshold.
[0029] For example, even if processor 11 determines that the person in the input image and the person in the registered image are the same person, they may actually be different people. Also, even if processor 11 determines that the person in the input image and the person in the registered image are different people, they may actually be the same person. Therefore, processor 11 can further determine whether the learning model was able to perform correct authentication for each input image, that is, whether it was able to accurately calculate a matching score for each input image. Thus, authentication accuracy can also be rephrased as an indicator of how accurately the learning model can calculate a matching score. An accurate matching score is a matching score at which processor 11's determination of whether the person in the input image is the same person or a different person from the person in the registered image, based on the predetermined first threshold described above, becomes correct.
[0030] Authentication accuracy can be determined, for example, by calculating the ratio of input images for which the learning model was able to calculate an accurate matching score, relative to the total number of input images input to the learning model.
[0031] For example, even if the calculated matching score of the input image is higher than the first threshold, if the person in the input image is actually a different person from the person in the registered image, this will reduce the authentication accuracy. Also, for example, even if the calculated matching score of the input image is lower than the first threshold, if the person in the input image is actually the same person from the registered image, this will reduce the authentication accuracy.
[0032] Furthermore, the processor 11 can calculate the authentication accuracy of the learning model for each attribute, as described later.
[0033] In this specification, terms such as "First" and "Second" are used merely for the purpose of distinguishing components for explanatory purposes, and are not intended to be interpreted as limiting the scope to any specific component.
[0034] Memory 12 holds the matching scores calculated by the learning model. Memory 12 may also hold matching scores calculated in past learning model selection processes. In other words, memory 12 may hold existing matching scores. This allows the processor 11 to determine the average matching score for each learning model.
[0035] Furthermore, memory 12 stores feature information for each learning model (for example, for each learning model identification information). The feature information for the learning model includes the matching score described above. The feature information for the learning model also includes performance information indicating the performance of the learning model. This performance information includes, for example, ROC curve information, or performance information for dark conditions, masks, blur, tilt at a predetermined angle (e.g., 45 degrees), backlighting, etc. Here, performance for dark conditions, masks, blur, tilt at a predetermined angle, or backlighting refers to the authentication accuracy for an input image that includes each of these attributes. An attribute is a feature included in the input image; for example, if a person in the input image is wearing glasses, the input image includes the attribute of glasses. In this embodiment, the attributes included in the input image are estimated by the learning model. The attributes that may be included in the input image may be set in advance by the user. For example, if a user has set four types of attributes—glasses, dark environment, mask, and blur—the learning model will estimate which of these four attributes is present in the input image. It is also possible that the input image may not contain any of these attributes.
[0036] Furthermore, the characteristic information of the learning model includes information about the learning model's related functions. Related functions include version information for the face detection function, a function for acquiring on-site authentication logs, and an anti-spoofing function. The characteristic information of the learning model also includes information about the operating environment in which the learning model is executed. The characteristic information of the learning model also includes information about the learning conditions of the learning model. This learning condition information may include information such as the distribution of the learning database, the type of neural network, the loss function, or the learning method. Memory 12 may also hold the learning model itself. The characteristic information of the learning model can be used, for example, in the graph creation process and table creation process described above.
[0037] Furthermore, memory 12 stores information regarding the selection criteria for the learning model. This information contributes to the selection of the learning model. The learning model selection criteria include at least one of the following: the operating environment conditions for the learning model, the processing time conditions, and the matching score conditions. In this embodiment, each item included in the selection criteria, such as the operating environment conditions for the learning model, the processing time conditions, and the matching score conditions, may be referred to as an item of the selection criteria. The operating environment includes the processor and other resources used to execute the learning model. For example, the operating environment includes the OS (e.g., Windows, Linux®), processor (e.g., GPU, CPU), memory (e.g., RAM), camera, or other device and system execution environment. OS stands for Operating System. As will be described later with reference to Figure 2, the learning model selection criteria can be set by the user of the support system 5. Memory 12 may, for example, store samples of the learning model selection criteria. This allows even users unfamiliar with the operating environment of a learning model to easily proceed with the learning model selection process, for example, by displaying sample selection criteria held in memory 12 on the display device 25. The samples may be pre-created by a user familiar with the learning model selection criteria.
[0038] The communication device 13 communicates various data or information according to a wired or wireless communication method. The communication method used by the communication device 13 may include, for example, a Local Area Network (LAN), a Wide Area Network (WAN), a mobile phone network, or power line communication. The communication device 13 may also communicate via a web application. The communication device 13 may communicate with the terminal 20 or other external devices.
[0039] The communication device 13 communicates various types of data or information with, for example, the terminal 20.
[0040] Terminal 20 is a PC (Personal Computer), smartphone, tablet, or mobile device, etc. Terminal 20 includes a processor 21, memory 22, communication device 23, input device 24, display device 25, etc. Terminal 20 may communicate various data or information with the server device 10, receive notifications or instructions of various information, perform actions in response to notifications or instructions, and issue various instructions in response to input such as operations on Terminal 20.
[0041] The processor 21 may be configured using, for example, a CPU, DSP, or GPU. The processor 21 may also be configured using various integrated circuits (for example, LSI or FPGA). The processor 21 implements various functions by executing programs held in memory 22. The processor 21 comprehensively controls each part of the terminal 20 and performs various processes.
[0042] Memory 22 includes primary storage (e.g., RAM or ROM). Memory 22 may also include secondary storage (e.g., HDD or SSD) or tertiary storage (e.g., optical disc or SD card). Furthermore, memory 22 may be an external storage medium and may be detachable from terminal 20. Memory 22 stores various data, information, or programs.
[0043] The communication device 23 communicates various data or information according to a wired or wireless communication method. The communication method used by the communication device 23 may include, for example, LAN, WAN, mobile phone network, or power line communication. The communication device 23 may also communicate via a web application. The communication device 23 may communicate with the server device 10 or other external devices.
[0044] The input device 24 may include various buttons, keys, a mouse, a keyboard, a touch panel, a microphone, or other input devices. The input device 24 accepts input of various data or information. The input device 24 may be operated by a user. The user is, for example, a user of terminal 20, or, for example, a person who selects a learning model.
[0045] The input device 24 accepts operations for making arbitrary inputs, such as selection, setting, specifying, or giving instructions regarding the selection of a learning model.
[0046] The display device 25 is, for example, a liquid crystal display or an organic EL display. The display device 25 displays various data or information. The display by the display device 25 may be confirmed, for example, by the user.
[0047] The display device 25 displays various screens to assist in the selection of a learning model. For example, the display device 25 displays the settings screen, the selection results screen, etc., which will be described later.
[0048] In the support system 5, the processor 11 of the server device 10 may acquire various types of information via the input device 24 of the terminal 20. Specifically, when the processor 11 of the server device 10 acquires information (input information) entered via the input device 24 of the terminal 20, the following operations may be performed: When the processor 21 of the terminal 20 acquires information entered by the user via the input device 24, it transmits it to the server device 10 via the communication device 23. The processor 11 of the server device 10 acquires the information entered by the user via the communication device 23. This specific operation is the same in subsequent processes, and its explanation may be omitted.
[0049] Furthermore, the processor 11 of the server device 10 may cause information to be displayed on the display device 25 of the terminal 20, that is, it may control the display of information. Specifically, when the processor 11 of the server device 10 causes information to be displayed on the display device 25 of the terminal 20, the following operations may be performed: The processor 11 of the server device 10 transmits the information to be displayed (display information) to the terminal 20 via the communication device 13. The information to be displayed may include a screen. The processor 21 of the terminal 20 obtains the information to be displayed via the communication device 23 and displays the information to be displayed on the display device 25. This specific operation is the same in subsequent processes, and their explanation may be omitted.
[0050] [Screen example] Next, we will describe the various screens that assist in the selection of a learning model. Hereafter, these various screens that assist in the selection of a learning model may be collectively referred to as support screens. Support screens are the UI screens of interactive tools. UI stands for User Interface. Support screens are displayed by the display device 25 of terminal 20, and input to support screens is performed by the input device 24 of terminal 20.
[0051] Figure 2 is a schematic diagram showing an example of a settings screen according to Embodiment 1. In Figure 2, settings screen G1 is displayed as a support screen. Settings screen G1 is a screen for setting the input images to be input to the learning model and the selection conditions for the learning model.
[0052] The settings screen G1 displays a directory selection area R1 for selecting the directory where the input image is stored. The directory selection area R1 may display directories within the server device 10, for example. The terminal 20 may configure the processor 21 to receive input from the user via the input device 24 and select which directory to select. The directory displayed in the directory selection area R1 may be a directory other than one within the server device 10, for example, a directory on a server in the cloud. In the example in Figure 2, directory 101 is selected.
[0053] Furthermore, the settings screen G1 displays an input image setting area R2 for setting the input images to be input to the learning model. The input image setting area R2 displays a list of images stored in the directory selected in the directory selection area R1. The terminal 20 may set which images to input to the learning model based on the input from the user via the input device 24, using the processor 21. Multiple images may be set to be input to the learning model. Also, images stored in each of multiple directories may be set as input images. For example, multiple images stored in directory 100 and multiple images contained in directory 101 may be set as input images. In the example in Figure 2, image b stored in directory 101 is selected. In the description of this embodiment, the selected state may be referred to as the selected state. The settings screen G1 displays an image display area R3 for displaying the selected image. In the example in Figure 2, the image display area R3 displays image b, which is in the selected state.
[0054] Furthermore, the settings screen G1 displays a selection condition setting area R4 for setting the selection criteria for the learning model. The selection condition setting area R4 may display, for example, the items of the selection criteria. Terminal 20's processor 21 may receive input from the user via the input device 24 and set specific numerical values for each item of the selection criteria. Examples of selection criteria to be set include authentication accuracy of 95% or higher, authentication accuracy of 90% or higher for input images containing the attribute "glasses", and processing speed of less than 1 second. Other examples of selection criteria to be set include conditions related to the operating environment, such as specifying the GPU or CPU to be used. Note that these are examples for illustrative purposes only, and the selection criteria are not limited to these. The selection condition setting area R4 may have, for example, the conditions for each item of the selection criteria pre-set. For example, a sample of pre-created selection criteria may be displayed in the selection condition setting area R4. For example, a user of the support system 5 can adjust the selection criteria to their desired conditions based on the sample selection criteria.
[0055] Furthermore, the settings screen G1 displays a results display settings area R5 for setting whether or not to display multiple candidate learning models after the selection of a learning model. Terminal 20's processor 21 may receive input from the user via the input device 24 and set whether or not to display multiple candidate learning models. The settings in the results display settings area R5 will be described later with reference to Figures 4 and 5.
[0056] On the settings screen G1, a "Start Model Selection" button b1 is displayed as a button to instruct the selection of a learning model. When the processor 21 receives a press of the "Start Model Selection" button b1 via the input device 24, it sends a learning model selection instruction to the server device 10 via the communication device 23. When the processor 11 of the server device 10 receives the learning model selection instruction via the communication device 13, it selects a learning model after evaluating the quality of the input image according to the instruction. The quality of the input image here is different from the quality of the input image (image quality) considered when setting the registration image, and means whether the input image has a certain degree of agreement with the registration image. Simply put, the quality of the input image here means the matching score of the input image. The matching score calculated during the learning model selection process and the estimation results of attributes contained in the input image may be stored in memory 12.
[0057] Figure 3 is a schematic diagram showing an example of the image evaluation result screen according to Embodiment 1. In Figure 3, the image evaluation result screen G2 and the image evaluation result screen G3 are displayed as support screens. The image evaluation result screen G2 is a screen that shows the evaluation result of the quality of the input image set in the settings screen G1.
[0058] When the processor 11 of the server device 10 receives a learning model selection instruction via the communication device 13, it performs a quality evaluation of the set input images before selecting a learning model. Specifically, the processor 11 inputs each of the set input images into an image quality evaluation model that calculates a matching score according to the image input, and obtains the calculated matching score. The image quality evaluation model may be prepared separately from the learning model, or a model suitable for evaluating the quality of input images may be set as the image quality evaluation model from among the learning models. For example, the learning model with the lowest average matching score may be set as the image quality evaluation model. The average matching score calculated for each learning model is obtained during the process of selecting a learning model in the past. Then, if the number of images among the set input images whose matching score calculated by the image quality evaluation model is above a predetermined second threshold is less than a predetermined number, the processor 11 displays the image evaluation result screen G2 on the display device 25 as the evaluation result of the input image quality.
[0059] The image evaluation results screen G2 displays information I1, which indicates the following: Information I1 indicates that there are insufficient images of sufficient quality to be used for selecting a learning model from the set input images, that the reliability of the selection result may be low if the learning model selection process proceeds as is, and that resetting the input images is recommended.
[0060] Furthermore, the image evaluation results screen G2 displays an input image selection area R6 for selecting a set input image. The image evaluation results screen G2 also displays an image display area R7 that displays the input image selected in the input image selection area R6. In the example in Figure 3, image e is selected in the input image selection area R6, and the selected image e is displayed in the image display area R7. The input image selection area R6 may also display a list of input images that do not have sufficient quality, that is, images whose matching score calculated by the image quality evaluation model is less than a predetermined second threshold.
[0061] Furthermore, on the image evaluation results screen G2, an input image reset button b2 is displayed as a button to instruct the user to reset the input image. When the input image reset button b2 is pressed, the screen displayed on the display device 25 transitions from the image evaluation results screen G2 to the settings screen G1. Also, on the image evaluation results screen G3, an input image non-reset button b3 is displayed as a button to instruct the user not to reset the input image. When the input image non-reset button b3 is pressed, the screen displayed on the display device 25 transitions from the image evaluation results screen G2 to the image evaluation results screen G3. Note that when the input image non-reset button b3 is pressed, the content that will be displayed on the image evaluation results screen G3 described later may be displayed below the input image reset button b2 and the input image non-reset button b3. In other words, the content that will be displayed on the image evaluation results screen G3 described later may be displayed consecutively within the image evaluation results screen G2.
[0062] The image evaluation results screen G3 displays information I2, which indicates the following: Information I2 indicates that a learning model that satisfies the selection criteria can be selected based on the past matching scores calculated by each learning model, and that selecting a learning model in this way may result in a more reliable selection result than performing the selection process using the currently set input image.
[0063] Furthermore, the image evaluation results screen G3 displays an input image use button b4, which instructs the system to select a model using the currently set input image. When the input image use button b4 is pressed, the learning model selection process is performed using the currently set input image. Additionally, the image evaluation results screen G3 displays an existing score use button b5, which instructs the system to select a model based on existing matching scores obtained in past learning model selection processes. When the existing score use button b5 is pressed, the learning model is selected, for example, as follows: The processor 11 may select the learning model that satisfies the set selection conditions and has the largest average value of previously calculated (in other words, output) matching scores. The processor 11 may then preferentially display the selected learning model on the selection results screen showing the learning model selection results. Note that the selection results screens shown in Figures 4, 5, and 6 below show the results of the selection process performed using the input image, without relying on past matching scores.
[0064] Figure 4 is a schematic diagram showing a first example of the selection results screen according to Embodiment 1. In Figure 4, the selection results screen G4 is displayed as a support screen. The selection results screen is a screen that shows the selection results of the learning model, and the selection results screen G4 shown in Figure 4 is a screen that displays the learning model that is best suited to the application site from among the selected learning models. If it is set not to display multiple learning model candidates in the result display setting area R5 of the setting screen G1, then, as shown in the selection results screen G4, one learning model that is best suited to the application site from among the selected learning models will be displayed.
[0065] The optimal learning model for a given application might be, for example, the model with the highest average value of the average matching scores for each tag (per person) of the input images captured at the application site. Alternatively, the optimal learning model might be the model with the highest average matching score for the input images containing the attribute with the highest proportion among the input images captured at the application site. Alternatively, the optimal learning model for a given application might be determined by weighting these factors. The priority for weighting may be predetermined by the user. For example, the priority of the average value of the average matching scores for each tag of the input images captured at the application site might be set highest. Alternatively, the priority of the average matching score for the input images containing the attribute with the highest proportion among the input images captured at the application site might be set highest. Note that what constitutes an optimal learning model for a given application site is not limited to the examples above and can be arbitrarily set by the user.
[0066] The selection results screen G4 displays information I3 indicating the optimal learning model for the application site. In the example in Figure 4, Model A is displayed as the optimal learning model for the application site on the selection results screen G4.
[0067] Furthermore, the selection results screen G4 displays a directory selection area R8, an input image selection area R9, an image display area R10, and a table T1 for checking the matching score for each input image calculated by the selected learning model. The directory selection area R8 displays a list of directories where the input images input to the learning model are stored. The input image selection area R9 displays a list of input images stored in the directories selected in the directory selection area R8. The image display area R10 displays the input images selected in the input image selection area R9. Table T1 shows the attributes and matching scores of the selected input images in tabular format. In the example in Figure 4, image b stored in directory 101 is displayed in the image display area R10, and the matching score of image b calculated by model A and the attributes contained in image b are shown in table T1. Note that, as will be explained later with reference to Figure 8, the attributes contained in the input image are not necessarily estimated by model A, but only need to be estimated by at least one of the learning models used in the selection process.
[0068] Furthermore, the selection results screen G4 displays Table T2, which shows in tabular format the average matching score for each attribute contained in the input images input to the learning model, and the average of the average matching scores for each attribute. For example, from Table T2 shown in Figure 4, it can be seen that the average matching score for images containing the attribute "mask" among the input images input to Model A is "500".
[0069] Here, the average of the average matching scores for each attribute may be equal to the average of the average matching scores for each tag (for each person). First, the average of the average matching scores for each attribute is calculated as follows: That is, processor 11 calculates the average matching score for a specific attribute of a specific tag (person). For example, processor 11 calculates the average matching score for the attribute "mask" of a certain tag (person). Processor 11 further calculates the average matching score for the relevant specific attribute (e.g., "mask") for each tag (for each person). This allows, for example, the average matching score for the attribute "mask" for each tag (for each person) to be calculated. Then, processor 11 can further calculate the average of the average matching scores for the attribute "mask" for each tag (for each person) (in the example in Figure 4, as shown in table T2, "500"). In the example in Figure 4, processor 11 calculates the average matching score for the attribute "blur" as "500", the average matching score for the attribute "glasses" as "600", and the average matching score for the attribute "front view" as "800". Then, in the example in Figure 4, processor 11 calculates the average of the average matching scores for each of these attributes as "600".
[0070] Thus, the processor 11 can determine the average matching score of a certain attribute by first calculating the average matching score of that attribute for each tag (for each person), and then further calculating the average of the calculated average matching scores for that attribute for each tag (for each person). The processor 11 can then further calculate the average of the average matching scores for each attribute.
[0071] If each input image contains some attribute, the average of the average matching scores for each attribute shown in Table T2 is equal to the average of the average matching scores for each tag (or person). However, if there are input images that do not contain any attributes, the average of the average matching scores for each attribute may not match the average matching scores for each tag (or person). This is because, when there are input images that do not contain any attributes, the matching score of those input images is not used in calculating the average matching score for each attribute. Table T2 may also be configured to show the average matching score for input images that do not contain any attributes. Furthermore, when there are input images that do not contain any attributes, Table T2 may be configured to show the average matching score for each tag (or person) in addition to the average matching score for each attribute. The same applies to Table T4 shown in Figure 6.
[0072] Note that the matching score values shown in Figure 4 are illustrative examples and do not limit the performance of the learning model. This also applies to the following explanation.
[0073] Furthermore, the selection results screen G4 displays attribute-specific performance information 60, which allows users to compare the performance of the learning model for each attribute. In the example in Figure 4, the attribute-specific performance information 60 is represented by a pentagon, showing the performance of Model A for each attribute: dark, mask, blur, 45-degree tilt, and backlight. Note that the performance for an attribute refers to the authentication accuracy for images containing that attribute. Which attribute's performance the attribute-specific performance information 60 shows may be set in advance by the user, for example. Alternatively, the attribute-specific performance information 60 may show the performance of attributes that appear frequently in the input image, in other words, the percentage of attributes that appear in the input image.
[0074] Furthermore, the selection results screen G4 displays processing speed information 70, which shows the processing speed of the learning model. In the example in Figure 4, the processing speed information 70 shows the processing speed when using the CPU and the processing speed when using the GPU for Model A in graph format. Note that the processing speed information 70 may also show the processing speed of the learning model in a format other than graph format (for example, in tabular format).
[0075] Furthermore, the selection results screen G4 displays Table T3, which shows information comparing the presence or absence of related functions in a tabular format. The information comparing the presence or absence of related functions may include, for example, version information of the face detection function of the learning model, information on the presence or absence of the field authentication log acquisition function, and information on the presence or absence of the impersonation prevention function.
[0076] Furthermore, the selection results screen G4 displays a selection reason explanation area R11 to explain why the learning model displayed on the selection results screen G4 was selected and is displayed preferentially. The selection reason explanation area R11 includes supplementary information SI that displays text explaining the reason for selecting the learning model. For example, if the proportion of input images containing a particular attribute is high, the system may be configured to preferentially select a learning model that satisfies the selection criteria and has a high average matching score for input images containing that particular attribute. In this case, the selection reason explanation area R11 indicates in the supplementary information SI that a learning model with high performance for that particular attribute was selected. The selection reason explanation area R11 may also display graphs showing the frequency of occurrence for each attribute. Hereinafter, the selection reason explanation area R11 may be referred to as the explanation section.
[0077] The processor 11 retrieves feature information of the learning model from memory 12 and creates a graph, for example, processing speed information 70, through a graph creation process. The processor 11 retrieves feature information of the learning model from memory 12 and creates a table, for example, table T1, through a table creation process.
[0078] In this specification, the process of having a learning model calculate a matching score for an input image, and the various processing using the calculated matching score (for example, calculating the average value of the matching score for each tag, calculating the average value of the average values of the matching score for each tag, calculating the average value of the matching score for each attribute, or creating various graphs or tables, etc.) may be referred to as the performance evaluation of the learning model. The selection result screen G4 consists of at least one performance evaluation result of the selected learning model. Furthermore, the performance evaluation result is not limited to the example shown in Figure 4. Therefore, the processor 11 may omit any of the performance evaluation results of Model A shown in Figure 4 (for example, Table T2, processing speed information 70, etc.) on the selection result screen G4, or may display further performance evaluation results.
[0079] The matching score may also be displayed as authentication accuracy. For example, a correspondence between the matching score and authentication accuracy may be defined, and the processor 11 may calculate the authentication accuracy from the matching score based on that correspondence. For example, the user may pre-configure the server device 10 to either display the matching score as a matching score or as authentication accuracy.
[0080] The selection results screen G4 displays a confirmation completion button b6. When the confirmation completion button b6 is pressed, the processor 11 saves the selection results of the learning model to memory 12. Then, it terminates the display of the selection results screen G4 on the display device 25.
[0081] On the selection results screen G4, a condition reset button b7 is displayed to instruct the user to reset the selection conditions. When the condition reset button b7 is pressed, the screen displayed by the display device 25 transitions from the selection results screen G4 to the settings screen G1. Then, the input image and selection conditions are reset.
[0082] Figure 5 is a schematic diagram showing a second example of the selection result screen according to Embodiment 1. In Figure 5, the selection result screen G5 is displayed as a support screen. The selection result screen G5 shown in Figure 5 is a screen that displays the selected learning models from among the selected learning models. In the explanation of the selection result screen G5, parts that overlap with the explanation of the selection result screen G4 may be omitted or simplified.
[0083] The selection results screen G5 displays information I5 indicating that candidate models can be recommended to the application site. A candidate model is one or more learning models that meet the selection criteria in the learning model selection process. There may be an upper limit on the number of candidate models.
[0084] The selection results screen G5 displays the candidate model selection area R12. The candidate model selection area R12 displays a list of candidate models.
[0085] If the setting screen G1 results display setting area R5 is set to display multiple candidate learning models, then the selection results screen G5 will display a list of candidate models as shown in the candidate model selection area R12.
[0086] In the candidate model selection region R12, the display order of each learning model may be determined, for example, based on the average of the average values of the matching scores calculated by each learning model for each input image. For example, each learning model may be displayed in a list from top to bottom in the candidate model selection region R12 in descending order of the average value of the matching scores calculated for each input image for each tag.
[0087] The selection results screen G5 displays the performance evaluation results and other information for the learning model selected in the candidate model selection region R12. In the example in Figure 5, Table T1, Table T2, Attribute-Specific Performance Information 60, Processing Speed Information 70, and Table T3 each show the performance evaluation results and information for Model A. However, if, for example, Model B is selected in the candidate model selection region R12, the performance evaluation results and information for Model B will be displayed. For example, if Model B is selected, Table T1 will show the matching score of image b calculated by Model B.
[0088] The selection results screen G5 displays a comparison setting area R13 for setting multiple learning models to be displayed for comparison. In the comparison setting area R13, the learning models to be displayed for comparison are set based on user input. In the example in Figure 5, Model A and Model B are added to the comparison setting area R13 and set as comparison targets. The selection results screen G5 displays a display change button b8. When the display change button b8 is pressed, the screen displayed on the display device 25 transitions from the selection results screen G5 to a screen that displays a comparison of the multiple learning models set in the comparison setting area R13. Figure 6 shows the screen that displays a comparison of multiple learning models.
[0089] Figure 6 is a schematic diagram showing a third example of the selection results screen according to Embodiment 1. In Figure 6, the selection results screen G6 is displayed as a support screen. The selection results screen G6 is a screen that compares and displays multiple learning models set by the user. In the explanation of the selection results screen G6, parts that overlap with the explanation of the selection results screen G4 or the selection results screen G5 may be omitted or simplified.
[0090] The selection results screen G6 displays multiple tabs for switching the displayed screen according to the user's selection. Specifically, the selection results screen G6 displays tab TB1 for displaying a screen that compares multiple learning models. In the example in Figure 6, tab TB1 is selected, and the selection results screen G6 comparing Model A and Model B is displayed. The selection results screen G6 also displays tabs for displaying screens that show the performance evaluation results of each of the learning models being compared (for example, selection results screen G5). In the example in Figure 6, tab TB2 for displaying the performance evaluation results of Model A and tab TB3 for displaying the performance evaluation results of Model B are displayed. For example, if tab TB2 is selected, the selection results screen G5 is displayed. The selection results screen G6 also displays tab TB4 for displaying the settings screen G1. The user can switch the displayed screen by selecting each tab via the input device 24. Even if the displayed screen is switched, each tab may remain displayed at the top of the screen. For example, even if tab TB2 is selected and the display screen transitions from selection results screen G6 to selection results screen G5, each tab may remain displayed on selection results screen G5, allowing the user to, for example, display selection results screen G6 again.
[0091] Furthermore, the selection results screen G6 displays Table T4, which shows in tabular format the average matching score for each attribute for each learning model being compared, and the average of the average matching scores for each attribute. In the example in Figure 6, Table T4 shows the average matching scores for each attribute for Model A and Model B, and the average of those averages. A user who views Table T4 can, for example, confirm that Model A has higher authentication accuracy than Model B for people wearing glasses.
[0092] Furthermore, the selection results screen G6 displays attribute-specific performance information 61, which allows users to check the performance of each learning model for each attribute. In the example in Figure 6, the attribute-specific performance information 61 is shown as a pentagon, and shows the performance of Model A and Model B for each attribute: dark conditions, mask, blur, 45-degree tilt, and backlight. A user who views the attribute-specific performance information 61 can, for example, confirm that Model A has higher authentication accuracy than Model B for people in input images taken in backlight.
[0093] Furthermore, the selection results screen G6 displays processing speed information 71, which shows the processing speed of each learning model being compared. In the example in Figure 6, the processing speed information 71 shows the processing speed when using the CPU and the processing speed when using the GPU for both Model A and Model B in graph format. A user who views the processing speed information 71 can, for example, confirm that using Model A allows for faster processing than using Model B.
[0094] Furthermore, the selection results screen G6 displays Table T5, which shows a comparison in tabular format of the presence or absence of related functions for each learning model being compared. In the example in Figure 6, Table T5 shows the presence or absence of related functions for both Model A and Model B. A user who views Table T5 can confirm that the version of the face detection function is newer for Model B than for Model A.
[0095] Furthermore, the selection results screen G6 displays a comparison reset button b9 for resetting the learning model to be compared. When the comparison reset button b9 is pressed, for example, the comparison setting area R13 and the display change button b8 may be displayed on the selection results screen G6, or the display screen of the display device 25 may transition from the selection results screen G6 to the selection results screen G5. The selection results screen G5 displays the comparison setting area R13 and the display change button b8, so the user can set the learning model to be compared.
[0096] [Operation Flow] Figure 7 is a flowchart showing the overall processing of the support system 5 according to Embodiment 1. It is assumed that the setting screen G1 is displayed on the display device 25 at the start of the flowchart in Figure 7.
[0097] First, the processor 11 of the server device 10 sets the input image and the selection conditions for the learning model based on user operation via the input device 24 (step S30). The user can set the input image and the selection conditions for the learning model, for example, in the directory selection area R1, input image setting area R2, and selection condition setting area R4 of the setting screen G1 shown in Figure 2.
[0098] The processor 11 determines, based on user operation via the input device 24, whether or not a button instructing the start of selection has been pressed, that is, whether or not the instruction to start selection has been received (step S31). The button instructing the start of selection is, for example, the model selection start button b1 on the settings screen G1. For example, if the processor 11 obtains information from the input device 24 from the terminal 20 indicating that the model selection start button b1 has been pressed, it determines that the model selection start button b1 has been pressed. If the processor 11 does not obtain information from the input device 24 from the terminal 20 indicating that the model selection start button b1 has been pressed, it determines that the model selection start button b1 has not been pressed.
[0099] If processor 11 determines that it has received the instruction to start the selection process (step S31: YES), it proceeds to step S33.
[0100] If processor 11 determines that it has not received the instruction to start selection (step S31: NO), it repeats the determination in step S31. In other words, processor 11 waits until it receives the instruction to start selection. At this time, processor 11 may return to the processing in step S30.
[0101] Processor 11 selects and sets a registered image from the input images (step S32). For example, processor 11 calculates a quality score for each input image and sets the input image with the highest quality score for each tag (for each person) as the registered image. This allows each model to calculate the matching score of the input image based on the registered image with the same tags as the input image, and the input image itself.
[0102] When the processor 11 receives an instruction to start the selection process, it inputs the input images set in step S30 into the image quality evaluation model and obtains the matching score for each input image calculated by the image quality evaluation model (step S33).
[0103] The processor 11 determines, based on the matching score of each input image acquired in step S33 and a predetermined second threshold for the matching score, whether the number of input images whose matching score is equal to or greater than the predetermined second threshold is equal to or greater than a predetermined number (step S34).
[0104] First, we will explain the case where the processor 11 determines that the number of input images whose matching score is equal to or greater than the predetermined second threshold is equal to or greater than the predetermined number.
[0105] If the processor 11 determines that the number of input images with a matching score equal to or greater than the predetermined second threshold is equal to or greater than the predetermined number (step S34: YES), it executes a model performance evaluation process (step S35). Details of the model performance evaluation process in step S35 will be described later with reference to Figure 8.
[0106] Processor 11 can obtain matching scores for each input image for each trained model through model performance evaluation processing. Processor 11 can also calculate the average matching score for each tag of each input image. Furthermore, Processor 11 can estimate the attributes contained in each input image. Processor 11 can also calculate the average matching score for each attribute of each input image. Furthermore, Processor 11 can calculate the proportion of each attribute contained in each input image. Finally, Processor 11 can obtain the processing time for each trained model through model performance evaluation processing.
[0107] The processor 11 narrows down and ranks the learning models based on the selection criteria for the learning models set in step S30 and the results of the model performance evaluation process (step S36). For example, if the selection criteria include a processing time condition and a matching score condition, the processor 11 narrows down the learning models to be selected to only those that satisfy the set processing time condition. Then, the processor 11 narrows down the narrowed-down learning models to only those that satisfy the set matching score condition. The matching score condition may be, for example, a condition that the average matching score of input images with the attribute "blur" is greater than or equal to a predetermined value. After narrowing down the learning models, the processor 11 ranks the narrowed-down learning models. The ranking of the learning models may be performed by weighting based on, for example, the average value of the matching score for each tag of each input image, the average of the average values of the matching score for each tag, the average value of the matching score for each input image for each attribute, and the proportion of each attribute included in each input image. Alternatively, the ranking of the learning models may be performed based on, for example, the average value of the average values of the matching score for each tag of each input image. The method for ranking the learning models may be set arbitrarily by the user in advance.
[0108] Processor 11 uses the results of narrowing down and ranking the learning models to perform a selection result output process (step S37). Details of the selection result output process in step S37 will be described later with reference to Figure 9. After step S37, Processor 11 terminates this processing flow.
[0109] Next, we will explain the case in step S34 when the processor 11 determines that the number of input images whose matching score is equal to or greater than the predetermined second threshold is not equal to or greater than the predetermined number.
[0110] If the processor 11 determines that the number of input images whose matching score is equal to or greater than the predetermined second threshold is not equal to or greater than the predetermined number (step S34: NO), it outputs the image evaluation result (step S38). Specifically, the processor 11 displays the image evaluation result screen G2 on the display device 25.
[0111] The processor 11 determines whether or not it has received an instruction to reset the input image (step S39). The processor 11 determines whether or not it has received an instruction to reset the input image based on whether the input image reset button b2 was pressed or the input image non-reset button b3 was pressed. If the processor 11 determines that the input image reset button b2 was pressed, it determines that it has received an instruction to reset the input image. If the processor 11 determines that the input image non-reset button b3 was pressed, it determines that it has not received an instruction to reset the input image.
[0112] If the processor 11 determines that it has received an instruction to reset the input image (step S39: YES), it returns to step S30 and repeats the process. At this time, the processor 11 displays the setting screen G1 on the display device 25.
[0113] If the processor 11 determines that it did not accept the instruction to reset the input image (step S39: NO), it determines whether it accepted the instruction to use an existing matching score (step S40). More specifically, if the processor 11 determines that it did not accept the instruction to reset the input image, it displays the image evaluation result screen G3 on the display device 25. Then, the processor 11 determines whether it accepted the instruction to use an existing matching score based on whether the input image use button b4 was pressed or the existing score use button b5 was pressed. If the input image use button b4 was pressed, the processor 11 determines that it did not accept the instruction to use an existing matching score. If the existing score use button b5 was pressed, the processor 11 determines that it accepted the instruction to use an existing matching score.
[0114] If processor 11 determines that it has received an instruction to use an existing matching score (step S40: YES), it proceeds to step S36. Then, processor 11 ranks the learning models that satisfy the selection criteria based on the matching scores that each learning model has calculated on average in the past.
[0115] If processor 11 determines that it did not accept the instruction to use an existing matching score (step S40: NO), it proceeds to step S35. Then, processor 11 performs a model performance evaluation process using the input image.
[0116] Next, the model performance evaluation process in step S35 will be described with reference to Figure 8. Figure 8 is a flowchart showing the model performance evaluation process by the support system 5 according to Embodiment 1.
[0117] The processor 11 of the server device 10 selects a learning model from the list of learning models for which the matching score calculation process and attribute estimation process are incomplete (step S41). The matching score calculation process is the process described in step S42, and the attribute estimation process is the process described in step S43.
[0118] The processor 11 inputs the input image to the learning model selected in step S41, calculates a matching score, and obtains the calculation result (step S42).
[0119] Furthermore, the processor 11 inputs the input image to the learning model selected in step S41, has it estimate the attributes contained in the input image, and obtains the estimation result (step S43).
[0120] The matching score calculation process and the attribute estimation process may be performed in parallel.
[0121] The processor 11 determines whether the calculation of matching scores and attribute estimation for all input images by the learning model selected in step S41 has been completed (step S44).
[0122] If the processor 11 determines that the calculation of matching scores and attribute estimation for all input images by the learning model selected in step S41 is not yet complete (step S44: NO), it returns to steps S42 and S43 and repeats the process. This allows the processor 11 to obtain the matching score for each input image by the learning model selected in step S41. The processor 11 can also obtain the estimated attribute results for each input image. This allows the processor 11 to associate each input image with the attributes it contains.
[0123] If the processor 11 determines that the calculation of matching scores and attribute estimation of input images by the learning model selected in step S41 is complete (step S44: YES), it calculates the average of the matching scores of all input images of the same person, i.e., all images with the same tag, calculated by the learning model, for each attribute and for each tag (step S45). In short, the processor 11 calculates the average of the matching scores of all input images for each tag (each person) for which matching scores are to be calculated. This calculates the average of the matching scores for the number of tag types (the number of people to be authenticated). In addition, the average of the matching scores for each attribute of the same tag is calculated. For example, the average matching score for the attribute "mask" or the average matching score for the attribute "glasses" of a certain person is calculated.
[0124] The processor 11 calculates the average matching score for each attribute of the input image, and further calculates the overall average of the matching scores (step S46). This allows, for example, the average of the average matching scores for each attribute "mask" of multiple people to be calculated. Furthermore, by calculating the average of the average matching scores for each tag (for each person) calculated in step S45, the overall average of the matching scores is calculated.
[0125] Note that the order of processing in steps S45 and S46 is not limited to the order shown in Figure 8.
[0126] The processor 11 calculates the proportion of each attribute present in the input image (step S47).
[0127] The processor 11 determines whether the matching score calculation process and attribute processing estimation process have been completed for all selected learning models (step S48).
[0128] If processor 11 determines that the matching score calculation process and attribute processing estimation process have not been completed for all selected learning models (step S48: NO), it returns to step S41 and repeats the process. In this way, processor 11 can calculate the average value of the matching score for each tag and attribute of all input images, the average value of the matching score for each tag of all input images, and the average value of the matching score for each attribute of all input images for each selected learning model.
[0129] Furthermore, if attribute estimation processing is performed once for all input images by one of the selected learning models, attribute estimation processing by other learning models may be omitted. In this embodiment, it is assumed that the attribute estimation results for the input images are the same for each learning model. Under this assumption, if attribute estimation processing is performed for all input images by a certain learning model, omitting attribute estimation processing by other learning models reduces the processing load and lowers costs such as time and power consumption. Also, if attribute estimation processing is performed once by a certain learning model and attribute estimation processing by other learning models is omitted, the processor 11 selects the learning model for which only the matching score calculation process is incomplete in the second and subsequent steps S41. Furthermore, in this case, the processor 11 omits step S43.
[0130] If the processor 11 determines that the matching score calculation process and attribute processing estimation process have been completed for all selected learning models (step S48: YES), it terminates this processing flow.
[0131] Next, with reference to Figure 9, the selection result output process in step S37 of the flowchart in Figure 7 will be explained. Figure 9 is a flowchart of the selection result output process by the support system 5 according to Embodiment 1.
[0132] The processor 11 of the server device 10 determines whether or not it is configured to display multiple candidate learning models (step S51). Although not shown in the flowchart of Figure 7, the processor 11 accepts a setting on the setting screen G1 to determine whether or not to display multiple candidate learning models in the result display setting area R5. Based on the setting in the result display setting area R5, the processor 11 determines whether or not it is configured to display multiple candidate learning models.
[0133] If the processor 11 determines that the setting to display multiple candidate learning models is not enabled (step S51: NO), it prioritizes displaying the evaluation result of the learning model with the highest average matching score among the learning models that meet the selection criteria on the display device 25 (step S52). The average matching score here may be the value obtained by further averaging each of the average matching scores for each tag (for each person) calculated in the process of step S45 of the flowchart shown in Figure 8, in step S46, that is, the average of the average matching scores for each tag. Alternatively, the average matching score here may be the average matching score of a pre-set tag among the average matching scores for each tag (for each person). Alternatively, the average matching score here may be the average of each of the average matching scores of multiple pre-set tags among the average matching scores for each tag (for each person). For example, if the matching score for a specific person is to be given priority consideration, the evaluation result of the learning model with the highest average matching score of the input images tagged with that specific person may be prioritized in displaying the result. Alternatively, the average matching score here may be the average matching score of a specific attribute that has been pre-configured.
[0134] For example, the processor 11 displays the selection results screen G4 on the display device 25. The selection results screen G4 displays, for example, the evaluation results of the learning model that was ranked 1st as a result of the ranking performed by the processor 11 in step S36 of the flowchart in Figure 7. After step S52, the processor 11 proceeds to step S56, which will be described later.
[0135] If processor 11 determines that it is set to display multiple candidate learning models, (Step S51: YES) The processor 11 displays a list of learning models that meet the selection criteria and the evaluation results of the selected learning models in that list on the display device 25 (Step S53). For example, the processor 11 displays a candidate model selection area R12 that displays a list of learning models and a selection result screen G5 on the display device 25 that shows the evaluation results of the learning models selected in the candidate model selection area R12. Then the processor 11 proceeds to step S54.
[0136] The processor 11 determines whether or not it has received an instruction to display a comparison of the evaluation results of the learning models (step S54). The processor 11 determines that it has received an instruction to display a comparison of the evaluation results of the learning models if the learning model to be compared is set in the comparison setting area R13 and the display change button b8 is pressed. The processor 11 determines that it has not received an instruction to display a comparison of the evaluation results of the learning models if the display change button b8 is not pressed.
[0137] If processor 11 determines that it has not received an instruction to display a comparison of the evaluation results of the learning model (step S54: NO), it proceeds to step S56.
[0138] If the processor 11 determines that it has received an instruction to display a comparison of the evaluation results of the learning models (step S54: YES), it displays the evaluation results of the multiple learning models set as the comparison target on the display device 25 (step S55). For example, the processor 11 displays the selection result screen G6 on the display device 25.
[0139] The processor 11 determines whether or not it has received an instruction to reset the input image or selection conditions (step S56). The processor 11 determines that it has received an instruction to reset the input image or selection conditions if the condition reset button b7 is pressed on the selection result screen G4, selection result screen G5, or G6. Alternatively, the processor 11 determines that it has received an instruction to reset the input image or selection conditions if the tab TB4 for displaying the settings screen G1 is selected on the selection result screen G6.
[0140] If the processor 11 determines that it has received an instruction to reset the input image or selection conditions (step S56: YES), it displays the setting screen G1 on the display device 25 (step S57). Then, the processor 11 terminates this processing flow.
[0141] If the processor 11 determines that it has not received instructions to reset the input image or selection conditions (step S56: NO), it accepts input indicating the end of the evaluation result confirmation (step S58). The processor 11 accepts input indicating the end of the evaluation result confirmation based on the press of the confirmation end button b6. Then, the processor 11 terminates this processing flow.
[0142] In this way, the processor 11 of the server device 10 of the support system 5 can display the setting screen G1 on the display device 25 of the terminal 20. The setting screen G1 is a screen for setting multiple images to be input to the learning model and the selection criteria for the learning model. The user can set multiple images to be input to the learning model and the selection criteria for the learning model on the setting screen G1. The processor 11 can evaluate the performance of each learning model based on the multiple images and selection criteria set on the setting screen G1. Specifically, for example, the processor 11 can obtain a matching score for each of the multiple images by inputting the set multiple images into each learning model. The processor 11 can then perform various processes based on the matching scores, such as calculating the average matching score for each tag (for each person), creating graphs or tables, etc. The processor 11 can display a selection result screen G4 (or selection result screen G5, selection result screen G6) containing at least one performance evaluation result on the display device 25. This allows the user of the support system 5 to visually confirm the selected learning model. Users can check the performance, processing speed, and other attributes of each learning model displayed on the display device 25, making it easy to select a learning model.
[0143] (Summary of Embodiment 1) The following technology is disclosed based on the description of Embodiment 1 above. Note that the components etc. in Embodiment 1 are examples, but are not limited to these.
[0144] (Technology 1) A learning model selection support system (e.g., support system 5) that assists in the selection of a learning model comprises a display device (e.g., terminal 20, display device 25) and a processing device (e.g., server device 10, processor 11). The processing device displays a setting screen (e.g., setting screen G1) on the display device for setting multiple images and selection conditions for a learning model. Based on the multiple images and selection conditions set on the setting screen, it evaluates the performance of each learning model and displays a selection result screen (e.g., selection result screen G4, selection result screen G5) on the display device that includes at least one performance evaluation result for each learning model (e.g., table T2, attribute-specific performance information 60, processing speed information 70).
[0145] This allows the learning model selection support system to display a selection results screen that includes the results of the learning model performance evaluation, and to provide the learning model selection results. As a result, users can visually confirm the performance of the learning model selected by the learning model support system, making it easier to select a learning model.
[0146] (Technology 2) In the learning model selection support system described in Technology 1, the learning model is a model that calculates and outputs a matching score indicating the similarity between an input image and a set registered image in response to an image input, the multiple images are images taken at the application site of the learning model, and the selection criteria may include at least one of the following: conditions for the operating environment of the learning model, conditions for processing time, and conditions for the matching score.
[0147] This allows the learning model selection support system to select, for example, a learning model suitable for the specific field where facial recognition processing is performed. Furthermore, the user can set at least one of the following conditions: the operating environment of the learning model, the processing time, and the matching score.
[0148] (Technology 3) In the learning model selection support system described in Technology 2, the processing device may set the registered images as targets for calculating the matching score and evaluate the performance of the learning model by calculating the average value of the matching score for each of the multiple images input to the learning model.
[0149] This allows the learning model selection support system to calculate an average matching score for multiple images for each learning model and for each target of matching score calculation, such as a person.
[0150] (Technology 4) In the learning model selection support system described in Technology 2 or 3, each of the multiple images has attributes, and the processing device may evaluate the performance of the learning model by calculating the average value of the matching scores for each attribute of the multiple images input to the learning model.
[0151] This allows the learning model selection support system to calculate the average matching score for each attribute of a set number of images for each learning model.
[0152] (Technology 5) In the learning model selection support system described in any one of technologies 2 to 4, the processing device may evaluate the performance of each learning model using multiple images from among the multiple images set on the settings screen, the matching score of which is equal to or greater than a predetermined threshold.
[0153] This allows the learning model selection support system to select a learning model using an image from among multiple pre-configured images that is of sufficient quality to evaluate the performance of the learning model.
[0154] (Technology 6) In the learning model selection support system described in Technology 3, the processing unit may preferentially display the performance evaluation results of the learning model that satisfies the selection criteria and has the highest average matching score on the selection results screen.
[0155] This allows the learning model selection support system to select a learning model that outputs a high matching score on average for a set number of images, and to provide the selection result to the user.
[0156] (Technology 7) In the learning model selection support system described in Technology 4, the processing device may calculate the proportion of each attribute included in multiple images and preferentially display the performance evaluation result of the learning model with the highest average matching score of the attribute with the highest calculated proportion on the selection result screen.
[0157] This allows the learning model selection support system to select a learning model with a high matching score for images containing specific attributes and provide the selection results to the user. Furthermore, the learning model selection support system can select a learning model that is more suitable for the application site.
[0158] (Technology 8) In the learning model selection support system described in Technology 5, if the number of images among the multiple images set on the settings screen whose matching score is equal to or greater than a predetermined threshold is less than a predetermined number, the processing device may preferentially display on the selection results screen the learning model that satisfies the selection conditions and has the largest average matching score output in the past.
[0159] As a result, when the number of images in a given set of images that are of sufficient quality to evaluate the performance of the learning model is small, the learning model selection support system can select a learning model that has historically produced a high matching score on average and provide the selection result to the user. This means that the learning model selection support system may be able to provide the user with a more reliable selection result than when selecting a learning model based on a set set of images.
[0160] (Technology 9) In the learning model selection support system described in any one of technologies 1 to 8, the selection result screen includes tabs (e.g., tab TB1, tab TB2, tab TB3) for switching between multiple selection result screens for comparing the performance evaluation results of each learning model, and a tab (e.g., tab TB4) for switching the display on the display device from the selection result screen to the settings screen, and the processing device may switch the selection result screen to one of the multiple selection result screens or the settings screen in response to input to the tabs.
[0161] This allows users to compare multiple learning models and return to the settings screen from the selection results screen.
[0162] (Technology 10) In the learning model selection support system described in Technology 7, the selection result screen may have an explanatory section (for example, selection reason explanation area R11) that visualizes the reason why the performance evaluation result of the learning model with the largest average matching score of the attribute with the highest calculated percentage is displayed preferentially.
[0163] This allows users to see why a particular learning model was selected.
[0164] (Technology 11) A terminal (e.g., terminal 20) that assists in the selection of a learning model includes a display device (e.g., display device 25) and a processor (e.g., processor 21). The processor displays a settings screen on the display device for setting multiple images and selection criteria for a learning model, and displays a selection results screen on the display device that includes at least one result of performance evaluation for each learning model based on the multiple images and selection criteria set on the settings screen.
[0165] This allows the device to achieve the same effect as in Technology 1.
[0166] (Technology 12) A method for assisting in the selection of a learning model includes: displaying a settings screen on a display device for setting multiple images and selection conditions for the learning model; performing a performance evaluation of each learning model based on the multiple images and selection conditions set on the settings screen; and displaying a selection results screen on the display device that includes at least one performance evaluation result for each learning model.
[0167] This allows the learning model selection support method to achieve the same effect as Technique 1.
[0168] While embodiments have been described above with reference to the attached drawings, this disclosure is not limited to such examples. It is clear to those skilled in the art that various modifications, alterations, substitutions, additions, deletions, and equivalents can be conceived within the scope of the claims, and these are also understood to fall within the technical scope of this disclosure. Furthermore, the components of the embodiments described above can be combined in any way without departing from the spirit of the invention. [Industrial applicability]
[0169] The technology disclosed herein is useful as a learning model selection support system, terminal, and learning model selection support method. [Explanation of Symbols]
[0170] 5. Support System 10 Server devices 11 processors 12 memory 13 Communication devices 20 devices 21 processors 22 memory 23 Communication devices 24 Input Devices 25 Display Devices G1 Settings Screen G2, G3 Image Evaluation Results Screen G4, G5, G6 selection result screen TB1, TB2, TB3, TB4 tabs
Claims
1. A learning model selection support system that assists in the selection of a learning model, Equipped with a display device and a processing device, The aforementioned processing apparatus is A settings screen for setting multiple images and selection criteria for the learning model is displayed on the display device. Based on the multiple images and selection conditions set in the aforementioned settings screen, the performance of each learning model is evaluated. The display device will show a selection results screen that includes at least one performance evaluation result for each of the aforementioned learning models. A system to support the selection of learning models.
2. The aforementioned learning model is a model that, in response to an image input, calculates and outputs a matching score indicating the similarity between the input image and a set registered image. The aforementioned multiple images are images taken at the application site of the learning model, The selection criteria include at least one of the following: the operating environment conditions of the learning model, the processing time conditions, and the matching score conditions. A learning model selection support system according to claim 1.
3. The aforementioned processing apparatus is The registered images are set for each target of the calculation of the matching score, The performance of the learning model is evaluated by calculating the average value of the matching score for each of the multiple images input to the learning model. The learning model selection support system according to claim 2.
4. Each of the aforementioned multiple images has an attribute, The processing device evaluates the performance of the learning model by calculating the average value of the matching scores for each attribute of the multiple images input to the learning model. The learning model selection support system according to claim 3.
5. The aforementioned processing apparatus is The performance of each learning model is evaluated using multiple images from among the multiple images set in the aforementioned settings screen, where the matching score is above a predetermined threshold. The learning model selection support system according to claim 2.
6. The aforementioned processing apparatus is The performance evaluation results of the learning model that satisfies the selection criteria and has the highest average matching score will be displayed preferentially on the selection results screen. The learning model selection support system according to claim 3.
7. The aforementioned processing apparatus is The proportion of each attribute included in the plurality of images is calculated. The performance evaluation results of the learning model with the highest average matching score for the attribute with the highest calculated proportion will be displayed preferentially on the selection results screen. The learning model selection support system according to claim 4.
8. The aforementioned processing apparatus is If, among the multiple images set on the settings screen, the number of images whose matching score is equal to or greater than the predetermined threshold is less than the predetermined number, the learning model that satisfies the selection conditions and has the highest average matching score output in the past will be preferentially displayed on the selection results screen. The learning model selection support system according to claim 5.
9. The selection results screen includes a tab for switching between multiple selection results screens for comparing the performance evaluation results of each learning model, and a tab for switching the display on the display device from the selection results screen to the settings screen. The processing device switches the selection result screen to one of the multiple selection result screens or the settings screen in response to input to the tab. A learning model selection support system according to claim 1.
10. The selection results screen includes an explanatory section that visualizes the reason why the performance evaluation results of the learning model with the highest average matching score for the attribute with the highest calculated percentage are displayed preferentially. The learning model selection support system according to claim 7.
11. A terminal that assists in the selection of a learning model, Equipped with a display device and a processor, The aforementioned processor, A settings screen for setting multiple images and selection criteria for the learning model is displayed on the display device. A selection results screen is displayed on the display device, which includes at least one of the performance evaluation results for each learning model based on the multiple images set in the settings screen and the selection conditions. Terminal.
12. A method for supporting the selection of a learning model, Displaying a settings screen on a display device for setting multiple images and selection criteria for the learning model, The performance of each learning model is evaluated based on the multiple images set in the settings screen and the selection conditions. The process includes displaying a selection results screen on the display device that includes at least one performance evaluation result for each of the aforementioned learning models, A method for supporting the selection of learning models.
Citation Information
Patent Citations
Evaluation device, evaluation method, and program
JP2022162454A