Input support device, input support method, and program
The input support device integrates machine learning-based estimation with user input to improve feature point positioning accuracy by providing feedback on discrepancies, addressing errors in existing technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NEC CORP
- Filing Date
- 2022-03-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies, such as those described in Patent Document 1, fail to effectively integrate user judgment with device estimation to accurately determine the position of feature points in images, leading to potential errors in feature point specification.
An input support device and method that utilizes a machine learning model to estimate feature point positions, assists users in specifying these points based on both estimated and user-specified locations, and provides feedback on discrepancies to enhance accuracy.
Enables more accurate determination of feature point positions by combining device estimation with user input, reducing errors and allowing users, even with little experience, to specify feature points with precision, particularly for images with individual variations.
Smart Images

Figure 0007859485000001 
Figure 0007859485000002 
Figure 0007859485000003
Abstract
Description
Technical Field
[0001] This disclosure relates to an input support device, an input support method, and a program.
Background Art
[0002] As a preparation for performing information processing using an image, a user may perform an operation of checking the image and designating the position of feature points in the image. For example, in order to create a database of a collation system that identifies a person using a face image, a user who is a database creator may designate the position of feature points for each of a plurality of known person's face images. Also, for example, in order to identify a person using a face image by a collation system, a user who is a search person may input the position of the feature points of the face of the person to be identified into the collation system. Also, for example, in order to prepare teacher data of a machine learning model, a user who is a teacher data creator may also designate the position of feature points of an image. Thus, for any purpose, an operation of a user designating the position of feature points in an image may be performed.
[0003] By the way, as a related technique, a contour detection device described in Patent Document 1 is known. According to this contour detection device, for an automatic detection result of feature points of the contour of a nail from an image, feature points with low reliability as a contour are specified.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] As shown in Patent Document 1, the estimation results obtained by the device may be incorrect. On the other hand, the user's (human) judgment may also be incorrect. Therefore, instead of focusing solely on errors in the device's estimation results, it is possible to determine the location of feature points more appropriately by focusing on the difference between the device's estimation results and the user's judgment. However, while the technology described in Patent Document 1 makes it easy to identify errors in the automatic detection results of feature points, it does not perform processing that focuses on the difference between the location estimated by the device for the feature points and the location specified by the user.
[0006] Therefore, one of the objectives that the embodiments disclosed in this specification aim to achieve is to provide an input support device, an input support method, and a program that can more appropriately determine the position of feature points. [Means for solving the problem]
[0007] The input support device according to the first aspect of this disclosure is An estimation means for estimating the position of feature points in an input image, Input support means assists the user in specifying the position of feature points in the input image based on the position estimated by the estimation means and the position specified by the input. It holds.
[0008] In the input support method relating to the second aspect of this disclosure, The position of feature points in the input image is estimated. The system assists the user in specifying the location of feature points in the input image, based on the estimated location and the location specified by the input.
[0009] The program relating to the third aspect of this disclosure is An estimation step to estimate the position of feature points in the input image, An input support step that assists the user in specifying the position of feature points in the input image based on the position estimated in the estimation step and the position specified by the input. Have the computer execute it. [Brief explanation of the drawing]
[0010] [Figure 1] This block diagram shows an example of the configuration of an input support device according to an outline of the embodiment. [Figure 2] This is a schematic diagram showing an example of an eye image. [Figure 3] This block diagram shows an example of the functional configuration of the input support device according to Embodiment 1. [Figure 4] This is a schematic diagram illustrating an example where the user interface displays feature points at estimated locations on the input image. [Figure 5] This block diagram shows an example of the hardware configuration of the input support device according to Embodiment 1. [Figure 6] This flowchart shows an example of the operation of the input support device according to Embodiment 1. [Figure 7] This is a schematic diagram illustrating input in an incorrect order. [Figure 8] This is a block diagram showing an example of the functional configuration of the input support device according to Embodiment 2. [Figure 9] This flowchart shows an example of the operation of the input support device according to Embodiment 2. [Modes for carrying out the invention]
[0011] <Overview of the Embodiment> First, an overview of the embodiment will be described. Figure 1 is a block diagram showing an example of the configuration of the input support device 1 according to an overview of the embodiment. As shown in Figure 1, the input support device 1 has an estimation unit 2 and an input support unit 3. The input support device 1 is a device used for the user to specify the position of feature points on an input image, which is an image input to the input support device 1. The input image is, for example, a face image showing a person's face, and the feature points are, for example, facial feature points, but the input image and feature points are not limited to these. For example, the input image may show any object such as an animal, a car, or a structure. Also, the feature points only need to be predefined for each object, and their type is not limited. The following describes each component of the input support device 1.
[0012] Estimation unit 2 estimates the positions of feature points in the input image. For example, estimation unit 2 may perform estimation using a pre-trained machine learning model. Here, the machine learning model is pre-trained using pairs of images and the positions of feature points of objects depicted in the image as training data. In this training data, the positions of the feature points are pre-specified, for example, by a skilled person who specializes in specifying the positions of feature points. In other words, in this training data, the positions of the feature points are specified to be the correct positions that match the definition of feature points.
[0013] The input support unit 3 assists the user in specifying the positions of feature points in the input image, based on the positions estimated by the estimation unit 2. Here, assisting the user in specifying the positions of feature points means performing processing that presents predetermined information to the user when the user specifies the positions of feature points, such as displaying the estimated positions, outputting warnings, or displaying the order in which positions are specified. In particular, the input support unit 3 assists the user's input based on the positions estimated by the estimation means and the positions specified by the input. For example, the input support unit 3 may assist the user's input by outputting a warning based on the difference between the positions estimated by the estimation unit 2 and the positions specified by the user's input.
[0014] According to the input support device 1, the input of the user is supported using the estimation result of the estimation unit 2 and the input result. Therefore, according to the input support device 1, it is possible to determine the position of the feature point after mutually complementing the estimation by the device and the judgment by the user. Therefore, according to the input support device 1, the position of the feature point can be determined more appropriately compared to the case where this device is not used.
[0015] Generally, since it is predefined which part of the object is to be the feature point, the user designates, as the position of the feature point, a position on the input image that conforms to this definition. However, not all users who know the definition of the feature point, that is, the criteria to be satisfied as the feature point, can plot the feature point at the same position. The position of the feature point designated by the user can vary depending on the user's experience or the ambiguity of the definition of the feature point. In particular, for feature points of an object with individual differences such as a human face, a general definition that ignores individual differences is made instead of a strict definition for each individual, so the definition of the feature point can be unclear for each individual. For this reason, it is not easy for the user to plot the feature point at an appropriate position. For example, assume that it is predefined to designate a point at the outer corner of the eye 90 (see FIG. 2) as the feature point of the eye. However, since the shape and size of the eyes are different for each person, it is difficult to accurately predefine in advance the position of the feature point of the outer corner of the eye that is common to all people. For this reason, especially for users with little experience, it is difficult to specify the position of the feature point at an appropriate position based only on the definition of the feature point. Therefore, when a user with little experience performs the operation, variations can occur in the position of the feature point of the outer corner of the eye.
[0016] On the other hand, when the user uses the input support device 1, the user can receive input support based on the positions of the feature points estimated by the estimation unit 2. For example, if an estimation is made using a model learned with teacher data specifying the appropriate positions of the feature points at the outer corners of the eyes, the appropriate positions of the feature points at the outer corners of the eyes can be estimated for most input images. Based on this estimation result, it becomes possible to prevent errors in the user's feature specification. Also, based on such an estimation result, the user can understand an appropriate position that is difficult to understand from the definition of general-purpose feature points. And a user who understands the appropriate position can specify the appropriate position even if an inappropriate position is estimated when a peculiar face image or the like is input to the estimation unit 2 (for example, a machine learning model). That is, even a beginner in the input operation of specifying the position of the feature point can specify the position of the feature point equally to an expert who is well aware of the standard of the position of the feature point.
[0017] In the above description, the input support device 1 having the configuration shown in FIG. 1 has been described, but the embodiments for obtaining the above effects are not limited to the device. For example, an input support method including the above-described processing of the input support device 1, a program that performs the above-described processing of the input support device 1, or a non-temporary computer-readable medium storing the program can also obtain the same effects.
[0018] Hereinafter, embodiments of this disclosure will be described in detail while referring to the drawings. In the following description and drawings, for clarity of explanation, appropriate omissions and simplifications are made. Also, in each drawing, the same reference numerals are assigned to the same components, and duplicate explanations are omitted as necessary.
[0019] <Embodiment 1> Figure 3 is a block diagram showing an example of the functional configuration of the input support device 100 according to Embodiment 1. The input support device 100 is the device corresponding to the input support device 1 in Figure 1. In this embodiment, the input support device 100 assists in specifying the position of feature points in a face image, for example, but the input support device 100 may also assist in specifying the position of feature points in an object other than a face. The user uses the input support device 100 to input the position of feature points in an object (face). The user is, for example, a beginner who is familiar with the definition of feature points but is not accustomed to the input work. However, the user of the input support device 100 is not limited to such a user. The user of the input support device 100 may be an expert in inputting the position of feature points, or a person who is not familiar with the definition of feature points.
[0020] As shown in Figure 3, the input support device 100 includes a model storage unit 101, an input image acquisition unit 102, an estimation unit 103, a user interface unit 104, a feature point data generation unit 105, and a feature point data storage unit 106. These will be described below.
[0021] The model memory unit 101 stores a machine learning model that estimates the positions of predetermined feature points of a predetermined object depicted in an image. In this embodiment, the predetermined object is a human face. The predetermined feature points are 19 predefined points corresponding to predetermined parts of the face. Specifically, these 19 feature points are 3 points for each eyebrow, 3 points for each eye, 4 points for the nose, and 3 points for the mouth. Note that the number of feature points is merely an example, and more or fewer feature points may be defined. Of course, the selection of which parts to use as feature points is also merely an example and is not limited to the above. This machine learning model is pre-trained using machine learning such as deep learning, using pairs of images and the positions of feature points of the object depicted in the image as training data. Specifically, the machine learning model stored in the model memory unit 101 is pre-trained using the positions of feature points specified by an expert in the task of specifying the positions of feature points as training data. In other words, this machine learning model is trained using training data that shows the correct positions that match the definition of feature points.
[0022] In the configuration shown in Figure 3, the model storage unit 101 is included in the input support device 100, but the model storage unit 101 may be implemented in another device that is communicably connected to the input support device 100 via a network or the like.
[0023] The input image acquisition unit 102 acquires an image that is input to the input support device 100, which should be used to specify the position of feature points. That is, the input image acquisition unit 102 acquires an image of a predetermined object (a person's face). Typically, the input image is an image taken by an imaging device such as a camera, but it does not necessarily have to be such an image; it may also be an image of the object represented by computer graphics. The input image acquisition unit 102 may acquire the input image by receiving the input image from another device, or by reading the input image from a storage device built into the input support device 100 or a storage device connected to the input support device 100.
[0024] The estimation unit 103 corresponds to the estimation unit 2 in Figure 1. The estimation unit 103 uses the machine learning model stored in the model storage unit 101 to estimate the positions of feature points in the input image acquired by the input image acquisition unit 102. In this case, the estimation unit 103 estimates the positions of 19 feature points in the face image.
[0025] The user interface unit 104 provides a user interface that accepts input from the user specifying the location of feature points in the input image acquired by the input image acquisition unit 102, and accepts such input from the user. For example, the user interface unit 104 displays the input image and displays a UI screen on the output device 150, which will be described later, where UI (user interface) components for accepting user input are arranged. In other words, the user interface unit 104 provides a GUI (Graphical User Interface) for accepting input from the user specifying the location of feature points. The user interface unit 104 then accepts the specification of the location of feature points input via the input device 151, which will be described later.
[0026] The user interface unit 104 may also be called the input support unit. The user interface unit 104 supports the user's input for specifying the position of feature points in the input image, based on the position estimated by the estimation unit 103. Specifically in this embodiment, the user interface unit 104 supports the user's input by displaying the position estimated by the estimation unit 103 prior to the user's input of the feature point position. More specifically, the user interface unit 104 displays the feature points on the input image at the position estimated by the estimation unit 103.
[0027] Figure 4 is a schematic diagram showing an example in which the user interface unit 104 displays feature points at estimated positions on the input image 91. As shown in Figure 4, the user interface unit 104 displays 19 feature points 92 at positions estimated by the estimation unit 103. Note that in Figure 4, only some of the 19 feature points are labeled with symbols to avoid making the diagram too complex.
[0028] In this case, the user makes an input specifying the position of each feature point, while referring to the position of the feature point displayed based on the estimation results. The input specifying the position of each feature point may be an input to correct the position of the feature point displayed based on the estimation results, or an input to approve the position of the feature point displayed based on the estimation results.
[0029] Furthermore, the user interface unit 104 may assist user input by outputting a warning based on the magnitude of the discrepancy between the position specified by user input and the position estimated by the estimation unit 103. In this case, the user interface unit 104 outputs a warning, for example, if the magnitude of the discrepancy between the two exceeds a predetermined threshold. Specifically, for example, the user interface unit 104 outputs a warning notifying the user that the position specified by the user may be incorrect. This warning may be displayed on the UI screen or output as audio. The user who receives the warning can specify an appropriate position for the feature point by correcting the position of the feature point as necessary. Also, if the user who receives the warning determines that the estimation result of the estimation unit 103 is incorrect, they may confirm the specified position despite the warning.
[0030] The feature point data generation unit 105 uses the positions specified by the user input, i.e., the positions determined according to the user input, as the positions of the feature points in the input image, and generates feature point data representing the feature points of the input image for each input image. The feature point data generation unit 105 stores the generated feature point data in the feature point data storage unit 106. The feature point data generation unit 105 may also associate the input image with the feature point data of this input image and store it in the feature point data storage unit 106.
[0031] The feature point data storage unit 106 stores the feature point data generated by the feature point data generation unit 105 based on user input. In the configuration shown in Figure 3, the feature point data storage unit 106 is included in the input support device 100, but the feature point data storage unit 106 may be implemented in another device that is communicably connected to the input support device 100 via a network or the like. Furthermore, the feature point data storage unit 106 may be configured as a database.
[0032] The data stored in the feature point data storage unit 106 can be used for any purpose. That is, the purpose of the user specifying the location of feature points in an input image is arbitrary and not limited to a specific purpose. For example, feature point data may be used for matching people using facial images. Specifically, the input support device 100 may be used to enable the identification of which of several known people a person to be identified corresponds to by matching the feature points of the facial image of the person to be identified with the feature points of the facial images of several known people. In this case, the input support device 100 may be used to pre-register the feature points of the facial images of known people in a database, or it may be used to identify the feature points of the facial image of the person to be compared with the feature points stored in the database. Furthermore, the use of feature point data is not limited to image matching. For example, feature point data may be collected to generate new data from feature point data. Specifically, feature point data may be collected to generate bounding box data surrounding a predetermined part (e.g., eyes). Also, feature point data may be collected to generate statistical data of feature point locations. Furthermore, feature point data may be collected for use as training data to create machine learning models. Thus, the purpose of the user specifying the location of feature points in the input image is optional.
[0033] Figure 5 is a block diagram showing an example of the hardware configuration of the input support device 100. As shown in Figure 5, the input support device 100 includes an output device 150, an input device 151, a storage device 152, a memory 153, and a processor 154.
[0034] The output device 150 is an output device such as a display that outputs information to the outside. The display may be a flat panel display such as a liquid crystal display, plasma display, or organic EL (Electro-Luminescence) display. The output device 150 may also include a speaker. The output device 150 displays the user interface provided by the user interface unit 104.
[0035] The input device 151 is a device for the user to input information via a user interface, and is, for example, an input device such as a pointing device or a keyboard. Examples of pointing devices include a mouse, trackball, touch panel, and pen tablet. The input device 151 and the output device 150 may be configured integrally as a touch panel.
[0036] The storage device 152 is a non-volatile storage device such as a hard disk or flash memory. The model storage unit 101 and the feature point data storage unit 106 described above are implemented by, for example, the storage device 152, but may be implemented by other storage devices.
[0037] Memory 153 is composed of, for example, a combination of volatile memory and non-volatile memory. Memory 153 is used to store software (computer programs) containing one or more instructions executed by the processor 154, and data used for various processes of the input assistance device 100.
[0038] The processor 154 reads and executes software (computer programs) from the memory 153 to perform the processing of the input image acquisition unit 102, estimation unit 103, user interface unit 104, and feature point data generation unit 105 described above. The processor 154 may be, for example, a microprocessor, MPU (Micro Processor Unit), or CPU (Central Processing Unit). The processor 154 may include multiple processors. Thus, the input support device 100 has the functionality of a computer.
[0039] The program, when loaded into a computer, includes a set of instructions (or software code) for causing the computer to perform one or more functions as described in the embodiments. The program may be stored on a non-temporary computer-readable medium or a physical storage medium. Examples, but not limited to, include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray® disc or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage or other magnetic storage devices. The program may be transmitted over a temporary computer-readable medium or a communication medium. Examples, but not limited to, include temporary computer-readable medium or a communication medium that includes electrically, optically, acoustically or otherwise propagating signals.
[0040] Next, the operation of the input support device 100 will be explained with reference to a flowchart. Figure 6 is a flowchart showing an example of the operation of the input support device 100. The following explanation of the operation of the input support device 100 will be given with reference to Figure 6.
[0041] In step S100, the input image acquisition unit 102 acquires an image showing a predetermined object (a person's face). Next, in step S101, the estimation unit 103 uses the machine learning model stored in the model storage unit 101 to estimate the positions of the feature points of the input image acquired in step S100.
[0042] Next, in step S102, the user interface unit 104 provides a user interface that receives input from the user to specify the position of feature points on the input image acquired in step S100, and accepts said input from the user. At that time, the user interface unit 104 accepts the position specification by the user while assisting the user's input using the estimation result obtained in step S101. Specifically, as described above, the user interface unit 104 displays the feature points at the estimated positions on the input image. In addition, the user interface unit 104 may output a warning if the difference between the position specified by the user input and the estimated position exceeds a threshold.
[0043] Next, in step S103, the feature point data generation unit 105 generates feature point data representing the feature points of the input image, using the positions determined according to the user's input as the positions of the feature points of the input image. Then, the feature point data generation unit 105 stores the generated feature point data in the feature point data storage unit 106. Next, in step S104, the input image acquisition unit 102 determines whether or not there is another input image. That is, the input image acquisition unit 102 determines whether or not there is another input image for which the position of feature points should be specified. If there is another input image, the process returns to step S100 and the process described above is repeated. On the other hand, if there is no other input image, the process ends.
[0044] The first embodiment has been described above. According to the input support device 100, user input is supported using the estimation results of a machine learning model. With this input support device 100, the user can easily specify an appropriate position as the location of the feature point in the image. In particular, in this embodiment, the user interface unit 104 supports user input by displaying the position estimated by the estimation unit 103 prior to the user's input of the feature point position. With this configuration, users can work while looking at the estimation results of a machine learning model that has been trained using training data created by experts who are well-versed in the criteria for the location of feature points. Therefore, even users with little experience in input work to specify the location of feature points can easily specify an appropriate position as the location of the feature point in the image. In other words, even beginners in input work to specify the location of feature points can specify the location of feature points at the same level as experts who are well-versed in the criteria for the location of feature points. Furthermore, it becomes possible to limit the user's work to only the correction of the estimated position, which can be expected to reduce the workload and enable efficient work.
[0045] Furthermore, as described above, the user interface unit 104 may output a warning based on the magnitude of the discrepancy between the position specified by user input and the position estimated by the estimation unit 103. With such a configuration, it is possible to suppress the user from mistakenly specifying an inappropriate position as the position of a feature point. In other words, it is possible to suppress the occurrence of human error. In particular, with such a configuration, since the position of the feature point is evaluated using the estimation result of the estimation unit 103 and the input result, an appropriate position that takes into account the estimation by the device and the judgment by the user can be set as the position of the feature point.
[0046] Furthermore, as mentioned above, the input image is a facial image, and the feature points may be facial feature points. By using the input support device 100 for such an input image and feature points, the user can appropriately specify the position of feature points even for faces, which are objects with individual differences.
[0047] <First modified example of Embodiment 1> Next, a first modification of Embodiment 1 described above will be explained. In the embodiment described above, the user interface unit 104 displayed the position estimated by the estimation unit 103 prior to user input. However, the user interface unit 104 does not have to display the position estimated by the estimation unit 103 prior to user input. In this case, the user interface unit 104 may support user input by evaluating the user's input based on the position specified by the user's input and the position estimated by the estimation unit 103. For example, the user interface unit 104 may evaluate the magnitude of the discrepancy between the position specified by the user's input and the position estimated by the estimation unit 103 and output the evaluation result. Specifically, for example, if the magnitude of the discrepancy is less than or equal to a predetermined threshold, the user interface unit 104 may output an evaluation result notifying that a position close to the position indicated by the model's estimation result has been specified. Note that this predetermined threshold may be the same as the threshold described in Embodiment 1, i.e., the threshold for outputting a warning based on the magnitude of the discrepancy. Furthermore, if the magnitude of the deviation exceeds a predetermined threshold, the user interface unit 104 may output a warning as an evaluation result. The user who receives the warning can specify an appropriate position for the feature point by correcting the position of the feature point as necessary. The evaluation result may be displayed on the UI screen or output as audio. With this configuration, the user can obtain information to determine whether the position they specified is appropriate or not, making it easy to specify an appropriate position for the feature point. In addition, this modification, in particular, reduces the amount of support displays compared to Embodiment 1. Therefore, it is possible to prevent the user from being bothered by such displays during input.
[0048] <Second Modification of Embodiment 1> Next, a second modification of the embodiment 1 described above will be explained. In the embodiment described above, the estimation unit 103 only estimated the positions of feature points. However, the estimation unit 103 may also estimate the order of multiple feature points in the input image. For example, if it is necessary to generate feature point data representing positional information in a predetermined order for each feature point, the order of these feature points is defined in advance. In this case, the user needs to specify the positions of multiple feature points of the object according to the predetermined order. To illustrate with a specific example, for instance, the order of 19 feature points in a face input image may be defined in advance, and the user must specify the position of each feature point according to this order. In this case, it is necessary to prevent the position of the feature points from being specified in an incorrect order.
[0049] Therefore, the estimation unit 103 may use a machine learning model to estimate the positions of multiple feature points in the input image, along with the order of the feature points. The user interface unit 104 may then assist the user in specifying the positions of the feature points in that order. This helps to prevent the position of feature points from being specified in an incorrect order. Such a machine learning model is pre-trained using machine learning, such as deep learning, with training data consisting of an image, the order of each feature point of an object depicted in the image, and the position of each feature point. In other words, the estimation unit 103 performs estimation processing using a machine learning model that has learned the position of each feature point along with the order information of each feature point.
[0050] The user interface unit 104 may assist user input by, for example, displaying the estimated position of each feature point and the order of each feature point on the input image. More specifically, prior to the user inputting the position of the feature points, the user interface unit 104 may display the feature points on the input image at the positions estimated by the estimation unit 103, as well as display information indicating the order of each feature point. Here, the information indicating the order of each feature point may be, for example, a number representing the order, but it may also be a mark such as an arrow indicating the order. With such a configuration, the user can work while looking at the estimation results of the machine learning model. Therefore, even users with little experience in inputting the position of feature points can specify the position of the feature points in the appropriate order.
[0051] Furthermore, the user interface unit 104 may assist user input by outputting a warning based on the difference between the order estimated by the machine learning model and the order in which the user specifies the locations of feature points. In this case, the user interface unit 104 outputs a warning if the user specifies the locations of feature points in an order different from the estimated order. Specifically, for example, the user interface unit 104 outputs a warning notifying the user that there is a possibility that the user is specifying the locations of feature points in an order different from a predetermined order. This warning may be displayed on the UI screen or output as audio. Figure 7 is a schematic diagram illustrating input in an incorrect order. Here, as indicated by the arrow 93 shown in Figure 7, the correct input order is assumed to be the input of the location of feature point 92a, which indicates the left end of the nose, then the input of the location of feature point 92b, which indicates the apex of the nose, and then the input of the location of feature point 92c, which indicates the right end of the nose. However, suppose the user inputs to plot feature point 92a, and then inputs to plot feature point 92d, which indicates the lower end of the nose. In this case, the user interface unit 104 outputs a warning. The user interface unit 104 determines that the input order is incorrect if, for example, the position of a feature point specified by the user is close to the estimated position of a feature point (feature point 92c) that is different from the next feature point (feature point 92b) to be input, as determined from the defined order. Here, the position of a feature point specified by the user being close to the estimated position means that the difference between the two positions is less than or equal to a predetermined threshold. This predetermined threshold may be the same as the threshold described in Embodiment 1, i.e., the threshold for outputting a warning based on the magnitude of the deviation. By outputting a warning about an order error in this way, it is possible to suppress the generation of feature point data in which the position information of each feature point is arranged in an order different from the predetermined order.
[0052] The second modification of Embodiment 1 has been described above, but the second modification described above may be combined with the first modification described above.
[0053] <Embodiment 2> Next, Embodiment 2 will be described. Embodiment 2 differs from Embodiment 1 in that the machine learning model is updated. Figure 8 is a block diagram showing an example of the functional configuration of the input support device 100a according to Embodiment 2. The input support device 100a according to Embodiment 2 differs from the input support device 100 according to Embodiment 1 in that it further includes a relearning unit 107. The processing of the relearning unit 107 is also performed, for example, by the processor 154 reading and executing software (computer program) from the memory 153.
[0054] The differences from Embodiment 1 will be explained in detail below, with redundant explanations omitted as appropriate. It should be noted that both the first modified example described above and the second modified example described above can be applied to Embodiment 2.
[0055] The retraining unit 107 performs machine learning on the machine learning model again by using the combination of the input image acquired by the input image acquisition unit 102 and the positions of feature points specified by the user for this input image as training data. In other words, the retraining unit 107 performs machine learning on the machine learning model again by using the feature point data generated by the feature point data generation unit 105 as training data.
[0056] The retraining unit 107 may use only the feature point data for some images generated by the user's instruction on the location of feature points, or it may use the feature point data for all images for retraining. In particular, the retraining unit 107 may use the feature point data generated when the user instructs to confirm the specified location as the feature point location, despite a warning that the magnitude of the discrepancy between the user-specified location and the location estimated by the estimation unit 103 exceeds a threshold. Since such feature data is feature data for unique images where the model's estimation results are incorrect, retraining using such feature data can update the machine learning model to a model that can make appropriate predictions even for such images. Furthermore, by using the feature point data for all images for retraining, the machine learning model is trained with more training data, thereby improving the stability of the machine learning model. The retraining unit 107 may also retrain using the training data initially used to generate the machine learning model.
[0057] Next, the operation of the input support device 100a according to Embodiment 2 will be described with reference to a flowchart. Figure 9 is a flowchart of an example of the operation of the input support device 100a according to Embodiment 2. As shown in Figure 9, the flowchart shown here differs from the flowchart shown in Figure 6 in that step S105 is added after step S104. The differences from the flowchart shown in Figure 6 will be explained below.
[0058] If it is determined in step S104 that there are no other input images, the process proceeds to step S105. In step S105, the retraining unit 107 reads the feature point data generated based on the series of processes from step S100 to step S104 from the feature point data storage unit 106 and retrains the machine learning model stored in the model storage unit 101. The retraining unit 107 then stores the retrained machine learning model in the model storage unit 101. As a result, in the next operation, estimation will be performed using the updated machine learning model.
[0059] The second embodiment has been described above. According to the input support device 100a, the machine learning model is retrained. As a result, the machine learning model is updated as needed, and the accuracy of the model can be improved. Therefore, the estimation by the estimation unit 103 can be performed with greater accuracy.
[0060] Although the present invention has been described above with reference to embodiments, the present invention is not limited thereto. Various modifications to the structure and details of the present invention can be made within the scope of the invention as can be understood by those skilled in the art. For example, various embodiments and variations can be combined as appropriate.
[0061] Some or all of the above embodiments may also be described as follows, but are not limited to the following: (Note 1) An estimation means for estimating the position of feature points in an input image, Input support means assists the user in specifying the position of feature points in the input image based on the position estimated by the estimation means and the position specified by the input. An input support device having the following features. (Note 2) The input support means further supports the input by displaying the estimated position prior to the user's input. The input support device described in Appendix 1. (Note 3) The input support means does not display the estimated position prior to the user's input, and supports the input by evaluating the input based on the position specified by the user's input and the estimated position. The input support device described in Appendix 1. (Note 4) The input support means provides support by outputting a warning based on the magnitude of the discrepancy between the position specified by the user's input and the estimated position. An input support device as described in any one of the items 1 to 3 of the appendix. (Note 5) The estimation means estimates the positions of multiple feature points in the input image along with the order of the feature points, The input support means assists the user in specifying the positions of feature points in the order described above. An input support device as described in any one of the items 1 to 4 of the appendix. (Note 6) The input support means provides support by displaying the estimated position of each feature point and the order of each feature point on the input image. The input support device described in Appendix 5. (Note 7) The input support means provides support by outputting a warning based on the difference between the estimated order and the order in which the user specifies the positions of the feature points. An input support device as described in Appendix 5 or 6. (Note 8) The estimation means estimates the position of feature points in the input image using a pre-trained machine learning model. The system further includes a retraining means for performing machine learning on the machine learning model again by using the combination of the input image and the position specified by the user's input as training data. An input support device as described in any one of the items 1 to 7 of the appendix. (Note 9) The input image is a facial image, and the feature points are facial feature points. An input support device as described in any one of the items 1 to 8 of the appendix. (Note 10) The position of feature points in the input image is estimated. The system assists the user in specifying the location of feature points in the input image, based on the estimated location and the location specified by the input. Input assistance methods. (Note 11) An estimation step to estimate the position of feature points in the input image, An input support step that assists the user in specifying the position of feature points in the input image based on the position estimated in the estimation step and the position specified by the input. A non-temporary, computer-readable medium containing a program that causes a computer to execute a program. [Explanation of Symbols]
[0062] 1. Input support device 2 Estimation part 3. Input Support Unit 100 Input support device 100a Input support device 101 Model Memory Unit 102 Input Image Acquisition Unit 103 Estimation part 104 User Interface Section 105 Feature Point Data Generation Unit 106 Feature Point Data Storage Unit 107 Relearning Section 150 Output device 151 Input device 152 Storage device 153 memory 154 processors
Claims
1. An estimation means for estimating the position of feature points in an input image, Input support means to assist the user in specifying the position of feature points in the input image, It has, The input support means does not display the estimated position prior to the user's input, and supports the input by evaluating the input based on the position specified by the user's input and the estimated position. Input assistance device.
2. An estimation means for estimating the positions of multiple feature points in an input image along with the order of the feature points, Input support means to assist the user in specifying the position of feature points in the input image, It has, The input support means assists the user in specifying the positions of feature points in the order described above by displaying the estimated position of each feature point and information indicating the order of each feature point on the input image. The information indicating the order of each of the aforementioned feature points is either a number representing the order or an arrow indicating the order. Input assistance device.
3. The input support means supports the user's specification of feature point locations in the order described above by outputting a warning based on the difference between the estimated order and the order in which the user specifies the feature point locations. The input support device according to claim 2.
4. The input support means further supports the input by displaying the estimated position prior to the user's input. The input support device according to claim 2 or 3.
5. The estimation means estimates the position of feature points in the input image using a pre-trained machine learning model. The system further includes a retraining means for performing machine learning on the machine learning model again by using the combination of the input image and the position specified by the user's input as training data. An input support device according to any one of claims 1 to 4.
6. The position of feature points in the input image is estimated. This system assists the user in specifying the position of feature points in the input image. In the aforementioned support, the estimated position is not displayed prior to the user's input, and the input is evaluated based on the position specified by the user's input and the estimated position, thereby supporting the input. Input assistance methods.
7. The positions of multiple feature points in the input image are estimated along with the order of the feature points. This system assists the user in specifying the position of feature points in the input image. The aforementioned support assists the user in specifying the positions of feature points in the aforementioned order by displaying the estimated position of each feature point and information indicating the order of each feature point on the input image. The information indicating the order of each of the aforementioned feature points is either a number representing the order or an arrow indicating the order. Input assistance methods.
8. The support is provided by outputting a warning based on the difference between the estimated order and the order in which the user specifies the positions of the feature points, in which the user specifies the positions of the feature points. The input support method according to claim 7.
9. An estimation step to estimate the position of feature points in the input image, An input support step that assists the user in specifying the position of feature points in the input image, and Have the computer run it, In the input support step, the estimated position is not displayed prior to the user's input, and the input is supported by evaluating the input based on the position specified by the user's input and the estimated position. program.
10. An estimation step that estimates the positions of multiple feature points in an input image along with the order of the feature points, An input support step that assists the user in specifying the position of feature points in the input image, and Have the computer run it, In the input support step, the user's specification of the positions of feature points in the order described above is supported by displaying the estimated position of each feature point and information indicating the order of each feature point on the input image. The information indicating the order of each of the aforementioned feature points is either a number representing the order or an arrow indicating the order. program.
11. The input support step supports the user's specification of the feature point locations in the order described above by outputting a warning based on the difference between the estimated order and the order in which the user specified the feature point locations. The program according to claim 10.