Machine learning programs, methods, apparatus, and estimation programs
By training a relative intensity model and a transformation model simultaneously with shared face images, the method addresses estimation errors in AU recognition due to individual differences, achieving accurate AU intensity estimation across varying facial muscle movements.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Conventional AU recognition technologies face significant estimation errors due to individual differences, particularly in the range of high intensity facial muscle movements.
A method involving a relative intensity model and a transformation model are trained simultaneously, using shared face images to determine a consistent measure of AU intensity, reducing memory consumption and minimizing estimation errors through a combined loss function.
Suppresses estimation errors caused by individual differences even in the high intensity range of facial muscle movements, ensuring accurate AU intensity estimation.
Smart Images

Figure 2026061260000001_ABST
Abstract
Description
[Technical Field]
[0001] The disclosed technologies relate to machine learning programs, machine learning methods, machine learning devices, and estimation programs. [Background technology]
[0002] Facial expressions play a crucial role in nonverbal communication, and facial recognition technology is one of the essential technologies for understanding and sensing people. An Action Unit (AU) is a type of facial expression that represents the movement of facial muscles, such as lowering the eyebrows or raising the cheeks. AU recognition technology, which estimates the intensity of such AUs, is a fundamental technology for facial recognition. By capturing the movement of facial muscles with AU recognition technology, it becomes possible to estimate the emotions and psychological states corresponding to those facial expressions.
[0003] The intensity of the AU (Auditory Characteristic) is defined to take values from 0 to 5, with 0 corresponding to a neutral expression. The intensity of the AU is defined based on features of a neutral face, such as the angle of the eyebrows and the prominence of the cheeks. Since the features of a neutral face differ from person to person, there are individual differences in the relationship between the AU intensity and facial features, and when recognizing an arbitrary person, estimation errors due to individual differences tend to be large. Therefore, in AU recognition, there is a need for techniques to suppress estimation errors caused by individual differences.
[0004] As a conventional technique to suppress estimation errors caused by individual differences, for example, a machine learning program has been proposed that has a computer perform the processes of generating a trained model and generating a third model. In this technique, in response to the input of training data including a pair of a first image and a second image, and a first label indicating which of the two images shows greater movement of the subject's facial muscles, the first image is input into the first model to obtain a first output value. This technique also generates a trained model by performing machine learning on the first model based on the first output value, a second output value obtained by inputting the second image into a second model that shares parameters with the first model, and the first label. Furthermore, this technique generates a third model by machine learning based on a third output value obtained by inputting a third image into the trained model, and a second label indicating the intensity or presence or absence of facial muscle movement of the subject contained in the third image. [Prior art documents] [Patent Documents]
[0005] [Patent Document 1] International Publication No. 2022 / 064660 [Overview of the project] [Problems that the invention aims to solve]
[0006] However, conventional technology has a problem in that, depending on the individual, estimation errors due to individual differences are small in the range of low intensity of facial muscle movement, but can become large in the range of high intensity.
[0007] One aspect of the disclosed technology is that it aims to suppress estimation errors caused by individual differences, even in the range of high intensity facial muscle movements. [Means for solving the problem]
[0008] In one embodiment, the disclosed technique uses first training data in which a pair of first and second images, each containing a person's face, is assigned a first ground truth value indicating the relative magnitude of the facial muscle movements between the first and second images. The disclosed technique calculates a first loss function for a first model based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value. The first model estimates the relative intensity of the facial muscle movements of a person when an image containing a person's face is input. The disclosed technique also determines a reference value for relative intensity based on the relative intensity for each of a plurality of images, each containing a person's face, obtained by inputting each of the plurality of images containing a person's face into the first model. Furthermore, the disclosed technique uses second training data in which a third image containing a person's face is assigned a second ground truth value for the facial muscle movement intensity. Furthermore, the disclosed technology calculates a second loss function for the second model based on the relative intensity obtained by inputting the third image into the first model, the determined reference value of the relative intensity, the intensity of the facial muscle movement obtained by inputting these values, and the second ground truth value. When the relative intensity and the reference value of the relative intensity are input, the second model converts the input relative intensity into the intensity of the facial muscle movement. The disclosed technology then trains the first and second models to minimize a third loss function, which is an integrated first and second loss function. [Effects of the Invention]
[0009] One aspect of this approach is that it can suppress estimation errors caused by individual differences, even in the range of high intensity facial muscle movements. [Brief explanation of the drawing]
[0010] [Figure 1] This figure illustrates why estimation errors occur in the range of high AU intensity in the comparison method. [Figure 2] This is a functional block diagram of the information processing device according to this embodiment. [Figure 3] This is a diagram to explain how to create a mini-batch. [Figure 4] A diagram for explaining the creation of mini - batches [Figure 5] A diagram for explaining the overall loss function [Figure 6] A block diagram showing the schematic configuration of a computer that functions as an information processing apparatus according to this embodiment [Figure 7] A flowchart showing an example of machine learning processing [Figure 8] A flowchart showing an example of mini - batch creation processing [Figure 9] A flowchart showing an example of parameter update processing [Figure 10] A flowchart showing an example of estimation processing [Figure 11] A diagram for explaining the effects of this embodiment
Mode for Carrying Out the Invention
[0011] Hereinafter, an example of an embodiment according to the disclosed technology will be described with reference to the drawings
[0012] First, before explaining the details of the embodiment, the comparison method and the estimation error caused by individual differences in the range of high movement intensity of facial muscles will be described
[0013] Here, as a comparison method, in AU recognition, a method focused on reducing the estimation error due to fluctuations in the standard of the movement intensity of facial muscles (hereinafter referred to as "AU intensity") caused by individual differences in faces will be described. The comparison method focuses on the fact that for the face of the same person, there is no fluctuation in the facial muscles at the time of expressionless face as a reference, and trains a relative intensity model and a conversion model composed of a neural network based only on the change in the characteristics of the face of the same person. Then, the comparison method estimates the AU intensity using the relative intensity model and the conversion model
[0014] The machine learning phase of the comparison method will be explained in detail. The comparison method focuses on the fact that, for images containing the region of a person's face (hereinafter referred to as "face image"), there is no change in the reference point for the relative magnitude of AU intensity for pairs of face images of the same person. Based on pairs of face images of the same person and the relative magnitude of their AU intensity, the comparison method trains a relative intensity model that estimates the relative intensity that captures the difference in AU intensity from the input face image. Furthermore, the AU intensity is based on the neutral expression of each person, and the neutral expression corresponds to the lower limit of the relative intensity for a series of face images of that person (hereinafter referred to as "face video"). Therefore, the comparison method trains a conversion model that converts relative intensity to AU intensity based on the relative intensity, the lower limit of relative intensity, and the ground truth value of AU intensity.
[0015] Next, we will explain the estimation phase of the comparison method in detail. The comparison method acquires facial images of the person for whom the AU intensity is to be estimated, and estimates the relative intensity for each facial image contained in the facial image using a relative intensity model. Based on the multiple estimated relative intensities, the comparison method determines a lower limit of the relative intensities. Then, using a transformation model, the comparison method converts the relative intensities for each facial image into AU intensities based on the lower limit of the relative intensities.
[0016] In the comparison method, the estimation error due to individual differences is small in the range of low AU intensity, but it may be large in the range of high AU intensity. This is because, in the comparison method, relative intensity is determined only by the relative magnitudes of AU intensity for the same person's face image, and therefore it is not a consistent measure independent of the person.
[0017] Therefore, even if two individuals happen to have almost the same relative intensity for the same AU intensity in the low AU intensity range, there may be a difference in their relative intensity for the same AU intensity in the high AU intensity range. Figure 1 schematically shows an example of the relative intensity of each face image in the face video with respect to the frame number or time information for person A and person B. In the example in Figure 1, person A and person B have almost the same relative intensity for the same AU intensity (e.g., 0) in the low AU intensity range (dashed ellipse in Figure 1). On the other hand, in the high AU intensity range, there is a difference in the relative intensity of person A and person B for the same AU intensity (e.g., 5).
[0018] As shown in the example in Figure 1, suppose a transformation model is trained to convert a relative intensity of 7 to an AU intensity of 5 using a facial image of person A (for example, with a relative intensity range of -23 to +7) as training data. In this case, when a facial image of person B (for example, with a relative intensity range of -23 to -5) is used as test data, the model cannot convert a relative intensity of -5 to an AU intensity of 5, and the estimation error increases. However, in the low range of AU intensity, the relative intensity is expected to correspond to an AU intensity of 0 and be close to the lower limit of relative intensity. Therefore, even if the correspondence between relative intensity and AU intensity differs to some extent, the estimation error does not increase, and the above problem does not occur.
[0019] In this embodiment, in order to suppress estimation errors due to individual differences even in the high AU intensity range, the relative intensity model and the transformation model are trained simultaneously, and the value of the AU intensity itself, which is a consistent measure independent of the person, is reflected in the relative intensity model via the transformation model. The information processing device according to this embodiment will be described in detail below.
[0020] As shown in Figure 2, the information processing device 10 functionally includes a machine learning unit 20 and an estimation unit 40. Furthermore, a relative intensity model DB (Database) 32, which stores the parameters of the relative intensity model, and a transformation model DB 34, which stores the parameters of the transformation model, are stored in a predetermined storage area of the information processing device 10. The machine learning unit 20 is an example of a "machine learning device" in the disclosed technology, the relative intensity model is an example of a "first model" in the disclosed technology, and the transformation model is an example of a "second model" in the disclosed technology.
[0021] Each of the relative intensity model and the transformation model is, for example, composed of a neural network. For instance, the relative intensity model may be a VGG (Visual Geometry Group)16, and the transformation model may be a fully connected network.
[0022] The machine learning unit 20 is a functional unit that operates during the learning phase. The machine learning unit 20 further includes a preprocessing unit 22, a creation unit 24, a calculation unit 26, and an update unit 28. The update unit 28 is an example of a “training unit” of the disclosed technology.
[0023] The preprocessing unit 22 acquires training face images for each of the multiple people input to the information processing device 10. A face image is a video that captures the region containing a person's face. Each frame included in the face image corresponds to a face image. In other words, a face image is a series of face images for each person. The training image is assigned identification information of the person being filmed (hereinafter referred to as "person ID"). In addition, the training face image is assigned a ground truth value for AU intensity for each face image included in the training face image. The preprocessing unit 22 normalizes the position, size, orientation, etc. of the face within the image for each face image of the acquired training face image using a method such as Procrustes analysis.
[0024] The creation unit 24 uses the multiple face images normalized by the preprocessing unit 22, and the ground truth values of the AU intensity of each face image, to create minibatches used for training the relative intensity model and the transformation model. A minibatch is a set of training data used for one parameter update.
[0025] In this embodiment, the relative intensity model and the transformation model are trained simultaneously. In the comparison method described above, it is assumed that all of a series of facial images of a person are used to determine the lower limit of the relative intensity. It is also assumed that minibatches for the relative intensity model and minibatches for the transformation model are created independently. In such a comparison method, if the relative intensity model and the transformation model are simply configured to be trained simultaneously, a large number of facial images necessary to determine the lower limit of the relative intensity, which is required to train the transformation model, will be loaded into memory simultaneously. For example, if facial video lasting several tens of seconds is used, the series of facial images for each person will amount to several hundred or more. In this case, if, for example, several GB of GPU memory is assumed to be used, the GPU memory will be insufficient, making it difficult to implement machine learning.
[0026] Therefore, in this embodiment, the number of face images used to determine the lower limit of relative intensity is reduced, and the face images used for training the relative intensity model and the transformation model are shared, thereby enabling accurate determination of the lower limit of relative intensity while suppressing memory consumption.
[0027] Specifically, the creation unit 24 randomly selects a predetermined number (for example, several dozen) of face images from a series of face images for each person. The creation unit 24 also shares at least one of the face images in a pair of face images for training the relative intensity model with the randomly selected face images. The two face images that make up the pair of face images are examples of the "first image" and "second image" of the disclosed technology.
[0028] More specifically, as shown in Figure 3A, the creation unit 24 sets one face image selected from a series of face images for each person as the face image for intensity learning. The creation unit 24 may also randomly select AU intensities for each of the multiple people to avoid bias in the distribution of AU intensities of the selected face images, and select the face image to which the correct value of the selected AU intensity is assigned as the face image for intensity learning. The face image for intensity learning is an example of the "third image" of the disclosed technology.
[0029] Furthermore, as shown in Figure 3B, the creation unit 24 sets a predetermined number of face images randomly selected from a series of face images for each person into a face image set for determining the lower limit of relative intensity. As shown in Figure 3C, the creation unit 24 uses the face image for intensity learning for one person, the correct value of the AU intensity assigned to that face image, and the face image set for determining the lower limit of relative intensity as a unit set for the conversion model.
[0030] Furthermore, as shown in Figure 3D, the creation unit 24 creates face image pairs, which are combinations of a face image for intensity learning and each of several face images selected from a set of face images for determining the lower limit of relative intensity. The creation unit 24 also assigns a correct value for the magnitude relationship of AU intensity to each face image pair, based on the correct value of AU intensity assigned to each face image pair. The correct value for the magnitude relationship may be information that identifies the face image with the larger AU intensity, or information that identifies the face image with the smaller AU intensity. Alternatively, an inequality sign indicating which AU intensity is larger may be used as the correct value for the magnitude relationship. As shown in Figure 3E, the creation unit 24 uses the multiple face image pairs and the correct value for the magnitude relationship assigned to each face image pair as a unit set for the relative intensity model.
[0031] The creation unit 24 creates a unit set for the relative intensity model and a unit set for the transformation model from each of the facial images of multiple people, and combines a predetermined number of unit sets to create a mini-batch for the relative intensity model and a mini-batch for the transformation model. The mini-batch for the relative intensity model is an example of the "first training data" of the disclosed technology, and the face images for intensity learning and the ground truth values of AU intensity included in the mini-batch for the transformation model are an example of the "second training data" of the disclosed technology.
[0032] Furthermore, as shown in Figure 4, the creation unit 24 may create multiple face image pairs by selecting two face images from a set of face images used to determine the lower limit of relative intensity. In the face image pair creation method shown in Figure 3, if the face images for intensity learning are selected so as not to have a bias in the distribution of AU intensity, the bias in the distribution of AU intensity is also reduced in the unit set for the relative intensity model, and machine learning proceeds efficiently. On the other hand, in the face image pair creation method shown in Figure 4, since face image pairs are created from a set of face images used to determine the lower limit of relative intensity that includes various facial expressions, a relative intensity model that can stably and accurately estimate for various facial expressions can be created.
[0033] As described above, in this embodiment, the face images used to determine the lower limit of relative intensity are randomly selected from a series of face images of the person in question. This is done to reduce memory usage, but in order to accurately determine the lower limit of relative intensity, it is better to select a larger number of face images. In either of the face image pair creation methods shown in Figures 3 and 4, since the face images are shared between the mini-batch for the relative intensity model and the mini-batch for the transformation model, memory usage during training can be reduced, and the number of face images selected can be increased accordingly.
[0034] The calculation unit 26 uses a minibatch for the relative intensity model created by the creation unit 24 to calculate the loss function of the relative intensity model, which estimates relative intensity when a face image is input. The loss function of the relative intensity model is an example of the "first loss function" of the disclosed technology. Specifically, the calculation unit 26 inputs two face images included in a pair of face images included in the minibatch for the relative intensity model into the relative intensity model, obtains the relative intensity of each, and obtains the magnitude relationship of the relative intensity. The calculation unit 26 calculates an output value of the relative intensity model that is smaller when the magnitude relationship of the obtained relative intensity is the same as the correct value of the magnitude relationship assigned to that pair of face images.
[0035] Furthermore, the calculation unit 26 determines the lower limit of relative intensity based on a plurality of relative intensities obtained by inputting each of the face images included in the face image set for determining the lower limit of relative intensity, which is included in the minibatch for the conversion model, into the relative intensity model. Note that the lower limit is an example of a reference value when converting relative intensity to AU intensity. The reference value is not limited to the lower limit and may be the median, etc., but in this embodiment, the case in which the reference of relative intensity is the lower limit will be described.
[0036] For example, the calculation unit 26 determines the q-quantile of multiple acquired relative intensities as the lower limit of the relative intensity. The q-quantile may be, for example, the 0.0 quantile, the 0.1 quantile, the 0.2 quantile, etc. Alternatively, the calculation unit 26 may determine multiple values, such as the 0.0 quantile, the 0.1 quantile, the 0.2 quantile, etc., as the lower limit of the relative intensity.
[0037] Furthermore, the calculation unit 26 uses a mini-batch for the transformation model to calculate a loss function for the transformation model, which outputs an estimated value obtained by converting the input relative intensity to AU intensity when relative intensity and a lower limit of relative intensity are input. The loss function for the transformation model is an example of the "second loss function" of the disclosed technology. Specifically, the calculation unit 26 inputs the intensity training face images included in the mini-batch for the transformation model into the relative intensity model to obtain the relative intensity of the intensity training face images. The calculation unit 26 inputs the obtained relative intensity of the intensity training face images and the determined lower limit of relative intensity into the transformation model to obtain an estimated value of AU intensity. Then, the calculation unit 26 calculates an output value of the transformation model loss function, which decreases as the degree of agreement between the obtained estimated value of AU intensity and the correct value of AU intensity assigned to the intensity training face images increases.
[0038] Furthermore, if the calculation unit 26 determines multiple lower limits for relative intensity, such as 0.0 quantile, 0.1 quantile, and 0.2 quantile, it inputs these multiple lower limits into the transformation model. In this case, the transformation model is trained to determine whether to selectively use one of the lower limits when estimating the AU intensity, or to determine the weighting of each lower limit.
[0039] Referring to FIG. 5, the loss functions of the relative intensity model and the conversion model calculated by the calculation unit 26 will be specifically described. Let the face image for reinforcement learning be x i , and the correct value of the AU intensity for the face image for reinforcement learning be y i . Also, let the index of the mini-batch for the conversion model be s = {(x i , v i )} i . v i is the index of the face image set for determining the lower limit of the relative intensity, and v i = {k} k . Also, let the index of the mini-batch for the relative intensity model be u = {(i, j)} i,j .
[0040] Also, let the relative intensity model be f, the relative intensity be ^z i = f(x i ), the function for determining the lower limit of the relative intensity be h, and the lower limit of the relative intensity be ^l i = h({^z j |j ∈ v i}). h may be, for example, the q-th quantile (q is, for example, 0.2). Also, let the conversion model be g and the estimated value of the AU intensity be ^y i = g(^z i , ^l i ). Note that in the figures and mathematical formulas, "^z" is denoted as "^ (hat) " above "z". The same applies to "^l" and "^y". Also, let the correct value of the magnitude relationship of the AU intensity assigned to the face image pair of face image x i and x j be r in the following formula (1) ij .
[0041]
Equation
[0042] In this case, for the loss function loss r of the relative intensity model, for face images x i and x jThe output value for the pair of face images is shown in equation (2) below. m is a parameter of the loss function of the relative intensity model, and here m=1. In this case, the calculation unit 26 calculates the output value of the loss function of the relative intensity model for the minibatch u for the relative intensity model loss using equation (3) below. R Calculate (u).
[0043]
number
[0044] Also, the loss function of the transformation model t The face image x i The output value for is shown in equation (4) below. In this case, the calculation unit 26 calculates the output value of the loss function of the transformation model for the minibatch s for the transformation model using equation (5) below. T Calculate (s).
[0045]
number
[0046] In the comparison method described above, the relative intensity model is trained by minimizing the loss function of the relative intensity model, and then the trained relative intensity model is fixed, and the transformation model is trained by minimizing the loss function of the transformation model. On the other hand, in this embodiment, by training the relative intensity model and the transformation model simultaneously, the parameters of the relative intensity model are updated via the transformation model so that the value of the AU intensity itself is reflected in the relative intensity model.
[0047] Specifically, the update unit 28 calculates the output value of the overall loss function, which is an integrated version of the loss function of the relative intensity model and the loss function of the transformation model. The overall loss function may be, for example, a weighted sum of the loss function of the relative intensity model and the loss function of the transformation model. The update unit 28 calculates the output value of the overall loss function loss using, for example, equation (6) below. A The following is calculated. λ is a parameter that represents the weighting of the two loss functions in the overall loss function.
[0048]
number
[0049] Furthermore, the overall loss function is not limited to the weighted sum of the loss function of the relative intensity model and the loss function of the transformation model described above, but may also be the product of the loss function of the relative intensity model and the loss function of the transformation model, etc.
[0050] The update unit 28 returns the output value of the overall loss function, loss. A The parameters of the relative intensity model and the transformation model are updated to minimize the loss, and this process is repeated until the termination condition for parameter updates is met. The termination condition is, for example, that the parameter updates have been performed a predetermined number of times, and the loss is terminated. A If the value falls below a predetermined value, the previously calculated loss A and the loss calculated this time A This may be the case when the difference between the two values falls below a predetermined value. The update unit 28 stores the parameters of the relative intensity model when the termination condition is met in the relative intensity model DB32 and stores the parameters of the conversion model in the conversion model DB34.
[0051] The estimation unit 40 is a functional unit that operates in the estimation phase. The estimation unit 40 further includes a preprocessing unit 42, a first estimation unit 44, and a second estimation unit 46.
[0052] The preprocessing unit 42 acquires a target face image of the person whose AU intensity is to be estimated, which is input to the information processing device 10. The target face image is the same as the training face image, except that the correct values for AU intensity are not attached. Similar to the preprocessing unit 22 of the machine learning unit 20, the preprocessing unit 42 normalizes the position, size, orientation, etc. of the face in the image for each face image of the acquired target face image using a method such as Procrustes analysis.
[0053] Furthermore, if the facial images in the target facial video contain multiple faces and it is desired to estimate the AU intensity for each face, the preprocessing unit 42 tracks the position of each person's face using face detection and tracking technologies and divides the image into multiple facial images, one for each person. The preprocessing unit 42 then performs the normalization preprocessing described above on the divided facial images. This makes it possible to estimate the AU intensity for each facial image in subsequent processing.
[0054] The first estimation unit 44 inputs each face image of the target video into a trained relative intensity model and estimates the relative intensity for each face image.
[0055] The second estimation unit 46 determines a lower limit of relative intensity based on the multiple relative intensities estimated by the first estimation unit 44. The second estimation unit 46 also inputs each of the relative intensities estimated by the first estimation unit 44 and the determined lower limit of relative intensity into the trained transformation model to estimate the AU intensity for each face image included in the target video, and outputs the estimated value of the estimated AU intensity.
[0056] The information processing device 10 may be implemented, for example, by a computer 50 as shown in Figure 6. The computer 50 includes a CPU (Central Processing Unit) 51, a GPU (Graphics Processing Unit) 52, a memory 53 as a temporary storage area, and a non-volatile storage device 54. The computer 50 also includes input / output devices 55 such as input devices and display devices, and an R / W (Read / Write) device 56 that controls the reading and writing of data to and from the storage medium 59. The computer 50 also includes a communication interface 57 that connects to a network such as the Internet. The CPU 51, GPU 52, memory 53, storage device 54, input / output devices 55, R / W device 56, and communication interface 57 are connected to each other via a bus 58.
[0057] The storage device 54 is, for example, an HDD (Hard Disk Drive), an SSD (Solid State Drive), or flash memory. The storage device 54, as a storage medium, stores a machine learning program 60 and an estimation program 70 that enable the computer 50 to function as an information processing device 10.
[0058] The machine learning program 60 includes a preprocessing process control instruction 62, a creation process control instruction 64, a calculation process control instruction 66, and an update process control instruction 68. The estimation program 70 includes a preprocessing process control instruction 72, a first estimation process control instruction 74, and a second estimation process control instruction 76. The storage device 54 also has an information storage area 80 in which information constituting the relative intensity model DB32 and the conversion model DB34 is stored.
[0059] The CPU 51 reads the machine learning program 60 from the storage device 54 and loads it into memory 53, and then sequentially executes the control instructions contained in the machine learning program 60. The CPU 51 operates as the preprocessing unit 22 shown in Figure 2 by executing the preprocessing process control instruction 62. The CPU 51 also operates as the creation unit 24 shown in Figure 2 by executing the creation process control instruction 64. The CPU 51 also operates as the calculation unit 26 shown in Figure 2 by executing the calculation process control instruction 66. The CPU 51 also operates as the update unit 28 shown in Figure 2 by executing the update process control instruction 68.
[0060] Furthermore, the CPU 51 reads the estimated program 70 from the storage device 54 and loads it into the memory 53, and sequentially executes the control instructions contained in the estimated program 70. The CPU 51 operates as the preprocessing unit 42 shown in Figure 2 by executing the preprocessing process control instruction 72. The CPU 51 also operates as the first estimation unit 44 shown in Figure 2 by executing the first estimation process control instruction 74. The CPU 51 also operates as the second estimation unit 46 shown in Figure 2 by executing the second estimation process control instruction 76.
[0061] Furthermore, the CPU 51 reads information from the information storage area 80 and loads the relative intensity model DB32 and the transformation model DB34 into memory 53. This allows the computer 50, which executed the machine learning program 60 and the estimation program 70, to function as an information processing device 10. The CPU 51 that executes the program is hardware. Also, a portion of the program may be executed by the GPU 52.
[0062] The functions implemented by the machine learning program 60 and the estimation program 70 may be implemented, for example, by semiconductor integrated circuits, more specifically by ASICs (Application Specific Integrated Circuits), FPGAs (Field-Programmable Gate Arrays), etc.
[0063] Next, the operation of the information processing device 10 according to this embodiment will be described. In the machine learning phase, when a training face image is input to the information processing device 10, the machine learning process shown in Figure 7 is executed in the information processing device 10. Furthermore, in the estimation phase, when an estimation target face image is input to the information processing device 10, the estimation process shown in Figure 10 is executed in the information processing device 10. The machine learning process is an example of a machine learning method of the disclosed technology.
[0064] First, let's explain the machine learning process shown in Figure 7.
[0065] In step S10, the update unit 28 initializes the parameters of the relative intensity model and the transformation model. For example, the update unit 28 sets the initial parameters of the relative intensity model, which is composed of VGG16, to values pre-trained on ImageNet. Also, for example, the update unit 28 sets the initial parameters of the transformation model, which is composed of a fully connected network, to random values.
[0066] Next, in step S20, the preprocessing unit 22 acquires training face images for each of the multiple people input to the information processing device 10. Then, the preprocessing unit 22 performs preprocessing on each face image of the acquired training face images, for example, by using a method such as Procrustes analysis to normalize the position, size, orientation, etc. of the face within the image.
[0067] Next, in step S30, the mini-batch creation process is executed. The mini-batch creation process will now be explained with reference to Figure 8.
[0068] In step S32, the creation unit 24 selects one target person from among the people whose face images have been input to the information processing device 10. Next, in step S34, the creation unit 24 acquires a face image associated with the target person's person ID, selects one face image from the multiple face images included in the acquired face image, and sets it as the face image for intensity learning. Next, in step S36, the creation unit 24 randomly selects a predetermined number of face images from the target person's face image and sets them as a face image set for determining the lower limit of relative intensity. Next, in step S38, the creation unit 24 creates the face image for intensity learning for the target person, the correct values of the AU intensity assigned to that face image, and the face image set for determining the lower limit of relative intensity as a unit set for the conversion model.
[0069] Next, in step S40, the creation unit 24 creates face image pairs, which are combinations of a face image for intensity learning and each of several face images selected from a set of face images for determining the lower limit of relative intensity. Then, the creation unit 24 assigns the correct values for the magnitude relationship of AU intensity to each face image pair, based on the correct values for AU intensity assigned to each face image pair. Next, in step S42, the creation unit 24 creates a unit set for the relative intensity model, which consists of the multiple face image pairs and the correct values for the magnitude relationship assigned to each face image pair.
[0070] Next, in step S44, the creation unit 24 determines whether a predetermined number of unit sets, which are the number of unit sets to be included in the mini-batch, have been created. If the number of unit sets has not reached the predetermined number, the process returns to step S32; if the predetermined number of unit sets have been created, the process proceeds to step S46.
[0071] In step S46, the creation unit 24 combines a predetermined number of unit sets for the relative intensity model to create a mini-batch for the relative intensity model, and combines a predetermined number of unit sets for the transformation model to create a mini-batch for the transformation model. Then the mini-batch creation process is completed, and the process returns to the machine learning process (Figure 7).
[0072] Next, in step S50, the parameter update process is executed. The parameter update process will now be explained with reference to Figure 9.
[0073] In step S52, the calculation unit 26 inputs each of the face images included in the face image set for determining the lower limit of relative intensity, which is included in the minibatch for the conversion model, into the relative intensity model to obtain multiple relative intensities. Next, in step S54, the calculation unit 26 determines the lower limit of relative intensity based on the multiple relative intensities obtained.
[0074] Next, in step S56, the calculation unit 26 inputs the face images for intensity learning included in the minibatch for the conversion model into the relative intensity model to obtain the relative intensity of the face images for intensity learning. Next, in step S58, the calculation unit 26 inputs the relative intensity of the face images for intensity learning obtained in step S56 and the lower limit of the relative intensity determined in step S54 into the conversion model to obtain an estimated value of the AU intensity of the face images for intensity learning.
[0075] Next, in step S60, the calculation unit 26 calculates the output value of the loss function of the transformation model, which decreases as the degree of agreement between the AU intensity obtained in step S58 and the correct value of the AU intensity assigned to the intensity learning face image increases.
[0076] Next, in step S62, the calculation unit 26 inputs two face images from a pair of face images included in a minibatch for the relative intensity model into the relative intensity model, obtains the relative intensity of each image, and obtains the magnitude relationship of the relative intensity. The calculation unit 26 then calculates the output value of the loss function of the relative intensity model, which is smaller when the magnitude relationship of the obtained relative intensity is the same as the correct magnitude relationship assigned to that pair of face images.
[0077] Next, in step S64, the update unit 28 calculates the output value of the overall loss function, which is an integrated version of the loss function of the relative intensity model and the loss function of the transformation model. The update unit 28 then updates the parameters of the relative intensity model and the transformation model so that the output value of the overall loss function becomes smaller. The parameter update process then ends, and the process returns to the machine learning process (Figure 7).
[0078] Next, in step S70, the update unit 28 determines whether the termination conditions for parameter updating have been met. If the termination conditions are met, the process proceeds to step S80; otherwise, it returns to step S30. In step S80, the update unit 28 stores the relative intensity model parameters at the time the termination conditions were met in the relative intensity model DB32, stores the transformation model parameters in the transformation model DB34, and the machine learning process ends.
[0079] Next, we will explain the estimation process shown in Figure 10.
[0080] In step S100, the first estimation unit 44 reads the parameters of the relative intensity model from the relative intensity model DB32, and the second estimation unit 46 reads the parameters of the transformation model from the transformation model DB34. Next, in step S102, the preprocessing unit 42 acquires the estimated target face video input to the information processing device 10. Then, the preprocessing unit 42 performs preprocessing on each face image of the acquired estimated target face video, for example, by a method such as Procrustes analysis to normalize the position, size, orientation, etc. of the face in the image.
[0081] Next, in step S104, the first estimation unit 44 inputs each face image of the target video into the trained relative intensity model and estimates the relative intensity for each face image. Then, in step S106, the second estimation unit 46 determines the lower limit of the relative intensity based on the estimated relative intensities.
[0082] Next, in step S106, the second estimation unit 46 inputs each of the relative intensities estimated in step S104 and the lower limit of the relative intensities determined in step S106 into the trained transformation model to obtain estimated values of the AU intensity for each face image included in the target video. Next, in step S110, the second estimation unit 46 outputs the estimated values of the estimated AU intensity, and the estimation process ends.
[0083] As described above, the information processing device according to this embodiment creates a minibatch for relative intensity by assigning the correct values of the magnitude relationship of AU intensity to pairs of face images. The information processing device also creates a minibatch for a transformation model that includes face images for intensity learning to which the correct values of AU intensity have been assigned, and a set of face images for determining the lower limit of relative intensity. Using the minibatch for relative intensity, the information processing device calculates the output value of the loss function of the relative intensity model based on the magnitude relationship of the relative intensity obtained by inputting each face image of the face image pair into the relative intensity model, and the correct values of the magnitude relationship. The information processing device also determines the lower limit of relative intensity based on the multiple relative intensities obtained by inputting each of the multiple face images included in the face image set into the relative intensity model. The information processing device also inputs the face images for intensity learning into the relative intensity model to obtain the relative intensity of the face images for intensity learning. The information processing device also calculates an estimated value of the loss function of the transformation model based on the estimated value of AU intensity obtained by inputting the obtained relative intensity and the determined lower limit of relative intensity into the transformation model, and the correct values of AU intensity. The information processing device then trains the relative intensity model and the transformation model to minimize the overall loss function, which is an integrated loss function combining the loss function of the relative intensity and the loss function of the transformation model. This makes it possible to suppress estimation errors caused by individual differences, even in the high intensity range of facial muscle movements.
[0084] In comparative methods, where the relative intensity model and the transformation model are trained separately, estimation errors due to individual differences can become large in the range of high AU intensity, as shown in the left figure of Figure 11. Note that the left figure of Figure 11 is a reproduction of Figure 1. On the other hand, in the method according to this embodiment (hereinafter also referred to as "this method"), the relative intensity model and the transformation model are trained simultaneously. This imposes a constraint on the relative intensity model to create a consistent, person-independent scale, such that for any person's face image, if the lower limit of relative intensity and AU intensity are the same, the relative intensity is also the same. Therefore, as shown in the right figure of Figure 11, even in the range of high AU intensity, the variation in relative intensity due to individuals for the same AU intensity can be suppressed, and the transformation model can accurately convert relative intensity to AU intensity based on the correspondence between relative intensity and AU intensity. As a result, estimation errors due to individual differences can be suppressed even in the range of high AU intensity.
[0085] Furthermore, if the comparison method simply incorporates a configuration that simultaneously trains the relative intensity model and the transformation model, the number of face images used simultaneously during training will be large, potentially leading to insufficient memory. In contrast, this method randomly selects a set of face images from a series of face images to determine the lower limit of relative intensity, and shares face images between the mini-batch for the relative intensity model and the mini-batch for the transformation model, as shown in Figures 3 to 5. This reduces memory usage.
[0086] In the above embodiment, the case in which the machine learning unit and the estimation unit are implemented on a single computer was described. However, the machine learning device including the functional unit of the machine learning unit and the estimation device including the functional unit of the estimation unit may be implemented on separate computers. Alternatively, the relative intensity model DB and the transformation model DB may be stored on an external device separate from the machine learning device. In this case, the machine learning device stores the parameters of the trained relative intensity model and the transformation model in each DB. The machine learning device can update the relative intensity model and the transformation model at any time. The machine learning device can then read the parameters of the trained relative intensity model and the transformation model from each DB when requested by the estimation device, and distribute them to the estimation device via the internet or the like. This allows the estimation device to be provided as an application for end users, and the machine learning device to be provided as a system for developers.
[0087] Furthermore, although the above embodiment was described assuming that there is only one type of AU, there are multiple types of AUs. This method is also applicable to multiple types of AUs. For example, a relative intensity model and a transformation model can be created for each type of AU, and estimation processing can be performed. Alternatively, for example, the relative intensity model and transformation model can be multi-labeled to allow estimation of multiple types of AU intensities at once. Specifically, the AU intensity can be represented by a vector in which the AU intensity for each type of AU is stored in each dimension, and the outputs of the relative intensity model and transformation model can also be vectors that match this.
[0088] Furthermore, although the above embodiment describes a case in which the entire series of target face images is acquired before determining the lower limit of the relative intensity in the estimation process, it is not limited to this. For example, the lower limit of the relative intensity may be determined for each predetermined frame of the target face image, assuming that the frequency of expressionless faces is higher than that of other expressions among the facial expressions shown in each face image included in the face image. This makes it possible to estimate the AU intensity in near real time.
[0089] Furthermore, in the above embodiment, the machine learning program and the estimation program are pre-stored (installed) in the storage device, but the invention is not limited thereto. The program relating to the disclosed technology may be provided in a form stored on a storage medium such as a CD-ROM, DVD-ROM, or USB memory.
[0090] The following additional information is disclosed regarding the embodiments described above.
[0091] (Note 1) Using first training data in which a pair of first and second images containing a person's face is assigned a first ground truth value indicating the relative magnitude of the facial muscle movements between the first and second images, a first model that estimates the relative intensity of the facial muscle movements of a person when an image containing a person's face is input calculates a first loss function based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value, Based on the relative intensity of each of the multiple images obtained by inputting each of the multiple images, including a person's face, a reference value for relative intensity is determined. Using second training data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and a reference value of relative intensity are input, a second loss function is calculated based on the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value, in the second model that converts the input relative intensity into the intensity of facial muscle movement. The first and second models are trained to minimize a third loss function obtained by integrating the first and second loss functions. A machine learning program that causes a computer to perform a process that includes the following.
[0092] (Note 2) The machine learning program described in Appendix 1, wherein the plurality of images for determining the reference value of the relative intensity are a predetermined number of frames randomly selected from a series of frames included in a video that captures a region including a person's face.
[0093] (Note 3) A machine learning program according to Appendix 1 or Appendix 2, wherein at least one of the first image and the second image is shared with the plurality of images for determining the reference value of the relative intensity.
[0094] (Note 4) The machine learning program described in Appendix 3, wherein the first training data is training data obtained by selecting multiple pairs of a first image and a second image from the multiple images for determining the reference value of the relative intensity to which the second correct value has been assigned, and assigning the first correct value to each of the pairs based on the second correct value assigned to each of the first image and the second image.
[0095] (Note 5) The machine learning program described in Appendix 3 is a machine learning program in which the first training data is obtained by creating multiple sets of training data in which the third image selected from a series of frames to which the second ground truth value is assigned is designated as the first image, and the image selected from the multiple images for determining the reference value of the relative intensity is designated as the second image, and the first ground truth value is assigned to each of the sets based on the second ground truth value assigned to each of the first and second images.
[0096] (Note 6) The machine learning program described in any one of the appendices 1 to 5, wherein the third loss function is a weighted sum of the first loss function and the second loss function.
[0097] (Note 7) The machine learning program described in any one of the appendices 1 to 6, wherein the reference value of the relative intensity is the lower limit or median of the relative intensity for each of the plurality of images.
[0098] (Note 8) The machine learning program described in Appendix 7, wherein the lower limit is a quantile of 1 or more relative intensity for each of the multiple images.
[0099] (Note 9) Using the first model and the second model, which have been trained by having a computer run the machine learning program described in any one of the appendices 1 to 8, Each of a series of images, including the face of a person whose facial muscle movement intensity is to be estimated, is input to the trained first model, and the relative intensity for each of the series of images is estimated. Based on the relative intensity of each of the aforementioned series of images, a reference value for relative intensity is determined. The second model, which has been trained, is used to estimate the intensity for each of the series of images by inputting the relative intensity for each of the series of images and the determined reference value for the relative intensity. A pre-programmed program that causes a computer to perform a process that includes the following.
[0100] (Note 10) Using first training data in which a pair of first and second images containing a person's face is assigned a first ground truth value indicating the relative magnitude of the facial muscle movements between the first and second images, a first model that estimates the relative intensity of the facial muscle movements of a person when an image containing a person's face is input calculates a first loss function based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value, Based on the relative intensity of each of the multiple images obtained by inputting each of the multiple images, including a person's face, a reference value for relative intensity is determined. Using second training data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and a reference value of relative intensity are input, a second loss function is calculated based on the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value, in the second model that converts the input relative intensity into the intensity of facial muscle movement. The first and second models are trained to minimize a third loss function obtained by integrating the first and second loss functions. A machine learning method in which a computer performs a process that includes [specific actions].
[0101] (Note 11) The machine learning method described in Appendix 10, wherein the plurality of images for determining the reference value of the relative intensity are a predetermined number of frames randomly selected from a series of frames included in a video that captures a region including a person's face.
[0102] (Note 12) The machine learning method according to Appendix 10 or Appendix 11, wherein at least one of the first image and the second image is shared with the plurality of images for determining the reference value of the relative intensity.
[0103] (Note 13) The machine learning method described in Appendix 12, wherein the first training data is training data obtained by selecting multiple pairs of a first image and a second image from the multiple images for determining the reference value of the relative intensity to which the second correct value has been assigned, and assigning the first correct value to each of the pairs based on the second correct value assigned to each of the first image and the second image.
[0104] (Note 14) The machine learning method described in Appendix 12, wherein the first training data is a training data obtained by creating multiple sets of images, each of which is selected as the first image, and each is selected as the second image, and each is selected as the first image, and each is selected as the first image, and each is selected as the second image, and the first ground truth value is assigned to each of the sets based on the second ground truth value assigned to each of the first image and the second image.
[0105] (Note 15) The machine learning method described in any one of the appendices 10 to 14, wherein the third loss function is a weighted sum of the first loss function and the second loss function.
[0106] (Note 16) The machine learning method described in any one of the appendices 10 to 15, wherein the reference value of the relative intensity is the lower limit or median of the relative intensity for each of the plurality of images.
[0107] (Note 17) The machine learning method described in Appendix 16, wherein the lower limit is a quantile of 1 or more relative intensity for each of the plurality of images.
[0108] (Note 18) Using the first and second models, which have been trained by a computer performing any one of the machine learning methods described in Appendix 10 to Appendix 17, Each of a series of images, including the face of a person whose facial muscle movement intensity is to be estimated, is input to the trained first model, and the relative intensity for each of the series of images is estimated. Based on the relative intensity of each of the aforementioned series of images, a reference value for relative intensity is determined. The second model, which has been trained, is used to estimate the intensity for each of the series of images by inputting the relative intensity for each of the series of images and the determined reference value for the relative intensity. An estimation method for how a computer will perform a process that includes the following.
[0109] (Note 19) Using first training data in which a pair of first and second images containing a person's face is assigned a first ground truth value indicating the relative magnitude of the facial muscle movement intensity between the first and second images, a first model is created to estimate the relative intensity of the facial muscle movement of a person when an image containing a person's face is input. A first loss function is calculated based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value. The first model is then created to estimate the relative intensity of the facial muscle movement of a person when an image containing a person's face is input. A calculation unit calculates a second loss function based on the relative intensity of each of the multiple images, using second learning data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and the reference value of relative intensity are input to a second model that converts the input relative intensity into the intensity of facial muscle movement, and the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value. A training unit that trains the first model and the second model to minimize a third loss function obtained by integrating the first loss function and the second loss function, A machine learning device that includes this.
[0110] (Note 20) Using the first and second models trained with the machine learning device described in Appendix 19, A first estimation unit inputs each of a series of images, including the face of a person whose facial muscle movement intensity is to be estimated, into the trained first model, and estimates the relative intensity for each of the series of images. A second estimation unit determines a reference value for relative intensity based on the relative intensity of each of the aforementioned series of images, and inputs the relative intensity of each of the aforementioned series of images and the determined reference value for relative intensity into the trained second model to estimate the intensity of each of the aforementioned series of images. An estimation device that includes this. [Explanation of Symbols]
[0111] 10 Information Processing Devices 20 Machine Learning Department 22 Pre-processing section 24 Creation Department 26 Calculation Section 28 Update section 32 Relative Intensity Model DB 34 Conversion Model DB 40 Estimation part 42 Pre-processing section 44 1st estimation part 46 Second estimation part 50 Computers 51 CPU 52 GPU 53 memory 54 Storage device 55 Input / Output Devices 56 R / W device 57 Communication I / F 58 Bus 59 Storage medium 60 Machine Learning Programs 62 Preprocessing process control instructions 64. Create process control instructions 66 Computation process control instructions 68 Update process control instructions 70 Estimated Programs 72 Preprocessing process control instructions 74 First Estimated Process Control Instruction 76 Second Estimated Process Control Instruction 80 Information storage area
Claims
1. Using first training data in which a pair of first and second images containing a person's face is assigned a first ground truth value indicating the relative magnitude of the facial muscle movements between the first and second images, a first model that estimates the relative intensity of the facial muscle movements of a person when an image containing a person's face is input calculates a first loss function based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value, Based on the relative intensity of each of the multiple images obtained by inputting each of the multiple images, including a person's face, a reference value for relative intensity is determined. Using second training data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and a reference value of relative intensity are input, a second loss function is calculated based on the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value, in the second model that converts the input relative intensity into the intensity of facial muscle movement. The first and second models are trained to minimize a third loss function obtained by integrating the first and second loss functions. A machine learning program that causes a computer to perform a process that includes the following.
2. The machine learning program according to claim 1, wherein the plurality of images for determining the reference value of the relative intensity are a predetermined number of frames randomly selected from a series of frames included in a video that captures a region including a person's face.
3. A machine learning program according to claim 1 or claim 2, wherein at least one of the first image and the second image is shared with the plurality of images for determining the reference value of the relative intensity.
4. The machine learning program according to claim 3, wherein the first training data is training data obtained by selecting multiple pairs of a first image and a second image from the plurality of images for determining the reference value of the relative intensity to which the second correct value has been assigned, and assigning the first correct value to each of the pairs based on the second correct value assigned to each of the first image and the second image.
5. The machine learning program according to claim 3, wherein the first training data is training data obtained by creating a plurality of sets of images, each of which is selected as the first image, and an image selected from the plurality of images for determining the reference value of the relative intensity, and each of the sets is selected as the first image, and the first correct value is assigned to each of the sets based on the second correct value assigned to each of the first image and the second image.
6. The machine learning program according to claim 1 or claim 2, wherein the third loss function is a weighted sum of the first loss function and the second loss function.
7. The machine learning program according to claim 1 or claim 2, wherein the reference value of the relative intensity is the lower limit or median of the relative intensity for each of the plurality of images.
8. The machine learning program according to claim 7, wherein the lower limit is a quantile of one or more relative intensity for each of the plurality of images.
9. Using the first model and the second model trained by having a computer execute the machine learning program described in claim 1 or claim 2, Each of a series of images, including the face of a person whose facial muscle movement intensity is to be estimated, is input to the trained first model, and the relative intensity for each of the series of images is estimated. Based on the relative intensity of each of the aforementioned series of images, a reference value for relative intensity is determined. The second model, which has been trained, is used to estimate the intensity for each of the series of images by inputting the relative intensity for each of the series of images and the determined reference value for the relative intensity. A pre-programmed program that causes a computer to perform a process that includes the following.
10. Using first training data in which a pair of first and second images containing a person's face is assigned a first ground truth value indicating the relative magnitude of the facial muscle movements between the first and second images, a first model that estimates the relative intensity of the facial muscle movements of a person when an image containing a person's face is input calculates a first loss function based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground truth value, Based on the relative intensity of each of the multiple images obtained by inputting each of the multiple images, including a person's face, a reference value for relative intensity is determined. Using second training data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and a reference value of relative intensity are input, a second loss function is calculated based on the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value, in the second model that converts the input relative intensity into the intensity of facial muscle movement. The first and second models are trained to minimize a third loss function obtained by integrating the first and second loss functions. A machine learning method in which a computer performs a process that includes [specific actions].
11. Using first training data in which a pair of first and second images containing a person's face is assigned a first ground value indicating the relative magnitude of the facial muscle movement intensity between the first and second images, a first model is used to estimate the relative intensity of the facial muscle movement of a person when an image containing a person's face is input. A first loss function is calculated based on the relative magnitude of the relative intensity obtained by inputting the first image and the relative intensity obtained by inputting the second image, and the first ground value. The first model is then used to estimate the relative intensity of the facial muscle movement of a person when an image containing a person's face is input. A calculation unit calculates a second loss function based on the relative intensity of each of the multiple images, using second learning data in which a second ground truth value of the intensity of facial muscle movement is assigned to a third image including a person's face, when relative intensity and the reference value of relative intensity are input to a second model that converts the input relative intensity into the intensity of facial muscle movement, and the relative intensity obtained by inputting the third image into the first model, the determined reference value of relative intensity, and the second ground truth value. A training unit that trains the first model and the second model to minimize a third loss function obtained by integrating the first loss function and the second loss function, A machine learning device that includes this.
Citation Information
Patent Citations
Machine learning program, machine learning method, and inference device
WO2022064660A1