Model Training Device, Model Training Method, and Program
The model training device and method address the challenge of maintaining image region classes during environmental conversion by using a trained image conversion model that updates parameters based on discrimination data and class information, resulting in effective environmental conversion without object disappearance.
Patent Information
- Application Number
- JP2023579963
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-02-10
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-02-10
AI Technical Summary
Existing image conversion models struggle to perform environmental conversion while maintaining the class of each image region, leading to issues like object disappearance during scene transformation.
A model training device and method that acquire a training dataset with class information for image regions, train an image conversion model to output images of a different environment while preserving image region classes, and update the model parameters using a calculated loss based on discrimination data and class information.
The approach enables effective environmental conversion of images while maintaining the class of each image region, preventing object disappearance and improving the robustness of image conversion models.
Smart Images

Figure 0007683751000012 
Figure 0007683751000013 
Figure 0007683751000014
Abstract
Description
Technical Field
[0001] The present disclosure relates to a technique for training a model that performs image conversion.
Background Art
[0002] A model that generates another image based on an input image, that is, a model that performs image conversion, has been developed. For example, Non-Patent Document 1 discloses a model that converts an input image into an image of another class, such as converting an image of a horse into an image of a zebra.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Non-Patent Document 2
Non-Patent Document 3
[0004] In Non-Patent Document 1, when an image is converted, the class of the object is converted. The present disclosure has been made in view of the above problems, and one of its objects is to provide a new technique for training a model for performing image conversion. [Means for Solving the Problems]
[0005] The model training device of the present disclosure includes an acquisition unit that acquires a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image, and a training execution unit that uses the first training dataset to train an image conversion model that outputs an image representing a scene in a second environment in response to an input of an image representing a scene in the first environment. The training execution means inputs the first training image into the image conversion model, inputs the first output image output from the image conversion model into the discrimination model, calculates a first loss using the discrimination data output from the discrimination model and the first class information, and updates the parameters of the image conversion model using the first loss. The discrimination data indicates, for each of a plurality of partial regions included in the image input to the discrimination model, whether the partial region is a fake image region, and indicates the class of the partial region when the partial region is not a fake image.
[0006] The model training method of the present disclosure is executed by a computer. The model training method includes an acquisition step of acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image, and a training execution step of training an image conversion model that outputs an image representing a scene in a second environment in response to an image representing a scene in the first environment being input, using the first training dataset. In the training execution step, the first training image is input into the image conversion model, the first output image output from the image conversion model is input into the discrimination model, a first loss is calculated using the discrimination data output from the discrimination model and the first class information, and the parameters of the image conversion model are updated using the first loss. The discrimination data indicates, for each of a plurality of partial regions included in the image input to the discrimination model, whether the partial region is a fake image region, and indicates the class of the partial region when the partial region is not a fake image.
[0007] The computer-readable medium of the present disclosure stores a program for causing a computer to execute the model training method of the present disclosure.
Advantages of the Invention
[0008] According to the present disclosure, a new technique for training a model for image conversion is provided.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Modes for Carrying Out the Invention
[0010] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. In each drawing, the same or corresponding elements are denoted by the same reference numerals, and redundant descriptions are omitted as necessary for clarity of explanation. Also, unless otherwise specified, predetermined values such as predetermined values and threshold values are stored in advance in a storage device or the like accessible from the device that uses the value. Furthermore, unless otherwise specified, the storage unit is configured by one or more arbitrary numbers of storage devices.
[0011] <Summary> FIG. 1 is a diagram illustrating an overview of an image conversion model trained by the model training apparatus of the present embodiment. The image conversion model 100 outputs an output image 20 in response to the input of an input image 10. The input image 10 is an image input to the image conversion model 100. The output image 20 is an image output from the image conversion model 100. For example, the image conversion model 100 is realized as any machine learning model (e.g., a neural network).
[0012] The image conversion model 100 is trained to perform the process of "when an image representing a scene in a first environment is input as the input image 10, output an image representing the scene in a second environment different from the first environment as the output image 20". Thereby, the image conversion model 100 can be made to pseudo-generate an image of a scene captured in another environment from an image of the scene captured in a certain specific environment.
[0013] For example, assume that the first environment is daytime and the second environment is nighttime. Also assume that the input image 10 is an image obtained by imaging a specific road with a camera. Here, the state of the road at night is different from that of the road during the day in terms of being generally dark, various lights such as car lights and streetlights being on, and the places illuminated by the lights being brighter compared to other places. The image conversion model 100 generates an image of the road at night from an image of the road during the day so as to pseudo-reproduce such characteristics of the road at night. Thereby, for example, as will be described later, data augmentation can be realized.
[0014] Note that the environment is not limited to time zones such as daytime or nighttime. For example, other examples of the environment include the environment related to the weather. For example, assume that the first environment is sunny and the second environment is rainy. In this case, the image conversion model 100 generates the output image 20 representing the scene under rainy weather from the input image 10 representing the scene under sunny weather. Note that weather such as snow can also be adopted instead of rain.
[0015] Furthermore, when generating the output image 20 from the input image 10, the image conversion model 100 is trained not to perform conversion of the class of each image region while performing the conversion of the environment from the first environment to the second environment. The class of the image region is represented by, for example, the type of object included in the image region. Therefore, for example, the conversion from the input image 10 to the output image 20 is performed so that the image region representing a car in the input image 10 also represents a car in the output image 20. By training the image conversion model 100 in this way, when performing the conversion from the input image 10 to the output image 20, it is possible to prevent predetermined types of objects such as cars from disappearing while performing the conversion of the environment. The importance of preventing the disappearance of objects will be described later.
[0016] The training of the image conversion model 100 is performed using an identification model. FIG. 2 is a diagram illustrating an overview of the identification model 200. For example, the identification model 200 is realized as an arbitrary machine learning model (for example, a neural network).
[0017] The identification model 200 identifies, for each of a plurality of image regions included in the input image 30, whether the image region is a true image region representing a scene in the second environment. Here, the true image region means an image region that is not an image region generated by the image conversion model 100 (that is, not a pseudo-generated image region). Further, the identification model 200 identifies the class of the true image region. Hereinafter, an image generated by the image conversion model 100 (that is, a pseudo image) and an image not generated by the image conversion model 100 are respectively referred to as a "false image" and a "true image". Further, an image region that is not a true image region is referred to as a "false image region".
[0018] The identification data 40 represents the result of identification by the identification model 200. For example, the identification data 40 indicates, for each of a plurality of image regions included in the input image 10, the probability of being a true image region belonging to each class and the probability of being a false image region. For example, it is assumed that n types from C1 to Cn are prepared as classes. In this case, the identification data 40 indicates, for each of a plurality of image regions included in the input image, a vector of (N + 1) dimensions (hereinafter, a score vector). The score vector indicates the probability that the corresponding image region is a true image region belonging to each of the classes C1 to CN, and the probability that the corresponding image region is a false image region. For example, the score vector indicates the probability that the corresponding image region is a true image region belonging to the class Ci (1 <= i <= n) as the i-th element, and the probability that the corresponding image region is a false image region as the (N + 1)-th element.
[0019] The image region to be identified by the identification model 200 may be one pixel or a region composed of a plurality of pixels. In the former case, the identification model 200 performs true / false identification and class identification for each pixel of the input image 10. On the other hand, in the latter case, for example, the identification model 200 divides the input image 10 into a plurality of image regions of a predetermined size, and performs true / false identification and class identification for each image region.
[0020] Based on the configurations of the above-described image conversion model 100 and discrimination model 200, an overview of the operation of the model training apparatus 2000 of the present embodiment will be described. FIG. 3 is a diagram illustrating an overview of the model training apparatus 2000 of the present embodiment. Here, FIG. 3 is a diagram for facilitating understanding of the overview of the model training apparatus 2000, and the operation of the model training apparatus 2000 is not limited to that shown in FIG. 1.
[0021] The model training apparatus 2000 acquires a first training dataset 50. The first training dataset 50 includes first training images 52 and first class information 54. The first training images 52 are images representing scenes in a first environment. The first class information 54 indicates the class of each of a plurality of image regions included in the first training images 52.
[0022] The model training apparatus 2000 inputs the first training Image 52 as input images 10 to the image conversion model 100, thereby obtaining output images 20 from the image conversion model 100. Further, the discrimination model 200 inputs these output images 20 to the discrimination model 200. As a result, the model training apparatus 2000 obtains discrimination data 40 representing discrimination results for each image region included in the output images 20.
[0023] Here, as described above, it is desirable that the image conversion model 100 performs environmental conversion but does not perform class conversion. Therefore, it is preferable to train the image conversion model 100 so that each image region of the output image 20 is discriminated by the discrimination model 200 as "being a true image region and belonging to the same class as the corresponding image region of the input image 10". That is, it is preferable to train the image conversion model 100 so that the class of each image region specified by the discrimination data 40 matches the class of each image region indicated by the first class information 54.
[0024] Therefore, the model training device 2000 calculates a first loss representing the magnitude of the difference between the identification data 40 and the first class information 54, and trains the image conversion model 100 to reduce the first loss. Specifically, the model training device 2000 updates the trainable parameters (for example, each weight of the neural network) included in the image conversion model 100 so as to reduce the first loss.
[0025] Note that the class of the image area specified by the identification data 40 is, for example, the class corresponding to the element with the maximum value in the above-described score vector. When the element with the maximum value in the score vector corresponds to a false image area, the score vector indicates that the corresponding image area is a false image area.
[0026] <Example of effects> In the method of Non-Patent Document 1, class conversion is performed on the entire image, such as converting an image of a horse into an image of a zebra. Therefore, in the method of Non-Patent Document 1, it is not possible to perform image conversion that maintains the class (for example, the type of object) of each image area while converting the environment of the scene represented by the entire image. As an example of such image conversion, image conversion of converting an image of a road during the day with a car running on it into an image of a road at night with a car running on it can be considered. In this image conversion, while converting the environment of the scene represented by the entire image from day to night, it is necessary that the image area representing a car in the image before conversion also represents a car in the image area after conversion.
[0027] In this regard, the model training device 2000 inputs the output image 20 obtained from the image conversion model 100 into the identification model 200, and trains the image conversion model 100 using the identification data 40 and the first class information 54 obtained from the identification model 200. Thereby, an image conversion model 100 having a function of "converting a scene in a first environment to a scene in a second environment while maintaining the class of each image area" can be obtained.
[0028] Hereinafter, the model training device 2000 of the present embodiment will be described in more detail.
[0029] <Example of Functional Configuration> FIG. 4 is a block diagram illustrating the functional configuration of the model training apparatus 2000 according to the present embodiment. The model training apparatus 2000 includes an acquisition unit 2020 and a training execution unit 2040. The acquisition unit 2020 acquires the first training dataset 50. The training execution unit 2040 trains the image conversion model 100 using the first training dataset 50. Specifically, the training execution unit 2040 inputs the first training image 52 into the image conversion model 100 to obtain the output image 20 from the image conversion model 100. Further, the training execution unit 2040 inputs the output image 20 into the discrimination model 200 to obtain the discrimination data 40 from the discrimination model 200. Then, the training execution unit 2040 calculates a first loss representing the magnitude of the difference between the discrimination data 40 and the first class information 54, and updates the image conversion model 100 using the first loss.
[0030] <Example of Hardware Configuration> Each functional component of the model training apparatus 2000 may be realized by hardware (e.g., a hard-wired electronic circuit, etc.) that realizes each functional component, or may be realized by a combination of hardware and software (e.g., a combination of an electronic circuit and a program that controls it, etc.). Hereinafter, the case where each functional component of the model training apparatus 2000 is realized by a combination of hardware and software will be further described.
[0031] FIG. 5 is a block diagram illustrating the hardware configuration of the computer 1000 that realizes the model training apparatus 2000. The computer 1000 is an arbitrary computer. For example, the computer 1000 is a stationary computer such as a PC (Personal Computer) or a server machine. In addition, for example, the computer 1000 is a portable computer such as a smartphone or a tablet terminal. The computer 1000 may be a dedicated computer designed to realize the model training apparatus 2000, or may be a general-purpose computer.
[0032] For example, by installing a predetermined application on computer 1000, each function of model training device 2000 is realized on computer 1000. The above application is composed of programs for realizing each functional component of model training device 2000. Note that the method for obtaining the above program is arbitrary. For example, the program can be obtained from a storage medium (such as a DVD disk or a USB memory) in which the program is stored. Among other things, for example, the program can be obtained by downloading the program from a server device that manages the storage device in which the program is stored.
[0033] Computer 1000 has a bus 1020, a processor 1040, a memory 1060, a storage device 1080, an input / output interface 1100, and a network interface 1120. Bus 1020 is a data transmission path for processor 1040, memory 1060, storage device 1080, input / output interface 1100, and network interface 1120 to transmit and receive data to and from each other. However, the method of connecting processor 1040 and the like to each other is not limited to bus connection.
[0034] Processor 1040 is various processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), or an FPGA (Field-Programmable Gate Array). Memory 1060 is a main storage device realized using, for example, a RAM (Random Access Memory). Storage device 1080 is an auxiliary storage device realized using, for example, a hard disk, an SSD (Solid State Drive), a memory card, or a ROM (Read Only Memory).
[0035] The input / output interface 1100 is an interface for connecting the computer 1000 and the input / output device. For example, an input device such as a keyboard and an output device such as a display device are connected to the input / output interface 1100.
[0036] The network interface 1120 is an interface for connecting the computer 1000 to a network. This network may be a LAN (Local Area Network) or a WAN (Wide Area Network).
[0037] The storage device 1080 stores a program (a program for realizing the aforementioned application) for realizing each functional component of the model training device 2000. The processor 1040 reads this program into the memory 1060 and executes it to realize each functional component of the model training device 2000.
[0038] The model training device 2000 may be realized by one computer 1000 or may be realized by a plurality of computers 1000. In the latter case, the configurations of the respective computers 1000 do not need to be the same and may be different from each other.
[0039] <Flow of processing> FIG. 6 is a flowchart illustrating the flow of processing executed by the model training device 2000 of the present embodiment. The acquisition unit 2020 acquires the first training dataset 50 (S102). The training execution unit 2040 inputs the first training image 52 to the image conversion model 100 (S104). The training execution unit 2040 inputs the output image 20 output from the image conversion model 100 to the identification model 200 (S106). The training execution unit 2040 calculates a first loss based on the magnitude of the difference between the identification data 40 output from the identification model 200 and the first class information 54 (S108). The training execution unit 2040 updates the image conversion model 100 using the first loss (S110).
[0040] Note that the model training device 2000 trains the image conversion model 100 by acquiring a plurality of first training data sets 50 and repeatedly updating the image conversion model 100 using the plurality of first training data sets 50.
[0041] <Example of Use of Image Conversion Model 100> To facilitate understanding of the usefulness of the model training device 2000, usage scenarios of the image conversion model 100 are exemplified. The usage scenarios described here are examples, and the usage scenarios of the model training device 2000 are not limited to the examples described below.
[0042] As a usage scenario, assume a case where video data obtained from a surveillance camera that images a road is used for vehicle surveillance. Vehicle surveillance is performed by detecting a vehicle from each video frame of the video data using a surveillance device. The surveillance device has a detection model that has been pre-trained to detect a vehicle from an image.
[0043] Here, the appearance of an object in an image (the image features of the object) can vary depending on the environment in which the object is imaged. For example, a vehicle imaged during the day and a vehicle imaged at night look different from each other. Also, a vehicle imaged on a sunny day and a vehicle imaged on a rainy day look different from each other.
[0044] The detection model used for vehicle surveillance is preferably robust to such environmental changes. That is, the detection model needs to be trained so that it can detect a vehicle from each video frame regardless of the time of day or the weather. For that purpose, it is necessary to train the detection model using images of roads taken under various environments as training images.
[0045] In this regard, the ease of obtaining training images can vary from environment to environment. For example, at night, the number of vehicles is less compared to daytime. Therefore, the number of images of vehicles on the road captured at night is less than that of images of vehicles on the road captured during the day that can be obtained from surveillance cameras. Also, in places with more sunny days, the number of images of vehicles on the road captured during non-sunny days such as rain or snow is less than that of images of vehicles on the road captured during sunny days that can be obtained from surveillance cameras. Due to the fact that the number of available images varies by environment, if the detection model is trained using only the images obtained from surveillance cameras, the detection accuracy of vehicles in environments such as at night or on rainy days will be low.
[0046] Therefore, by using the image conversion model 100 trained by the model training device 2000 to perform data augmentation using images in an environment where acquisition is easy, images in an environment where acquisition is difficult are pseudo-generated. For example, assume that the model training device 2000 has been pre-trained such that when an image of a vehicle on a road during the day is input as the input image 10, an image of a vehicle on a road at night is output as the output image 20. FIG. 7 is a diagram illustrating the effect of data augmentation using the image conversion model 100.
[0047] The upper part of FIG. 7 shows a case where the detection model is trained using only the images obtained from the surveillance camera without performing data augmentation by the image conversion model 100. In this case, since the number of training images of vehicles captured at night is insufficient, the detection accuracy of vehicles at night becomes low.
[0048] On the one hand, the lower part of FIG. 7 illustrates a case where data augmentation is performed by the image conversion model 100. The user inputs an image of a car on a daytime road obtained from a surveillance camera into the image conversion model 100 to obtain an image that pseudo-represents a car on a nighttime road. By doing so, it is possible to obtain the same number of images of cars on a nighttime road as those on a daytime road. By using the images obtained using the image conversion model 100 as training images to train the detection model, it is possible to generate a detection model that can accurately detect cars at night. That is, it is possible to generate a detection model that is robust to environmental changes.
[0049] Here, in order to train the detection model, in addition to the training images, information indicating which part of the training image the car is located in is also required. This information can be regarded as class information indicating which of the two classes, car and others, each image region included in the training image belongs to. However, when the detection model is to be able to detect not only cars but also other types of objects (such as people and roads), those types are also indicated by the class information.
[0050] Here, if the generation of class information has to be done manually for the images generated using the image conversion model 100, a lot of time will be required for data augmentation (generation of the training dataset) using the image conversion model 100. In this regard, if the class of each image region of the input image 10 matches the class of each image region of the output image 20, the class information of the input image 10 can be directly reused as the class information of the output image 20. Therefore, the time required for data augmentation using the image conversion model 100 can be greatly reduced. Thus, as described above, the image conversion model 100 is trained not to perform class conversion although it performs environmental conversion.
[0051] <Regarding Classes> The types of classes of image regions handled by the model training device 2000 can be arbitrarily set according to characteristics of the scene represented by the images handled by the image conversion model 100, etc. For example, the image regions can be classified into two classes: a predetermined object that can be included in the images handled by the image conversion model 100, and the rest. For example, when the predetermined object is a car, the first class information 54 indicates the class of "car" for the image region representing a car, and the class of "other than car" for the image region representing other than a car.
[0052] As the predetermined object, multiple types of objects may be handled. For example, it is conceivable to further classify cars in more detail. Specifically, it is conceivable to provide classes such as "ordinary car", "bus", "truck", "motorcycle", and "bicycle". In addition, for example, classes other than cars such as "road", "building", or "person" may be further provided. When providing the class of road, the road may be further classified according to the traveling direction of the car.
[0053] <Configuration of the Image Conversion Model 100> For example, the image conversion model 100 is configured to extract features from the input image 10 and generate the output image 20 based on the extracted features. FIG. 8 is a diagram illustrating the configuration of the image conversion model 100. The image conversion model 100 includes two models: a feature extraction model 110 and an image generation model 120. The feature extraction model 110 is configured to extract a feature map from the input image 10. Here, the feature map extracted from the image is a set of feature amounts obtained from each of a plurality of partial regions included in the image. The image generation model 120 is configured to generate the output image 20 from the feature map.
[0054] Both the feature extraction model 110 and the image generation model 120 are configured as arbitrary types of machine learning models. For example, both the feature extraction model 110 and the image generation model 120 are configured by neural networks.
[0055] Note that the image conversion model 100 may use the class information corresponding to the input image 10 to generate the output image 20. In this case, for example, when generating the output image 20 from the first training image 52, the image conversion model 100 further uses the first class information 54. For example, the first training image 52 is input to the image generation model 120. Here, as a technique for using class information in a model for generating an image, for example, the technique disclosed in Non-Patent Document 2 can be used.
[0056] <Acquisition of the first training dataset 50: S102> The acquisition unit 2020 acquires the first training dataset 50 (S102). There are various methods for the acquisition unit 2020 to acquire the first training dataset 50. For example, the first training dataset 50 is stored in an arbitrary storage device in a manner that can be acquired from the model training device 2000 in advance. In this case, the acquisition unit 2020 reads out the first training dataset 50 from the storage device. In addition, for example, the acquisition unit 2020 may acquire the first training dataset 50 by receiving the first training dataset 50 transmitted from another device.
[0057] <Training of the image conversion model 100: S104~S110> The training execution unit 2040 trains the image conversion model 100 using the first training dataset 50. As described above, the training execution unit 2040 inputs the first training image 52 to the image conversion model 100 (S104), and inputs the output image 20 output from the image conversion model 100 to the identification model 200 (S106). Further, the training execution unit 2040 calculates a first loss representing the magnitude of the difference between the identification data 40 output from the identification model 200 and the first class information 54, and updates the image conversion model 100 using the first loss. Note that various existing methods can be used as a specific method for updating the parameters of the model based on the loss.
[0058] Here, as the loss function for calculating the first loss (hereinafter referred to as the first loss function), various functions that can represent the magnitude of the difference between the identification data 40 and the first class information 54 can be used. For example, as the first loss function, the following formula (1) can be used.
Equation
[0059] Here, x1 and t1 represent the first training image 52 and the first class information 54, respectively. L1(x1,t) represents the first loss calculated using the first training image x1 and the first class information t1. c represents the class identifier. N represents the total number of classes. α_c represents the weight given to the class with the identifier c. Note that the calculation method of this weight is disclosed in Non-Patent Document 3. Also, the symbol “_” represents a subscript. i represents the identifier of the image region to be identified. M represents the total number of image regions included in the output image 20. For example, when each pixel is treated as an image region and the number of pixels in the vertical and horizontal directions in the output image 20 is H and W, respectively, M = H * W. t1_i,c indicates 1 when the class of the image region i in the first class information t1 is c, and indicates 0 when the class of the image region i in the first class information t1 is not c. G(x1) represents the output image 20 generated by inputting the first training image x1 into the image conversion model 100. Note that when the first Class information is also input, G(x1,t1) is used instead of G(x1). D(G(x1)) represents the identification data 40 output by the identification model 200 in response to the input of the output image 20. D(G(x1))_i,c is the value indicated by the score vector corresponding to the image region i in the identification data 40 for the class c. That is, it represents the probability that the class of the image region i of the output image 20 calculated by the identification model 200 is c.
[0060] The training execution unit 2040 may further calculate a loss based on the difference between the first training image 52 and the output image 20, and update the image conversion model 100 based on both the loss and the aforementioned first loss. For example, in this case, the training execution unit 2040 calculates an overall loss as a weighted sum of these two losses, and updates the image conversion model 100 so that the overall loss becomes smaller.
[0061] As the loss based on the difference between the first training image 52 and the output image 20, for example, the patchwise contrastive loss disclosed in Non-Patent Document 1, the cycle consistency loss disclosed in Non-Patent Document 4, etc. can be used. However, the loss based on the difference between the first training image 52 and the output image 20 is not limited to those disclosed in these non-patent documents. Also, when using the patchwise contrastive loss, the following contrivance may be applied.
[0062] Note that the loss may be calculated collectively for a plurality of first training data sets 50. In this case, the loss for training the image conversion model 100 can be generalized by, for example, the following formula.
Equation
[0063] <Regarding the identification model 200> As described above, the discrimination model 200 discriminates the authenticity and class of each of a plurality of image regions included in the input image. Here, if "a false image region" is treated as one class, the discrimination model 200 can be regarded as a model that discriminates the class of each of a plurality of image regions included in the input image, that is, a model that performs semantic segmentation. Therefore, as the discrimination model 200, various models capable of realizing semantic segmentation can be adopted. As such a model, for example, a model composed of an encoder and a decoder can be adopted in the same manner as the OASIS discriminator disclosed in Non-Patent Document 3.
[0064] FIG. 9 is a diagram illustrating the configuration of the discrimination model 200. The encoder 210 receives the input image 30 as an input and generates a feature map of the input image 30. The decoder 220 receives the feature map output from the encoder 210 as an input and calculates discrimination data 40 from the feature map. For example, both the encoder 210 and the decoder 220 are each composed of a plurality of resblocks, in the same manner as the OASIS discriminator. Further, a skip connection may be provided between the encoder 210 and the decoder 220 so that the intermediate output of the encoder 210 can also be utilized by the decoder 220.
[0065] The discrimination model 200 may be pre-trained or may be trained together with the image conversion model 100. In the latter case, for example, the model training device 2000 trains the image conversion model 100 and the discrimination model 200 by training an adversarial generation network composed of the image conversion model 100 and the discrimination model 200. This case will be further described below.
[0066] The acquisition unit 2020 acquires a second training dataset 60 and a third training image 70 for use in training the identification model 200. The second training dataset 60 includes a second training image 62 and second class information 64. The second training image 62 is a true image representing a scene in a second environment. For example, the second training image 62 is generated by actually imaging a scene in the second environment with a camera. The second class information 64 indicates the class of each image region included in the second training image 62. The third training image 70 is an image representing a scene in a first environment.
[0067] The second training dataset 60 is used to obtain an identification model 200 that can correctly identify the classes of true image regions. The training execution unit 2040 obtains identification data 40 by inputting the second training image 62 into the identification model 200. Then, the training execution unit 2040 calculates a second loss using this identification data 40 and the second class information 64.
[0068] Here, since the second training image 62 is a true image, it is desirable for the identification model 200 to be able to correctly identify the classes of each image region included in the second training image 62. That is, for all image regions, it is preferable that the class indicated by the second class information 64 and the class specified by the identification data 40 match each other. Thus, for example, the second loss is made smaller as the class indicated by the identification data 40 for each image region matches the class indicated by the second class information 64.
[0069] On the other hand, the third training image 70 is used to obtain an identification model 200 that can correctly identify false image regions. The training execution unit 2040 obtains an output image 20 by inputting the third training image 70 into the image conversion model 100. Further, the training execution unit 2040 obtains identification data 40 by inputting the output image 20 into the identification model 200. Then, the training execution unit 2040 calculates a third loss using this identification data 40.
[0070] In addition, when the image conversion model 100 uses class information to generate the output image 20, the acquisition unit 2020 further acquires the class information corresponding to the third training image 70. Then, the training execution unit 2040 inputs the third training image 70 and this class information into the image conversion model 100 to obtain the output image 20.
[0071] Here, since the output image 20 input to the discrimination model 200 is a fake image, it is preferable that the discrimination model 200 can discriminate that each image region included in the third training image 70 is a fake image region. That is, it is preferable that the discrimination data 40 obtained using the third training image 70 indicates that all image regions are fake image regions. Therefore, for example, the higher the probability that the discrimination data 40 indicates a fake image region for each image region, the smaller the third loss becomes.
[0072] In view of the above, the training execution unit 2040 updates the trainable parameters of the discrimination model 200 using the second loss calculated using the second training dataset 60 and the third loss calculated using the third training image 70. For example, the training execution unit 2040 calculates the weighted sum of the second loss and the third loss, and updates the trainable parameters of the discrimination model 200 so as to minimize the weighted sum. For example, this weighted sum can be expressed by the following formula (3).
Equation
[0073] Unless otherwise specified, among the symbols included in Equation (3), the symbols also included in Equation (1) have the same meaning as in Equation (1). x2, t2, and x3 represent the second training image 62, the second class information 64, and the third training image 70, respectively. L_D(x2, t2, x3) represents the loss for training the discrimination model 200 calculated using the second training image x2, the second class information t2, and the third training image x3. L2(x2, t2) represents the second loss calculated using the second training image x2 and the second class information t2. L3(x3) represents the third loss calculated using the third training image x3. γ represents the weight given to the third loss. t2_i,c represents 1 when the class of the image region i in the second class information t2 is c, and represents 0 when the class of the image region i in the second class information t2 is not c. D(x2) represents the discrimination data 40 output by the discrimination model 200 in response to the input of the second training image x2. D(x2)_i,c represents the probability that the class of the image region i indicated by this discrimination data 40 is c.
[0074] G(x3) represents the output image 20 output by the image conversion model 100 in response to the input of the third training image x3. D(G(x3)) represents the discrimination data 40 output by the discrimination model 200 in response to the input of this output image 20. D(G(x3))_i,c = N + 1 represents the probability that the image region i is a fake image region indicated by this discrimination data 40. Here, the score vector of the discrimination data 40 indicates the probability that the target image region is a fake image region in the (N + 1)-th element.
[0075] Note that, similar to the loss L_G for training the image conversion model 100, the loss L_D for training the discrimination model 200 may also be calculated collectively for a plurality of second training data sets 60 and third training images 70. In this case, the loss L_D can be generalized as follows.
Equation
[0076] The training execution unit 2040 improves the accuracy of both the image conversion model 100 and the discrimination model 200 by repeatedly performing both the training of the image conversion model 100 and the training of the discrimination model 200. For example, the training execution unit 2040 alternately repeats the training of the image conversion model 100 and the training of the discrimination model 200. Additionally, for example, the training execution unit 2040 may alternately repeat the training of the image conversion model 100 for a predetermined number of times and the training of the discrimination model 200 for a predetermined number of times. However, the number of times of training of the image conversion model 100 and the number of times of training of the discrimination model 200 may be different from each other.
[0077] <Output of processing results> The model training device 2000 outputs, as a processing result, information (hereinafter, output information) that can identify the trained image conversion model 100. The output information includes at least the parameter group of the image conversion model 100 obtained by training. In addition to this, the output information may include a program that realizes the image conversion model 100. Further, the output information may further include the parameter group of the discrimination model 200 and a program that realizes the discrimination model 200.
[0078] The output mode of the output information is arbitrary. For example, the model training device 2000 stores the output information in an arbitrary storage unit. Additionally, for example, the model training device 2000 transmits the output information to another device (for example, a device used for the operation of the image conversion model 100).
[0079] <Considerations in calculating patchwise contrastive loss> Here, for the case of using patchwise contrastive loss in the training of the image conversion model 100, the points of consideration in its calculation will be described. First, the patchwise contrastive loss will be briefly described.
[0080] FIG. 10 is a diagram illustrating a method for calculating a patch-wise contrastive loss. The training execution unit 2040 obtains an output image 20 by inputting the first training image 52 into the image conversion model 100. Further, the training execution unit 2040 obtains a first feature map 130, which is a feature map of the first training image 52 calculated by the feature extraction model 110. Furthermore, the training execution unit 2040 obtains a second feature map 140, which is a feature map of the output image 20, by inputting the output image 20 into the feature extraction model 110. The training execution unit 2040 calculates a patch-wise contrastive loss using the first feature map 130 and the second feature map 140.
[0081] More specifically, the training execution unit 2040 extracts feature amounts corresponding to the positive example patches and one or more negative example patches of the first training image 52 from the first feature map 130. The training execution unit 2040 also extracts a feature amount corresponding to the positive example patch of the output image 20 from the second feature map 140.
[0082] Here, the positive example patch and the negative example patch will be described. FIG. 11 is a diagram illustrating the positive example patch and the negative example patch. Both the positive example patch 522 and the negative example patch 524 are partial image regions of the first training image 52. The positive example patch 22 is an image region representing the same location as the location represented by the positive example patch 522 among partial image regions of the output image 20. In this way, the image regions for which feature amounts are to be extracted in both the first training image 52 and the output image 20 are called positive example patches. On the other hand, the image regions for which feature amounts are to be extracted only in the first training image 52 are called negative example patches. Hereinafter, the combination of the positive example patch 522, the negative example patch 524, and the positive example patch 22 is called a patch set.
[0083] As shown in FIG. 11, among the feature amounts included in the first feature map 130, there are feature amounts corresponding to each image region of the first training image 52. Therefore, the training execution unit 2040 extracts the feature amounts corresponding to the positive example patch 522 and the negative example patch 524 from the first feature map 130. Similarly, the training execution unit 2040 extracts the feature amount corresponding to the positive example patch 22 from the second feature map 140.
[0084] The training execution unit 2040 generates one or more patch sets for the pair of the first training image 52 and the output image 20. Then, for each patch set, the training execution unit 2040 extracts feature amounts from the first feature map 130 and the second feature map 140.
[0085] Here, in Non-Patent Document 1, the position of the positive example patch is randomly selected. In this regard, for example, the training execution unit 2040 is devised to extract positive example patches mainly from image regions belonging to a specific class (hereinafter referred to as specific regions). Here, the term "mainly" means that the case where the positive example patch 522 is extracted from the specific region is more frequent than the case where the positive example patch 522 is extracted from other partial regions. By mainly extracting the positive example patch 522 from the specific region in this way, the features of the image regions belonging to a specific class (for example, the features of a specific type of object) can be mainly learned by the image conversion model 100. Therefore, the image conversion model 100 can accurately convert the image regions of a specific class in the first environment into the image regions in the second environment.
[0086] For example, it is assumed that the image conversion model 100 is used to perform data augmentation on the training data of the detection model illustrated using FIG. 7. In this case, it is preferable that the image conversion model 100 can accurately convert the features of the car in the first environment into the features of the car in the second environment. Therefore, by mainly using the image region of the car as the positive example patch, the features of the car are mainly learned by the image conversion model 100.
[0087] Note that the specific method of focusing on using the image regions of a specific class as positive example patches will be described later. First, the method of calculating the patch-wise contrastive loss will be described in more detail.
[0088] The training execution unit 2040 calculates the patch-wise contrastive loss using the feature amounts corresponding to the positive example patches 522, the feature amounts corresponding to the negative example patches 524, and the feature amounts corresponding to the positive example patches 22 obtained for each patch set. The loss for one patch set is calculated as the cross-entropy loss represented by, for example, the following formula (5).
Equation
[0089] When there is one patch set, the patch-wise contrastive loss is calculated by the above formula (5). On the other hand, considering the case where there are multiple patch sets, the patch-wise contrastive loss can be generalized as in the following formula (6).
Equation
[0090] The feature extraction model 110 may be configured to perform multi-stage feature extraction. For example, such a feature extraction model 110 may include a convolutional neural network having a plurality of convolutional layers. In a convolutional neural network having a plurality of convolutional layers, the n-th convolutional layer outputs the n-th feature map by performing a convolution operation of the (n - 1)-th filter on the (n - 1)-th feature map output from the (n - 1)-th convolutional layer (n is an integer of 2 or more).
[0091] When multi-stage feature extraction is performed in this way, not only the first feature map 130 and the second feature map 140, which are the finally obtained feature maps, but also the feature maps obtained in the intermediate stages can be used for calculating the patch-wise contrastive loss. That is, a plurality of feature maps obtained from the first training image 52 and a plurality of feature maps obtained from the output image 20 can be used for calculating the patch-wise contrastive loss.
[0092] For example, when the feature extraction model 110 is an n-layer convolutional neural network, n feature maps can be obtained by obtaining feature maps from each layer. And the feature amounts corresponding to the positive example patch 522, the negative example patch 524, and the positive example patch 22 can be extracted from each of the n feature maps. Therefore, the training execution unit 2040 extracts the feature amounts corresponding to the positive example patch 522, the negative example patch 524, and the positive example patch 22 from each of the n feature maps, and calculates the patch-wise contrastive loss using the extracted feature amounts.
[0093] When calculating the patch-wise contrastive loss using a plurality of feature maps obtained from each of the first training image 52 and the output image 20, for example, the patch-wise contrastive loss is represented by the following formula (7).
Equation
[0094] Also, as described above, the patch-wise contrastive loss may be calculated collectively for a plurality of first training images 52. In this case, the patch-wise contrastive loss can be generalized by the following formula (8).
Equation
[0095] The training execution unit 2040 calculates the first loss and the patch-wise contrastive loss using one or more first training data sets 50, and updates the image conversion model 100 using the comprehensive loss calculated using these. For example, this comprehensive loss is represented by the above-described formula (2).
[0096] <<Regarding the generation of patch sets>> The training execution unit 2040 generates patch sets for the first training image 52 and the output image 20. As described above, one patch set includes one positive example patch 522, one or more negative example patches 524, and one positive example patch 22. For example, after the training execution unit 2040 extracts the positive example patch 522 from the first training image 52, it performs a process of extracting one or more negative example patches 524 from the regions other than the positive example patch 522 in the first training image 52, and a process of extracting the positive example patch 22 from the output image 20.
[0097] As described above, the positive example patch 522 is preferably extracted intensively from a specific region. Therefore, the training execution unit 2040 detects a specific region from the first training image 52 for use in extracting the positive example patch 522. Here, existing techniques can be used for the technique of detecting an image region of a specific class from the first training image 52. Hereinafter, this "specific class" is referred to as the "target class".
[0098] The target class may be predetermined or may be specifiable by the user. In the latter case, the training execution unit 2040 acquires information representing the target class and detects the image region of the target class indicated in the information as the specific region. The information representing the target class is obtained, for example, as a result of user input.
[0099] Hereinafter, several examples of a method for extracting the positive example patch 522 based on the detected specific region will be exemplified.
[0100] <<Method 1>> First, the training execution unit 2040 determines whether to extract the positive example patch 522 from inside or outside the specific region. This determination is made so that the number of positive example patches 522 extracted from inside the specific region is larger than the number of positive example patches 522 extracted from outside the specific region. By doing so, the positive example patch 522 is extracted intensively from the specific region.
[0101] For example, the above determination is made probabilistically. As a method of probabilistically selecting one of the two options, for example, a method of sampling a value from a Bernoulli distribution and making a determination based on the sampled value can be considered. More specifically, for example, when the sampled value is 1, the positive example patch 522 is extracted from within the specific region, and when the sampled value is 0, the positive example patch 522 is extracted from outside the specific region. At this time, by making the probability that the sampled value is 1 greater than 50%, the number of positive example patches 522 extracted from within the specific region can be made probabilistically larger than the number of positive example patches 522 extracted from outside the specific region.
[0102] After determining whether to extract the positive example patch 522 from inside or outside the specific region, the training execution unit 2040 extracts the positive example patch 522 based on this determination. Here, the size of the positive example patch 522 (hereinafter referred to as the patch size) is determined in advance. When extracting the positive example patch 522 from within the specific region, the training execution unit 2040 extracts a region of the patch size from an arbitrary location within the specific region and treats this region as the positive example patch 522. On the other hand, when extracting the positive example patch 522 from outside the specific region, the training execution unit 2040 selects a region of the patch size from an arbitrary location outside the specific region and determines the selected region as the positive example patch 522. Note that existing techniques can be used for the technique of arbitrarily selecting a region of a predetermined size from within a certain region.
[0103] Note that when extracting the positive example patch 522 from within the specific region, a part of the positive example patch 522 may be outside the specific region. For example, in this case, the positive example patch 522 is extracted so as to satisfy the condition that "a predetermined ratio or more of the positive example patch 522 is within the specific region".
[0104] <<Method 2>> The training execution unit 2040 extracts the positive example patches 522 such that the probability of extraction as the positive example patches 522 is higher for the regions with a greater overlap with the specific region. To this end, for example, the training execution unit 2040 generates an extraction probability map indicating a higher extraction probability as the overlap rate with the specific region is higher. For example, the extraction probability map is generated as a probability distribution indicating the probability that a region of the patch size with that pixel as the starting point (for example, the upper left corner of the positive example patch 522) is extracted as the positive example patch 522 for each pixel of the first training image 52. In order to increase the extraction probability as the overlap rate with the specific region is higher, the extraction probability map is generated such that for each pixel, the higher the degree of overlap between the region of the patch size with that pixel as the starting point and the specific region, the higher the extraction probability. Note that it can also be said that the extraction probability map indicates the probability that each partial region of the patch size included in the first training image 52 is extracted as the positive example patch 522. And the extraction probability of each partial region is set higher as the degree of overlap between that partial region and the specific region is higher.
[0105] To generate such an extraction probability map, for example, first, the training execution unit 2040 sets a value representing the degree of overlap between the region of the patch size with each pixel of the extraction probability map as the starting point and the specific region. Then, the training execution unit 2040 changes the value of each pixel of the extraction probability map to a value obtained by dividing the value by the sum of the values of all the pixels.
[0106] FIG. 12 is a diagram illustrating the extraction probability map. In this example, the size of the positive example patch 522 is 2x2. Also, the size of the specific region 410 is 4x3. Each pixel of the extraction probability map 400 indicates a higher extraction probability as the degree of overlap between the positive example patch 522 and the specific region is greater when the positive example patch 522 is extracted with that pixel as the upper left corner. Here, in FIG. 12, the pixels with a higher extraction probability are represented by darker dots. Therefore, in FIG. 12, it shows that the higher the pixel represented by a darker dot, the higher the probability that the positive example patch 522 is extracted with that pixel as the starting point.
[0107] The training execution unit 2040 samples the coordinates of pixels from the probability distribution represented by the extraction probability map, and extracts a region of the patch size starting from the sampled coordinates as the positive example patch 522.
[0108] <<Method 3>> When the target class represents the class of an object, the object may be further classified into finer sub-classifications, and based on the sub-classifications, the extraction probability of each pixel of the above-mentioned extraction probability map may be determined. For example, when the target class is a car, the sub-classifications may include types such as passenger cars, trucks, or buses. Hereinafter, the class on the sub-classification to which the object included in the first training image 52 belongs is referred to as a sub-class.
[0109] When considering sub-classifications, among the objects belonging to the target class, the importance in the training of the image conversion model 100 may vary for each sub-class. For example, for an object of a class that appears less frequently in the first training image 52, since it is necessary to enable the image conversion model 100 to learn its features with less training, it can be said that it is an important object in training.
[0110] As a specific example, assume that an image representing the state of a road during the day is used as the input image 10, and the image conversion model 100 is trained so that an output image 20 representing the state of the road at night is generated from the input image 10. Here, assume that on the road where the first training image 52 was captured, the appearance frequency of trucks is lower than that of passenger cars. In this case, the opportunity for the image conversion model 100 to learn the features of trucks is less than the opportunity to learn the features of passenger cars. Therefore, it is necessary to enable the image conversion model 100 to learn the features of trucks with less training.
[0111] Therefore, for example, the lower the appearance frequency of a subclass, the higher its importance in training. More specifically, the training execution unit 2040 generates an extraction probability map such that in the first training image 52, the extraction probability of a specific region representing an object belonging to a subclass with a lower appearance frequency is higher. For this purpose, for each subclass, a higher weight is set as its appearance frequency is lower.
[0112] For each pixel of the extraction probability map, the training execution unit 2040 sets a value obtained by multiplying the degree of overlap between the pixel and the specific region by the weight corresponding to the subclass of the object represented by the specific region. Then, the training execution unit 2040 changes the value of each pixel to a value obtained by dividing it by the sum of the values of all pixels.
[0113] The training execution unit 2040 samples the coordinates of pixels from the probability distribution represented by this extraction probability map, and extracts a region of patch size starting from the sampled coordinates as a positive example patch 522.
[0114] Here, the weight of each subclass may be predetermined or determined by the training execution unit 2040. In the latter case, for example, before extracting the positive example patch 522, the training execution unit 2040 performs a process of detecting an object of the target class for each first training image 52 acquired by the acquisition unit 2020, and counts the number of detected objects for each subclass. Thereby, the number of appearances of each subclass in the training image group is specified. The training execution unit 2040 determines the weight of each subclass based on the number of appearances of each subclass. This weight is determined such that the weight of a subclass with a smaller number of appearances is larger.
[0115] <<Method for Extracting Negative Example Patch 524>> The training execution unit 2040 randomly extracts a region of the patch size from regions in the first training image 52 other than the region extracted as the positive example patch 522 among the regions included in the first training image 52, and uses that region as the negative example patch 524. As described above, one patch set may include a plurality of negative example patches 524. The number of negative example patches 524 included in one patch set is determined in advance.
[0116] <<Method for Extracting Positive Example Patch 22>> The training execution unit 2040 extracts the positive example patch 22 from the position on the output image 20 corresponding to the position on the first training image 52 from which the positive example patch 522 was extracted. That is, the coordinates of the pixel serving as the starting point for extracting the positive example patch 22 are the same as the coordinates used as the starting point for extracting the positive example patch 522.
[0117] <Other Methods> In the above-described model training apparatus 2000, by extracting the positive example patch 522 intensively from the image region of the target class, the features of the object of the target class are learned with particularly high accuracy. However, the method of learning the features of the object of the target class with high accuracy is not limited to the method of intensively extracting the positive example patch 522 from a specific region.
[0118] For example, in addition to or instead of intensively extracting the positive example patch 522 from a specific region, the model training apparatus 2000 calculates a patch-wise contrastive loss so that the influence of the loss (for example, the cross-entropy loss described above) calculated using the features corresponding to the positive example patch 522 extracted from the specific region is greater than the influence of the loss calculated using the feature amounts corresponding to the positive example patches 522 extracted from other regions. Note that when the method of intensively extracting the positive example patch 522 from a specific region is not adopted, for example, the positive example patch 522 is extracted with the same probability from any location in the first training image 52.
[0119] Next, a method for determining the influence degree of the loss based on the feature amount corresponding to the positive example patch 522 will be described according to whether the positive example patch 522 is extracted from inside or outside the specific region.
[0120] For example, the training execution unit 2040 calculates the patch-wise contrastive loss using the following formula (9).
Equation
[0121] In Equation (7), for the loss calculated for each patch set, when the positive example patch 522 included in the patch set is extracted from inside the specific region, the weight a is multiplied, while when the positive example patch 522 included in the patch set is extracted from outside the specific region, the weight b is multiplied. Since a > b > 0, the influence of the loss when the positive example patch 522 is extracted from inside the specific region is greater than the influence of the loss when the positive example patch 522 is extracted from outside the specific region.
[0122] Note that the same applies when calculating the patch-wise contrastive loss using the aforementioned Equations (7) and (8). That is, when feature maps are obtained from multiple layers of the feature extraction model 110, the above-described weighting is performed on the loss calculated for the feature maps obtained from each layer.
[0123] Also, weights similar to w_s may be used for calculating the first loss, the second loss, and the third loss. In this case, for example, these losses can be calculated using the following formula (10).
Equation
[0124] When feature maps are obtained from multiple layers, weights may be set for each layer or only for a specific layer based on the relationship between the size of the partial region of the input image corresponding to one cell of the feature map and the patch size. This method will be described below.
[0125] When feature maps are obtained from multiple layers, the size of the partial region of the input image corresponding to one cell of the feature map is different for each feature map (each layer). For example, assume that convolution processing of a filter of size 3x3 is performed in each layer. In this case, one cell of the first feature map corresponds to a partial region of size 3x3 in the input image. Also, one cell of the second feature map corresponds to a set of cells of size 3x3 in the first feature map. From this, one cell of the second feature map corresponds to a region of size 9x9 in the input image. For the same reason, one cell of the third feature map corresponds to a region of size 27x27 in the input image. Thus, the feature maps generated by the subsequent layers correspond to larger partial regions of the input image.
[0126] In this regard, in the multiple feature maps generated from different layers for the first training image 52, it is considered that the feature maps in which the size of the partial region of the first training image 52 corresponding to one cell is closer to the patch size represent the features of the positive example patch 522 more accurately. The same applies to the negative example patch 524 and the positive example patch 22.
[0127] Therefore, for example, the training execution unit 2040 calculates the patch-wise contrastive loss so that larger weights are assigned to the feature amounts extracted from the feature maps in which the size of the partial region of the first training image 52 corresponding to one cell is closer to the patch size. The same applies to the positive example patch 22 and the negative example patch 524. In this case, for example, the patch-wise contrastive loss is calculated using the following formula (11).
Equation
[0128] Note that, by attaching a weight greater than 1 only to the layer \(l\) with the smallest difference between \(z_p\) and \(z_l\) and not attaching weights to other layers, weights may be attached only to the layer where the size of the partial region of the input image corresponding to the cell of the feature map is closest to the patch size. Also, a method may be adopted in which weights greater than 1 are attached only to a predetermined number of top layers in ascending order of the small difference between \(z_p\) and \(z_l\).
[0129] As described above, the present invention has been described with reference to the embodiments, but the present invention is not limited to the above embodiments. Various changes that can be understood by those skilled in the art can be made to the configuration and details of the present invention within the scope of the present invention.
[0130] In the above example, when the program is loaded into a computer, it includes a set of instructions (or software code) for causing the computer to perform one or more functions described in the embodiments. The program may be stored in a non-transitory computer-readable medium or a tangible storage medium. By way of example and not limitation, the computer-readable medium or tangible storage medium includes random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD), or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray (registered trademark) disc, or other optical disc storage, magnetic cassette, magnetic tape, magnetic disk storage, or other magnetic storage devices. The program may also be transmitted on a transient computer-readable medium or a communication medium. By way of example and not limitation, the transient computer-readable medium or communication medium includes electrical, optical, acoustic, or other forms of propagated signals.
[0131] Some or all of the above embodiments may be described as follows, but are not limited thereto. (Appendix 1) An acquisition means for acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; Training execution means for training an image conversion model that outputs an image representing a scene in a second environment in response to an image representing a scene in the first environment being input, using the first training dataset; and having, The training execution means inputs the first training image into the image conversion model, inputs a first output image output from the image conversion model into an identification model, calculates a first loss using the identification data output from the identification model and the first class information, and updates the parameters of the image conversion model using the first loss. The identification data indicates, for each of a plurality of partial regions included in the image input to the identification model, whether the partial region is a false image region, and indicates the class of the partial region when the partial region is not a false image. A model training device. (Appendix 2) The first loss is smaller as the number of image regions in which the class indicated by the identification data matches the class indicated by the first class information is larger. The model training device according to Appendix 1. (Appendix 3) The training execution means calculates the first loss by giving a larger weight to an image region belonging to a specific class than to an image region not belonging to the specific class. The model training device according to Appendix 2. (Appendix 4) The image conversion model includes a feature extraction model that extracts a feature map from the input image. The training execution means inputs the first training image into the image conversion model, and obtains from the image conversion model the first output image and a first feature map that is the feature map of the first training image. inputs the first output image into the feature extraction model, and obtains from the feature extraction model a second feature map that is the feature map of the first output image. updates the parameters of the image conversion model using both the feature loss calculated using the first feature map and the second feature map and the first loss. The model training device according to any one of Appendices 1 to 3. (Appendix 5) The training execution means generates one or more patch sets that are a set of a first positive example patch and a first negative example patch that are partial regions of the first training image, and a second positive example patch that is a partial region at a position corresponding to the first positive example patch in the first output image. extracts feature amounts corresponding to the first positive example patch and the first negative example patch respectively from the first feature map, extracts the feature amount corresponding to the second positive example patch from the second feature map, and calculates the feature loss using each of the extracted feature amounts. The training execution means In the generation of the patch set, among the regions included in the first training image, the first positive example patch is extracted preferentially from a specific region belonging to a specific class, or The model training device according to Supplementary Note 4, wherein the feature loss is calculated such that the influence of the loss calculated for the patch set including the first positive example patch extracted from the specific region is greater than the influence of the loss calculated for the patch set including the first positive example patch extracted from outside the specific region. (Supplementary Note 6) The acquisition means acquires a second training dataset including a second training image representing a scene in the first environment and second class information indicating the class of each of a plurality of image regions included in the second training image, and a third training image representing a scene in the second environment. The training execution means Inputs the second output image obtained by inputting the second training image into the image conversion model into the identification model, and calculates a second loss using the identification data output from the identification model and the second class information. Inputs the third training image into the identification model, and calculates a third loss using the identification data output from the identification model. The model training device according to any one of Supplementary Notes 1 to 5, wherein the parameters of the identification model are updated using the second loss and the third loss. (Supplementary Note 7) A model training method executed by a computer, comprising: An acquisition step of acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; A training execution step of training an image conversion model that outputs an image representing a scene in a second environment in response to an input of an image representing a scene in the first environment using the first training dataset. In the training execution step, input the first training image into the image conversion model, input the first output image output from the image conversion model into the discrimination model, calculate a first loss using the discrimination data output from the discrimination model and the first class information, and update the parameters of the image conversion model using the first loss. The discrimination data indicates, for each of a plurality of partial regions included in the image input to the discrimination model, whether the partial region is a false image region, and indicates the class of the partial region when the partial region is not a false image. A model training method. (Appendix 8) The first loss is smaller as the number of image regions where the class indicated by the discrimination data and the class indicated by the first class information match is larger. The model training method according to Appendix 7. (Appendix 9) In the training execution step, for an image region belonging to a specific class, a weight larger than that for an image region not belonging to the specific class is given to calculate the first loss. The model training method according to Appendix 8. (Appendix 10) The image conversion model includes a feature extraction model that extracts a feature map from the input image. In the training execution step, Input the first training image into the image conversion model, and obtain from the image conversion model the first output image and the first feature map that is the feature map of the first training image. Input the first output image into the feature extraction model, and obtain from the feature extraction model the second feature map that is the feature map of the first output image. Update the parameters of the image conversion model using both the feature loss calculated using the first feature map and the second feature map and the first loss. The model training method according to any one of Appendices 7 to 9. (Appendix 11) In the training execution step, Generate one or more patch sets, which are sets of a first positive example patch and a first negative example patch that are partial regions of the first training image, and a second positive example patch that is a partial region at a position corresponding to the first positive example patch in the first output image. Extract feature amounts corresponding to the first positive example patch and the first negative example patch respectively from the first feature map, extract the feature amount corresponding to the second positive example patch from the second feature map, and calculate the feature loss using each of the extracted feature amounts. In the training execution step In the generation of the patch set, either extract the first positive example patch intensively from a specific region belonging to a specific class among the regions included in the first training image, or Calculate the feature loss such that the influence of the loss calculated for the patch set including the first positive example patch extracted from the specific region is greater than the influence of the loss calculated for the patch set including the first positive example patch extracted from outside the specific region. The model training method according to Supplementary Note 10. (Supplementary Note 12) In the acquisition step, acquire a second training dataset including a second training image representing a scene in the first environment and second class information indicating the class of each of a plurality of image regions included in the second training image, and a third training image representing a scene in the second environment. In the training execution step Input a second output image obtained by inputting the second training image into the image conversion model into the discrimination model, and calculate a second loss using the discrimination data output from the discrimination model and the second class information. Input the third training image into the discrimination model, and calculate a third loss using the discrimination data output from the discrimination model. Update the parameters of the discrimination model using the second loss and the third loss. The model training method according to any one of Supplementary Notes 7 to 11. (Supplementary Note 13) To a computer An acquisition step of acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; A training execution step of training an image conversion model that outputs an image representing a scene in a second environment in response to an image representing a scene in the first environment being input, using the first training dataset, is stored. In the training execution step, the first training image is input to the image conversion model, the first output image output from the image conversion model is input to an identification model, a first loss is calculated using the identification data output from the identification model and the first class information, and the parameters of the image conversion model are updated using the first loss. The identification data indicates, for each of a plurality of partial regions included in the image input to the identification model, whether the partial region is a fake image region, and indicates the class of the partial region if the partial region is not a fake image, a non-transitory computer-readable medium. (Appendix 14) The first loss is smaller as the number of image regions where the class indicated by the identification data and the class indicated by the first class information match is larger, a computer-readable medium according to Appendix 13. (Appendix 15) In the training execution step, a larger weight is given to the image regions belonging to a specific class than to the image regions not belonging to the specific class to calculate the first loss, a computer-readable medium according to Appendix 14. (Appendix 16) The image conversion model includes a feature extraction model that extracts a feature map from the input image. In the training execution step, The first training image is input to the image conversion model, and the first output image and a first feature map that is the feature map of the first training image are obtained from the image conversion model. Input the first output image into the feature extraction model, and obtain, from the feature extraction model, a second feature map that is the feature map of the first output image. A computer-readable medium according to any one of Appendices 13 to 15, wherein the image conversion model is updated with parameters using both a feature loss calculated using the first feature map and the second feature map and the first loss. (Appendix 17) In the training execution step, Generate one or more patch sets, which are sets of a first positive example patch and a first negative example patch that are partial regions of the first training image, and a second positive example patch that is a partial region at a position corresponding to the first positive example patch in the first output image. Extract feature amounts corresponding to the first positive example patch and the first negative example patch respectively from the first feature map, extract a feature amount corresponding to the second positive example patch from the second feature map, and calculate the feature loss using each of the extracted feature amounts. In the training execution step, In the generation of the patch set, either extract the first positive example patch intensively from a specific region belonging to a specific class among the regions included in the first training image, or A computer-readable medium according to Appendix 16, wherein the feature loss is calculated such that the influence of the loss calculated for the patch set including the first positive example patch extracted from within the specific region is greater than the influence of the loss calculated for the patch set including the first positive example patch extracted from outside the specific region. (Appendix 18) In the acquisition step, acquire a second training dataset including a second training image representing a scene in the first environment and second class information indicating the class of each of a plurality of image regions included in the second training image, and a third training image representing a scene in the second environment. In the training execution step, Input the second output image obtained by inputting the second training image into the image conversion model into the identification model, calculate a second loss using the identification data output from the identification model and the second class information, Input the third training image into the identification model, and calculate a third loss using the identification data output from the identification model, Update the parameters of the identification model using the second loss and the third loss. The computer-readable medium according to any one of appendices 13 to 17.
Explanation of symbols
[0132] 10 Input image 20 Output image 22 Positive example patch 30 Input image 40 Identification data 50 First training dataset 52 First training image 54 First class information 60 Second training dataset 62 Second training image 64 Second class information 70 Third training image 100 Image conversion model 110 Feature extraction model 120 Image generation model 130 First feature map 140 Second feature map 200 Identification model 210 Encoder 220 Decoder 400 Extraction probability map 410 Specific area 522 Positive example patch 524 Negative example patch 1000 Computer 1020 Bus 1040 Processor 1060 Memory 1080 Storage device 1100 Input / output interface 1120 Network Interface 2000 Model Training Device 2020 Acquisition Unit 2040 Training Execution Unit
Claims
1. An acquisition means for acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; Training execution means for training an image conversion model that outputs an image representing a scene in a second environment in response to an input of an image representing a scene in the first environment using the first training dataset, and having: The training execution means inputs the first training image into the image conversion model, inputs a first output image output from the image conversion model into an identification model, calculates a first loss using the identification data output from the identification model and the first class information, and updates the parameters of the image conversion model using the first loss. The identification data indicates, for each of a plurality of sub-regions included in the image input to the identification model, whether the sub-region is a fake image region, and indicates the class of the sub-region if the sub-region is not a fake image. A model training device.
2. The model training device according to claim 1, wherein the first loss is smaller as the number of image regions in which the class indicated by the identification data and the class indicated by the first class information match is larger.
3. The training execution means calculates the first loss by giving a larger weight to an image region belonging to a specific class than to an image region not belonging to the specific class. The model training device according to claim 2.
4. The image conversion model includes a feature extraction model that extracts a feature map from an input image. The training execution means Inputs the first training image into the image conversion model, and obtains the first output image and a first feature map that is the feature map of the first training image from the image conversion model. Inputs the first output image into the feature extraction model, and obtains a second feature map that is the feature map of the first output image from the feature extraction model. The model training device according to any one of claims 1 to 3, wherein the parameters of the image conversion model are updated using both a feature loss calculated using the first feature map and the second feature map and the first loss.
5. The training execution means Generate one or more patch sets, which are sets of a first positive example patch and a first negative example patch that are partial regions of the first training image, and a second positive example patch that is a partial region at a position corresponding to the first positive example patch in the first output image. Extract feature amounts corresponding to the first positive example patch and the first negative example patch respectively from the first feature map, extract the feature amount corresponding to the second positive example patch from the second feature map, and calculate the feature loss using each of the extracted feature amounts. The training execution means In generating the patch set, either extract the first positive example patch intensively from a specific region belonging to a specific class among the regions included in the first training image, or The model training device according to claim 4, wherein the feature loss is calculated such that the influence of the loss calculated for the patch set including the first positive example patch extracted from the specific region is greater than the influence of the loss calculated for the patch set including the first positive example patch extracted from outside the specific region.
6. The acquisition means acquires a second training dataset including a second training image representing a scene in the first environment and second class information indicating the class of each of a plurality of image regions included in the second training image, and a third training image representing a scene in the second environment. The training execution means Input the second output image obtained by inputting the second training image into the image conversion model into the identification model, and calculate a second loss using the identification data output from the identification model and the second class information. Input the third training image into the identification model, and calculate a third loss using the identification data output from the identification model. The model training device according to any one of claims 1 to 5, wherein the parameters of the identification model are updated using the second loss and the third loss.
7. A model training method executed by a computer, comprising: An acquisition step of acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; A training execution step of training an image conversion model that outputs an image representing a scene in a second environment in response to an image representing a scene in the first environment being input, using the first training dataset. In the training execution step, input the first training image into the image conversion model, input the first output image output from the image conversion model into the discrimination model, calculate a first loss using the discrimination data output from the discrimination model and the first class information, and update the parameters of the image conversion model using the first loss. The discrimination data indicates, for each of a plurality of partial regions included in the image input to the discrimination model, whether the partial region is a fake image region, and indicates the class of the partial region when the partial region is not a fake image. A model training method. Claim 8 The first loss is smaller as the number of image regions where the class indicated by the discrimination data matches the class indicated by the first class information is larger. The model training method according to claim 7. Claim 9 On a computer, an acquisition step of acquiring a first training dataset including a first training image representing a scene in a first environment and first class information indicating the class of each of a plurality of image regions included in the first training image; a training execution step of training an image conversion model that outputs an image representing a scene in a second environment in response to an image representing a scene in the first environment being input, using the first training dataset; In the training execution step, input the first training image into the image conversion model, input the first output image output from the image conversion model into the discrimination model, calculate a first loss using the discrimination data output from the discrimination model and the first class information, and update the parameters of the image conversion model using the first loss. The discrimination data indicates, for each of a plurality of partial regions included in the image input to the discrimination model, whether the partial region is a fake image region, and indicates the class of the partial region when the partial region is not a fake image. A program. Claim 10 The first loss is smaller as the number of image regions where the class indicated by the discrimination data matches the class indicated by the first class information is larger. The program according to claim 9.
Citation Information
Patent Citations
Information processing apparatus, information processing method and program
JP2018163444A
Device and method of generating teacher data for machine learning
JP2019028876A
X-ray image object recognition system
JP2020014799A
Machine learning system, domain conversion device and machine learning method
JP2020095364A