Learning device and learning method
By using a learning device including a first neural network, a second neural network and a learning assisted neural network in the object mark recognition, the problem of difficulty in extracting robust features in the prior art is solved by reducing the influence of biased features by counter-learning, and a more accurate and stable object mark recognition is achieved.
Patent Information
- Application Number
- CN202210066481.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-02-18
- Filing Date
- 2022-01-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-01-20
AI Technical Summary
The prior art is difficult to adaptively extract robust features relative to domains in object mark recognition, and the method using a huge and varied data set is expensive. The method using a single domain data set is easily affected by biased features, resulting in inaccurate recognition results.
Using a learning device including a first neural network, a second neural network and a learning assisted neural network, the second feature and the third feature are approached by adversarial learning, and the frequency of occurrence of the third feature in the first feature is reduced, thereby reducing the influence of the biased feature.
It realizes the adaptive extraction of robust features relative to the domain in object mark recognition, and improves the accuracy and stability of the recognition results.
Smart Images

Figure CN115019116B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a learning device and a learning method. Background Art
[0002] In recent years, there has been known a technique for inputting an image captured by a camera into a deep neural network (DNN) and recognizing an object in the image through inference processing of the DNN.
[0003] In order to improve the robustness of object identification performed by DNN, it is necessary to implement learning (training) using large and varied data sets from different domains. Through learning using large and varied data sets, DNN can extract robust image features that are not inherent to the domain, but from the perspective of data collection costs and huge processing costs, such methods are mostly difficult.
[0004] On the other hand, a technique for extracting robust features by learning DNN using a dataset from one domain has been studied. For example, in a DNN used for object identification, sometimes, in addition to the features that should be focused on, features that are different from the features that should be focused on (biased features) are also considered for learning. In this case, when new image data is subjected to recognition processing, it is sometimes affected by the biased features and cannot output a correct recognition result (i.e., it cannot extract robust features).
[0005] In order to solve such a problem, the following technology is proposed in non-patent document 1: use a model (DNN) that can easily extract local features of an image to extract biased features of the image (texture features in non-patent document 1), and use the HSIC (Hilbert-Schmidt Independence Criterion) benchmark to remove the biased features from the features of the image.
[0006] Prior art literature
[0007] Non-patent literature
[0008] Non-patent literature 1: Hyojin Bahng, 4 others, “Learning De-biased Representations with Biased Representations”, arXiv:1910.02806v2[cs.CV], March 2, 2020 Summary of the invention
[0009] Problems to be solved by the invention
[0010] In the technology proposed in non-patent document 1, a specific model for extracting texture features is determined by design, assuming that the biased features are texture features. That is, in non-patent document 1, a technology dedicated to the case of processing texture features as biased features is proposed. In addition, in non-patent document 1, the HSIC benchmark is used to remove the biased features, and other methods for removing the biased features are not considered.
[0011] The present invention is made in view of the above-mentioned problems, and an object of the present invention is to provide a technology capable of extracting robust features in a domain-adaptive manner in object recognition.
[0012] Means used to solve problems
[0013] According to the present invention, a learning device is provided, the learning device comprising a processing mechanism, wherein the processing mechanism comprises:
[0014] A first neural network extracts a first feature of an object in the image data;
[0015] A second neural network, which uses a network structure different from that of the first neural network to extract a second feature of the object in the image data; and
[0016] learning an auxiliary neural network that extracts a third feature from the first feature extracted by the first neural network,
[0017] The second feature and the third feature are features offset relative to the object,
[0018] The processing mechanism causes the learning auxiliary neural network to learn so that the second feature extracted by the second neural network is close to the third feature extracted by the learning auxiliary neural network, and the processing mechanism causes the first neural network to learn so that the third feature appearing in the first feature extracted by the first neural network is reduced.
[0019] In addition, according to the present invention, a learning device is provided, the learning device comprising a first neural network, a second neural network, a learning auxiliary neural network and a loss output unit, characterized in that:
[0020] The first neural network extracts features of the image data from the image data,
[0021] The second neural network, whose network structure is smaller than that of the first neural network, extracts features of the image data from the image data.
[0022] The learning auxiliary neural network extracts features including a bias factor of the image data from features of the image data extracted by the first neural network,
[0023] The loss output unit compares the features extracted by the second neural network with the features including the bias factor extracted by the learning assistance neural network and outputs the loss.
[0024] Furthermore, according to the present invention, a learning device is provided, the learning device comprising a processing mechanism, wherein the processing mechanism comprises:
[0025] A first neural network extracts features of objects in the image data and classifies the objects;
[0026] A learning auxiliary neural network that learns to extract the biased feature from features that are originally to be focused on for classifying the object and features that are different from the features that are originally to be focused on, included in the features extracted by the first neural network; and
[0027] a second neural network that extracts biased features of the object within the image data,
[0028] The processing mechanism causes the learning auxiliary neural network to learn so that the difference between the biased feature extracted by the learning auxiliary neural network and the biased feature extracted by the second neural network becomes smaller, and the processing mechanism causes the first neural network to learn so that a feature that increases the difference is extracted from the image data as a result extracted by the learning auxiliary neural network.
[0029] In addition, according to the present invention, there is provided a learning method, which is executed in a learning device including a processing mechanism, and is characterized in that:
[0030] The processing mechanism includes: a first neural network that extracts a first feature of an object in image data; a second neural network that uses a network structure different from that of the first neural network to extract a second feature of the object in the image data; and a learning auxiliary neural network that extracts a third feature from the first feature extracted by the first neural network, wherein the second feature and the third feature are features that are biased relative to the object.
[0031] The learning method has a processing step, in which the learning auxiliary neural network is made to learn by the processing mechanism so that the second feature extracted by the second neural network is close to the third feature extracted by the learning auxiliary neural network, and the first neural network is made to learn by the processing mechanism so that the third feature appearing in the first feature extracted by the first neural network is reduced.
[0032] Effects of the Invention
[0033] According to the present invention, in object recognition, robust features can be extracted adaptively with respect to the domain. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a block diagram showing an example of the functional configuration of the information processing server according to the first embodiment.
[0035] Figure 2 This is a diagram for explaining the problem of feature extraction including biased features (features of bias factors) in object recognition processing.
[0036] Figure 3A This is a diagram illustrating a configuration example of a deep neural network (DNN) of a model processing unit according to the first embodiment during a learning phase.
[0037] Figure 3B This is a diagram illustrating a configuration example of a deep neural network in the inference stage of a model processing unit according to the first embodiment.
[0038] Figure 3C This is a diagram showing an example of the output of the model processing unit involved in the first embodiment.
[0039] Figure 4 This is a diagram showing an example of learning data according to the first embodiment.
[0040] Figure 5A , Figure 5B This is a flowchart showing a series of operations of the processing in the learning phase in the model processing unit according to the first embodiment.
[0041] Figure 6 This is a flowchart showing a series of operations of the processing in the inference phase in the model processing unit according to the first embodiment.
[0042] Figure 7 This is a block diagram showing an example of a functional configuration of a vehicle according to the second embodiment.
[0043] Figure 8 This is a diagram showing a main configuration for vehicle travel control according to the second embodiment.
[0044] Description of Reference Numerals
[0045] 100: information processing server; 113: image data acquisition unit; 114: model acquisition unit; 310: DNN_R; 311: DNN_E; 312: DNN_B; 313: difference calculation unit. DETAILED DESCRIPTION
[0046] Hereinafter, the embodiments are described in detail with reference to the accompanying drawings. It should be noted that the following embodiments do not limit the invention involved in the technical solution. In addition, all combinations of features described in the embodiments are not limited to the content necessary for the invention. Two or more of the multiple features described in the embodiments may also be combined arbitrarily. In addition, the same or similar configurations are marked with the same reference numerals, and repeated descriptions are omitted.
[0047] (Implementation Method 1)
[0048] <Configuration of Information Processing Server>
[0049] Next, refer to Figure 1 The functional configuration example of the information processing server is described. It should be noted that the various functional blocks described with reference to the following figures can be combined or separated, and the functions described can also be implemented by other blocks. In addition, the components described as hardware can be implemented by software, and vice versa.
[0050] The control unit 104 includes, for example, a CPU 110, a RAM 111, and a ROM 112, and controls the operation of each part of the information processing server 100. The control unit 104 performs the functions of each part of the control unit 104 by having the CPU 110 load and execute a computer program stored in the ROM 112 or the storage unit 103 in the RAM 111. In addition to the CPU 110, the control unit 104 may include a GPU or dedicated hardware suitable for executing machine learning processing or neural network processing.
[0051] The image data acquisition unit 113 acquires image data transmitted from an external device such as an information processing device operated by a user or a vehicle. The image data acquisition unit 113 stores the acquired image data in the storage unit 103. The image data acquired by the image data acquisition unit 113 can be used as learning data described later, or can be input into a learned model at the inference stage in order to obtain an inference result based on new image data.
[0052] The model processing unit 114 includes the learning model involved in this embodiment, and performs the processing of the learning phase and the processing of the inference phase of the learning model. The learning model performs the operation of the deep learning algorithm using the deep neural network (DNN) described later, and performs the processing of recognizing the object included in the image data. The object may include pedestrians, vehicles, two-wheeled vehicles, signs, logos, roads, lines drawn in white or yellow on the road, etc.
[0053] The DNN becomes a learned state by performing the processing of the learning stage described later. By inputting new image data into the learned DNN, it is possible to identify the object target for the new image data (processing of the inference stage). When the inference processing using the learned model is performed in the information processing server 100, the processing of the inference stage is performed. It should be noted that the information processing server 100 can execute the learned learned model on the information processing server 100 side and send the inference result to an external device such as a vehicle and an information processing device, or it can perform the inference stage processing based on the learning model in the vehicle or information processing device as needed. When the inference stage processing based on the learning model is performed in the vehicle or information processing device, the model providing unit 115 provides information of the learned model to an external device such as a vehicle or an information processing device.
[0054] When the inference processing using the learned model is performed in the vehicle or the information processing device, the model providing unit 115 sends the information of the learned model learned in the information processing server 100 to the vehicle or the information processing device. For example, when the vehicle receives the information of the learned model from the information processing server 100, the learned model in the vehicle is updated to the latest learned model, and the latest learned model is used to perform the object recognition processing (inference processing). The information of the learned model includes the version information of the learned model, the information of the weight coefficient of the learned neural network, etc.
[0055] It should be noted that the information processing server 100 is generally able to use more abundant computing resources than vehicles and the like. In addition, by receiving and accumulating image data captured by various vehicles, it is possible to collect learning data in a variety of versatile situations, and to perform learning corresponding to more situations. Therefore, if the learned model learned using the learning data collected on the information processing server 100 can be provided to the vehicle or an external information processing device, the inference results for the images in the vehicle and the information processing device become more robust.
[0056] The learning data generation unit 116 generates learning data using the image data stored in the storage unit 103 based on access from an external predetermined information processing device operated by a user who manages the learning data. For example, the learning data generation unit 116 receives information on the type and position of the object included in the image data stored in the storage unit 103 (i.e., a label indicating the correct answer of the object to be recognized), and stores the received label in the storage unit 103 in association with the image data. The label associated with the image data is stored in the storage unit 103 as learning data in the form of a table, for example. This will be referred to later. Figure 4 Describe the details of the learning data.
[0057] The communication unit 101 is, for example, a communication device including a communication circuit, and communicates with external devices such as vehicles and information processing devices through a network such as the Internet. In addition to receiving actual images sent from external devices such as vehicles and information processing devices, the communication unit 101 also sends information about a learning model that has been learned at a predetermined time or cycle to the vehicle. The power supply unit 102 supplies power to each part in the information processing server 100. The storage unit 103 is a non-volatile memory such as a hard disk and a semiconductor memory. The storage unit 103 stores learning data described later, programs executed by the CPU 110, other data, etc.
[0058] <Example of Learning Model in Model Processing Unit>
[0059] Next, an example of a learning model in the model processing unit 114 according to the present embodiment will be described. Figure 2 , the problem of feature extraction including the bias factor in object recognition processing is explained. Figure 2 In the example, the color becomes a bias factor when the feature that should be focused on in the object recognition process is the shape. For example, Figure 2 The DNN shown is a DNN that infers whether the object in the image data is a truck or a passenger car, and is learned using image data of a black truck and image data of a red passenger car. That is, in addition to the features of the shape that should be focused on, the DNN also considers the features of the color that are different from the features that should be focused on (biased features) for learning. In such a DNN, when the image data of a black truck and the image data of a red passenger car are input in the inference stage, the correct inference result (truck or passenger car) can be output. Regarding such inference results, there are cases where the correct inference results are output according to the features that should be focused on, and there are also cases where the inference results are output according to the features of the color that are different from the features that should be focused on.
[0060] In the case where the DNN outputs an inference result based on the color feature, if the image data of a red truck is input to the DNN, the inference result is a passenger car, and if the image data of a black passenger car is input to the DNN, the inference result is a truck. In addition, when an image of a vehicle of an unknown color that is neither black nor red is input, it is unclear what classification result will be obtained.
[0061] On the other hand, when the DNN outputs an inference result based on the shape feature, if the image data of a red truck is input to the DNN, the inference result is a truck, and if the image data of a black passenger car is input to the DNN, the inference result is a passenger car. In addition, when an image of a truck of an unknown color that is neither black nor red is input, the inference result is a truck. In this way, when the DNN is learned with biased features, it is impossible to output a correct inference result when performing inference processing on new image data (that is, it is impossible to extract robust features).
[0062] In order to reduce the influence of such biased features and learn the features that should be focused on, in this embodiment, the model processing unit 114 is composed of Figure 3A Specifically, the model processing unit 114 includes DNN_R310, DNN_E311, DNN_B312 and a difference calculation unit 313.
[0063] DNN_R310 is a deep neural network (DNN) composed of one or more deep neural networks (DNNs) that extracts features from image data and outputs inference results of objects contained in the image data. Figure 3A In the example shown, DNN_R310 has two DNNs inside, namely DNN321 and DNN322. DNN321 is a DNN of an encoder that encodes the features of image data, and outputs features extracted from the image data (for example, set to z). The feature z includes the feature f that should be focused on and the biased feature b. DNN322 is a classifier that classifies objects based on the feature z extracted from the image data (ultimately becoming z→f through learning).
[0064] DNN_R310 example output Figure 3C The data of the inference result shown as an example in FIG. Figure 3C The data of the inference result shown in the output image include the presence or absence of an object mark (for example, if an object mark exists, it is set to 1, and if an object mark does not exist, it is set to 0), the center position and size of the object mark area. In addition, the probability of each object mark category is included. For example, the probability that the recognized object mark is a truck, a passenger car, a forklift, etc. is output in the range of 0 to 1.
[0065] It should be noted that Figure 3C The example of the data shown shows a case where one object mark is detected for the image data, but the data may include the probability of the object category according to the presence or absence of the object mark for each predetermined area.
[0066] In addition, DNN_R310 can, for example, Figure 4 The data and image data shown are used as learning data to carry out the processing in the learning phase. Figure 4 The data shown includes, for example, an identifier for identifying image data and a corresponding label. The label indicates the correct answer for the object included in the image data indicated by the image ID. The label indicates, for example, the category of the object included in the corresponding image data (e.g., truck, passenger car, forklift, etc.). In addition, the learning data may include data on the center position and size of the object. DNN_R310 is learned as when the image data of the learning data is input and the image data of the learning data is output Figure 3C When the inference result data shown is compared with the labels of the learning data, the error of the inference result is minimized. However, the learning of DNN_R310 is constrained to maximize the loss function of the feature described later.
[0067] DNN_E311 is a DNN that extracts the biased feature b from the feature z (z = feature f that should originally be focused on + biased feature b) output by DNN_R310. DNN_E311 functions as a learning auxiliary neural network that assists the learning of DNN_R310. DNN_E311 is learned to be able to extract the biased feature b with higher accuracy by learning in an adversarial manner with DNN_R310 during the learning phase. On the other hand, DNN_R310 is able to remove the biased feature b and extract the feature f that should originally be focused on with higher accuracy by learning in an adversarial manner with DNN_E311. That is, the feature z output from DNN_R310 is infinitely close to f.
[0068] DNN_E311 has a well-known GRL (Gradient reversal layer) inside that can perform adversarial learning. GRL is a layer that reverses the sign of the gradient for DNN_E311 when the weight coefficients of DNN_E311 and DNN_R310 are changed based on back propagation. Therefore, in adversarial learning, the gradient of the weight coefficient of DNN_E311 is changed in association with the gradient of the weight coefficient of DNN_R310, so that both neural networks can learn simultaneously.
[0069] DNN_B312 is a DNN that inputs image data and infers classification results based on biased features. DNN_B312 is learned to perform the same inference task as DNN_R310 (e.g., classification of object targets). That is, DNN_B312 is learned to minimize the same target loss function as that used by DNN_R310 (e.g., a loss function that minimizes the difference between the inference result of the object target and the learning data).
[0070] However, it is learned to extract biased features inside DNN_B312 and output the best classification result based on the extracted features. In this embodiment, image data is input to DNN_B312 in a learned state, and DNN_B312 extracts the biased features b' extracted inside.
[0071] DNN_B312 has completed its learning before DNN_R310 and DNN_E311 are made to learn. Therefore, DNN_B312 performs the following functions: during the learning process of DNN_R310 and DNN_E311, the correct bias factor (biased feature b') contained in the image data is extracted and provided to DNN_E311. DNN_B312 has a network structure different from that of DNN_R310, and is configured to extract features different from those extracted by DNN_R310. For example, DNN_B312 is configured as a neural network having a network structure smaller in scale (fewer parameters and lower complexity) than that of the neural network possessed by DNN_R, and extracts surface features (bias factors) of image data. The structure of DNN_E311 can also be set to a structure that processes image data with a lower resolution than DNN_R310, or to a structure with fewer layers than DNN_R310. In DNN_E311, for example, the main color in the image is extracted as a biased feature. Alternatively, in order to extract the texture features within the image as biased features, the kernel size may be smaller than that of DNN_R310, so that DNN_B312 is constructed in a manner to extract local features of the image data.
[0072] It should be noted that although Figure 3A Although not explicitly stated in the example, DNN_B312 may have two DNNs internally, similar to the example of DNN_R310. For example, it may include a DNN of an encoder that extracts a biased feature b', and a DNN of a classifier that infers a classification result based on the biased feature b'. In this case, the encoder DNN of DNN_B312 is configured to extract different features (from the encoder DNN of DNN_R310) from the image data through a network structure different from that of the encoder DNN of DNN_R310.
[0073] The difference calculation unit 313 compares the biased feature b' output from DNN_B 312 with the biased feature b output from DNN_E 311 to calculate the difference. The difference calculated by the difference calculation unit 313 is used to calculate the loss function of the feature.
[0074] In this embodiment, DNN_E311 is made to learn so as to minimize the loss function of the feature based on the difference of the difference calculation unit 313. Therefore, DNN_E311 is made to learn so as to make the biased feature b extracted by DNN_E311 close to the biased feature b' extracted by DNN_B312. That is, DNN_E311 is made to learn so as to extract the biased feature b with higher accuracy from the feature z extracted by DNN_R310.
[0075] On the other hand, DNN_R310 is advanced in learning so as to maximize the loss function of the feature based on the difference of the difference calculation unit 313 and minimize the target loss function of the inference task (e.g., classification of the object). In other words, in the present embodiment, explicit constraints are imposed in learning so that the feature z extracted by DNN_R310 maximizes the feature f that should be focused on and minimizes the bias factor b. In particular, in the learning method of the present embodiment, DNN_R310 and DNN_E311 are made to learn adversarially, so that DNN_R310 and DNN_E311 learn the parameters of DNN_R310 in the direction of extracting feature z such that it is difficult for DNN_E311 that extracts the biased feature b to extract the biased feature b (including DNN_E311).
[0076] In the present embodiment, such adversarial learning is illustrated by taking the case where the GRL included in DNN_E311 is used to simultaneously update DNN_R310 and DNN_E311 as an example, but DNN_R310 and DNN_E311 may also be updated alternately. For example, first, on the basis of fixing DNN_R310, DNN_E311 is updated to minimize the loss function of the features based on the differences of the difference calculation unit 313. Next, on the basis of fixing DNN_E311, DNN_R310 is updated to maximize the loss function of the features based on the differences of the difference calculation unit 313 and minimize the target loss function of the inference task (such as the classification of the object). Through such learning, DNN_R310 can extract the feature f that should be paid attention to with high precision and can extract robust features.
[0077] When the learning phase of DNN_R310 is completed through the above-mentioned adversarial learning, DNN_R310 becomes a learning-completed model and can be used in the inference phase. Figure 3BAs shown, the image data is only input to DNN_R 310, and DNN_R 310 only outputs the inference result (object classification result). That is, DNN_E 311, DNN_B 312 and difference calculation unit 313 do not operate in the inference stage.
[0078]
[0079] Next, refer to Figure 5A as well as Figure 5B , a series of actions in the learning phase in the model processing unit 114 are described. It should be noted that this process is implemented by the CPU 110 of the control unit 104 loading and executing the program stored in the ROM 112 or the storage unit 103 in the RAM 111. It should be noted that the various DNNs of the model processing unit 114 of the control unit 104 are not completely learned, but become a state of complete learning through this process.
[0080] In S501, the control unit 104 causes the DNN_B312 of the model processing unit 114 to learn. DNN_B312 can use the same learning data as the learning data used to learn DNN_R310 for learning. Image data of the learning data is input to DNN_B312, and a classification result is calculated from DNN_B312. As described above, DNN_B312 is learned to minimize the loss function obtained based on the difference between the classification result and the label of the learning data. As a result, DNN_B312 is learned to extract features that are biased internally. Although the description is simplified in this flowchart, in the learning of DNN_B312, iterative processing corresponding to the number of learning data and the number of generations is also performed.
[0081] In S502, the control unit 104 reads the image data associated with the learning data from the storage unit 103. Here, the learning data includes the reference Figure 4 Describing data.
[0082] In S503 , the model processing unit 114 applies the weight coefficient of the current neural network to the read image data, and outputs the extracted feature z and the inference result.
[0083] In S504, the model processing unit 114 inputs the feature z extracted in DNN_R310 to DNN_E311 to extract the biased feature b. Furthermore, in S505, the model processing unit 114 inputs the image data to DNN_B312 to extract the biased feature b' from the image data.
[0084] In S506, the model processing unit 114 calculates the difference (absolute value of the difference) between the biased feature b and the biased feature b' through the difference calculation unit 313. In S507, the model processing unit 114 calculates the above-mentioned target loss function (L f In S508, the model processing unit 114 calculates the above-mentioned feature loss function (L b ) loss.
[0085] In S509, the model processing unit 114 determines whether the above-mentioned processing of S502 to S508 has been executed for all the predetermined learning data. If the model processing unit 114 determines that the processing has been executed for all the predetermined learning data, the processing proceeds to S510. If not, the processing returns to S502 in order to execute the processing of S502 to S508 using further learning data.
[0086] In S510, the model processing unit 114 changes the weight coefficient of DNN_E311 so that the characteristic loss function (L b ) is reduced (i.e., the biased feature b is extracted more accurately from the feature z extracted by DNN_R310). On the other hand, in S511, the model processing unit 114 changes the weight coefficient of DNN_R so that the feature loss function (L b ) increases the sum of the losses and makes the objective loss function (L f ) is reduced. That is, the model processing unit 114 is learned to maximize the feature f that should be focused on while making the feature z extracted by DNN_R310 minimize the bias factor b.
[0087] In S512, the model processing unit 114 determines whether the processing of the predetermined number of generations has been completed. In other words, it determines whether the processing of S502 to S511 has been repeated a predetermined number of times. By repeating the processing of S502 to S511, changes are made so that the weight coefficients of DNN_R310 and DNN_E311 gradually converge to the optimal value. If the model processing unit 114 determines that the predetermined number of generations has not been completed, the processing returns to S502. If not, the series of processing is terminated. In this way, when a series of actions in the learning stage of the model processing unit 114 are completed, each DNN in the model processing unit 114 (especially DNN_R310) becomes a state of completed learning.
[0088]
[0089] Next, refer to Figure 6, a series of actions in the inference stage in the model processing unit 114 are explained. This process is a process for outputting the classification result of the object target for the image data actually captured by the vehicle or the information processing device (that is, unknown image data without a correct answer). It should be noted that this process is implemented by the CPU 110 of the control unit 104 loading and executing the program stored in the ROM 112 or the storage unit 103 in the RAM 111. In addition, this process is a state in which the DNN_R310 of the model processing unit 114 has been pre-learned. That is, the weight coefficient is determined in such a way that the DNN_R310 maximizes the detection of the feature f that should be paid attention to.
[0090] In S601, the control unit 104 inputs the image data acquired from the vehicle or the information processing device to the DNN_R310. In S602, the model processing unit 114 performs object recognition processing based on the DNN_R310 and outputs the inference result. When the inference processing is completed, the control unit 104 ends the series of actions involved in this processing.
[0091] As described above, in this embodiment, the information processing server includes: DNN_R, which extracts the features of the object in the image data; DNN_B, which uses a network structure different from DNN_R to extract the features of the object in the image data; and DNN_E, which extracts the biased features from the features extracted in DNN_R. Then, DNN_E311 is made to learn so that the biased features extracted in DNN_B312 are close to the biased features extracted in DNN_E311, and DNN_R310 is made to learn so that the biased features appearing in the features extracted by DNN_R310 are reduced. As a result, in object recognition, robust features can be extracted adaptively with respect to the domain.
[0092] (Implementation Method 2)
[0093] Next, embodiment 2 involved in the present invention is described. In the above-mentioned embodiment, the case where the processing of the learning phase and the processing of the inference phase of the neural network are executed in the information processing server 100 is described as an example. However, this embodiment is not limited to the case where the processing of the learning phase is executed in the information processing server, and can also be applied to the case where the processing of the learning phase is executed in the vehicle. That is, the learning data provided by the information processing server 100 can also be input into the model processing unit of the vehicle to make the neural network learn in the vehicle. In addition, the processing of the inference phase can also be executed using the learned neural network. Below, an example of the functional configuration of a vehicle in such an embodiment is described.
[0094] In the following example, the control unit 708 is described as a control mechanism assembled in the vehicle 700, but an information processing device having the structure of the control unit 708 may be mounted in the vehicle 700. That is, the vehicle 700 may be a vehicle equipped with an information processing device having the structure of the CPU 710 included in the control unit 708, the model processing unit 714, and the like.
[0095] <Vehicle Configuration>
[0096] First, refer to Figure 7 , an example of the functional configuration of the vehicle 700 involved in this embodiment is described. It should be noted that the various functional blocks described with reference to the following figures can be combined or separated, and the functions described can also be implemented by other blocks. In addition, the components described as hardware can be implemented by software, and vice versa.
[0097] The sensor unit 701 includes a camera (photographing mechanism) that outputs a photographed image obtained by photographing the front (or, further, the rear direction, the surroundings) of the vehicle. The sensor unit 701 may also include a Lidar (Light Detection and Ranging) that outputs a distance image obtained by measuring the distance in front (or, further, the rear direction, the surroundings) of the vehicle. The photographed image is used, for example, for inference processing of object identification in the model processing unit 714. In addition, various sensors that output acceleration, position information, steering angle, etc. of the vehicle 700 may also be included.
[0098] The communication unit 702 is, for example, a communication device including a communication circuit, and communicates with the information processing server 100, a surrounding transportation system, etc., via, for example, LTE, LTE-Advanced, etc., or mobile communication standardized as so-called 5G. The communication unit 702 acquires learning data from the information processing server 100. In addition, the communication unit 702 receives part or all of map data, traffic information, etc. from other information processing servers and surrounding transportation systems.
[0099] The operation unit 703 includes, in addition to operation components such as buttons and a touch panel installed in the vehicle 700, components such as a steering wheel and a brake pedal that receive input for driving the vehicle 700. The power supply unit 704 includes a battery such as a lithium-ion battery, and supplies power to various components in the vehicle 700. The power unit 705 includes, for example, an engine or a motor that generates power for driving the vehicle.
[0100] The driving control unit 706 controls the driving of the vehicle 700, for example, by maintaining driving in the same lane or following the driving of the vehicle in front, based on the result of the inference processing output from the model processing unit 714 (e.g., the result of object recognition). It should be noted that in the present embodiment, the driving control can be performed using a known method. It should be noted that in the description of the present embodiment, the driving control unit 706 is illustrated as a different configuration from the control unit 708, but it can also be included in the control unit 708.
[0101] The storage unit 707 includes a nonvolatile large-capacity storage device such as a semiconductor memory. The actual image output from the sensor unit 701 and various sensor data output from the sensor unit 701 are temporarily stored. In addition, the learning data acquisition unit 713 described later stores learning data received from the external information processing server 100 via the communication unit 702 for learning by the model processing unit 714.
[0102] The control unit 708 includes, for example, a CPU 710, a RAM 711, and a ROM 712, and controls the operation of each part of the vehicle 700. In addition, the control unit 708 acquires image data from the sensor unit 701 and performs the above-mentioned inference processing including object recognition processing, and also uses the image data received from the information processing server 100 to perform the processing of the learning phase of the model processing unit 714. The control unit 708 loads the computer program stored in the ROM 712 into the RAM 711 through the CPU 710 and executes it, so that the functions of each part of the control unit 708, such as the model processing unit 714, are exerted.
[0103] The CPU 710 includes one or more processors. The RAM 711 is composed of a volatile storage medium such as a DRAM, and functions as a working memory of the CPU 710. The ROM 712 is composed of a non-volatile storage medium, and stores computer programs executed by the CPU 710, setting values for operating the control unit 708, and the like. It should be noted that in the following embodiments, the case where the CPU 710 executes the processing of the model processing unit 714 is used as an example for explanation, but the processing of the model processing unit 714 may also be executed by one or more other processors (e.g., a GPU) not shown in the figure.
[0104] The learning data acquisition unit 713 acquires image data and Figure 4 The data shown is regarded as learning data and is stored in the storage unit 707. The learning data is used when the model processing unit 714 is made to learn in the learning phase.
[0105] The model processing unit 714 has the same structure as in the first embodiment. Figure 3AThe model processing unit 714 performs the processing of the learning phase and the processing of the inference phase using the learning data acquired by the learning data acquisition unit 713. The processing of the learning phase and the processing of the inference phase performed by the model processing unit 714 can be performed in the same manner as the processing shown in Implementation Example 1.
[0106] <Main components for vehicle driving control>
[0107] Next, refer to Figure 8 , the main components for driving control of the vehicle 700 are described. The sensor unit 701, for example, captures the front of the vehicle 700 and outputs the captured image data at a predetermined number of frames per second. The image data output from the sensor unit 701 is input to the model processing unit 714 of the control unit 708. The image data input to the model processing unit 714 is used for object recognition processing (processing in the inference stage) for controlling the driving of the vehicle at the current time point.
[0108] The model processing unit 714 receives the image data output from the sensor unit 701 to perform object recognition processing, and outputs the classification result to the driving control unit 706. The classification result can be the same as that in the first embodiment. Figure 3C The output shown is the same.
[0109] The driving control unit 706 outputs a control signal to the power unit 705, for example, to control the vehicle 700 based on the result of the object recognition and various sensor information such as the acceleration and steering angle of the vehicle obtained from the sensor unit 701. As described above, the vehicle control performed by the driving control unit 706 can be performed using a known method, so a detailed description is omitted in this embodiment. The power unit 705 controls the generation of power according to the control signal of the driving control unit 706.
[0110] The learning data acquisition unit 713 acquires the learning data, namely, the image data and Figure 4 The data shown. The acquired data is used to enable the DNN of the model processing unit 714 to learn.
[0111] The vehicle 700 can use the learning data stored in the storage unit 707 and Figure 5A , Figure 5B The process shown in FIG. 1 also performs a series of processes in the learning phase. Figure 6 The illustrated processing similarly performs a series of processing in the inference stage.
[0112] As described above, in the present embodiment, deep neural network learning for object identification is used in the model processing unit 714 in the vehicle 700. That is, the vehicle has: DNN_R, which extracts features of object targets in image data; DNN_B, which uses a network structure different from that of DNN_R to extract features of object targets in image data; and DNN_E, which extracts biased features from features extracted in DNN_R. Then, DNN_E311 is made to learn so that the biased features extracted in DNN_B312 are close to the biased features extracted in DNN_E311, and DNN_R310 is made to learn so that the biased features appearing in the features extracted by DNN_R310 are reduced. Thus, in object identification, robust features can be extracted adaptively with respect to the domain.
[0113] It should be noted that in the above-mentioned embodiment, the information processing server as an example of the learning device and the vehicle as an example of the learning device are executed. Figure 3A However, the learning device is not limited to the information processing server and the vehicle, and can also be executed by other devices. Figure 3A Processing of the DNN shown.
[0114] The present invention is not limited to the above-described embodiments, and various modifications and changes can be made within the scope of the gist of the invention.
Claims
1. A learning device, comprising a processing mechanism, characterized in that: The processing mechanism comprises: A first neural network extracts a first feature of an object in the image data; A second neural network, which uses a network structure different from that of the first neural network to extract a second feature of the object in the image data; and learning an auxiliary neural network that extracts a third feature from the first feature extracted by the first neural network, The second feature and the third feature are features offset relative to the object, The processing mechanism causes the learning auxiliary neural network to learn so that the second feature extracted by the second neural network is close to the third feature extracted by the learning auxiliary neural network, and the processing mechanism causes the first neural network to learn so that the third feature appearing in the first feature extracted by the first neural network is reduced.
2. The learning device according to claim 1, characterized in that The scale of the network structure of the second neural network is smaller than that of the network structure of the first neural network.
3. The learning device according to claim 1, characterized in that: The first neural network and the second neural network have kernels for extracting local features of an image, The size of the kernel of the second neural network is smaller than the size of the kernel of the first neural network.
4. The learning device according to claim 1, characterized in that: The first neural network is a neural network that classifies the object by extracting the first feature of the object in the image data. The second neural network is a neural network that classifies the object by extracting the second feature of the object in the image data.
5. The learning device according to claim 4, characterized in that: When the first neural network is made to learn so that the third feature appearing in the first features extracted by the first neural network is reduced, the processing unit makes the first neural network learn so that the difference between the classification result for the object and the learning data becomes smaller.
6. The learning device according to claim 1, characterized in that The processing mechanism uses a gradient reversal layer to vary the weight coefficients of the first neural network in association with the weight coefficients of the learning assistance neural network.
7. The learning device according to claim 1, characterized in that: The second neural network is a neural network that has been previously learned so as to extract the second feature of the object in the image data.
8. The learning device according to any one of claims 1 to 7, characterized in that: The learning device is an information processing server.
9. The learning device according to any one of claims 1 to 7, characterized in that: The learning device is a vehicle.
10. A learning device, comprising a first neural network, a second neural network, a learning auxiliary neural network and a loss output unit, characterized in that: The first neural network extracts features of the image data from the image data, The second neural network, whose network structure is smaller than that of the first neural network, extracts features of the image data from the image data. The learning auxiliary neural network extracts features including a bias factor of the image data from features of the image data extracted by the first neural network, The loss output unit compares the features extracted by the second neural network with the features including the bias factor extracted by the learning auxiliary neural network and outputs the loss. The first neural network is caused to learn so that features extracted by the learning-assisted neural network including a bias factor of the image data are reduced.
11. A learning device, comprising a processing mechanism, characterized in that: The processing mechanism comprises: A first neural network extracts features of objects in the image data and classifies the objects; A learning auxiliary neural network that performs learning so as to extract the biased feature from features that are originally to be focused on for classifying the object and features that are different from the features that are originally to be focused on, included in the features extracted by the first neural network; as well as a second neural network that extracts biased features of the object within the image data, The processing mechanism causes the learning auxiliary neural network to learn so that the difference between the biased feature extracted by the learning auxiliary neural network and the biased feature extracted by the second neural network becomes smaller, and the processing mechanism causes the first neural network to learn so that a feature that increases the difference is extracted from the image data as a result of extraction by the learning auxiliary neural network.
12. A learning method, the learning method being executed in a learning device comprising a processing mechanism, characterized in that: The processing mechanism includes: a first neural network that extracts a first feature of an object in image data; a second neural network that uses a network structure different from that of the first neural network to extract a second feature of the object in the image data; and a learning auxiliary neural network that extracts a third feature from the first feature extracted by the first neural network, wherein the second feature and the third feature are biased features relative to the object. The learning method has a processing step, in which the learning auxiliary neural network is made to learn by the processing mechanism so that the second feature extracted by the second neural network is close to the third feature extracted by the learning auxiliary neural network, and the first neural network is made to learn by the processing mechanism so that the third feature appearing in the first feature extracted by the first neural network is reduced.
Citation Information
Patent Citations
Data processing system and data processing method
CN111630530A
Learning device, identification device, and program
CN111771216A