Image processing device, image processing system, and image processing method

The image processing device and method address the instability in generating additional learning images by using feature amount similarities to select images, thereby improving accuracy and optimizing image selection for autonomous driving systems.

WO2025094442A1PCT designated stage expired Publication Date: 2025-05-08ASTEMO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/021209
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-30
Filing Date
2024-06-11
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing image processing technologies for autonomous driving rely heavily on developer experience and know-how to generate additional learning images, leading to unstable accuracy and suboptimal selection of images for relearning.

Method used

An image processing device and method that utilize an inference unit to calculate feature amounts from both target and candidate images, and a learning image generation unit to select learning images based on similarity between these feature amounts, thereby generating additional learning images without relying on developer experience.

Benefits of technology

This approach improves the accuracy and optimizes the selection of additional learning images, enhancing the learning accuracy of image recognition models used in autonomous driving systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024021209_08052025_PF_FP_ABST
    Figure JP2024021209_08052025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an image processing device which uses AI technology, is capable of generating an additional training image independent of experience and know-how, and can improve the accuracy of the additional training image. This invention comprises: an inference unit that, on the basis of an intermediate layer output of an image recognition model obtained by inference using a target image and a candidate image group as input to the image recognition model, derives a feature amount related to the target image and a feature amount related to a candidate image included in the candidate image group, and outputs the feature amounts; and a training image generation unit that generates, in accordance with a feature amount similarity, a training image to be used for learning the target image, the feature amount similarity being the degree of similarity between the feature amount related to the target image and the feature amount related to the candidate image output from the inference unit.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing device, image processing system, and image processing method

[0001] The present invention relates to an image processing device, an image processing system, and an image processing method.

[0002] Image processing technology is known for recognizing the external environment using a camera mounted on a vehicle and for use in autonomous driving, etc. The image processing technology described above requires accurate recognition of objects from the obtained images, and therefore AI technology is used, and there is a demand for improved re-learning technology for recognizing the external environment.

[0003] In AI-based external recognition technology for in-vehicle cameras, the method of collecting and generating additional training images (padded images) to improve recognition errors during evaluation of a trained AI model is largely dependent on the experience and know-how of the developer. As a result, the accuracy of the additional training images is unstable, and improvements in accuracy are desired.

[0004] In other words, a technology is needed that can collect and generate additional training images without relying on the experience or know-how of the developer, and that can improve accuracy.

[0005] Patent Document 1 describes a technology that generates a derived image from a training candidate image, calculates the similarity between the training image and the derived image (using the G component of an RGB image), compares the calculated similarity with a judgment threshold, and stores training candidate images whose similarity is smaller than the judgment threshold as training images.

[0006] Patent document 2 describes a technology that matches a template image area with a candidate image to be added to training, sets a threshold for the matching results, selects all candidate images to be added to training whose similarity to the template image area exceeds the set threshold, and generates training images.

[0007] Patent Document 1: WO2017 / 109854 Patent Document 2: JP 2022-182149 A

[0008] The similarity between the training image and the derived image described in Patent Document 1 is different from the similarity in AI technology, and it is difficult to apply the technology described in Patent Document 1 to a device that selects training candidate images using AI technology, and there is a possibility that sufficient improvement in accuracy will not be obtained.

[0009] Furthermore, since images are selected based solely on a similarity threshold, the number of selected images may be large, which is not the optimal number of images for re-learning, making it difficult to improve accuracy.

[0010] Furthermore, the similarity described in Patent Document 2 is different from the similarity in AI technology, and like the technology described in Patent Document 1, it is difficult to apply the technology described in Patent Document 2 to a device that selects learning candidate images using AI technology, and there is a possibility that sufficient improvement in accuracy will not be obtained.

[0011] Furthermore, since images are selected based solely on a similarity threshold, the number of selected images may be large, which is not the optimal number of images for re-learning, making it difficult to improve accuracy.

[0012] An object of the present invention is to provide an image processing device and an image processing method using AI technology that can generate additional training images without relying on experience or know-how and can improve the accuracy of the additional training images.

[0013] In order to achieve the above object, the present invention is configured as follows.

[0014] The image processing device includes an inference unit that calculates feature quantities related to the target image and feature quantities related to candidate images included in the group of candidate images based on an intermediate layer output of an image recognition model obtained by inference using a target image and a group of candidate images as inputs to the image recognition model, and outputs the feature quantities; and a training image generation unit that generates training images to be used for training the target image in accordance with the feature similarity, which is the degree of similarity between the feature quantities related to the target image and the feature quantities related to the candidate images output from the inference unit.

[0015] According to the present invention, it is possible to provide an image processing device and an image processing method using AI technology that can generate additional training images without relying on experience or know-how and can improve the accuracy of the additional training images.

[0016] FIG. 1 is a block diagram showing the functional configuration of an image processing device according to Example 1. FIG. 2 is a block diagram showing the functional configuration of an image processing device according to Example 2. FIG. 3 is a diagram showing the characteristics of the number of images required to be added depending on the similarity with an image to be subjected to accuracy improvement. FIG. 4 is a block diagram showing the functional configuration of an image processing device according to Example 3. FIG. 5 is a diagram explaining a method for determining the number of images to be selected depending on the similarity and reliability. FIG. 6 is a block diagram showing the functional configuration of an image processing device according to a first modified example. FIG. 7 is a block diagram showing the functional configuration of an image processing device according to a second modified example. FIG. 8 is a system configuration diagram showing the second modified example. FIG. 9 is a hardware configuration diagram of an image processing system including an image processing device and a server according to one embodiment of the present invention.

[0017] Hereinafter, embodiments of the present invention will be described with reference to the drawings. In addition, in the examples other than Example 1 of the present invention, the same parts as those in Example 1 are designated by the same reference numerals, and detailed description thereof will be omitted.

[0018] First Embodiment FIG. 1 is a block diagram showing the functional configuration of an image processing apparatus 100 and a server 500 according to a first embodiment of the present invention.

[0019] The image processing device 100 shown in FIG. 1 includes an inference unit 1, a training image generation unit 2, and a memory unit 3. The training image generation unit 2 includes a similarity calculation unit 21 and a similar image selection unit 22. The memory unit 3 includes a training candidate image group storage unit 31, a similarity threshold storage unit 32, an accuracy improvement target image storage unit 33, and a training image storage unit 34. The training candidate image group storage unit 31, the similarity threshold storage unit 32, the accuracy improvement target image storage unit 33, and the training image storage unit 34 may be different memory areas of the same storage unit, or may be different storage units. For example, as described below, if the image processing device 100 is an on-board ECU (electronic control unit) mounted on a vehicle, the training candidate image group storage unit 31 and the similarity threshold storage unit 32 may be included in the on-board ECU, and the accuracy improvement target image storage unit 33 and the training image storage unit 34 may be included in a server connected to and communicating with the on-board ECU. The hardware configuration of the image processing device 100 will be described later.

[0020] 1 includes a learning unit 510 and a storage unit 520. The storage unit further includes a model storage unit 530. The server 500 is connected to an image processing device via a network.

[0021] The inference unit 1 executes a recognition process that uses an image recognition model to recognize a predetermined object or information from image data. The inference unit 1 receives as input one accuracy improvement target image stored in an accuracy improvement target image storage unit 31 (described later) and multiple learning candidate images stored in a learning candidate image group storage unit 33. The image recognition model 30 may be, for example, a classification AI model, but is not limited to this. The number of accuracy improvement target images input to the inference unit 1 is limited to one, but may be multiple.

[0022] The inference unit 1 receives an image to be improved and a candidate image for training as input, performs computation using an AI model, and calculates an intermediate layer output. The intermediate layer output is, for example, the final convolution layer of a classification AI model, but is not limited to this and may be any output value in the intermediate layer of the AI ​​model. This intermediate layer output is hereinafter referred to as a "feature amount." The inference unit 1 inputs the feature amount obtained by computation using the image to be improved as input and the feature amount obtained by computation using the candidate image for training as input to the training image generation unit 2.

[0023] The learning image generation unit 2 includes a similarity calculation unit 21 and a similar image selection unit 22. The similarity calculation unit 21 calculates the similarity (hereinafter referred to as feature similarity) between the feature corresponding to the image to be subjected to accuracy improvement and the feature corresponding to the learning candidate image. Here, an example of a method for calculating feature similarity is shown. The method for calculating feature similarity is not limited to the following example.

[0024] Assume that the feature amounts corresponding to the target image for accuracy improvement calculated by the inference unit 1 are A1, A2, A3, ... A64. Assume that the feature amounts of the candidate images for learning calculated by the inference unit 1 are B1, B2, B3, ... B64. In this case, the feature similarity is calculated by the following formula (1).

[0025] Feature similarity = (A1 - B1)^2 + (A2 - B2)^2 + (A3 - B3)^2 + ... + (A64 - B64)^2 ... (Equation 1) Next, the similar image selection unit 22 references a similarity threshold stored in advance in the similarity threshold storage unit 33 and selects training images from the multiple training candidate images based on the similarity threshold. The training image generation unit 2 selects all or some of the multiple training candidate images whose feature similarity exceeds the similarity threshold as training images. Here, one similarity threshold stored in the similarity threshold storage unit 33 is used, but multiple similarity thresholds may also be used. An embodiment using multiple similarity thresholds will be described in Example 2. Details of the similarity threshold will be described later.

[0026] The storage unit 3 includes a learning candidate image group storage unit 31, an accuracy improvement target image storage unit 32, a similarity threshold storage unit 33, and a learning image storage unit 34. The learning candidate image group storage unit 31 is a storage unit that stores learning candidate images, which are any group of images that are candidates for images used in training an AI model (image recognition model) for accuracy improvement target images. The learning candidate images may be images automatically generated using a known image generation technology, or may be images captured by a camera. For example, the learning candidate images may be multiple images captured by a camera mounted on a vehicle of a traffic scene or environment while the vehicle is traveling.

[0027] The accuracy improvement target image storage unit 32 is a storage unit that stores accuracy improvement target images, which are target images for improving the accuracy of recognition processing by an image recognition model. The accuracy improvement target image stored in the accuracy improvement target image storage unit 32 may be one or more. The accuracy improvement target image may be, for example, an image that has been recognized incorrectly during evaluation of an arbitrary trained image recognition model.

[0028] The similarity threshold storage unit 33 is a storage unit that stores a similarity threshold that serves as a criterion when the learning image generation unit 2 selects learning images from among the learning candidate images. An example of a method for determining the similarity threshold is shown below. Note that the step of determining the similarity threshold stored in the similarity threshold storage unit 33 may be executed in the image processing device 100, may be executed by a different computer having the same functional configuration as the image processing device 100, or may be partially executed in the image processing device 100.

[0029] First, the feature similarity, which is the degree of similarity between the feature corresponding to the image to be improved determined by the inference unit 1 and the feature corresponding to multiple candidate images for training, is determined, and the candidate images for training are sorted in descending order of feature similarity (i.e., in descending order of the value determined from Equation 1) (Step 1).

[0030] Next, an arbitrary similarity threshold is set as a provisional similarity threshold (step 2).

[0031] Based on the set provisional similarity threshold and the feature similarity, one or more learning images are selected from the plurality of learning candidate images (step 3).

[0032] The image recognition model is retrained based on the one or more selected candidate learning images and the trained images used to train the image recognition model, and a retrained image recognition model is generated (step 4).

[0033] Finally, the inference unit 1 inputs an evaluation image for evaluating the recognition accuracy of the retrained image recognition model and an image to be targeted for accuracy improvement into the retrained image recognition model, and performs inference (step 5).

[0034] If the inference unit 1 determines that the recognition accuracy of the target image for accuracy improvement and the evaluation image exceeds the threshold, the provisional similarity threshold is set as the formal similarity threshold. If the inference unit 1 determines that the recognition accuracy of the target image for accuracy improvement and the evaluation image falls below the threshold, the process returns to step 2, resets the provisional similarity threshold, and repeats steps 3 to 5. This process is repeated until the inference unit 1 determines that the recognition accuracy of the target image for accuracy improvement and the evaluation image exceeds the threshold (step 5).

[0035] Returning to FIG. 1, the learning image generation unit 2 stores the learning candidate images selected from the learning candidate images in the learning image storage unit 34 as learning images to be used for training the image recognition model.

[0036] The learning unit 510 reads out the learning images stored in the learning image storage unit 34. Then, the learning unit 510 re-learns the image recognition model stored in the model storage unit 530 using the trained images and the read-out learning images, and stores the re-learned image recognition model in the model storage unit.

[0037] The model storage unit 530 stores the image recognition model used for inference in the inference unit 1. The image processing device 100 stores the re-trained image recognition model re-trained by the learning unit 510 in the memory unit 3, and can use it for inference in the inference unit 1.

[0038] According to the first embodiment described above, it is possible to generate training images that can improve the training accuracy of an image recognition model without relying on experience or know-how. That is, by generating training images based on the similarity (feature similarity) of intermediate layer outputs obtained by executing arithmetic processing using an image recognition model, rather than the similarity of image features of the image itself between an accuracy improvement target image and a group of candidate training images, it is possible to generate training images that can improve the training accuracy of an image recognition model.

[0039] Second Embodiment Next, a second embodiment of the present invention will be described.

[0040] FIG. 2 is a functional block diagram of an image processing apparatus 101 according to the second embodiment.

[0041] The difference between the first embodiment shown in FIG. 1 and the second embodiment shown in FIG. 2 is that a generation number adjustment unit 4 is added.

[0042] 2 , the similarity calculation unit 21 calculates the similarity of the features (feature similarity) between the learning candidate image and the accuracy improvement target image based on the feature output from the inference unit 1. The similarity calculation unit 21 outputs the calculated feature similarity to the similar image selection unit 22 and also to the generation number adjustment unit 4. The similar image selection unit 22 selects learning candidate images whose feature similarity exceeds a similarity threshold, and outputs the selected images to the generation number adjustment unit 4.

[0043] The generation number adjustment unit 4 determines the number of images or the amount of data to be selected according to the feature similarity from the candidate learning images output from the learning image generation unit 2. Then, the generation number adjustment unit 4 stores the learning images in the learning image storage unit 34 according to the determined number of images or the amount of data.

[0044] FIG. 3 shows the characteristics of the number of images required for additional processing depending on the similarity to the image to be subjected to accuracy improvement, and explains a method for determining the number of images to be selected depending on the feature similarity.

[0045] FIG. 3 shows the relationship between feature similarity and the required number of images. When the target image for accuracy improvement and the candidate training images are 100% similar (i.e., identical), one additional image is sufficient for retraining the image recognition model. When the target image for accuracy improvement and the candidate training images are completely dissimilar (0% similar), theoretically, even if an infinite number of images are added, no retraining effect for the image recognition model is achieved. Here, the specific value of the required number of images (required data volume) is related to the number of trained images used in training the image recognition model. The larger the number of trained images, the higher the relationship between feature similarity and the required number of additional images shifts in parallel upward on the graph. For example, to prevent overtraining for a specific target image for accuracy improvement, the number of additional training images to be added is approximately 5% of the number of trained images used in training the image recognition model.

[0046] As shown in FIG. 3 , the generation number adjustment unit 4 sets the number of training images or the data amount of the training images so that the number of images required decreases as the average feature similarity (average feature similarity) between the feature corresponding to the accuracy improvement target image and the multiple feature similarities corresponding to the multiple training candidate images increases. The setting of the number of training images or the data amount is achieved, for example, by changing the similarity threshold. If the number of training images or the data amount for one accuracy improvement target image is to be greater than those for other accuracy improvement target images, the similarity threshold for the one accuracy improvement target image is set lower than the similarity thresholds for the other accuracy improvement target images. In other words, different similarity thresholds can be set for each accuracy improvement target image. The method for setting the number of training images or the data amount is not limited to the method described above. The generation number adjustment unit 4 can use not only the average of multiple feature similarities, but also any index (such as the median) statistically determined from the feature similarities.

[0047] In Example 2, in addition to being able to obtain the same effects as in Example 1, it is possible to generate more appropriate training images. That is, by adjusting the number of training images or the amount of data using the generation number adjustment unit 4, it is possible to prevent overtraining from occurring when relearning an image recognition model using the generated training images. Details of the effects are described below.

[0048] In Example 1, training images are selected from the training candidate images based on a single similarity threshold. However, it is generally assumed that there are multiple accuracy improvement target images, and training images must be selected from the training candidate images for each accuracy improvement target image. Assume that the feature similarity levels (distributions) of each accuracy improvement target image and the training candidate image differ. In this case, if training images for each accuracy improvement target image are selected based on a single similarity threshold, a large number of training images may be generated for a specific accuracy improvement target image. As a result, the retrained image recognition model retrained by the training unit 510 using the training images output from the training image generation unit 2 can accurately recognize the accuracy improvement target images from which a large number of training images have been extracted. However, overtraining due to the generation of a large number of training images may result in a decrease in the recognition accuracy of the evaluation images that were recognized before the retraining. In other words, the recognition accuracy of the evaluation images by the image recognition model retrained by the training unit 510 may be lower than the recognition accuracy of the evaluation images by the image recognition model before the retraining by the training unit 510.

[0049] Therefore, in a second embodiment of the present invention, when there are multiple images to be subjected to accuracy improvement, an average value of feature similarities, which is the degree of similarity between the feature values ​​of the training images and the feature values ​​of each of the images to be subjected to accuracy improvement, is calculated, and a constraint is imposed on the number of training images or the amount of data to be generated based on the relationship shown in Fig. 3. This makes it possible to prevent overtraining in the image recognition model after retraining. Therefore, it is possible to generate training images that can improve the training accuracy of the image recognition model.

[0050] Third Embodiment Next, a third embodiment of the present invention will be described.

[0051] FIG. 4 is a diagram showing the configuration of an image processing apparatus 102 according to the third embodiment.

[0052] The difference between Example 2 shown in FIG. 2 and Example 3 shown in FIG. 4 is that the inference unit 1 uses the result of inference using the target image for accuracy improvement as input as a reliability, and the generation number adjustment unit 4 determines the number of training images to be generated or the amount of data based on not only the feature similarity but also the reliability. Here, reliability refers to the image recognition model's recognition confidence for the generated recognition result when the target image for accuracy improvement is inferred using the image recognition model. The higher the reliability, the higher the confidence of the image recognition model and the higher the possibility that the recognition result will match the correct value. This is also called a confidence value. For example, if the reliability is 50% or less, the reliability is considered low. If the reliability of the inference result by the inference unit 1 is 50% or less, the image on which the inference was performed (the image input to the inference unit 1) becomes the target image for accuracy improvement.

[0053] The inference unit 1 outputs the feature amount (intermediate layer output) to the learning image generation unit 2, and further outputs the result of inference using the image to be improved as input to the number of images to be generated adjustment unit 4 as the reliability of the image to be improved.

[0054] The generation number adjustment unit 4 determines the number of candidate learning images output from the learning image generation unit 2 to be selected as learning images in accordance with the feature similarity and reliability.

[0055] FIG. 5 is a diagram illustrating the characteristics of the number of images required depending on the feature similarity and reliability of the images to be subjected to accuracy improvement, and explains a method for determining the number of learning images or the amount of data to be selected depending on the feature similarity and reliability.

[0056] 5, the generation number adjustment unit 4 sets the number of training images or the amount of data so that the higher the feature similarity, the fewer images are required, and the higher the reliability, the fewer images are required. For example, if the reliability of a specific accuracy improvement target image is half that of other accuracy improvement target images, the generation number adjustment unit 4 sets the required number of images or the required amount of data so that the required number of training images to be generated for the specific accuracy improvement target image is twice that of the other accuracy improvement target images.

[0057] In Example 3, it is possible to generate a more appropriate number of training images than in Example 2. In other words, by having the training unit 510 retrain an image recognition model using training images generated in accordance with the number of images or amount of data determined by the generation number adjustment unit 4, it is possible to prevent overtraining from occurring in the image recognition model after retraining. Details of the effects will be described below.

[0058] As described above, the higher the reliability, the higher the confidence of the AI ​​model, and the higher the likelihood that the recognition result will match the correct value. Conversely, the lower the reliability, the lower the confidence of the AI ​​model, and the lower the likelihood that the recognition result will match the correct value. Here, an accuracy improvement target image with a low reliability may be an image with features that the image recognition model has never learned or has only learned a small amount. Therefore, for an accuracy improvement target image with a low reliability, a large number of training images with features similar to those of the accuracy improvement target image are required when retraining the image recognition model. On the other hand, for an accuracy improvement target image with a high reliability, a large number of training images similar to the accuracy improvement target image may not be required when retraining the image recognition model. If training images are provided when an accuracy improvement target image with a high reliability is used as an image recognition model, overfitting may occur, and the recognition accuracy of the evaluation image by the image recognition model after retraining may be reduced compared to the recognition accuracy of the evaluation image by the image recognition model before retraining.

[0059] Therefore, in Example 3 according to the present invention, when there are multiple images to be subjected to accuracy improvement, the reliability of each image to be subjected to accuracy improvement is calculated from the result of inference by the inference unit 3, and a restriction is imposed on the number of training images or the amount of data to be generated based on the relationship shown in Fig. 5. This makes it possible to prevent overtraining in the image recognition model after relearning generated by the training unit 510.

[0060] As described above, according to the present invention, a configuration is made such that training images to be used for training a target image are generated in accordance with feature similarity, which is the degree of similarity between features related to a target image and features related to candidate images. This makes it possible to provide an image processing device, image processing system, and image processing method that use AI technology, which can generate additional training images without relying on experience or know-how and can improve the accuracy of the additional training images.

[0061] (Modification 1) Next, a first modification of the image processing device 100 will be described.

[0062] Fig. 6 is a configuration diagram showing a first modified example of the image processing device 100. As shown in Fig. 6, the learning unit 510 and the model storage unit 530 may be included in the image processing device 100. This allows the image processing device 100 to re-learn the image recognition model for a desired target image for accuracy improvement without communicating with a server, thereby improving the recognition accuracy of the image recognition model.

[0063] (Second Modification) Next, a second modification of the image processing system including the image processing device 100 and the server 500 will be described.

[0064] FIG. 7 is a configuration diagram showing a second modified example of an image processing system including an image processing device 100 and a server 500. As shown in FIG. 7, the accuracy improvement target image storage unit 32 and the learning image storage unit 34 may be included in a storage unit 520 of the server 500. Furthermore, the image processing device 100 may be included in a vehicle 900 or may be communicatively connected to a camera 99 mounted on the vehicle 900. In this case, the inference unit 1 can use images captured by the camera 99 as learning candidate images instead of the learning candidate images stored in the learning candidate image group storage unit 31, or in combination with the learning candidate images stored in the learning candidate image group storage unit 31. The camera 99 may be a single monocular camera, a stereo camera consisting of a pair of left and right cameras, or a multi-camera system with multiple cameras installed at different positions on the vehicle 900.

[0065] FIG. 8 is a diagram showing an example of a system configuration in which the image processing device 100 is provided in a vehicle 900. As shown in FIG.

[0066] In the vehicle 900, the image processing device 100 is connected to a camera 99, a vehicle control device 90, a brake control device 94, a steering control device 95, and the like via a bus 96.

[0067] The vehicle control device 90 is configured by a microcomputer that combines, for example, a CPU (Central Processing Unit) that executes calculations, a ROM (Read Only Memory) as a secondary storage device that records programs for the calculations, and a RAM (Random Access Memory) as a temporary storage device that saves the calculation progress and temporary control variables, and by executing the stored programs, it realizes each function of a recognition unit 91, a judgment unit 92, a control unit 93, etc.

[0068] The recognition unit 91 reads out the inference result of the inference unit 1 of the image processing device 100 that uses the captured image acquired by the camera 99 as input, and recognizes the external environment based on the inference result.

[0069] The determination unit 92 determines the action plan and driving route of the vehicle 900 based on the recognition result of the external environment by the recognition unit 91, and outputs the determination result to the control unit 93.

[0070] The control unit 93 calculates control command values ​​for the brake control device 94, the steering control device 94, etc. based on the judgment result by the judgment unit 92, and controls the vehicle 900 by outputting them to the brake control device 94, the steering control device 94, etc.

[0071] With this configuration, the server 500 communicates with many vehicles 900, distributes the desired accuracy improvement target images stored in the accuracy target image storage unit 32 to many vehicles, and uses a large number of images captured by the cameras 99 mounted on many vehicles as candidate images for learning. This allows for efficient re-learning of the desired image recognition model.

[0072] (Hardware Configuration) Finally, a description will be given of the hardware configuration of the image processing device 100 and the server 500. Here, the hardware configuration will be described using the image processing device 100 and the server according to Example 1 as an example, but the same applies to the image processing devices according to Example 2, Example 3, Modification Example 1, and Modification Example 2.

[0073] FIG. 8 is a diagram illustrating an example of the hardware configuration of the image processing device 100.

[0074] The image processing apparatus 100 realizes an image processing method in which the blocks cooperate to perform the above-described processes by causing a computer to execute a program.

[0075] The image processing apparatus 100 includes a CPU (Central Processing Unit) 200 , a ROM (Read Only Memory) 201 , a RAM (Random Access Memory) 202 , a non-volatile storage 203 , and a transmission / reception unit 204 , all of which are connected to a bus 205 .

[0076] The CPU 200 reads out program code of software that realizes each function according to this embodiment from the ROM 201, loads it into the RAM 202, and executes it. Variables, parameters, etc. generated during the calculation processing of the CPU 200 are temporarily written to the RAM 202, and these variables, parameters, etc. are read out by the CPU 202 as appropriate. However, an MPU (Micro Processing Unit) or a GPU (Graphics Processing Unit) may be used instead of the CPU 200, or the CPU 200 and a GPU (Graphics Processing Unit) may be used together. For example, the functions of the inference unit 1 and the learning image generation unit 2 are realized by the CPU 200, the ROM 201, and the RAM 202.

[0077] The non-volatile storage 203 may be, for example, a hard disk drive (HDD), a solid state drive (SSD), a flexible disk, an optical disk, a magneto-optical disk, a CD-ROM, a CD-R, a magnetic tape, or a non-volatile memory. The non-volatile storage 203 stores an operating system (OS), various parameters, and programs for operating the image processing device 100. The ROM 201 and the non-volatile storage 203 store programs and data necessary for the CPU 200 to operate, and are used as examples of computer-readable, non-transitory storage media that store programs executed by the image processing device 100. For example, the functions of the learning candidate image group storage unit 31, the similarity threshold storage unit 32, the accuracy improvement target image storage unit 33, and the learning image storage unit 34 are realized by the non-volatile storage 203.

[0078] The transmitter / receiver 508 may be, for example, a network interface card (NIC), which allows various data to be transmitted and received between devices via wireless communication connected to the terminal of the NIC. The server 500 is also configured to be able to transmit and receive various data to and from a server or the like via a network such as a LAN or the Internet, or a dedicated line. The server 500 also has a similar hardware configuration, and a computer executes a program to realize an information processing method in which the above-described processes are performed in cooperation with each block.

[0079] Although the above-described second modification is an example applied to the recognition of images captured by an in-vehicle camera, the present invention can also be applied to other image recognition devices. For example, the present invention can be applied to a recognition device for images captured by a traffic volume survey camera, a security camera, or a camera for inspecting processed products in a factory.

[0080] The above-described configurations, functions, processing units, processing means, etc. may be partly or entirely realized in hardware by, for example, designing them as integrated circuits, etc. Furthermore, the above-described configurations, functions, etc. may be realized in software by a processor interpreting and executing a program that realizes each function.

[0081] The present invention is not limited to the above-described embodiments, and various other applications and modifications are possible without departing from the spirit and scope of the present invention as defined in the claims. For example, the above-described embodiments provide detailed and specific descriptions of the device and system configurations to clearly explain the present invention, and are not necessarily limited to those including all of the described configurations. Furthermore, it is possible to replace part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add the configuration of another embodiment to the configuration of one embodiment. It is also possible to add, delete, or replace part of the configuration of each embodiment with other configurations. Furthermore, the control lines and information lines shown are those considered necessary for the explanation, and not all control lines and information lines are necessarily shown in the product. In reality, it can be assumed that almost all components are interconnected.

[0082] Inference unit 1, learning image generation unit 2, similarity calculation unit 21, similar image selection unit 22, memory unit 3, learning candidate image group memory unit 31, accuracy target image memory unit 32, similarity threshold memory unit 33, learning image memory unit 34, server 500, learning unit 510, server memory unit 520, model storage unit 530, image processing device 100

Claims

1. An image processing device comprising: an inference unit that determines features for a target image and features for candidate images included in a group of candidate images based on an intermediate layer output of an image recognition model obtained by inference in which a target image and a group of candidate images are input to the image recognition model, and outputs the features; and a learning image generation unit that generates learning images to be used for learning the target image in accordance with feature similarity, which is the degree of similarity between the features for the target image and the features for the candidate images output from the inference unit.

2. In the image processing device described in claim 1, the learning image generation unit outputs the feature similarity and the generated learning images, and the image processing device is characterized in that it is equipped with a generation number adjustment unit that determines the amount of learning images to be generated based on the feature similarity and the learning images output from the learning image generation unit.

3. In the image processing device described in claim 1, the learning image generation unit has a similar image selection unit that selects a similar image from the candidate images included in the candidate image group whose feature similarity with the target image exceeds a threshold, and the learning image is generated using the similar image selected by the similar image selection unit.

4. An image processing device according to claim 2, wherein the generation number adjustment unit reduces the number of learning images generated as the feature similarity increases.

5. An image processing device as described in claim 1, wherein the inference unit outputs the feature and further outputs the result of the inference as a reliability of the target image, the training image generation unit outputs the feature similarity and the generated training images, and the image processing device further comprises a generation number adjustment unit which receives as input the feature similarity output from the training image generation unit, the training images, and the reliability output from the inference unit, and determines the amount of training images to be generated based on the feature similarity and the reliability.

6. An image processing device according to claim 5, wherein the generation number adjustment unit reduces the number of learning images generated as the reliability increases.

7. An image processing device according to claim 1, wherein the inference unit recognizes images captured by an in-vehicle camera.

8. An image processing system including the image processing device according to claim 1 and a server communicatively connected to said image processing device via a network, wherein said server is provided with a learning unit that performs learning of said image recognition model based on said training images, and a model storage unit that stores said trained model.

9. An image processing method which calculates features for a target image and features for candidate images included in a group of candidate images based on the intermediate layer output of an image recognition model obtained by inference in which a target image and a group of candidate images are input into the image recognition model, outputs the features, and generates a learning image to be used for learning the target image according to the feature similarity, which is the degree of similarity between the output features for the target image and the features for the candidate images.

Citation Information

Patent Citations

  • Object detection device, detection model generator, program, and method capable of learning based on search result

    JP2018169972A

  • Learning data collection device, learning device, learning data collection method, and program

    JP2022038941A