Image recognition processing AI learning device and learning method

By using XAI to set and train AI systems on appropriate gaze regions, the method addresses the instability of image recognition AI, enhancing accuracy and reliability for autonomous driving systems.

JP2026054928APending Publication Date: 2026-03-30ASTEMO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-17
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Existing image recognition AI systems face challenges in accurately and reliably recognizing objects due to limited training data and instability in recognition performance, particularly when faced with unknown scenarios or environmental changes, leading to potential false detections and missed detections, which can compromise the safety of autonomous driving systems.

Method used

A learning device and method that utilizes explainable AI (XAI) to determine an appropriate normative gaze region for each label and situation, setting a normative label and training the AI to focus on this region, thereby improving recognition accuracy and reliability by visualizing and adjusting the gaze area based on heatmaps and retraining the model.

Benefits of technology

The method enhances the accuracy and reliability of image recognition AI by ensuring it focuses on the appropriate gaze areas, reducing reliance on irrelevant information and improving stability across varying environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026054928000001_ABST
    Figure 2026054928000001_ABST
Patent Text Reader

Abstract

To obtain a learning device and learning method for an image recognition processing AI that can improve accuracy and reliability. [Solution] The learning device and learning method for an image recognition processing AI of the present invention group multiple images in an image dataset according to their features, determine the gaze region with the highest appropriateness rating in each group as the normative gaze region A, assign a normative label to each normative gaze region A, and set a new gaze region A' based on the normative gaze region A for other images in the same group that do not have the normative label assigned. Furthermore, the difference between the gaze region during AI learning and the new gaze region A' is evaluated, and the AI ​​learning model of the image recognition processing AI is retrained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0002] ,

[0004] , , , , ,

[0003]

[0001] The present invention relates to a learning device and a learning method for image recognition processing AI.

Background Art

[0002] Camera systems mounted on automobiles and the like use technologies that analyze the surrounding situation through images by a computer and utilize it for autonomous driving. In recent years, with the introduction of AI (artificial intelligence), it has become possible to recognize complex and changing road conditions in more detail and make quick and accurate judgments. Since this technology for recognizing the surrounding environment of autonomous driving is related to human lives, it is not only necessary to be accurate, but also to satisfy reliability and safety at the same time.

[0003] Image processing by AI exhibits very high performance in enabling a vehicle to understand the surrounding environment. However, there is a problem that it is difficult to perfectly grasp all situations because the data learned by AI is limited. In an actual driving environment, since the AI may face unknown scenarios that it has not learned in advance, there is a possibility that false detections and missed detections may occur. Such incomplete recognition has a significant impact on the safety of autonomous driving, so the recognition processing system of AI is constantly evaluated and it is necessary to ensure high reliability and robustness. In order for the AI to accurately understand various scenes that occur in the actual environment, it is required to continuously learn based on the data collected in the actual environment and improve the recognition ability of the AI.

[0004] To improve AI's image processing capabilities, it's necessary to collect data from external sources, such as new objects or scenes that previously experienced recognition errors, and use this data for AI training. By training the AI ​​with this collected data, it can become able to accurately recognize objects and scenes that it previously failed to recognize correctly. A challenge in this training process is that even when the AI ​​becomes able to recognize correctly, the results often include unstable recognition outcomes. Specifically, it may only recognize a part of the target object, or make decisions based on irrelevant information. This instability means that recognition ability deteriorates even with slight environmental changes, such as changes in the angle of the target object or when the recognition area is obscured by other objects. Therefore, to develop a highly reliable image AI, it is necessary to select data from the collected dataset that the AI ​​recognizes particularly poorly and to focus the training on that data. To achieve this, it is necessary to utilize explainable AI (XAI) technology to evaluate which areas the AI ​​uses as the basis for its recognition decisions and to perform training based on that information.

[0005] Various techniques have been explored to improve the reliability of AI recognition. One such technique proposes expanding the gaze region, which is the basis for the AI's judgment, in order to increase accuracy. Patent document 1 describes an image processing method in which the gaze region is visualized using XAI and then blacked out, and the gaze region is then used to train the AI ​​to expand the gaze region. By using such a technique, it becomes possible to correct the behavior of AI that only uses a portion of the recognition object as its basis. [Prior art documents] [Patent Documents]

[0006] [Patent Document 1] Japanese Patent Publication No. 2023-84981 [Overview of the project] [Problems that the invention aims to solve]

[0007] In the technology described in Patent Document 1, if the gaze area occupies a large portion of the recognition target, resulting in an image that is inherently highly reliable, the majority of the area may be painted black, making it unsuitable as training data. Therefore, determining whether or not to process the image, and which pixels to process, requires visual confirmation and trial and error to verify the effect, which presents challenges in terms of significant human cost and time. Furthermore, processing the gaze area may cause the AI ​​to learn information from unnecessary areas such as the background, potentially leading to unstable AI recognition performance. In addition, it may be necessary to set an appropriate gaze area according to the type, situation, and characteristics of the recognition target, for example, by checking for the presence or absence of a cargo bed when distinguishing between a regular vehicle and a truck.

[0008] The present invention has been made in view of the above problems, and its objective is to provide a method and apparatus for learning an image recognition processing AI that can improve accuracy and reliability by setting an appropriate gaze area for each label, feature, and situation of the object to be recognized during the learning of the image recognition processing AI, and training the image recognition processing AI so that it can obtain a gaze area that matches the appropriate gaze area. [Means for solving the problem]

[0009] The learning device for image recognition processing AI of the present invention, which solves the above problems, A learning device for an image recognition processing AI that recognizes objects captured in an image, The learning device is A normative label setting unit determines a normative gaze region from among the gaze regions that the AI ​​learning model of the image recognition processing AI focuses on in multiple images, and assigns a normative label to the normative gaze region. A normative gaze region setting unit sets a new gaze region based on the normative gaze region for other images that have gaze regions to which the normative label has not been assigned, It is characterized by having the following features. Furthermore, the learning method for the image recognition processing AI of the present invention, which solves the above problems, is A method for training an image recognition processing AI that uses a computer to recognize objects captured in an image, A gaze region calculation step in which the AI ​​learning model of the aforementioned image recognition processing AI calculates gaze regions from multiple images, A gaze area evaluation step for evaluating the appropriateness of the gaze area, Based on the evaluation of the appropriateness of the gaze region, a normative gaze region is determined from among the multiple gaze regions to serve as a norm for the gaze region that the AI ​​learning model of the image recognition processing AI focuses on, and a normative label is assigned to the normative gaze region in a normative label setting step. A normative gaze region setting step, which sets a new gaze region based on the normative gaze region for other images that have gaze regions to which the normative label has not been assigned, The method is characterized by including a retraining step which compares the new gaze region set for the other image by the normative gaze region setting step with the gaze region calculated by the gaze region calculation step before the new gaze region was set for the other image, and retrains the AI ​​learning model so that the gaze region matches the new gaze region. [Effects of the Invention]

[0010] According to the present invention, in a learning method and learning apparatus for an image recognition processing AI, an appropriate normative gaze region is determined for each label, feature, and situation of the recognition target, and the image recognition processing AI is trained to obtain a gaze region that matches that normative gaze region, thereby improving the accuracy and reliability of image recognition by the image recognition processing AI.

[0011] Further features related to the present invention will become apparent from the description herein and the accompanying drawings. Problems, configurations, and effects not described above will be revealed by the following description of embodiments. [Brief explanation of the drawing]

[0012] [Figure 1]A diagram showing an example of the hardware configuration of a learning device for an image recognition processing AI according to the first embodiment of the present invention. [Figure 2] A functional block diagram of a learning device for an image recognition processing AI according to the first embodiment of the present invention. [Figure 3] A flowchart showing a learning method for an image recognition processing AI according to the first embodiment of the present invention. [Figure 4] A diagram showing an example of classifying and grouping recognition targets by feature. [Figure 5] A diagram showing a part of a list of the original image of group 3 in which recognition targets on the left rear surface of a vehicle are collected and the attention area heat map. [Figure 6] A diagram showing an example of a procedure for setting a new attention area in another image based on a standard attention area. [Figure 7] A flowchart showing a learning process procedure for updating weights by evaluation based on a standard attention area. [Figure 8] A flowchart showing a learning method for an image recognition processing AI according to the second embodiment of the present invention. [Figure 9] A diagram showing an example of an occlusion area in an image. [Figure 10] A flowchart showing a procedure for generating a processed image based on attention area information executed in step S61 of FIG. 8. [Figure 11] A diagram showing a method of replacing the pixel color of the pixel coordinates of the area outside the attention area and the pixel color of the pixel coordinates of the overlapping area D where the attention area evaluation value is greater than or equal to the threshold with black. [Figure 12] A flowchart showing a learning method for an image recognition processing AI according to the third embodiment of the present invention. [Figure 13] A graph showing the relationship between the score distribution before and after learning and the presence or absence of improvement.

Embodiments for Carrying Out the Invention

[0013] <0000Embodiments of the present invention will be described below with reference to the drawings. Each embodiment is illustrative for explaining the present invention, and has been omitted and simplified as appropriate for clarity of explanation. The present invention can also be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.

[0014] The positions, sizes, shapes, and ranges of the components shown in the drawings may not represent their actual positions, sizes, shapes, and ranges in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the positions, sizes, shapes, and ranges disclosed in the drawings. When there are multiple components with the same or similar functions, they may be described using the same reference numeral but with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description.

[0015] In each embodiment, the processing performed by executing the program may be described. Here, the computer executes the program using a processor (e.g., CPU, GPU) and performs the processing defined in the program using memory resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the main entity performing the processing by executing the program may be the processor. Similarly, the main entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The main entity performing the processing by executing the program may be an arithmetic unit, and may include a dedicated circuit that performs a specific processing. Here, a dedicated circuit is, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a CPLD (CompleX Programmable Logic Device).

[0016] The program may be installed on the computer from the program source. The program source may be, for example, a program distribution server or a storage medium readable by the computer. If the program source is a program distribution server, the program distribution server includes a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to other computers. In addition, in this embodiment, two or more programs may be implemented as one program, or one program may be implemented as two or more programs.

[0017] As autonomous driving technology advances, AI image programs using machine learning are being developed to enable vehicles to accurately perceive their surroundings. Unlike rule-based algorithms, these programs present a challenge in that it is difficult to address problems when they occur because the reasoning process used by the AI ​​is not transparent.

[0018] To address this challenge, a technique called explainable AI (XAI) has been proposed, which visualizes the gaze region—the area that the AI ​​uses as the basis for its judgment in image processing. For example, even if an AI evaluates in its inference that it recognizes the entire vehicle, in reality, it may only be focusing on a part of the vehicle, and if that part is obscured, it may not be able to recognize it accurately. Therefore, in order to develop a highly reliable AI, it is necessary to select data from the collected dataset in which the AI's recognition is particularly unstable, and to focus the training on that data. To achieve this, it is necessary to utilize XAI technology to evaluate which areas the AI ​​uses as the basis for its judgment and to perform training based on that information.

[0019] This invention improves the accuracy and reliability of recognition by visualizing the region of interest that serves as the basis for AI's judgment when recognizing objects in an image using XAI, and by performing AI learning based on that information. The following embodiments describe an example of its application to image recognition processing AI in an in-vehicle ECU for autonomous driving (AD). This invention is applicable not only to advanced driver assistance systems (ADAS) and autonomous driving systems (ADS), but can also be applied to image recognition processing AI applications in a wide range of fields.

[0020] —First Embodiment— Hereinafter, with reference to Figures 1 to 7, a learning device and learning method for an image recognition processing AI according to the first embodiment of the present invention will be described.

[0021] Figure 1 shows an example of the hardware configuration of a learning device 1 for an image recognition processing AI according to the first embodiment of the present invention (hereinafter referred to as the "learning device"). Learning device 1 is a computer such as a server or PC, and as shown in Figure 1, it is composed of an arithmetic unit 102, a storage device 103, an input device 104, an output device 105, and a communication interface (communication I / F) 106, which are connected to each other via a system bus 101.

[0022] The computing unit 102 is composed of a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), and the like. It controls the operation of each device and controller connected to the system bus 101 in the learning device 1. It also performs calculations such as training the image recognition processing AI.

[0023] The storage device 103 is a recording medium capable of storing various programs and data non-temporarily or temporarily, and is configured using, for example, ROM (Read Only Memory), RAM (Random Access Memory), HDD (Hard Disk Drive), etc. It stores various types of data such as images and analysis / learning results.

[0024] The input device 104 consists of a keyboard, mouse, touch panel, microphone, etc. This is a means for the user to input information into the learning device 1. The output device 105 consists of a display, projector, printer, speaker, etc. This is a device that outputs data and information to the user.

[0025] The communication interface 106 consists of communication devices for connecting to a network. It has the function of sending and receiving data between the learning device 1 and external information equipment (not shown) via a communication network (not shown).

[0026] Figure 2 is a functional block diagram of a learning device according to the first embodiment of the present invention. The learning device 1 acquires the image recognition processing AI 200 and the training image dataset (input image dataset) 201, retrains the AI ​​learning model of the image recognition processing AI 200, and updates the image recognition processing AI 200 to image recognition processing AI 200'.

[0027] The learning device 1 clusters and groups the features of the training image dataset 201, and uses the image recognition processing AI 200 and the training image dataset 201 to calculate the gaze region of each cluster and visualize the gaze region using XAI. Then, it determines a normative gaze region A for each group from among the multiple gaze regions and sets a normative label (annotation using visualization information). Finally, it updates the AI ​​learning model by evaluating the appropriateness of the gaze region using the normative gaze region A in each group.

[0028] The training image dataset 201 is a collection of multiple images in which the object to be recognized is captured. In this embodiment, the training image dataset 201 is an image dataset consisting of multiple images of a vehicle captured and collected by an in-vehicle camera, and is stored, for example, in the storage device 103. Each image may be the entire captured image, or it may be a part of the captured image, such as the area in which the object to be recognized is captured.

[0029] The learning device 1 is realized when the learning program in the storage device 103 is executed by the arithmetic unit 102, and each part's function is realized. It includes a grouping unit 202, a gaze area calculation unit 203, a gaze area evaluation unit 204, a norm label setting unit 205, a norm gaze area setting unit 206, and a retraining unit 207.

[0030] The grouping unit 202 acquires a training image dataset 201, divides (clusters) the recognition targets in the training image dataset 201 according to their type, characteristics, and situations, and groups them. In this embodiment, the recognition target is a vehicle, and the images are classified and grouped by imaging surface, such as the rear or left side of the vehicle, as clusters. By grouping, it is possible to efficiently find appropriate gaze area candidates and learning conditions individually.

[0031] The gaze region calculation unit 203 calculates gaze regions and generates heatmaps. The gaze region calculation unit 203 acquires the images grouped by the grouping unit 202 and the image recognition processing AI 200 to be retrained, and calculates gaze regions from the grouped images using the image recognition processing AI 200. The gaze region is the region that the image recognition processing AI 200 used as evidence when recognizing the object to be recognized, and is calculated for each object to be recognized from multiple image data (original images) included in the training image dataset 201. The gaze region is calculated so that the value increases as the degree of prediction evidence increases for each pixel, and the heatmap is color-set so that the color gradually changes according to the calculated gaze region values. The gaze region calculation unit 203 visualizes the gaze regions by superimposing the created heatmap onto the image.

[0032] The gaze area evaluation unit 204 evaluates the appropriateness of the gaze area calculated for each recognized object. For example, the score distribution of the visualization result, which visualizes the basis for predicting the gaze area, is used to evaluate the appropriateness of the gaze area. The gaze area evaluation unit 204 evaluates so that the gaze area evaluation value increases as the score distribution of the gaze area increases. The gaze area evaluation unit 204 may also evaluate the appropriateness of the gaze area based on the score distribution, size, location, and at least one of the distributions of the gaze area.

[0033] The norm label setting unit 205 sets a norm gaze region A for each group from among multiple gaze regions, which will serve as the norm for the region that the AI ​​learning model gazes upon, and assigns a norm label to that norm gaze region A. In this embodiment, the gaze region with the highest appropriateness rating for each group is designated as the norm gaze region A, and a norm label is assigned to that norm gaze region A, while the other gaze regions become gaze regions to which no norm labels have been assigned.

[0034] The normative gaze area setting unit 206 sets a new gaze area A' based on normative gaze area A for other images in the same group that have gaze areas that have not been assigned a normative label. This setting of normative gaze area A and the setting of the new gaze area A' is performed for each group.

[0035] The new gaze region A' is a region that the image recognition processing AI 200 should focus on, and it is a gaze region that is evaluated as more appropriate compared to gaze region A, which does not have a normative label. The normative gaze region setting unit 206 adds the coordinate information of the new gaze region A' to the information of the recognition target included in other images in the same group in the training image dataset 201 that have gaze regions that do not have a normative label.

[0036] The retraining unit 207 retrains the AI ​​learning model of the image recognition processing AI 200 using the training image dataset 201 to which the coordinate information of a new gaze region A' has been added to the recognition target. The retraining unit 207 evaluates the difference (error) between the gaze region without a reference label and the new gaze region A', and updates the weights so that the gaze region appears across the entire new gaze region A' when image recognition processing is performed.

[0037] Figure 3 is a flowchart showing the learning method of the learning device 1 according to the first embodiment of the present invention, and Figure 4 is a diagram showing an example in which recognition targets are classified and grouped according to their features.

[0038] In step S10, the learning device 1 acquires the image recognition processing AI 200 and the training image dataset 201. The image recognition processing AI 200 is a pre-trained model using a DNN (Deep Neural Network). Note that the model used for the image recognition processing AI 200 is not limited to a DNN (Deep Neural Network); other known technologies such as support vector machines or Transformers may also be used.

[0039] The training image dataset 201 is used as training data when training the image recognition processing AI 200. In this embodiment, the open dataset BDD100K (a dataset of vehicle images collected by an in-vehicle camera) was used. This dataset contains images and the correct labels (class, coordinates) of the recognition targets present within them.

[0040] In step S20, the grouping unit 202 groups the recognition targets in the training image dataset 201. The grouping is performed for each label and feature of the training data using clustering. The grouping unit 202 performs clustering of the recognition targets in the training image dataset 201 using unsupervised learning for each correct label, thereby dividing and grouping the recognition targets in the training image dataset 201 according to their respective features.

[0041] In this embodiment, as shown in Figure 4, clustering was performed with vehicles as the target of recognition, and the results were grouped based on the vehicle body angle in the images. Group 1 was a collection of images of the rear of the vehicle, Group 2 was a collection of images of the left front of the vehicle, Group 3 was a collection of images of the left rear of the vehicle, and Group 4 was a collection of images of the left side of the vehicle. The method for this grouping step is not limited to any particular means, and for example, the grouping may be done by visual confirmation or automatic grouping by feature extraction.

[0042] In step S30, the gaze region calculation unit 203 calculates the gaze region and generates a heat map. The gaze region calculation unit 203 calculates the gaze region of the target to be recognized from each image in the training image dataset using the AI ​​learning model of the image recognition processing AI 200 (gaze region calculation step), and generates a heat map. In this specification, the heat map is drawn so that the larger the calculated gaze region value, the more black is set, and the smaller the value, the more white is set. By superimposing this color setting onto the image, the gaze region can be visualized (see the gaze region heat map in Figure 5). In this embodiment, D-RISE, a known technology, is used for calculating the gaze region and generating the heat map, but the invention is not limited to this, and other known technologies such as Grad-CAM may be used.

[0043] In step S40, the gaze area evaluation unit 204 evaluates the appropriateness of the gaze area (gaze area evaluation step). Here, the gaze area calculated in step S30 is evaluated using the score distribution of the visualization results. The score distribution of the visualization results of the gaze area is calculated by integrating the numerical values ​​of each pixel in the gaze area of ​​the entire region where the recognition target exists, and dividing it by the area of ​​the region where the recognition target exists. In this embodiment, the integrated value of the heatmap is used to evaluate the appropriateness of the gaze area, but other indicators such as the area of ​​each color in the heatmap (size of the gaze area) may also be used. Furthermore, the appropriateness of the gaze area can be evaluated not only based on the score distribution and size of the gaze area, but also on its location and distribution. For example, if the goal is to distinguish between a passenger car and a truck, and the goal is to recognize the truck bed as the gaze area, the appropriateness of the gaze area can be evaluated based on its location and distribution. The appropriateness of the gaze area can be evaluated based on at least one of the score distribution, size, location, and distribution of the gaze area, and can also be a combination of two or more factors.

[0044] In step S50, the norm label setting unit 205 performs the process of setting a norm label in norm gaze area A. The norm label setting unit 205 determines the gaze area with the highest appropriateness evaluation within each group as norm gaze area A and assigns a norm label to that norm gaze area A (norm label setting step). The norm label setting unit 205 assigns a norm label to the gaze area with the highest evaluation value calculated in step S40 for each group. As a specific example, the case of applying this to group 3, which consists of recognition targets on the left rear of a vehicle, will be described later with reference to Figure 5. Note that here, the norm label is set to the gaze area with the highest evaluation value, but other means such as visual confirmation by a person or setting a norm label to match a specified gaze area pattern (location or distribution) may also be used.

[0045] In step S60, a new gaze region A' is set based on the norm gaze region A for other images that have gaze regions that do not have a norm label (norm gaze region setting step). The norm gaze region setting unit 206 sets gaze regions that have a norm label within each group as the norm gaze region A. Then, for other images in the same group that have gaze regions that do not have a norm label, a new gaze region A' is set based on the norm gaze region A. In this embodiment, a gaze region having the same coordinates as the norm gaze region A of the same group is set as the new gaze region A' for the other images. In the other images, a gaze region has already been set by the gaze region calculation unit 203 through the processing in step S30, but that gaze region is changed to the new gaze region A'.

[0046] The normative gaze area setting unit 206 adds coordinate information for a new gaze area A' to the information of other images in the training image dataset 201. For a specific example, refer to Figure 6 and describe the case where it is applied to group 3 of images containing recognition targets on the left rear of a vehicle. Note that the normative gaze area A may be set, for example, based on human judgment or on known information about the characteristics of the target (such as the cargo bed of a truck).

[0047] In step S70, the new gaze region set for other images in step S60 is compared with the gaze region calculated in step S30 before the new gaze region was set for other images, and the AI ​​learning model is retrained so that the gaze region matches the new gaze region (retraining step). The retraining unit 207 retrains the AI ​​learning model using the training image dataset 201, to which the information of the new gaze region A' was added in step S60. Here, the gaze region is calculated during training, the difference (error) between the gaze region without a reference label and the new gaze region A' is evaluated, and the weights are updated so that the gaze region appears throughout the entire new gaze region A'. Details of this learning method will be described later with reference to the flowchart in Figure 7.

[0048] Figure 5 is a partial list of the images and gaze area heatmaps for Group 3, which consists of original images of the left rear of the vehicle. In this figure, the larger the calculated gaze area value, the more black is set, and the smaller the value, the more white is set. This list summarizes the relationship between the gaze area heatmap and the gaze area evaluation value calculated and evaluated in steps S30 and S40. For simplicity, in this example, we will explain using images (1), (2), and (3) from the list. Among these, image (2) has the largest gaze area evaluation value, so we set a reference label for the gaze area of ​​image (2) as the target for the most robust detection.

[0049] Figure 6 shows an example of the settings for each reference label in Group 3, which consists of original images of the left rear of a vehicle. Since the gaze region of image (2) in Figure 6 has a reference label, the coordinates of the region where this gaze region of the recognition target exists are obtained, and that region is set as the reference gaze region A. Then, a new gaze region A' with the same coordinates as the reference gaze region A is set in the other images (1) and (3) that have gaze regions without reference labels. The coordinates of this new gaze region A' are added to the annotation information, such as the correct label, of the data of images (1), (2), and (3) in the training image dataset 201 that belong to this group.

[0050] Figure 7 is a flowchart showing the steps of the learning method performed in step S70 of Figure 3. In step S701, a training image dataset 201 containing information on the normative gaze region A and the image recognition processing AI 200 to be trained are acquired. The training image dataset 201 contains information on the new gaze region A' set in step S60.

[0051] In step S702, the AI ​​learning model of the image recognition processing AI200 is trained using the training image dataset 201. Once training has started, batch processing is used to acquire images and related information (ground truth labels, reference gaze regions, etc.) equal to the batch size, and input them into the image recognition processing AI200.

[0052] In step S703, the learning is evaluated using the loss function and the weights are updated. In general, machine learning evaluates the accuracy of the training data and predicted labels for each batch size using the loss function and updates the weights. However, in this embodiment, the gaze region is calculated for each batch size and evaluated using the loss function, including the difference from the normative gaze region.

[0053] In this embodiment, the difference between the gaze region calculated by the gaze region calculation unit 203 and a new gaze region A' based on the reference gaze region A is evaluated using the known IoU (Intersection over Union) loss. This is obtained by dividing the intersection of the calculated gaze region and the reference gaze region A by their combined portion and subtracting the resulting score from 1. The evaluation method is not limited to this, and other known techniques such as cross-entropy loss may be used.

[0054] In step S704, it is determined whether processing has been completed for all batches of the divided training data.

[0055] In step S705, based on the specified conditions, the program proceeds to the next epoch if the conditions are not met, and terminates learning if the conditions are met. In this embodiment, the learning termination condition is when there are no more weight updates. Other known termination conditions, such as setting an upper limit on the number of epochs or a threshold for the loss function, may also be used as examples.

[0056] The learning method for the image recognition processing AI 200 using the learning device 1 is a learning method for an image recognition processing AI that recognizes objects to be recognized that are captured in the input image. This learning method involves a step S30 (gaze area calculation step) in which the AI ​​learning model of the image recognition processing AI200 calculates the gaze area from multiple images, Step S40 (gaze area evaluation step) is to evaluate the appropriateness of the gaze area, Step S50 (normative label setting step) involves determining a normative gaze region A from among multiple gaze regions based on an evaluation of the appropriateness of the gaze region, which will serve as a norm for the gaze region that the AI ​​learning model of the image recognition processing AI200 will focus on, and assigning a normative label to the normative gaze region A. Step S60 (normative gaze area setting step) sets a new gaze area A' based on normative gaze area A for other images that have gaze areas that have not been assigned normative labels, The process includes step S70 (retraining step), which compares the new gaze region A' set for the other image by step S60 with the gaze region calculated by step S30 before the new gaze region A' was set for the other image, and retrains the AI ​​learning model so that the gaze region matches the new gaze region A'.

[0057] In this learning method, the computing unit 102 classifies and groups the training image dataset by label, feature, and situation (S20), and calculates and evaluates the gaze region for each data (S30, S40). Then, the target with the highest evaluation of the appropriateness of the gaze region in each group is designated as the normative gaze region A (S50), and a new gaze region A' based on that normative gaze region A is set for the other images in the same group (S60). The gaze region during training and the new gaze region A' are compared in the other images, and the weights are updated so that the gaze region during training matches the new gaze region A' (S70). Therefore, the appropriateness of the gaze region calculated by the image recognition processing AI 200' can be improved, enabling correct inference regardless of which region of the recognition target is viewed, thereby improving accuracy and reliability.

[0058] Steps S20, S50, and S60 implement a function that automatically sets the gaze area according to the vehicle body angle and other conditions captured in the image. This reduces manual work by humans and enables the rapid and efficient setting of a highly robust gaze area that should serve as a standard.

[0059] Through step S70, the AI ​​learning model continuously evaluates the appropriateness of the gaze area during the learning process and updates the model. This enables continuous adjustment of the gaze area through learning, resulting in a more efficient learning process.

[0060] According to the learning method for the image recognition processing AI 200 using the learning device 1, in the learning of the image recognition processing AI, an appropriate gaze region (new gaze region A') can be set by grouping each label, feature, and situation of the recognition target, and the image recognition processing AI can be trained to obtain a gaze region that matches this new gaze region A', thereby improving accuracy and reliability.

[0061] —Second Embodiment— A learning method and learning apparatus for an image recognition processing AI according to a second embodiment of the present invention will be described.

[0062] In the first embodiment, the image recognition processing AI for recognizing objects to be recognized that are captured in the input image may experience occlusion, a state in which objects in the foreground obscure objects behind them, because there are many objects to be recognized in a single image. In such cases, it may not be possible to properly evaluate the appropriateness of the gaze region.

[0063] Furthermore, in the first embodiment, a normative gaze region A was set, and the weights of the image recognition processing AI were updated by evaluating it with a loss function under training, thereby improving the appropriateness of the gaze region. In this process, it is thought that the appropriateness of the gaze region can be efficiently improved by actively training the AI ​​on areas of the normative gaze region A that are not currently recognized.

[0064] In this embodiment, in addition to the first embodiment, a method for evaluating the gaze region when a recognition target in an image is hidden by other recognition targets (occlusion occurs), and a method for generating a processed image using the gaze region will be described. The hardware configuration of the learning device 1 according to this embodiment is the same as the hardware configuration of the learning device 1 shown in Figure 1 described in the first embodiment.

[0065] In the learning device 1 of this embodiment, the gaze area evaluation unit 204 evaluates the appropriateness of the gaze area of ​​the recognition target by excluding the occlusion area B, which is hidden by other objects within the gaze area of ​​the recognition target, from the target area in the image when there is an image in which the recognition target is partially hidden by other objects due to occlusion.

[0066] The retraining unit 207 performs image processing on occlusion regions to reduce the degree of gaze of the AI ​​learning model. Then, within the new gaze region A', the retraining unit 207 performs image processing on overlapping regions D that overlap with the gaze region calculated by the gaze region calculation unit 203 to reduce the degree of gaze of the AI ​​learning model. Finally, the retraining unit 207 performs image processing on regions C outside the gaze region, excluding the new gaze region A' of other images, to reduce the degree of gaze of the AI ​​learning model.

[0067] Figure 8 is a flowchart showing the learning method of the learning device 1 according to the second embodiment of the present invention. The learning method of the learning device 1 differs from the learning method described in the first embodiment in that it further includes steps S41 and S61. Steps S10 to S30 perform the same processing as in the first embodiment.

[0068] In step S41, the appropriateness of the gaze area is evaluated in the area excluding the occlusion region. Here, it is determined whether occlusion occurs for each recognition object. This is done by checking for overlaps between the coordinate information of the correct label of the recognition object being evaluated and the coordinate information of other recognition objects present in the same image.

[0069] If the objects to be recognized overlap in the image, the object whose base is relatively lower in the image, as shown in Figure 9, is considered the object in the foreground. If the object to be judged is located behind the object (further away from the foreground object), occlusion is determined, and the overlapping occlusion region B is excluded from the area of ​​the object to be recognized. The same gaze region evaluation calculation as in step S40 of the first embodiment is then performed. If there are multiple overlapping objects to be recognized, the occlusion region B (see, for example, Figure 9) is excluded for all objects in the foreground. Specific examples of occlusion determination and occlusion region exclusion will be described later with reference to Figure 9.

[0070] Steps S50 and S60 perform the same processes as in the first embodiment, respectively.

[0071] In step S61, a processed image is generated using the gaze region information calculated in S30 and the new gaze region A' information set in S60. The norm gaze region setting unit 206 obtains the pixel coordinates of region C outside the norm gaze region A from the area of ​​the recognition target in the image having the norm gaze region A, and applies a mask by replacing the color of the pixels in region C with black. Similarly, for other images with gaze regions that do not have a reference label, the color of pixels in region C outside the new gaze region A' from the region to be recognized is replaced with black. Then, the color of pixels in the overlapping region D between the new gaze region A' and the calculated gaze region is replaced with black and a mask is applied. The procedure for this processing method will be described later with reference to the flowchart in Figure 10.

[0072] Such processed images are generated and added to the training image dataset 201 along with the correct label information. As a result, the training image dataset 201 consists of the original image data and the processed image data from this step. In this embodiment, the degree of attention is reduced by blacking out the image using a mask, but other processing methods such as blurring may be used. Step S70 is performed in the same way as in the first embodiment.

[0073] Figure 9 shows the evaluation of the gaze region excluding the occlusion region, which is performed in step S41 of Figure 8. This figure will be used to explain an example of occlusion determination and overlapping region exclusion. In this figure, half of the body of car (1) on the other side is hidden by car (2) in the foreground. Even though half of the body of car (1) is hidden in this way, the coordinates of the correct label as the object to be recognized are defined by the target region, which is a black rectangle with a size that encloses the entire body of car (1). In this case, occlusion region B is defined as the area where car (1) overlaps with car (2). To determine which car is hidden in occlusion region B, it is determined that the object whose respective base edges are relatively lower relative to the image is the object in the foreground. As a result, it is determined that occlusion region B is hidden by car (2) for car (1), and the appropriateness of the gaze region 91 is evaluated for the target region of car (1) excluding occlusion region B.

[0074] Figure 10 is a flowchart showing the procedure for generating a processed image based on gaze area information performed in step S61 of Figure 8, and Figure 11 is a diagram showing a specific example, illustrating a method for replacing the pixel color of the pixel coordinates in the area outside the gaze area and the pixel color of the pixel coordinates in the overlapping area D where the gaze area evaluation value is greater than or equal to a threshold with black.

[0075] In step S901, the training image dataset, gaze region evaluation values, normative gaze regions A and A', and normative label settings created in steps S10 to S60 in Figure 8 are obtained. In step S902, the pixel coordinates of region C outside the newly defined gaze region A' are obtained from the region to be recognized (see Figure 11(1)). In step S903, a mask is applied by replacing the color of the pixels in region C with black (see Figure 11(3)).

[0076] In step S904, it is determined whether a reference label has been assigned to the gaze region of the image to be recognized. If a reference label has been assigned (NO in S904), the processing is terminated. On the other hand, if a reference label has not been assigned (YES in S904), the process proceeds to S905. In step S905, for the image to be recognized that has a gaze region without a reference label, the newly set gaze region A' and the pixel coordinates of the overlapping region D of the gaze region calculated by the gaze region calculation unit 203 (pixel coordinates of the overlapping region D where the gaze region evaluation value is greater than or equal to the threshold) are obtained (see Figure 11(2)). In step S906, a mask is applied by replacing the pixel color of the obtained pixel coordinates of the overlapping region D with black (see Figure 11(3)). Here, since the overlapping region D is already gazeable and recognized as a gaze region by the gaze region calculation unit 203, the mask is applied to reduce the degree of gaze.

[0077] According to the second embodiment of the present invention described above, the following effects are achieved. Step S41 allows for the determination of occluded recognition objects within the image and the exclusion of the hidden occluded region B from the recognition area of ​​the car (1), thereby enabling the evaluation of an appropriate gaze region 91.

[0078] Step S61 masks the area C outside the gaze region, which is unnecessary as a gaze region, and the overlapping area D, which can already be gazed on because the gaze region evaluation value is above the threshold. By generating an image with areas C and D removed from the image to be recognized and training the AI ​​200, it is possible to efficiently learn information about a new gaze region A' that the AI ​​200 has not yet recognized. In this embodiment, the case in which both the processing in step S41 and the processing in step S61 are performed has been explained as an example, but either one of them may be performed.

[0079] —Third Embodiment— Next, a learning method and learning apparatus for an image recognition processing AI according to a third embodiment of the present invention will be described. In the first and second embodiments, the same normative label settings, normative gaze areas, and processing conditions were used for all groups, but in this case, it may not necessarily lead to improvement for all groups.

[0080] For example, the appropriate processing conditions may differ depending on the size of the object to be recognized. Therefore, we will describe an example in which each group is determined to improve through learning, and the conditions for the groups that have not improved are changed and learning is performed. Note that the hardware configuration of the learning device 1 according to this embodiment is the same as the hardware configuration of the learning device 1 in Figure 1 described in the first embodiment. Accordingly, the following description will explain the learning method of the learning device 1 of this embodiment.

[0081] In the learning device 1 of this embodiment, if the relearning unit 207 determines that the gaze area evaluation value (evaluation of the appropriateness of the gaze area) has not improved due to the relearning of the AI ​​learning model, the norm label setting unit 205 changes the gaze area to which the norm label is assigned. Then, the norm gaze area setting unit 206 sets a new gaze area A' based on the changed norm gaze area A. The relearning unit 207 changes the processing conditions for image processing.

[0082] Figure 12 is a flowchart showing the learning method of the learning device 1 for the image recognition processing AI 200 according to the third embodiment of the present invention. The learning method of the learning device 1 differs from the learning method described in the second embodiment in that it further includes steps S80 to S140.

[0083] Steps S10 to S70 perform the same processing as in the second embodiment. In step S80, the gaze region calculation performed in S30 is performed using the AI ​​learning model of the retrained image recognition processing AI200'. In step S90, the suitability of the gaze region evaluation performed in S40 is performed using the AI ​​learning model of the retrained image recognition processing AI200'.

[0084] Step S100 compares the average gaze region evaluation values ​​of the group before and after relearning. If the pre-relearning value is higher than the post-relearning value, it can be said that the gaze region of this group tended to worsen after relearning. If the post-relearning value is higher than the pre-relearning value, it can be said that the gaze region of this group tended to improve after relearning. The improvement or deterioration is determined for each group, and this information is saved. Note that the method of comparing gaze region evaluation values ​​is not limited to the average value; methods such as determining the median or a baseline may also be used.

[0085] Figure 13 is a graph showing an example of the score distribution before and after learning. For example, if both the pre-training score and the post-training score are low, or both are high, it is judged that there is no improvement. If the post-training score is higher than the pre-training score, it is judged that there is a tendency towards improvement. If the pre-training score is higher than the post-training score, it is judged that there is a tendency towards deterioration.

[0086] Step S110 is the termination condition for relearning, and in this embodiment, the condition is improvement in the evaluation value for all groups. If the evaluation value improves for all groups (YES in S110), this flow is terminated (End). On the other hand, for groups that do not improve (NO in S110), the processing conditions are changed in S61 (S120), the reference attention area is changed in S60 (S130), and the reference label is changed in S50 (S140).

[0087] In step S120, the processing conditions for the group that did not improve are changed. For the group whose focus area was worsened, it is assumed that the processing in the generation of the processed image in S61 had a negative impact, and conditions for no processing are set. Note that the processing conditions are not limited to whether or not processing is performed; for example, the processing intensity or type may also be changed.

[0088] In step S130, the normative attention area of ​​the group that did not improve is modified. Here, the new attention area A' set according to the previous learning conditions is modified. For the new attention area A' of the group whose attention area has worsened, if occlusion due to overlapping recognition targets does not occur, the new attention area A' is considered insufficient, and the new attention area A' is set to cover the entire target area. If occlusion occurs, the occlusion area B hidden by occlusion is excluded using the method performed in step S41, since the system has learned to focus on areas that should not be focused on, and the new attention area A' is set. Note that the modification of the new attention area A' is not limited to any particular method; for example, it may be modified based on judgment by human visual confirmation or based on known information about the characteristics of the target (e.g., the back of a truck).

[0089] In step S140, the normative labels set according to the previous learning conditions are changed for groups that did not improve. For groups where the fixation area worsened, it is considered that the default label setting was inappropriate, and the default label is set for images with a high score distribution in the fixation area, second only to the original default label setting target. Note that the means by which the normative labels are changed are not limited; for example, they may be changed by human visual confirmation.

[0090] According to the third embodiment of the present invention described above, the following effects are achieved. Step S100 compares the gaze region evaluation values ​​of each group before and after relearning. For groups where no improvement is observed, the conditions can be changed to further improve the gaze region. Steps S120, S130, and S140 allow for the modification of processing conditions, norm gaze regions, and norm labels for groups where the gaze region has deteriorated, and relearning is performed to improve the appropriateness of the gaze region and the accuracy of learning. Furthermore, the relearning termination condition set in step S110 allows relearning to continue until the gaze region evaluation values ​​improve for all groups, thereby further enhancing the reliability of the image recognition processing AI. Overall, this embodiment makes the learning process of the image recognition processing AI more efficient and effective, resulting in the provision of an image recognition processing AI with higher accuracy and reliability.

[0091] Although embodiments of the present invention have been described in detail above, the present invention is not limited to the embodiments described above, and various design modifications can be made without departing from the spirit of the invention as described in the claims. For example, the embodiments described above are described in detail in order to explain the present invention in an easy-to-understand manner, and are not necessarily limited to those having all the configurations described. Furthermore, it is possible to replace a part of the configuration of one embodiment with the configuration of another embodiment, and it is also possible to add a configuration of another embodiment to the configuration of one embodiment. Moreover, it is possible to add, delete, or replace a part of the configuration of each embodiment with other configurations. [Explanation of Symbols]

[0092] 1...Learning device, 200, 200'...Image recognition processing AI, 201...Training image dataset, 202...Grouping unit, 203...Looking region calculation unit, 204...Looking region evaluation unit, 205...Reference label setting unit, 206...Reference looking region setting unit, 207...Retraining unit, A...Reference looking region, A'...New looking region, B...Occlusion region, C...Region outside the looking region, D...Overlapping region

Claims

1. A learning device for an AI image recognition process that recognizes objects captured in an image, The learning device is A normative label setting unit determines a normative gaze region from among the gaze regions that the AI ​​learning model of the image recognition processing AI focuses on in multiple images, and assigns a normative label to the normative gaze region. A normative gaze region setting unit sets a new gaze region based on the normative gaze region for other images that have gaze regions to which the normative label has not been assigned, A learning device for image recognition processing AI, characterized by comprising the following:

2. A gaze region calculation unit calculates the gaze region from each of the multiple images using the AI ​​learning model of the image recognition processing AI, It includes a gaze area evaluation unit that evaluates the appropriateness of the gaze area, The learning device for image recognition processing AI according to claim 1, characterized in that the norm label setting unit determines the norm gaze area based on an evaluation of the appropriateness of the gaze area.

3. The learning device for an image recognition processing AI according to claim 2, characterized in that the gaze area evaluation unit evaluates the appropriateness of the gaze area based on at least one of the score distribution, size, location, and distribution of the gaze area.

4. The learning device for an image recognition processing AI according to claim 3, further comprising a retraining unit that compares the gaze region calculated by the gaze region calculation unit with the new gaze region set by the norm gaze region setting unit, and retrains the AI ​​learning model so that the gaze region matches the new gaze region.

5. The learning device for an image recognition processing AI according to claim 4, characterized in that the retraining unit evaluates the learning using a loss function that includes the difference between the gaze region calculated by the gaze region calculation unit and the new gaze region, and updates the weights.

6. The learning device for an image recognition processing AI according to claim 4, characterized in that the retraining unit performs image processing to reduce the degree of gaze of the AI ​​learning model in overlapping regions within the new gaze region that overlap with the gaze region calculated by the gaze region calculation unit.

7. The gaze region evaluation unit evaluates the appropriateness of the gaze region by excluding the occlusion region within the gaze region where the recognized object is partially obscured by another object, if there is a gaze region among the multiple gaze regions in which occlusion occurs. The learning device for image recognition processing AI according to feature 4.

8. The learning device for an image recognition processing AI according to claim 7, characterized in that the retraining unit performs image processing on the occlusion region to reduce the degree of gaze of the AI ​​learning model.

9. The retraining unit performs image processing on the areas outside the gaze region of the other image, excluding the newly gazed region, to reduce the degree of gaze of the AI ​​learning model. The learning device for image recognition processing AI according to feature 4.

10. The system includes a grouping unit that classifies and groups the multiple images according to the characteristics of the recognition target, The learning device for image recognition processing AI according to claim 1, characterized in that the norm label setting unit sets the gaze area with the highest appropriateness rating within each group as the norm gaze area.

11. The aforementioned relearning unit, The learning device for an image recognition processing AI according to claim 6, characterized in that it determines whether the evaluation of the appropriateness of the gaze region has improved by retraining the AI ​​learning model, by comparing the aforementioned gaze region with the new gaze region, and comparing the evaluation of the appropriateness of the gaze region using the AI ​​learning model after retraining the AI ​​learning model with the evaluation of the appropriateness of the gaze region using the AI ​​learning model before retraining.

12. If it is determined that the evaluation of the appropriateness of the gaze area has not improved by the retraining of the AI ​​learning model by the retraining unit, The learning device for image recognition processing AI according to claim 11, characterized in that the norm label setting unit changes the norm gaze area to the gaze area of ​​the other image.

13. If it is determined that the evaluation of the appropriateness of the gaze area has not improved by the retraining of the AI ​​learning model by the retraining unit, The learning device for image recognition processing AI according to claim 11, characterized in that the norm gaze area setting unit changes the norm gaze area to a gaze area to which the norm label is not assigned.

14. If it is determined that the evaluation of the appropriateness of the gaze area has not improved by the retraining of the AI ​​learning model by the retraining unit, The learning device for an image recognition processing AI according to claim 11, characterized in that the retraining unit changes the processing conditions for the image processing.

15. A method for training an AI for image recognition processing that uses a computer to recognize objects that appear in an image, A gaze region calculation step in which the AI ​​learning model of the aforementioned image recognition processing AI calculates gaze regions from multiple images, A gaze area evaluation step for evaluating the appropriateness of the gaze area, Based on the evaluation of the appropriateness of the gaze region, a normative gaze region is determined from among the plurality of gaze regions to serve as a norm for the gaze region that the AI ​​learning model of the image recognition processing AI focuses on, and a normative label is assigned to the normative gaze region in a normative label setting step. A normative gaze region setting step, which sets a new gaze region based on the normative gaze region for other images that have gaze regions to which the normative label has not been assigned, A method for learning an image recognition processing AI, comprising: a retraining step which compares the new gaze region set for the other image by the normative gaze region setting step with the gaze region calculated by the gaze region calculation step before the new gaze region was set for the other image, and retrains the AI ​​learning model so that the gaze region matches the new gaze region.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, and program

    JP2023084981A