Analysis method and analysis device for image recognition processing
The method automates the adjustment of disturbances in image recognition processing to enhance efficiency in analyzing large datasets by optimizing visualization of attention areas, addressing the inefficiencies in existing parameter adjustment processes.
Patent Information
- Application Number
- JP2023223650
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2025-07-10
AI Technical Summary
Existing image recognition technologies for vehicle surroundings analysis require significant human effort and time for parameter adjustment to ensure high-quality visualization of attention areas, especially in complex environments, leading to inefficiencies in analyzing large datasets.
An automated method and apparatus that generate multiple disturbed images by superimposing disturbances on input images, evaluate recognition results, and adjust disturbances based on influence degrees to optimize visualization of attention areas, thereby improving analysis efficiency.
Significantly enhances the analysis efficiency of large-scale datasets by automatically adjusting disturbances for accurate visualization of attention areas in image recognition processing.
Smart Images

Figure 2025105234000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an analysis method and an analysis apparatus for image recognition processing.
Background Art
[0002] Conventionally, in vehicles such as automobiles, a technique is known in which a computer recognizes the surrounding situation of a vehicle based on an image captured by a camera and performs automatic driving of the vehicle using the recognition result. For the surrounding recognition processing for realizing such automatic driving, not only high recognition performance but also strong requirements for reliability and safety are required. Therefore, in recent years, higher recognition performance has been realized by recognizing the surrounding situation of a vehicle using image recognition processing applying AI (artificial intelligence).
[0003] In image recognition processing by AI, the recognition performance varies greatly depending on the learning degree of AI. However, it is difficult to prepare in advance learning data covering all situations that can occur in the actual environment during the automatic driving of a vehicle. Therefore, it is known that there are infinitely many unknown scenes that are not included in the learning data of AI in the actual environment, and recognition errors unintended by the designer may occur. In order to reduce the occurrence of such recognition errors as much as possible and ensure reliability and safety, in image recognition processing for automatic driving, data indicating the surrounding situation of a vehicle in the actual environment is acquired as external data, and by repeatedly performing re-learning of AI using this external data, it is strongly required to continuously evolve AI.
[0004] In order to appropriately evolve the AI for image recognition processing, for example, it is preferable to collect data such as objects that the AI is newly desired to recognize and scenes where recognition errors have occurred in the past as external data, and use the collected external data for the re-learning of the AI. As a result, the AI after re-learning can correctly recognize objects and scenes that the AI before re-learning could not correctly recognize until now. However, among the objects and scenes that the AI before re-learning could correctly recognize, there are many unstable ones that the AI got correct answers accidentally even without confidence, or that cannot be correctly recognized even with a slight environmental change (such as changes in illuminance, chroma, noise, etc.). It is said that the occurrence rate of such unstable recognition results is higher than the probability of recognition errors. That is, in order to realize an AI for image recognition processing with truly high reliability, it can be said that it is necessary to extract more learning data from various data collected so far and the existing data sets in hand, for which the recognition results of the AI become unstable.
[0005] Regarding methods for quantifying the stability, confidence, etc. of image recognition processing by AI, various existing technologies have been studied so far. One of them is a technology for visualizing the inference basis of AI. For example, in Patent Document 1, a technology for visualizing the attention area of AI is disclosed using information on how the result of executing inference by adding disturbance to the input image of AI has changed with respect to the inference result without disturbance. By utilizing the information of this attention area, it becomes possible to extract data for which the recognition of AI becomes unstable.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0007] In many techniques for visualizing the attention area of inference, like the technique of Patent Document 1, the designer visually checks whether the visualized result is of sufficiently high quality. If it is determined that the quality is poor, trial and error is required to reset various parameters used for visualization and perform visualization again. In that case, it is necessary to manually adjust the setting parameters of the disturbance according to the object size of the recognition target and the difficulty of recognition. Therefore, a large amount of human cost and time are required for the analysis of the collected large dataset, and it is difficult to improve the analysis efficiency.
Means for Solving the Problems
[0008] The method for analyzing image recognition processing according to the present invention is a method for analyzing image recognition processing for recognizing an object to be recognized reflected in an input image, and a computer generates a plurality of disturbed images in which disturbances are superimposed on the input image, and the computer obtains evaluation values for the recognition results of the object obtained by respectively executing the image recognition processing on the plurality of disturbed images, and the computer calculates the disturbance influence degree of the disturbance on the image recognition processing based on the evaluation values obtained for each disturbed image, and the computer adjusts the disturbance based on the disturbance influence degree. The analysis apparatus according to the present invention is an apparatus for analyzing image recognition processing for recognizing an object to be recognized reflected in an input image, and includes a disturbance superimposing unit that generates a plurality of disturbed images in which disturbances are superimposed on the input image, an evaluation value obtaining unit that obtains evaluation values for the recognition results of the object obtained by respectively executing the image recognition processing on the plurality of disturbed images generated by the disturbance superimposing unit, a disturbance ratio evaluation unit that calculates the disturbance influence degree of the disturbance on the image recognition processing based on the evaluation values obtained for each disturbed image by the evaluation value obtaining unit, and a parameter adjusting unit that adjusts the disturbance based on the disturbance influence degree calculated by the disturbance ratio evaluation unit.
Effects of the Invention
[0009] According to the present invention, it is possible to automatically adjust the disturbance necessary for visualizing the attention area in inference in image recognition processing, and significantly improve the analysis efficiency of a large-scale dataset.
Brief Description of Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each embodiment is an exemplification for explaining the present invention, and for the sake of clarity of explanation, appropriate omissions and simplifications have been made. The present invention can be implemented in various other forms. Unless otherwise particularly limited, each component may be singular or plural.
[0012] The positions, sizes, shapes, ranges, etc. of the respective components shown in the drawings may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings. When there are a plurality of components having the same or similar functions, they may be described with the same reference numeral and different subscripts. Also, when it is not necessary to distinguish these plurality of components, the description may be made with the omission of subscripts.
[0013] In each embodiment, the processes performed by executing a program may be described. Here, a computer executes a program by a processor (e.g., a CPU or a GPU), and performs the processes defined by the program while using a storage resource (e.g., a memory) and an interface device (e.g., a communication port), etc. Therefore, the entity performing the processes performed by executing the program may be the processor. Similarly, the entity performing the processes performed by executing the program may be a controller, a device, a system, a computer, or a node having a processor. The entity performing the processes performed by executing the program may be an arithmetic unit and may include a dedicated circuit for performing specific processes. Here, the dedicated circuit is, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), a CPLD (Complex Programmable Logic Device), etc.
[0014] The program may be installed in a computer from a program source. The program source may be, for example, a program distribution server or a storage medium readable by a computer. When the program source is a program distribution server, the program distribution server includes a processor and a storage resource for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Also, in an embodiment, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0015] In recent years, for the realization of advanced autonomous driving, the spread of a peripheral recognition application using an image recognition program applying a deep neural network (DNN), which is one type of AI using machine learning, has been progressing. Programs using AI such as DNN have a problem that it is difficult to handle when there is a defect because it is unclear why such an inference was made for the recognition result obtained by the inference of the AI compared to a rule-based algorithm.
[0016] To solve such problems, techniques for visualizing the image regions that have attracted attention in AI inference (such as Grad-CAM and RISE) have been reported. As a specific example, consider a situation where the entire vehicle, which is the object to be recognized in an image, is clearly shown, and the AI inference result for this image accurately infers the position of the vehicle and has a very high inference confidence score. However, when attempting to continuously detect the vehicle, frequent undetected events occur. As a result of visualizing the attention area for the vehicle captured in the image where this event occurred, it may be found that only a part of the vehicle (for example, a part like the windshield) is being attended to in the inference. Such a situation can be said to correspond to conditions that are likely to result in undetected by AI when the relevant part is blocked by other objects or when the relevant part becomes unclear due to the amount of light or noise during shooting.
[0017] As described above, visualizing the attention area in AI inference can be effectively utilized to identify the causes of machine learning defects. However, the structure of the neural network used in DNN and the characteristics of the object to be recognized in the input image (size, orientation, difficulty of recognition, etc.) vary from situation to situation and are not uniform. In visualizing the attention area, the optimal values of the parameters for obtaining high-quality visualization results differ according to these differences. Therefore, it is necessary to visually confirm the quality of the visualized attention area and, in the case of low quality, adjust the parameters and visualize again. That is, when analyzing the AI image recognition results for a large-scale dataset, the man-hours of trial and error required for parameter adjustment become an issue. In the visualization of the attention area in an application where conditions such as the size, orientation, and difficulty of recognition of the object to be recognized are almost constant, such parameter adjustment work is often unnecessary. However, in applications that require recognizing various types of objects in various scenes, such as the external recognition of a vehicle, parameter adjustment is particularly necessary.
[0018] The present invention automatically adjusts parameters when visualizing an attention area in inference by AI such as DNN, thereby efficiently realizing the analysis of the image recognition results of the AI for a large-scale dataset. In the following embodiments, an example in which the present invention is applied to an object recognition DNN in an in-vehicle ECU for vehicle control, for example, an Advanced Driver Assistance System (ADAS) or Autonomous Driving (AD), will be described. However, the present invention is not limited to in-vehicle ECUs for ADAS and AD, and can be applied to all applications using object detection AI.
[0019] - First Embodiment - Hereinafter, with reference to FIGS. 1 to 7, an analysis method and an analysis apparatus for image recognition processing according to the first embodiment of the present invention will be described.
[0020] FIG. 1 is a diagram showing a hardware configuration example of an image recognition processing analysis apparatus (hereinafter simply referred to as "analysis apparatus") 1 according to the first embodiment of the present invention. The analysis apparatus 1 is a computer such as a server or a PC, and a processor 101, a storage device 102, an input device 103, an output device 104, and a communication interface (communication IF) 105 are connected to each other via a communication bus 106.
[0021] The processor 101 controls the operation of the analyzer 1 by executing a predetermined program, and causes a computer such as a server or a PC to function as the analyzer 1. The storage device 102 is a recording medium capable of storing various programs and data non-temporarily or temporarily, and is configured using, for example, a ROM (Read Only Memory), a RAM (Random Access Memory), an HDD (Hard Disk Drive), a flash memory, or the like. The input device 103 is a device that receives data input from the user to the analyzer 1, and is configured using, for example, a keyboard, a mouse, a touch panel, a microphone, or the like. The output device 104 is a device that outputs data to the user, and is configured using, for example, a display, a printer, a speaker, or the like. The communication IF 105 transmits and receives data between the analyzer 1 and an external information device (not shown) via a communication network (not shown).
[0022] FIG. 2 is a functional block diagram showing a functional configuration example of the analyzer 1 according to the first embodiment of the present invention. The analyzer 1 includes, as functions for performing image recognition processing by AI and analyzing its inference results, an input image acquisition unit 11, a teacher data acquisition unit 12, a disturbance setting value storage unit 13, a disturbance generation unit 14, a disturbance superposition unit 15, an image recognition unit 16, an evaluation value acquisition unit 17, a target information generation unit 18, a target information integration unit 19, a drawing unit 20, a disturbance ratio evaluation unit 21, a parameter adjustment unit 22, and a GUI (Graphical User Interface) 23.
[0023] Specifically, the input image acquisition unit 11, the teacher data acquisition unit 12, the disturbance generation unit 14, the disturbance superposition unit 15, the image recognition unit 16, the evaluation value acquisition unit 17, the target information generation unit 18, the target information integration unit 19, the drawing unit 20, the disturbance ratio evaluation unit 21, the parameter adjustment unit 22, and the GUI 23 are realized by causing the processor 101 to execute a program stored in the storage device 102 shown in FIG. 1, for example. Specifically, the disturbance setting value storage unit 13 is realized by the storage device 102 shown in FIG. 2, for example.
[0024] The input image acquisition unit 11 acquires an input image that is the target of the image recognition process performed by the analysis device 1. The teacher data acquisition unit 12 acquires a correct label indicating the correct answer of the recognition result for the input image as the teacher data of the input image. The input image acquisition unit 11 and the teacher data acquisition unit 12 can acquire the input image and the correct label respectively based on the information stored in the storage device 102 or the information received via the communication network by the communication IF 105.
[0025] The disturbance generation unit 14 generates a disturbance that becomes noise when performing the image recognition process on the input image. In the disturbance setting value storage unit 13, setting values of various parameters when the disturbance generation unit 14 generates a disturbance are stored. For example, information such as the type, intensity and resolution of the disturbance, the analysis method of the disturbance, and the threshold value used for the evaluation of the disturbance is stored in the disturbance setting value storage unit 13 as the setting value of the parameters used for the generation of the disturbance. The disturbance generation unit 14 generates a plurality of disturbances with different patterns according to the values of the parameters stored in the disturbance setting value storage unit 13, and outputs them to the disturbance superimposing unit 15 and the attention information generation unit 18.
[0026] The disturbance superimposing unit 15 generates a disturbed image in which the disturbance generated by the disturbance generation unit 14 is superimposed on the input image acquired by the input image acquisition unit 11. Here, as described above, in the disturbance generation unit 14, a plurality of disturbances with different patterns are generated. The disturbance superimposing unit 15 can generate a plurality of disturbed images for the same input image by superimposing disturbances with different patterns on the input image respectively.
[0027] The image recognition unit 16 executes image recognition processes using the learned DNN on the plurality of disturbed images generated by the disturbance superimposing unit 15 respectively. In the image recognition process executed by the image recognition unit 16, among various objects reflected in the input image, objects to be recognized in ADAS or AD, such as surrounding vehicles, obstacles, and various road signs, are recognized. Note that when the image recognition unit 16 executes the image recognition process using the DNN, it is possible to utilize well-known arithmetic processes.
[0028] The evaluation value acquisition unit 17 acquires an evaluation value for the recognition result of the object obtained by the image recognition unit 16 performing image recognition processing on each disturbance image. For example, the evaluation value acquisition unit 17 compares the recognition result of the object for the disturbance image with the correct label acquired by the teacher data acquisition unit 12, and sets the evaluation value for the recognition result of the object for each disturbance image such that the higher the degree of coincidence, the higher the evaluation value. Details of the method for acquiring the evaluation value by the evaluation value acquisition unit 17 will be described later.
[0029] The attention information generation unit 18 generates attention information regarding the degree of attention in the image recognition process when the image recognition unit 16 recognizes an object in the disturbance image, by applying the disturbance generated by the disturbance generation unit 14 to the evaluation value acquired for each disturbance image by the evaluation value acquisition unit 17. Details of the method for generating the attention information by the attention information generation unit 18 will be described later.
[0030] The attention information integration unit 19 integrates the attention information for each disturbance image generated by the attention information generation unit 18 for the plurality of disturbance images generated by the disturbance superimposition unit 15. By this process, it is possible to integrate the attention information obtained for the plurality of disturbance images and generate information for visualizing the attention region that was focused on during inference in the image recognition process for the input image.
[0031] The drawing unit 20 draws the information on the attention region obtained by the attention information integration unit 19 integrating the attention information, superimposed on the input image acquired by the input image acquisition unit 11. Thereby, an output image in which the position of the attention region in the input image is shown in a map form is generated. The output image generated by the drawing unit 20 is presented to the user via the GUI 23. The user can visually confirm the degree to which the disturbance affects the result of the image recognition process for the input image by viewing this output image.
[0032] The external disturbance ratio evaluation unit 21 calculates an external disturbance influence degree indicating the degree of influence of external disturbances on the image recognition process based on the evaluation values obtained for each external disturbance image by the evaluation value acquisition unit 17. Then, based on the calculated external disturbance influence degree, it evaluates the ratio of the area where the external disturbance generated by the external disturbance generation unit 14 hides the input image (hereinafter referred to as the "external disturbance ratio"), and notifies the evaluation result to the parameter adjustment unit 22.
[0033] The external disturbance ratio evaluation unit 21 includes an aggregation unit 211, a statistical processing unit 212, and a determination unit 213. The aggregation unit 211 aggregates the evaluation values for each external disturbance image obtained by the evaluation value acquisition unit 17. The statistical processing unit 212 statistically processes the evaluation values aggregated by the aggregation unit 211 to obtain the distribution of the evaluation values in a plurality of external disturbance images. The determination unit 213 calculates the external disturbance influence degree based on the distribution of the evaluation values obtained by the statistical processing unit 212, and determines whether the external disturbance ratio is appropriate by comparing this external disturbance influence degree with a predetermined determination threshold. Then, the obtained determination result is notified to the parameter adjustment unit 22 as the evaluation result of the external disturbance ratio. Note that the determination threshold used for comparison with the external disturbance influence degree in the determination unit 213 can be set by the user via the GUI 23.
[0034] The parameter adjustment unit 22 adjusts the parameters stored in the external disturbance setting value storage unit 13 so that the external disturbance ratio becomes an appropriate value based on the evaluation result of the external disturbance ratio by the external disturbance ratio evaluation unit 21. The specific method of parameter adjustment by the parameter adjustment unit 22 will be described later.
[0035] Figure 3 is a flowchart showing an analysis method of the image recognition process according to the first embodiment of the present invention. In this embodiment, the analysis device 1 can analyze the image recognition result by the image recognition unit 16 by executing the processing shown in the flowchart of Figure 3 by the processor 101.
[0036] Steps S10 to S50 are for initial settings. In step S10, a test image dataset is acquired. Here, the input image acquisition unit 11 and the teacher data acquisition unit 12 can acquire, as a test image dataset, a combination of an arbitrary input image including an object to be recognized in the image recognition process executed by the image recognition unit 16 and the correct label, for example, within a range of several to several tens of sets.
[0037] In step S20, initial parameters of the disturbance are specified. Here, for various parameters such as the type, intensity, and resolution of the disturbance, the analysis method of the disturbance, and the threshold value used for the evaluation of the disturbance, the values input by the user via the GUI 23 are stored in the disturbance setting value storage unit 13 as the initial parameters.
[0038] In step S30, a disturbance evaluation process is executed using the test image dataset acquired in step S10. Here, the disturbance superimposing unit 15 superimposes the disturbance generated by the disturbance generation unit 14 on the input image to generate a disturbed image including noise, and the image recognition unit 16 executes an image recognition process on this disturbed image. Then, the evaluation value acquisition unit 17 acquires an evaluation value for the recognition result of the object in the image recognition process. Details of this disturbance evaluation process will be described later with reference to the flowchart of FIG. 4.
[0039] In step S40, the evaluation result of the disturbance by the disturbance evaluation process executed in step S30 is output. Here, for example, the evaluation value acquired in the disturbance evaluation process is superimposed on the input image and presented to the user via the GUI 23, so that the evaluation result of the disturbance by the set initial parameters is visualized and confirmed by the user.
[0040] In step S50, it is determined whether the evaluation result of the disturbance output in step S40 is appropriate. Here, for example, based on the content of the input operation performed by the user via the GUI23 on the evaluation result of the disturbance, it is determined whether the evaluation result of the disturbance is appropriate. As a result, if it is determined that the evaluation result of the disturbance is appropriate, the process proceeds to step S60. On the other hand, if it is determined that the evaluation result of the disturbance is not appropriate, the process returns to step S20. After resetting the initial parameters of the disturbance, the disturbance evaluation process in step S30 is executed again, and it is determined whether the evaluation result is appropriate. Thereby, until an appropriate disturbance evaluation result is obtained for the test image data set, the setting of the initial parameters of the disturbance is repeatedly performed, and the value of the finally set initial parameters is used as the parameters for generating the disturbance in the subsequent processing.
[0041] Steps S60 to S90 are processes for adjusting the disturbance. In step S60, an image data set to be analyzed is acquired. Here, similar to step S10, the input image acquisition unit 11 and the teacher data acquisition unit 12 can acquire, as the image data set to be analyzed, a combination of an arbitrary input image including an object to be recognized in the image recognition process executed by the image recognition unit 16 and the correct label. At this time, the same combination of the input image and the correct label as in step S10 may be acquired as the image data set to be analyzed, or a different combination may be acquired. Also, more image data sets than the test image data set may be acquired as the image data set to be analyzed, or fewer may be acquired. However, in order to perform learning of the image recognition unit 16 using a large amount of image data, it is preferable to acquire more image data sets than the test image data set as the image data set to be analyzed.
[0042] In step S70, using the image data set to be analyzed acquired in step S60, the same disturbance evaluation process as in step S30 is executed.
[0043] In step S80, disturbance ratio adjustment processing is executed based on the evaluation result of the disturbance obtained by the disturbance evaluation processing executed in step S70. Here, the evaluation values for each disturbance image acquired in the disturbance evaluation processing are aggregated by the disturbance ratio evaluation unit 21 and statistically processed to calculate the degree of disturbance influence that the disturbance has on the image recognition processing. Then, it is determined whether the disturbance ratio is appropriate based on the calculated degree of disturbance influence, and the determination result is output to the parameter adjustment unit 22. In the parameter adjustment unit 22, based on the determination result of the disturbance ratio input from the disturbance ratio evaluation unit 21, among the disturbance parameters stored in the disturbance setting value storage unit 13, the value of the parameter related to the disturbance ratio is adjusted. The details of this disturbance ratio adjustment processing will be described later with reference to the flowchart of FIG. 5.
[0044] In step S90, it is determined whether the adjustment of the disturbance parameters is completed based on the determination result of the disturbance ratio obtained by the disturbance ratio adjustment processing executed in step S80. Here, if in the immediately preceding disturbance ratio adjustment processing, a determination result indicating that the disturbance ratio is appropriate is obtained and thus the adjustment of the parameter value is not performed, it is determined that the adjustment of the disturbance parameters is completed, and the processing shown in the flowchart of FIG. 3 is completed. On the other hand, if in the immediately preceding disturbance ratio adjustment processing, a determination result indicating that the disturbance ratio is not appropriate is obtained and thus the adjustment of the parameter value is performed, it is determined that the adjustment of the disturbance parameters is not completed and the process returns to step S70. In this case, using the adjusted parameter value, the disturbance evaluation processing in step S70 and the disturbance ratio adjustment processing in step S80 are executed again. Thereby, the adjustment of the disturbance parameters is repeated until an appropriate disturbance ratio is obtained for the image data set to be analyzed.
[0045] FIG. 4 is a flowchart showing the procedure of the disturbance evaluation processing executed in steps S30 and S70 of FIG. 3.
[0046] In step S101, the disturbance generation unit 14 generates a disturbance for the input image. Here, based on the parameters stored in the disturbance setting value storage unit 13, a disturbance is generated that randomly masks the images of each region obtained by dividing the input image into predetermined ranges. Specifically, for example, for each region obtained by dividing the input image into predetermined ranges, a disturbance value of "0" or "1" is randomly set. The size of the region at this time, the appearance ratio of "0" and "1" in the disturbance value, etc. are determined according to the parameter values stored in the disturbance setting value storage unit 13.
[0047] In step S102, the disturbance superimposing unit 15 synthesizes the disturbance generated in step S101 onto the input image to generate a disturbed image. Here, for each sub-image corresponding to each region of the input image, the disturbance value of "0" or "1" set in step S101 is multiplied respectively. As a result, in each region where the disturbance value is "0", the sub-image corresponding to that part in the input image is hidden by the disturbance, while in each region where the set disturbance value is "1", the sub-image corresponding to that part in the input image is not hidden by the disturbance, and thus a disturbed image with the disturbance superimposed on the input image can be generated.
[0048] In step S103, the image recognition unit 16 makes an inference on the disturbed image generated in step S102. In this inference, by executing an image recognition process using a learned DNN, the object to be recognized included in the input image before the disturbance is superimposed is extracted, and the type of the object and the range of the region occupied by the object in the image are recognized. At this time, the certainty of the obtained inference result may be obtained. Also, the image recognition unit 16 may be provided outside the analyzer 1, and the result of the inference executed in this image recognition unit 16 may be acquired by the analyzer 1 in step S103.
[0049] In step S104, the evaluation value acquisition unit 17 calculates an evaluation value for the inference result of step S103. For example, the object region obtained by inference is compared with the object region indicated by the correct label, and the degree of overlap (IoU: Intersection over Union) between them is calculated as the evaluation value (score) for the inference result. In this case, if the inference result and the correct label completely do not match, the evaluation value becomes 0, and if they completely match, the evaluation value becomes 1. At this time, instead of the correct label, the inference result for the input image without disturbance may be obtained, and this may be compared with the inference result of step S103 to calculate the evaluation value. Also, the evaluation value may be calculated in consideration of the recognition class (type of object) in the inference result. For example, when the recognition class in the inference matches the type of object indicated by the correct label, the value of IoU is used as the evaluation value, and when they are different, the evaluation value is set to 0 regardless of the value of IoU. Alternatively, even if the recognition class and the correct label are different but in the same category, the evaluation value is adjusted by multiplying the value of IoU by a predetermined penalty rate (for example, 0.5). In addition to this, the evaluation value for the inference result of step S103 can be calculated by any arbitrary method.
[0050] In step S105, the attention information generation unit 18 generates attention information regarding the degree of attention in the inference in step S103 based on the evaluation value calculated in step S104. Specifically, for example, the attention information can be generated by calculating the product of the disturbance superimposed on the input image by the disturbance superimposing unit 15 in step S102 and the evaluation value calculated in step S104 for each partial image corresponding to each region of the input image on which the disturbance is set. In this attention information, for each partial image corresponding to each region in the disturbance image where the disturbance value is "1", that is, each region not hidden by the disturbance, the magnitude of the evaluation value for the inference is represented. That is, the attention information represents, together with the similarity of the inference results, the regions of the input image for which it is possible to obtain the same or similar inference results by using the images of the parts not hidden by the disturbance regardless of the presence or absence of the disturbance. Such information corresponds to the degree of attention in the inference, and it can be said that the higher the value of the attention information (similarity of the inference results), the higher the degree of attention in the inference. Conversely, an image region with a small value of the attention information indicates that the influence of the part hidden by the disturbance is large, and even if the image recognition unit 16 looks at the partial image within that region when performing inference on the disturbance image, it cannot reproduce the same inference result as the original input image without the superimposed disturbance.
[0051] In step S106, the attention information integration unit 19 further integrates the attention information generated in step S105 with the integration result of the attention information obtained in the previous processing.
[0052] In step S107, it is determined whether or not the processing of steps S101 to S106 has been completed a predetermined number of times (for example, N times). If the number of executions of the processing of steps S101 to S106 is less than N times, the process returns to step S101. After regenerating the disturbance in step S101, the processing after step S102 is performed to repeat the integration of the attention information. On the other hand, if the number of executions of the processing of steps S101 to S106 reaches N times, the disturbance evaluation processing shown in the flowchart of FIG. 4 is terminated.
[0053] In the disturbance evaluation process of steps S30 and S70 in FIG. 3, as described above, the processes of steps S101 to S107 are repeatedly executed N times. As a result, the target information is accumulated each time the number of executions is increased. Since the disturbance superimposed on the input image has a high degree of randomness, in the integration result of the target information after N executions, the value of the degree of attention of the part corresponding to the target area becomes large, and the values of the other parts become relatively small. Therefore, the target area can be represented by the integration result of the target information.
[0054] FIG. 5 is a flowchart showing the procedure of the disturbance ratio adjustment process executed in step S80 of FIG. 3.
[0055] In step S201, the aggregation unit 211 aggregates the N evaluation values (scores) calculated by the evaluation value acquisition unit 17 in the disturbance evaluation process of step S70 in FIG. 3. Here, for example, the number of occurrences of each evaluation value in the calculation results for N times is counted respectively, and the count values for each obtained evaluation value are histogrammed. As a result, for the N disturbed images obtained by randomly superimposing disturbances on the input image, the distribution of the evaluation values for the inference results of the image recognition unit 16 can be obtained.
[0056] In step S202, the statistical processing unit 212 calculates a statistical value according to the distribution of the evaluation values aggregated in step S201. Here, for the aggregation result by evaluation value based on the calculation results for N times, by calculating the average value Sa and the variance Sv, a statistical value according to the distribution of the evaluation values can be calculated. Note that the statistical value calculated in step S202 represents the magnitude of the influence of the random disturbance superimposed on the input image on the image recognition process, that is, the above-mentioned disturbance influence degree. In other words, in step S202, the statistical information in the distribution of the evaluation values for the inference results of the image recognition unit 16 is calculated as the disturbance influence degree.
[0057] In step S203, the determination unit 213 determines the degree of disturbance influence based on the statistical value calculated in step S202. Here, for example, using a determination threshold value previously set by the user via the GUI 23, it is determined whether the degree of disturbance influence is appropriate, or too large or too small, and the obtained determination result is output to the parameter adjustment unit 22. Note that the method for determining the degree of disturbance influence in step S203 will be described later with reference to FIGS. 6 and 7.
[0058] In step S204, based on the determination result of the degree of disturbance influence input from the determination unit 213 in step S203, the parameter adjustment unit 22 determines which of "appropriate", "small influence degree", or "large influence degree" the degree of disturbance influence corresponds to. Then, if it is determined that the degree of disturbance influence is "appropriate", it proceeds to step S205; if it is determined that the influence degree is "small", that is, if the degree of disturbance influence is too small, it proceeds to step S206; and if it is determined that the influence degree is "large", that is, if the degree of disturbance influence is too large, it proceeds to step S207 respectively.
[0059] In step S205, the parameter adjustment unit 22 determines that the parameter adjustment is completed and ends the disturbance ratio adjustment process shown in the flowchart of FIG. 5. In this case, the parameter adjustment is not performed. Therefore, in the determination process of step S90 executed following the disturbance ratio adjustment process of step S80 in the flowchart of FIG. 3, it is determined that the adjustment of the disturbance parameter is completed.
[0060] In step S206, the parameter adjustment unit 22 adjusts the parameter value so as to increase the strength of the disturbance by increasing the disturbance ratio with respect to the parameter stored in the disturbance setting value storage unit 13 compared to the current value. As a result, when it is considered that the degree of disturbance influence is too small with the current parameter value and the influence on the inference result is limited even when the disturbance is superimposed on the input image, the determination result of the degree of disturbance influence obtained in step S203 is fed back, and the parameter adjustment is performed so as to automatically increase the ratio of the disturbance superimposed on the input image.
[0061] In step S207, the parameter adjustment unit 22 adjusts the parameter value so as to weaken the intensity of the disturbance by reducing the disturbance ratio with respect to the parameter stored in the disturbance setting value storage unit 13 compared to the current state. Thereby, when the disturbance influence degree is too large with the current parameter value and it becomes difficult to obtain a correct inference result when the disturbance is superimposed on the input image, the determination result of the disturbance influence degree obtained in step S203 is fed back, and parameter adjustment is performed so as to automatically reduce the disturbance ratio of the disturbance superimposed on the input image.
[0062] In step S206 or S207, when the parameter of the disturbance setting value storage unit 13 is adjusted so that the disturbance ratio changes according to the disturbance influence degree, the disturbance ratio adjustment process shown in the flowchart of FIG. 5 is terminated.
[0063] FIG. 6 is a flowchart showing an example of a method for determining the disturbance influence degree performed in step S203 of FIG. 5.
[0064] In step S301, the average value Sa among the statistical values calculated in step S202 of FIG. 5 is compared with an upper threshold value Sa_thu of the preset average value. As a result, if the average value Sa is greater than the upper threshold value Sa_thu, the process proceeds to step S304, and if it is less than or equal to the upper threshold value Sa_thu, the process proceeds to step S302.
[0065] In step S302, the average value Sa is compared with a lower threshold value Sa_thl of the preset average value. As a result, if the average value Sa is less than the lower threshold value Sa_thl, the process proceeds to step S303, and if it is greater than or equal to the lower threshold value Sa_thl, it is determined that the disturbance influence degree is appropriate and the determination of the disturbance influence degree is terminated.
[0066] In step S303, the variance Sv among the statistical values calculated in step S202 of FIG. 5 is compared with a preset threshold Sv_th. As a result, if the variance Sv is less than the threshold Sv_th, it is determined that the disturbance influence degree is too large and the evaluation value is biased toward the low evaluation side, and the determination of the disturbance influence degree is terminated. On the other hand, if the variance Sv is greater than or equal to the threshold Sv_th, it is determined that the disturbance influence degree is appropriate and the determination of the disturbance influence degree is terminated.
[0067] In step S304, the variance Sv is compared with the threshold Sv_th. As a result, if the variance Sv is less than the threshold Sv_th, it is determined that the disturbance influence degree is too small and the evaluation value is biased toward the high evaluation side, and the determination of the disturbance influence degree is terminated. On the other hand, if the variance Sv is greater than or equal to the threshold Sv_th, it is determined that the disturbance influence degree is appropriate and the determination of the disturbance influence degree is terminated.
[0068] FIG. 7 is a diagram showing an example of the distribution of evaluation values (scores) for each disturbance influence degree. In FIG. 7, case (1) shown on the left side shows an example of the distribution of evaluation values (scores) in the target information obtained by repeating the processes of steps S101 to S107 of FIG. 4 N times when the disturbance influence degree is too large, and an example of the visualization result of the target region output in step S40 of FIG. 3. In this case, a correct inference result cannot be obtained due to the disturbance, and thus it can be seen that the evaluation value is concentrated at 0 in the target information and the target region cannot be visualized. Therefore, in the process of FIG. 6, in steps S301 and S302, it is determined that the average value Sa is less than the lower threshold Sa_thl, and in the subsequent step S303, it is determined that the variance Sv is less than the threshold Sv_th, so that it is determined that the disturbance influence degree is too large. As a result, in step S207 of FIG. 5, the parameter value is adjusted to reduce the disturbance ratio.
[0069] In FIG. 7, the case (2) shown in the center shows a distribution example of evaluation values (scores) in the attention information obtained by repeating the processes of steps S101 to S107 in FIG. 4 N times when the disturbance influence degree is appropriate, and an example of the visualization result of the attention area output in step S40 of FIG. 3. In this case, when a specific attention area (for example, a person) in the input image is not hidden by the disturbance, a correct inference result can be obtained. On the other hand, when the attention area is hidden, a correct inference result cannot be obtained. Therefore, in the attention information, the evaluation values are moderately dispersed, and it can be seen that the attention area is easily visible and visualized. Therefore, in the process of FIG. 6, in steps S301 and S302, it is determined that the average value Sa is equal to or greater than the lower threshold Sa_thl and equal to or less than the upper threshold Sa_thu, so it is determined that the disturbance influence degree is appropriate. As a result, in step S205 of FIG. 5, it is determined that the parameter adjustment is completed, and the adjustment of the parameter value is not performed.
[0070] In FIG. 7, the case (3) shown on the right shows a distribution example of evaluation values (scores) in the attention information obtained by repeating the processes of steps S101 to S107 in FIG. 4 N times when the disturbance influence degree is too small, and an example of the visualization result of the attention area output in step S40 of FIG. 3. In this case, even if the disturbance is superimposed on the input image, the inference result hardly changes. Therefore, it is not known where the attention area exists in the input image, and it can be seen that the visualization quality of the attention area has deteriorated. Therefore, in the process of FIG. 6, in step S301, it is determined that the average value Sa is greater than the upper threshold Sa_thu, and in the subsequent step S304, it is determined that the variance Sv is less than the threshold Sv_th, so it is determined that the disturbance influence degree is too small. As a result, in step S206 of FIG. 5, the parameter value is adjusted to increase the disturbance ratio.
[0071] According to the first embodiment of the present invention described above, the following operational effects are achieved.
[0072] (1) The analysis method of the image recognition process by the analysis device 1 is a method for analyzing the image recognition process to recognize the object to be recognized reflected in the input image. In this analysis method, a plurality of disturbed images with disturbances superimposed on the input image are generated by the processor 101 which is a computer (step S102), and evaluation values for the recognition results of the objects obtained by respectively executing the image recognition process on the plurality of disturbed images are acquired (step S104). Then, based on the evaluation values acquired for each disturbed image, the disturbance influence degree of the disturbance on the image recognition process is calculated (step S202), and the disturbance is adjusted based on the calculated disturbance influence degree (steps S206, S207). By doing so, since the disturbance necessary for visualizing the attention area in the inference of the image recognition process can be automatically adjusted, the analysis efficiency of a large-scale dataset can be significantly improved.
[0073] (2) In step S202, the disturbance influence degree is calculated based on the distribution of the evaluation values acquired for each disturbed image. Specifically, statistical information (average value Sa, variance Sv) in the distribution of the evaluation values acquired for each disturbed image is calculated, and this statistical information is used as the disturbance influence degree. By doing so, the disturbance influence degree of the disturbance on the image recognition process can be appropriately and quantitatively calculated.
[0074] (3) In the adjustment of the disturbance, when the distribution of the evaluation values is biased towards the high evaluation side (steps S301, S304: Yes), the intensity of the disturbance is adjusted to be increased (step S206), and when the distribution of the evaluation values is biased towards the low evaluation side (steps S302, S303: Yes), the intensity of the disturbance is adjusted to be decreased (step S207). By doing so, the intensity of the disturbance can be automatically adjusted to an optimal value according to the distribution of the evaluation values.
[0075] (4) The analysis device 1 is a device that analyzes image recognition processing for recognizing an object to be recognized reflected in an input image. The analysis device 1 includes a disturbance superimposing unit 15 that generates a plurality of disturbed images obtained by superimposing a disturbance on an input image, an evaluation value acquisition unit 17 that acquires an evaluation value for the recognition result of an object obtained by the image recognition unit 16 executing image recognition processing on each of the plurality of disturbed images generated by the disturbance superimposing unit 15, a disturbance ratio evaluation unit 21 that calculates the degree of disturbance influence that the disturbance exerts on the image recognition processing based on the evaluation values acquired for each disturbed image by the evaluation value acquisition unit 17, and a parameter adjustment unit 22 that adjusts the disturbance based on the degree of disturbance influence calculated by the disturbance ratio evaluation unit 21. By doing so, by using the analysis device 1, it is possible to automatically adjust the disturbance necessary for visualizing the attention area in the inference in the image recognition processing, so that the analysis efficiency of a large-scale dataset can be significantly improved.
[0076] - Second Embodiment - Next, an analysis method and an analysis device for image recognition processing according to the second embodiment of the present invention will be described. In the first embodiment, an example of adjusting the disturbance ratio among the disturbance parameters was described, but in this embodiment, in addition to this, an example of adjusting the resolution of the disturbance will be described. Note that the hardware configuration of the analysis device 1A according to this embodiment is the same as the hardware configuration of FIG. 1 described in the first embodiment. Therefore, in the following description, the analysis device 1A of this embodiment will be described using the hardware configuration of FIG. 1.
[0077] FIG. 8 is a functional block diagram showing a functional configuration example of the analysis device 1A according to the second embodiment of the present invention. The analysis device 1A is different from the analysis device 1 described in the first embodiment in that it further has a disturbance resolution evaluation unit 24.
[0078] The external disturbance resolution evaluation unit 24 includes a target area distribution analysis unit 241 and a determination unit 242. The target area distribution analysis unit 241 analyzes the distribution of the target areas that were focused on in the image recognition process when the image recognition unit 16 recognized an object in the external disturbance image, based on the information of the target areas obtained by the target information integration unit 19 integrating the target information. The determination unit 242 determines whether the resolution of the external disturbance is appropriate based on the analysis result of the target area distribution by the target area distribution analysis unit 241. Then, the obtained determination result is notified to the parameter adjustment unit 22 as the evaluation result of the external disturbance resolution.
[0079] In addition to adjusting the external disturbance ratio described in the first embodiment, the parameter adjustment unit 22 adjusts the external disturbance resolution. That is, based on the evaluation result of the external disturbance resolution by the external disturbance resolution evaluation unit 24, the parameters stored in the external disturbance setting value storage unit 13 are adjusted so that the external disturbance resolution becomes an appropriate value.
[0080] FIG. 9 is a flowchart showing an analysis method of the image recognition process according to the second embodiment of the present invention. In this embodiment, the analysis device 1A can analyze the image recognition result by the image recognition unit 16 by executing the process shown in the flowchart of FIG. 9 by the processor 101.
[0081] In steps S10 to S80, the same processes as those in the first embodiment are respectively performed. After performing the process of step S80, in subsequent step S81, a disturbance resolution adjustment process based on the evaluation result of the disturbance obtained by the disturbance evaluation process executed in step S70 is executed. Here, using the information of the target area generated in the disturbance evaluation process, that is, the integration result of the target information by the target information integration unit 19, the distribution of the target area is analyzed by the disturbance resolution evaluation unit 24. Then, it is determined whether the disturbance resolution is appropriate based on the obtained analysis result, and the determination result is output to the parameter adjustment unit 22. In the parameter adjustment unit 22, based on the determination result of the disturbance resolution input from the disturbance resolution evaluation unit 24, among the disturbance parameters stored in the disturbance setting value storage unit 13, the value of the parameter related to the disturbance resolution is adjusted. The details of this disturbance resolution adjustment process will be described later with reference to the flowchart of FIG. 10.
[0082] In step S90, based on the determination result of the disturbance ratio obtained by the disturbance ratio adjustment process executed in step S80 and the determination result of the disturbance resolution obtained by the disturbance resolution adjustment process executed in step S81, it is determined whether the adjustment of the disturbance parameters is completed. Here, in the immediately preceding disturbance ratio adjustment process and disturbance resolution adjustment process, determination results indicating that the disturbance ratio and the disturbance resolution are appropriate have been respectively obtained. Therefore, if the parameter value adjustment has not been performed, it is determined that the adjustment of the disturbance parameters is completed, and the process shown in the flowchart of FIG. 9 is completed. On the other hand, if a determination result indicating that the disturbance ratio is not appropriate is obtained in the immediately preceding disturbance ratio adjustment process, or if a determination result indicating that the disturbance resolution is not appropriate is obtained in the immediately preceding disturbance resolution adjustment process, and thus the parameter value adjustment has been performed, it is determined that the adjustment of the disturbance parameters is incomplete, and the process returns to step S70. In this case, using the adjusted parameter value, the disturbance evaluation process in step S70, the disturbance ratio adjustment process in step S80, and the disturbance resolution adjustment process in step S81 are executed again. Thereby, the adjustment of the disturbance parameters is repeatedly performed until an appropriate disturbance ratio and disturbance resolution are obtained for the image data set to be analyzed.
[0083] Figure 10 is a flowchart showing the procedure of the disturbance resolution adjustment process executed in step S81 of FIG. 9.
[0084] In step S401, the attention area distribution analysis unit 241 obtains the distribution of the visualization result of the attention area generated by the attention information integration unit 19 integrating the attention information for N times in the disturbance evaluation process of step S70 in FIG. 9, respectively in the x direction (the horizontal direction of the input image) and the y direction (the vertical direction of the input image). When a predetermined condition is satisfied, for example, when the size ratio in the x direction and the y direction of the object area indicated by the correct label acquired by the teacher data acquisition unit 12 is equal to or greater than a predetermined value (for example, 2:1), the distribution of the visualization result of the attention area may be obtained only for the direction with the larger size. In that case, the processes after step S402 may be implemented only for that direction.
[0085] In step S402, the determination unit 242 determines whether the distribution of the attention area acquired in step S401 exceeds the object area indicated by the correct label acquired by the teacher data acquisition unit 12. Here, for example, when the value of the attention area is normalized in the range of 0 to 1, if the portion having a value of 0.8 or more in the distribution of the attention area acquired in step S401 is within the object area, it is determined that the distribution of the attention area does not exceed the object area, and the process proceeds to step S403. On the other hand, when the portion protrudes from the object area, it is determined that the distribution of the attention area exceeds the object area, and the process proceeds to step S406.
[0086] In step S403, the attention area distribution analysis unit 241 obtains the spatial frequency of the distribution of the attention area acquired in step S401. Here, for example, for each of the x direction and the y direction, the spatial frequency of the distribution of the attention area can be obtained by obtaining the number of peaks of the distribution of the attention area within the object area.
[0087] In step S404, the determination unit 242 determines whether the spatial frequency of the distribution of the target region acquired in step S403 is appropriate. Here, for example, as described above, when the number of peaks in the distribution of the target region is acquired as the spatial frequency, if the number of peaks is equal to or greater than a predetermined number (for example, 2), it is determined that the spatial frequency of the distribution of the target region is appropriate, and the process proceeds to step S405. At this time, when the number of peaks is equal to or greater than the predetermined number in both the x-direction and the y-direction, it may be determined that the spatial frequency of the distribution of the target region is appropriate, or if the number of peaks is equal to or greater than the predetermined number in at least either one, it may be determined that the spatial frequency of the distribution of the target region is appropriate. On the other hand, if the number of peaks is less than the predetermined number, it is determined that the spatial frequency of the distribution of the target region is insufficient, and the process proceeds to step S406.
[0088] In step S405, the parameter adjustment unit 22 determines that the parameter adjustment is completed and ends the disturbance resolution adjustment process shown in the flowchart of FIG. 10. In this case, assuming that the disturbance resolution is appropriate, the parameter adjustment is not performed.
[0089] In step S406, the parameter adjustment unit 22 adjusts the parameter value of the parameter stored in the disturbance setting value storage unit 13 so as to increase the resolution of the disturbance more than the current level. Thereby, the determination results for the disturbance resolution performed in steps S402 and S404 are fed back, and the parameter adjustment is performed so as to automatically increase the resolution of the disturbance superimposed on the input image.
[0090] In step S406, when the parameter of the disturbance setting value storage unit 13 is adjusted to increase the insufficient disturbance resolution, the disturbance resolution adjustment process shown in the flowchart of FIG. 10 ends.
[0091] FIG. 11 is a diagram showing an example of the visualization result and distribution of the region of interest for each disturbance resolution. In FIG. 11, in the upper column, an example of visualizing the region of interest obtained when the disturbance resolution is low and an example of the distribution of this region of interest in the y direction are shown. In this case, it can be seen that there is only one peak in the distribution of the region of interest in the y direction due to the low disturbance resolution. Therefore, in the process of FIG. 10, in step S404, it is determined that the spatial frequency of the distribution of the region of interest is insufficient. As a result, in step S406, the parameter value is adjusted to increase the disturbance resolution.
[0092] In FIG. 11, in the middle column, an example of visualizing the region of interest obtained when the disturbance resolution is medium and an example of the distribution of this region of interest in the y direction are shown. In this case, it can be seen that there are two peaks in the distribution of the region of interest in the y direction. Therefore, in the process of FIG. 10, in step S404, it is determined that the spatial frequency of the distribution of the region of interest is appropriate. As a result, in step S405, it is determined that the parameter adjustment is completed and the parameter value is not adjusted.
[0093] In FIG. 11, in the lower column, an example of visualizing the region of interest obtained when the disturbance resolution is high and an example of the distribution of this region of interest in the y direction are shown. In this case, it can be seen that there are three peaks in the distribution of the region of interest in the y direction. Therefore, also in this case, as in the case of medium disturbance resolution, in the process of FIG. 10, in step S404, it is determined that the spatial frequency of the distribution of the region of interest is appropriate. As a result, in step S405, it is determined that the parameter adjustment is completed and the parameter value is not adjusted.
[0094] As described above, in the analyzer 1A of the present embodiment, it is determined that medium to high is appropriate as the disturbance resolution, and when the disturbance resolution is low, parameter adjustment is performed to increase the disturbance resolution.
[0095] Note that since the resolution of the disturbance is not too fine to cause problems, in the analyzer 1A of the present embodiment, it is preferable to adjust the disturbance parameter in the direction of increasing the insufficient resolution. However, generally, the lower the disturbance resolution, that is, the coarser the disturbance superimposed on the input image, the easier it is for the evaluation result of the disturbance ratio to converge. Therefore, if the image recognition process is performed with a high disturbance resolution from the beginning, the evaluation result of the disturbance ratio may not converge. Therefore, when the adjustment of the disturbance ratio does not converge indefinitely (for example, it does not converge even after performing it five times), the disturbance resolution may be lowered once and then the disturbance ratio may be adjusted again.
[0096] According to the second embodiment of the present invention described above, the processor 101, which is a computer, evaluates the resolution of the disturbance based on the distribution of the attention regions of the image recognition process when an object is recognized in the disturbance image (steps S402, S404). The disturbance is adjusted based on the evaluation result of the resolution of the disturbance (step S406). By doing so, it is possible to automatically adjust the resolution of the disturbance to an optimal value according to the distribution of the attention regions.
[0097] - Third Embodiment - Next, an analysis method and an analyzer for image recognition processing according to the third embodiment of the present invention will be described. In the first and second embodiments, an example in which only one object to be recognized in the image recognition process is reflected in the input image has been described. However, in the present embodiment, an example in which disturbances are applied to a plurality of objects and parameter adjustment is performed when a plurality of objects to be recognized are reflected in the input image will be described. Note that the hardware configuration of the analyzer 1B according to the present embodiment is the same as the hardware configuration of FIG. 1 described in the first embodiment. Therefore, in the following description, the analyzer 1B of the present embodiment will be described using the hardware configuration of FIG. 1.
[0098] FIG. 12 is a functional block diagram showing a functional configuration example of the analyzer 1B according to the third embodiment of the present invention. The analyzer 1B is different from the analyzer 1A described in the second embodiment in that it further includes an analysis region determination unit 25 and an analysis region quality determination unit 26.
[0099] For each of the plurality of recognition objects reflected in the input image, the analysis region determination unit 25 determines an analysis region corresponding to the object within the input image based on the correct label acquired by the teacher data acquisition unit 12. In the present embodiment, for each part corresponding to the analysis region set for each recognition object by the analysis region determination unit 25 in the input image, based on the set value of the parameter stored in the disturbance setting value storage unit 13, the disturbance generation unit 14 generates a disturbance that becomes noise when performing the image recognition process, and the disturbance superimposing unit 15 superimposes the disturbance.
[0100] The analysis region quality determination unit 26 determines whether the analysis region for each recognition object determined by the analysis region determination unit 25 is appropriate based on the evaluation result of the disturbance ratio by the disturbance ratio evaluation unit 21 and the evaluation result of the disturbance decomposition by the disturbance resolution evaluation unit 24. As a result, when it is determined that the analysis region is not appropriate, an instruction is given to the analysis region determination unit 25 to re-set the analysis region.
[0101] FIG. 13 is a flowchart showing an analysis method of image recognition processing according to the third embodiment of the present invention. In the present embodiment, the analysis device 1B can analyze the image recognition result by the image recognition unit 16 by executing the processing shown in the flowchart of FIG. 13 by the processor 101.
[0102] In steps S10 to S60, the same processing as in the first embodiment is respectively performed. After performing the processing of step S60, in the subsequent step S61, the analysis region determination unit 25 sets analysis regions for a plurality of objects reflected in the input image included in the image data set of the analysis target acquired in step S60. Here, for example, for the region of each object indicated by the correct label included in the image data set of the analysis target acquired in step S60, the object region is expanded by a predetermined margin rate (for example, 10%, 15%, 20%, etc.) to set a plurality of analysis regions with different margin rates for each object.
[0103] In step S62, the analysis region pass / fail determination unit 26 calculates the overlap between the analysis regions set in step S61. Here, it is determined whether the analysis regions of each object set for each margin rate overlap on the input image.
[0104] In step S63, based on the overlap between the analysis regions calculated in step S62, the analysis region pass / fail determination unit 26 extracts combinations of non-overlapping analysis regions. Here, among the combinations of the analysis regions of each object set for each margin rate in step S61, combinations of analysis regions that do not overlap are extracted as the combination with the highest margin rate. This makes it possible to set an optimal analysis region for each object.
[0105] In step S64, one of the analysis regions in the combination of analysis regions extracted in step S63 is selected. In steps S70 to S90, which are performed after step S64, the same processing as that described in the first and second embodiments is performed on the partial image within the analysis region selected in step S64.
[0106] In step S91, it is determined whether all of the analysis regions extracted in step S63 have been selected in step S64. If there are unselected analysis regions, the process returns to step S64. After selecting one of the unselected analysis regions in step S64, the processing in steps S70 to S90 is repeated for the partial image within that analysis region. If all of the analysis regions have been selected, the processing shown in the flowchart of FIG. 13 is completed.
[0107] FIG. 14 is a diagram showing an example of analysis regions set for each object. In the example shown in FIG. 14, for objects 141 and 142 reflected in the input image, analysis regions 151 and 152 are set by expanding these ranges so that they do not overlap each other.
[0108] In addition, in the present embodiment, in order to confirm whether the combination of analysis regions extracted in step S63 is appropriate, the processing after step S64 may be performed for analysis regions with different margin ratios. For example, by setting the number of times of generating the disturbance image to be smaller than the aforementioned predetermined number N, visualizing the regions of interest obtained for each analysis region, and presenting them to the user, the user can confirm that there is no change in these regions of interest.
[0109] According to the third embodiment of the present invention described above, the processor 101, which is a computer, sets an analysis region for each of a plurality of objects reflected in the input image (steps S61 to S63), calculates the disturbance influence degree for each of these analysis regions, and adjusts the disturbance (steps S64 to S81). By doing so, even when a plurality of objects to be recognized are reflected in the input image, the processing time can be shortened.
[0110] In addition, in the third embodiment of the present invention described above, the analysis device 1B may not have the disturbance resolution evaluation unit 24. In this case, the analysis device 1B only adjusts the disturbance ratio as in the first embodiment, and does not perform the adjustment of the disturbance resolution.
[0111] - Fourth Embodiment - Next, an analysis method and an analysis device for image recognition processing according to the fourth embodiment of the present invention will be described. In the third embodiment, an example was described in which when a plurality of objects to be recognized in the image recognition processing are reflected in the input image, an analysis region is set for each object and parameter adjustment is performed. However, such a parameter adjustment method cannot be applied well when the objects are densely packed. Therefore, in the present embodiment, an example will be described in which the objects reflected in the input image are grouped, the images are reconstructed for each group of objects in the entire dataset, and parameter adjustment is performed to efficiently analyze the entire dataset. The hardware configuration of the analysis device 1C according to the present embodiment is the same as the hardware configuration of FIG. 1 described in the first embodiment. Therefore, in the following description, the analysis device 1C of the present embodiment will be described using the hardware configuration of FIG. 1.
[0112] FIG. 15 is a functional block diagram showing a functional configuration example of the analyzer 1C according to the fourth embodiment of the present invention. The analyzer 1C is different from the analyzer 1B described in the third embodiment in that it further has an image division / assignment unit 27.
[0113] The image division / assignment unit 27 divides the input image acquired by the input image acquisition unit 11 into objects reflected in the input image, thereby obtaining a divided image for each object from the input image. Then, the divided images obtained for the entire dataset are appropriately combined and synthesized to reconstruct the image to be analyzed. In the present embodiment, for the image reconstructed by the image division / assignment unit 27, based on the set value of the parameter stored in the disturbance setting value storage unit 13, a disturbance that becomes noise when performing image recognition processing is generated by the disturbance generation unit 14 and superimposed by the disturbance superimposition unit 15.
[0114] FIG. 16 is a flowchart showing an analysis method of image recognition processing according to the fourth embodiment of the present invention. In the present embodiment, the analyzer 1C can analyze the image recognition result by the image recognition unit 16 by executing the processing shown in the flowchart of FIG. 16 by the processor 101.
[0115] In steps S10 to S60, the same processing as in the first embodiment is respectively performed. After the processing of step S60 is performed, in subsequent step S65, the image division / assignment unit 27 selects any one of the image dataset for analysis acquired in step S60, and reads the combination of the input image and the correct label in the dataset.
[0116] In step S66, the image segmentation and assignment unit 27 groups each object reflected in the input image included in the dataset read in step S61 according to its size. Here, for example, based on the size of the area of each object indicated by the correct label included in the dataset, each object is assigned to one of three groups: "large", "medium", and "small". Note that the number of groups for assignment is not limited to three, and can be arbitrarily set according to the specifications of the image recognition process performed by the image recognition unit 16 and the like.
[0117] In step S67, it is determined whether or not the grouping of the objects in the input images included in all the datasets to be analyzed acquired in step S60 has been completed. If the grouping process in step S66 has been performed for all the recognition target objects reflected in the input images of all the datasets, the process proceeds to step S68. On the other hand, if there are unclassified objects for which the process in step S66 has not been performed, the process returns to step S65 to reselect the dataset, and then the process in step S66 is performed again to continue the grouping of the objects.
[0118] In step S68, the image segmentation and assignment unit 27 generates a composite image to be analyzed by cutting out the images of the objects grouped in step S66 from the input image and synthesizing them by size. For example, when each object is classified into the three groups of "large", "medium", and "small" described above, partial images within the analysis area set by the analysis area determination unit 25 for the objects in each group are cut out from the input image in which the object is reflected. Then, the partial images are synthesized in different numbers for each group to generate a composite image.
[0119] Specifically, for example, for each object classified into the "large" group, with the number of partial images assigned per composite image being 1, the partial image of each object is used as the composite image to be analyzed as it is. On the other hand, for example, for each object classified into the "medium" group, with the number of partial images assigned per composite image being 2, two partial images are combined to form one composite image. In this case, the two regions of the composite image (for example, the left and right regions) are respectively assigned to different objects. Also, for example, for each object classified into the "small" group, with the number of partial images assigned per composite image being 4, four partial images are combined to form one composite image. In this case, the four regions of the composite image (for example, the upper left, upper right, lower left, and lower right regions) are respectively assigned to different objects. Note that the number of partial images assigned for each group in the composite image is not limited to this and can be any number.
[0120] In step S69, among the partial images of each object included in each composite image generated in step S68, one partial image in one of the composite images is selected. In steps S70 to S90 implemented after step S69, for the partial image selected in step S69, the same processing as described in the first and second embodiments is respectively performed.
[0121] In step S92, it is determined whether all the partial images in all the composite images generated in step S68 have been selected in step S69. If there are unselected composite images or partial images, return to step S69, select one of them in step S69, and then repeat the processing of steps S70 to S90 for the partial image. If all the partial images in all the composite images have been selected, complete the processing shown in the flowchart of FIG. 16.
[0122] According to the fourth embodiment of the present invention described above, the processor 101, which is a computer, cuts out partial images corresponding to the analysis region from the input image for each object (steps S65 to S67), generates a composite image by combining a plurality of partial images cut out from different input images (step S68), and calculates the degree of disturbance influence for each partial image of this composite image to adjust the disturbance (steps S69 to S81). By doing so, when a plurality of objects to be recognized are reflected in the input image, the analysis of the data set can be performed more efficiently.
[0123] Note that also in the fourth embodiment of the present invention described above, similar to the third embodiment described above, the analysis device 1C may not have the disturbance resolution evaluation unit 24. In this case, the analysis device 1C only adjusts the disturbance ratio as in the first embodiment and does not perform the adjustment of the disturbance resolution.
[0124] Note that in each of the first to fourth embodiments described above, the following modification examples may be applied.
[0125] (Modification example) In each of the flowcharts of FIGS. 3, 9, 13, and 16, it may be evaluated whether the disturbance is appropriate during the disturbance evaluation process of FIG. 4 executed in steps S30 and S70, and it may be determined whether to continue generating the disturbance image. Specifically, for example, in the disturbance evaluation process of FIG. 4, each time a disturbance image is generated until the number of generated disturbance images reaches a predetermined upper limit number (the number of times N described above), the evaluation value of the disturbance image is obtained in step S104. Here, when the number of generated disturbance images is less than the upper limit number N, the distribution of the evaluation values of the disturbance images obtained so far is tabulated in the disturbance ratio evaluation unit 21, and when the bias of the distribution exceeds a predetermined bias state (for example, when the variance Sv is less than the above-mentioned threshold Sv_th), the generation of the disturbance image is terminated. Then, after changing the disturbance parameter by the parameter adjustment unit 22 to adjust the disturbance, the disturbance evaluation process is restarted to start over the generation of the disturbance image. By doing so, it is possible to prevent the disturbance evaluation process from continuing to be performed by an inappropriate disturbance, so that the processing time can be further shortened.
[0126] Note that the present invention is not limited to the various embodiments and modifications described above, and includes various other modifications. For example, the above-described embodiments are specifically described for the purpose of easily explaining the present invention, and are not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment can be replaced with a part of the configuration of another embodiment. Also, the configuration of another embodiment can be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, it can be deleted, a part of another configuration can be added, and a part of another configuration can be replaced.
[0127] The present invention is not limited to the above-described embodiments, and various changes can be made without departing from the spirit of the present invention.
Description of Reference Numerals
[0128] 1, 1A, 1B, 1C: Image recognition processing analysis device (analysis device), 11: Input image acquisition unit, 12: Teacher data acquisition unit, 13: Noise setting value storage unit, 14: Noise generation unit, 15: Noise superposition unit, 16: Image recognition unit, 17: Evaluation value acquisition unit, 18: Attention information generation unit, 19: Attention information integration unit, 20: Drawing unit, 21: Noise ratio evaluation unit, 211: Aggregation unit, 212: Statistical processing unit, 213: Determination unit, 22: Parameter adjustment unit, 23: GUI (Graphical User Interface), 24: Noise resolution evaluation unit, 241: Attention area distribution analysis unit, 242: Determination unit, 25: Analysis area determination unit, 26: Analysis area quality determination unit, 27: Image division / assignment unit
Claims
1. A method for analyzing an image recognition process for recognizing an object to be recognized reflected in an input image, comprising: generating, by a computer, a plurality of disturbed images obtained by superimposing a disturbance on the input image; obtaining, by the computer, an evaluation value for the recognition result of the object obtained by respectively executing the image recognition process on the plurality of disturbed images; calculating, by the computer, a disturbance influence degree of the disturbance on the image recognition process based on the evaluation value obtained for each of the disturbed images; adjusting, by the computer, the disturbance based on the disturbance influence degree, the method for analyzing an image recognition process.
2. The method for analyzing an image recognition process according to claim 1, wherein the disturbance influence degree is calculated based on a distribution of the evaluation value obtained for each of the disturbed images, the method for analyzing an image recognition process.
3. The method for analyzing an image recognition process according to claim 2, wherein statistical information in the distribution of the evaluation value is calculated and used as the disturbance influence degree, the method for analyzing an image recognition process.
4. The method for analyzing an image recognition process according to claim 2 or 3, wherein when the distribution of the evaluation value is biased toward the high evaluation side, the intensity of the disturbance is adjusted to be increased, when the distribution of the evaluation value is biased toward the low evaluation side, the intensity of the disturbance is adjusted to be decreased, the method for analyzing an image recognition process.
5. The method for analyzing an image recognition process according to claim 2 or 3, wherein until the number of generated disturbed images reaches a predetermined upper limit number, an evaluation value of each disturbed image is obtained each time a disturbed image is generated, when the number of generated disturbed images is less than the upper limit number and the deviation of the distribution of the evaluation value exceeds a predetermined deviation state, generation of the disturbed images is terminated, the disturbance is adjusted, and then generation of the disturbed images is restarted from the beginning, the method for analyzing an image recognition process.
6. The method for analyzing an image recognition process according to claim 1, wherein the computer evaluates a resolution of the disturbance based on a distribution of a region of interest of the image recognition process when the object is recognized in the disturbed image, the computer adjusts the disturbance based on an evaluation result of the resolution of the disturbance, the method for analyzing an image recognition process.
7. The method for analyzing an image recognition process according to claim 1, wherein An analysis method for image recognition processing, which sets an analysis region for each of the plurality of objects reflected in the input image, calculates the disturbance influence degree for each analysis region, and adjusts the disturbance.
8. In the analysis method for image recognition processing according to Claim 7, each partial image corresponding to the analysis region is cut out from the input image for each object, a composite image is generated by combining a plurality of the partial images cut out from different input images respectively, and the disturbance influence degree is calculated for each partial image with respect to the composite image to adjust the disturbance. An analysis method for image recognition processing.
9. An apparatus for analyzing an image recognition process for recognizing an object to be recognized reflected in an input image, comprising: a disturbance superimposing unit that generates a plurality of disturbance images in which disturbances are superimposed on the input image; an evaluation value acquisition unit that acquires an evaluation value for the recognition result of the object obtained by respectively executing the image recognition process on the plurality of disturbance images generated by the disturbance superimposing unit; a disturbance ratio evaluation unit that calculates a disturbance influence degree exerted by the disturbance on the image recognition process based on the evaluation value acquired for each disturbance image by the evaluation value acquisition unit; and a parameter adjustment unit that adjusts the disturbance based on the disturbance influence degree calculated by the disturbance ratio evaluation unit.
Citation Information
Patent Citations
Explanatory visualizations for object detection
US20210357644A1