Analysis method and analysis device for image recognition processing
The method automates the adjustment of disturbances in image recognition systems to optimize parameter settings, enhancing dataset analysis efficiency and recognition accuracy by visualizing attention areas.
Patent Information
- Application Number
- PCT/JP2024/028153
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-08-06
- Publication Date
- 2025-07-03
AI Technical Summary
Existing image recognition systems, particularly in autonomous driving, face challenges in handling variations in environmental conditions and require significant manual parameter adjustment for visualizing attention areas, leading to inefficiencies in analyzing large datasets.
An automated method and apparatus that generates disturbed images by superimposing disturbances on input images, evaluates recognition results, and adjusts disturbances based on influence degrees to optimize parameter settings for improved visualization of attention areas.
Significantly enhances the analysis efficiency of large-scale datasets by automatically adjusting disturbances for accurate visualization of attention areas, reducing manual effort and improving recognition performance.
Smart Images

Figure JP2024028153_03072025_PF_FP_ABST
Abstract
Description
Image recognition processing analysis method and analysis device
[0001] The present invention relates to an analysis method and an analysis device for image recognition processing.
[0002] Conventionally, a technology has been known in which a computer recognizes the surroundings of a vehicle based on images captured by a camera, and the recognition results are used to automatically drive the vehicle. The surrounding recognition processing required to realize such automatic driving requires not only high recognition performance, but also reliability and safety. Therefore, in recent years, image recognition processing using AI (artificial intelligence) has been used to recognize the surroundings of a vehicle, thereby achieving higher recognition performance than ever before.
[0003] In image recognition processing by AI, recognition performance varies significantly depending on the AI's learning level. However, it is difficult to prepare training data in advance that covers all possible situations that may occur in the real environment during autonomous driving. As a result, there are an infinite number of unknown scenes in the real environment that are not included in the AI's training data, and it is known that recognition errors unintended by the designer may occur. In order to minimize the occurrence of such recognition errors and ensure reliability and safety, there is a strong demand for image recognition processing for autonomous driving to continuously evolve by acquiring data indicating the vehicle's surroundings in the real environment as external data and repeatedly re-training the AI using this external data.
[0004] In order to appropriately evolve AI for image recognition processing, it is preferable to collect data on external environments, such as new objects that the AI is to recognize and scenes where recognition errors have occurred in the past, as external data, and use the collected external data for relearning the AI. This allows the AI to correctly recognize objects and scenes that the AI was previously unable to correctly recognize. However, among the objects and scenes that the AI correctly recognized before relearning, there are many that the AI recognized correctly by chance even when it was not confident, or that are unstable and cannot be correctly recognized even with slight changes in the environment (such as changes in illuminance, saturation, or noise). It is said that the occurrence rate of such unstable recognition results is higher than the probability of recognition errors. In other words, to achieve truly reliable AI for image recognition processing, it is necessary to extract more training data that will cause the AI's recognition results to be unstable from the various data collected to date and existing datasets.
[0005] Various existing technologies have been studied to date as methods for quantifying the stability and confidence of image recognition processing by AI. One of these is a technology that visualizes the basis of AI inference. For example, Patent Literature 1 discloses a technology that visualizes the area of interest of AI using information on how the results of inference performed by adding disturbances to an input image of the AI change compared to the inference results without the disturbance. By utilizing this information on the area of interest, it is possible to extract data that makes AI recognition unstable.
[0006] US Patent Application Publication No. 2021 / 0357644
[0007] In many technologies for visualizing an area of interest for inference, such as the technology in Patent Document 1, a designer must visually check whether the visualized results are of sufficiently high quality. If the designer determines that the quality is poor, the designer must reset various parameters used for visualization and revisit the visualization, a process of trial and error. In this case, the disturbance setting parameters must be manually adjusted depending on the size of the object to be recognized and the difficulty of recognition. As a result, analyzing a large amount of collected data sets requires a lot of human resources and time, making it difficult to improve analysis efficiency.
[0008] The method for analyzing image recognition processing according to the present invention is a method for analyzing image recognition processing for recognizing an object to be recognized that is reflected in an input image, in which a computer generates a plurality of disturbance images by superimposing a disturbance on the input image, the computer obtains an evaluation value for the recognition results of the object obtained by performing the image recognition processing on each of the plurality of disturbance images, the computer calculates the degree of disturbance influence that the disturbance has on the image recognition processing based on the evaluation value obtained for each of the disturbance images, and the computer adjusts the disturbance based on the degree of disturbance influence. The analysis device according to the present invention is a device for analyzing image recognition processing for recognizing a target object that is reflected in an input image, and includes a disturbance superposition unit that generates a plurality of disturbance images by superimposing a disturbance on the input image, an evaluation value acquisition unit that acquires an evaluation value for the recognition result of the object obtained by executing the image recognition processing on each of the plurality of disturbance images generated by the disturbance superposition unit, a disturbance ratio evaluation unit that calculates the degree of disturbance influence that the disturbance has on the image recognition processing based on the evaluation value acquired for each disturbance image by the evaluation value acquisition unit, and a parameter adjustment unit that adjusts the disturbance based on the degree of disturbance influence calculated by the disturbance ratio evaluation unit.
[0009] According to the present invention, it is possible to automatically adjust disturbances required for visualizing regions of interest in inference in image recognition processing, thereby significantly improving the efficiency of analyzing large-scale data sets.
[0010] FIG. 1 is a diagram illustrating an example of the hardware configuration of an image recognition processing analysis device according to a first embodiment of the present invention. FIG. 2 is a functional block diagram illustrating an example of the functional configuration of an image recognition processing analysis device according to the first embodiment of the present invention. FIG. 3 is a flowchart illustrating an analysis method for image recognition processing according to the first embodiment of the present invention. FIG. 4 is a flowchart illustrating the procedure of disturbance evaluation processing. FIG. 5 is a flowchart illustrating the procedure of disturbance ratio adjustment processing. FIG. 6 is a flowchart illustrating an example of a method for determining a disturbance influence level. FIG. 7 is a diagram illustrating an example distribution of evaluation values for each disturbance influence level. FIG. 8 is a functional block diagram illustrating an example of the functional configuration of an image recognition processing analysis device according to a second embodiment of the present invention. FIG. 9 is a flowchart illustrating an analysis method for image recognition processing according to the second embodiment of the present invention. FIG. 10 is a flowchart illustrating the procedure of disturbance resolution adjustment processing. FIG. 11 is a diagram illustrating an example of visualization results and distribution of attention areas for each disturbance resolution. FIG. 12 is a functional block diagram illustrating an example of the functional configuration of an image recognition processing analysis device according to a third embodiment of the present invention. FIG. 13 is a flowchart illustrating an analysis method for image recognition processing according to the third embodiment of the present invention. FIG. 14 is a diagram illustrating an example of an analysis area set for each object. FIG. 15 is a functional block diagram illustrating an example of the functional configuration of an image recognition processing analysis device according to a fourth embodiment of the present invention. FIG. 16 is a flowchart illustrating an analysis method for image recognition processing according to the fourth embodiment of the present invention.
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each embodiment is an example for explaining the present invention, and for clarity of explanation, appropriate omissions and simplifications have been made. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.
[0012] The position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings. When there are multiple components having the same or similar functions, they may be described using the same reference numeral with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description.
[0013] In each embodiment, processing performed by executing a program may be described. Here, a computer executes the program using a processor (e.g., a CPU or a GPU) and performs processing defined by the program using storage resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the entity performing the processing by executing the program may be the processor. Similarly, the entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The entity performing the processing by executing the program may be any computing unit, and may include a dedicated circuit that performs specific processing. Here, the dedicated circuit may be, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a CPLD (Complex Programmable Logic Device).
[0014] A program may be installed on a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server may include a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. In addition, in the embodiments, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0015] In recent years, in order to realize advanced autonomous driving, peripheral recognition applications using image recognition programs that apply deep neural networks (DNNs), a type of AI that uses machine learning, have become increasingly popular.Compared to rule-based algorithms, programs using AI such as DNNs have the problem that it is difficult to respond when a malfunction occurs, because it is unclear why the AI made the inference it did in response to the recognition results obtained through AI inference.
[0016] To address these challenges, technologies that visualize image regions focused on by AI inference (e.g., Grad-CAM and RISE) have been reported. As a specific example, consider a situation in which an image clearly captures the entire vehicle (the target of recognition), and the AI inference results for this image accurately infer the vehicle's position and achieve a very high reliability score. However, when attempting to continuously detect the vehicle, frequent non-detections occur. Visualizing the focus area for a vehicle captured in an image where this phenomenon occurred may result in the inference focusing on only a portion of the vehicle (e.g., a portion such as the windshield). This situation can be considered a condition that makes AI inference prone to non-detection, as it may be due to the fact that the relevant portion is occluded by another object or obscured due to the amount of light or noise during capture.
[0017] As described above, visualizing the region of interest in AI inference can be effectively used to identify the cause of machine learning problems. However, the structure of the neural network used in DNN and the characteristics of the object to be recognized in the input image (size, orientation, recognition difficulty, etc.) vary depending on the situation and are not uniform. When visualizing the region of interest, the optimal parameter values for obtaining high-quality visualization results vary depending on these differences. Therefore, it is necessary to visually check the quality of the visualized region of interest, and if the quality is low, adjust the parameters and revisit the visualization. In other words, when analyzing AI image recognition results for large datasets, the trial-and-error labor required for parameter adjustment becomes an issue. When visualizing the region of interest in an application where the size, orientation, recognition difficulty, etc. of the object to be recognized are nearly constant, such parameter adjustment work is often unnecessary. However, parameter adjustment is particularly necessary in applications that require the recognition of various types of objects in various scenes, such as automotive external environment recognition.
[0018] The present invention automatically adjusts parameters when visualizing regions of interest in inference by AI such as DNN, thereby efficiently analyzing the image recognition results of the AI for large data sets. In the following embodiments, an example in which the present invention is applied to an object recognition DNN in an in-vehicle ECU for vehicle control, for example, an Advanced Driver Assistance System (ADAS) or Autonomous Driving (AD), is described. However, the present invention is not limited to in-vehicle ECUs for ADAS and AD, and can be applied to general applications using object detection AI.
[0019] -First Embodiment- Hereinafter, an analysis method and an analysis device for image recognition processing according to a first embodiment of the present invention will be described with reference to FIGS.
[0020] 1 is a diagram showing an example of the hardware configuration of an image recognition processing analysis device (hereinafter simply referred to as "analysis device") 1 according to a first embodiment of the present invention. The analysis device 1 is a computer such as a server or a PC, and is configured by connecting a processor 101, a storage device 102, an input device 103, an output device 104, and a communication interface (communication IF) 105 to one another via a communication bus 106.
[0021] The processor 101 executes a predetermined program to control the operation of the analysis apparatus 1 and to cause a computer such as a server or PC to function as the analysis apparatus 1. The storage device 102 is a recording medium capable of non-temporarily or temporarily storing various programs and data, and is configured using, for example, a read-only memory (ROM), a random access memory (RAM), a hard disk drive (HDD), or a flash memory. The input device 103 is a device that accepts data input from a user to the analysis apparatus 1, and is configured using, for example, a keyboard, a mouse, a touch panel, a microphone, etc. The output device 104 is a device that outputs data to the user, and is configured using, for example, a display, a printer, a speaker, etc. The communication IF 105 transmits and receives data between the analysis apparatus 1 and an external information device (not shown) via a communication network (not shown).
[0022] 2 is a functional block diagram showing an example of the functional configuration of the analysis device 1 according to the first embodiment of the present invention. The analysis device 1 has, as functions for performing image recognition processing using AI and analyzing the inference results, an input image acquisition unit 11, a teacher data acquisition unit 12, a disturbance setting value storage unit 13, a disturbance generation unit 14, a disturbance superposition unit 15, an image recognition unit 16, an evaluation value acquisition unit 17, an attention information generation unit 18, an attention information integration unit 19, a drawing unit 20, a disturbance ratio evaluation unit 21, a parameter adjustment unit 22, and a GUI (Graphical User Interface) 23.
[0023] The input image acquisition unit 11, teacher data acquisition unit 12, disturbance generation unit 14, disturbance superposition unit 15, image recognition unit 16, evaluation value acquisition unit 17, attention information generation unit 18, attention information integration unit 19, drawing unit 20, disturbance ratio evaluation unit 21, parameter adjustment unit 22, and GUI 23 are specifically realized by causing the processor 101 to execute programs stored in, for example, the storage device 102 shown in Fig. 1. The disturbance setting value storage unit 13 is specifically realized by, for example, the storage device 102 shown in Fig. 2.
[0024] The input image acquisition unit 11 acquires an input image to be subjected to image recognition processing performed by the analysis device 1. The teacher data acquisition unit 12 acquires a correct answer label indicating the correct answer of the recognition result for the input image as teacher data for the input image. The input image acquisition unit 11 and the teacher data acquisition unit 12 can respectively acquire the input image and the correct answer label based on information stored in the storage device 102 and information received by the communication IF 105 via a communication network.
[0025] The disturbance generator 14 generates disturbances that become noise when performing image recognition processing on an input image. The disturbance setting value storage unit 13 stores setting values of various parameters used when the disturbance generator 14 generates disturbances. For example, information such as the type, intensity, and resolution of the disturbance, an analysis method for the disturbance, and a threshold value used for evaluating the disturbance is stored in the disturbance setting value storage unit 13 as setting values of the parameters used to generate the disturbances. The disturbance generator 14 generates a plurality of disturbances with different patterns according to the parameter values stored in the disturbance setting value storage unit 13, and outputs the generated disturbances to the disturbance superimposition unit 15 and the attention information generator 18.
[0026] The disturbance superimposing unit 15 generates a disturbance image in which the disturbance is superimposed on the input image by superimposing the disturbance generated by the disturbance generating unit 14 on the input image acquired by the input image acquiring unit 11. Here, as described above, the disturbance generating unit 14 generates a plurality of disturbances with different patterns. By superimposing disturbances with different patterns on the input image in this way, the disturbance superimposing unit 15 can generate a plurality of disturbance images for the same input image.
[0027] The image recognition unit 16 performs image recognition processing using a trained DNN on each of the multiple disturbance images generated by the disturbance superimposition unit 15. In the image recognition processing performed by the image recognition unit 16, objects that should be recognized by the ADAS or AD, such as surrounding vehicles, obstacles, and various road signs, are recognized from among various objects reflected in the input image. Note that when the image recognition unit 16 performs image recognition processing using a DNN, it is possible to use well-known arithmetic processing.
[0028] The evaluation value acquisition unit 17 acquires an evaluation value for the object recognition result obtained by the image recognition unit 16 executing the image recognition process on each disturbance image. For example, the evaluation value acquisition unit 17 compares the object recognition result for the disturbance image with the correct label acquired by the teacher data acquisition unit 12, and sets an evaluation value for the object recognition result for each disturbance image so that the higher the degree of match between them, the higher the evaluation value. Note that the method of acquiring the evaluation value by the evaluation value acquisition unit 17 will be described in detail later.
[0029] The attention information generation unit 18 generates attention information related to the attention level in the image recognition process when the image recognition unit 16 recognizes an object in the disturbance image by applying the disturbance generated by the disturbance generation unit 14 to the evaluation value acquired for each disturbance image by the evaluation value acquisition unit 17. Details of the method for generating attention information by the attention information generation unit 18 will be described later.
[0030] The attention information integrating unit 19 integrates the attention information for each disturbance image generated by the attention information generating unit 18 for the plurality of disturbance images generated by the disturbance superimposing unit 15. This processing integrates the attention information obtained for the plurality of disturbance images, and makes it possible to generate information for visualizing the attention area that was focused on during inference in the image recognition processing for the input image.
[0031] The drawing unit 20 draws information about the attention area obtained by integrating the attention information by the attention information integration unit 19, superimposing it on the input image acquired by the input image acquisition unit 11. As a result, an output image is generated in which the positions of the attention areas in the input image are indicated in a map-like manner. The output image generated by the drawing unit 20 is presented to the user via the GUI 23. By looking at this output image, the user can visually confirm the extent to which disturbances have affected the results of the image recognition processing for the input image.
[0032] The disturbance ratio evaluation unit 21 calculates a disturbance influence degree, which indicates the degree of influence that a disturbance has on the image recognition processing, based on the evaluation value acquired for each disturbance image by the evaluation value acquisition unit 17. Then, based on the calculated disturbance influence degree, it evaluates the ratio of the area where the disturbance generated by the disturbance generation unit 14 hides the input image (hereinafter referred to as the "disturbance ratio") and notifies the parameter adjustment unit 22 of the evaluation result.
[0033] The disturbance ratio evaluation unit 21 includes a counting unit 211, a statistical processing unit 212, and a determination unit 213. The counting unit 211 counts the evaluation value for each disturbance image acquired by the evaluation value acquisition unit 17. The statistical processing unit 212 statistically processes the evaluation values counted by the counting unit 211 to determine the distribution of evaluation values for multiple disturbance images. The determination unit 213 calculates a disturbance influence degree based on the distribution of evaluation values determined by the statistical processing unit 212 and compares this disturbance influence degree with a predetermined determination threshold to determine whether the disturbance ratio is appropriate. The obtained determination result is then notified to the parameter adjustment unit 22 as the evaluation result of the disturbance ratio. Note that the determination threshold used for comparison with the disturbance influence degree in the determination unit 213 can be set by the user via the GUI 23.
[0034] The parameter adjustment unit 22 adjusts the parameters stored in the disturbance set value storage unit 13 so that the disturbance ratio becomes an appropriate value, based on the evaluation result of the disturbance ratio by the disturbance ratio evaluation unit 21. Note that a specific method of parameter adjustment by the parameter adjustment unit 22 will be described later.
[0035] 3 is a flowchart showing an analysis method of image recognition processing according to the first embodiment of the present invention. In this embodiment, the analysis device 1 can analyze the image recognition results obtained by the image recognition unit 16 by causing the processor 101 to execute the processing shown in the flowchart of FIG.
[0036] Steps S10 to S50 are processes for initial setting. In step S10, a test image dataset is acquired. Here, the input image acquisition unit 11 and the training data acquisition unit 12 can acquire, as a test image dataset, a number ranging from several pairs to several tens of pairs of combinations of arbitrary input images containing objects to be recognized in the image recognition process executed by the image recognition unit 16 and correct labels.
[0037] In step S20, initial parameters of the disturbance are specified. Here, values of various parameters, such as the type, intensity, and resolution of the disturbance, the method of analyzing the disturbance, and the threshold value used for evaluating the disturbance, input by the user via the GUI 23 are stored as initial parameters in the disturbance setting value storage unit 13.
[0038] In step S30, a disturbance evaluation process is performed using the test image data set acquired in step S10. Here, the disturbance superimposing unit 15 superimposes the disturbance generated by the disturbance generator 14 on the input image to generate a disturbance image containing noise, and the image recognition unit 16 performs image recognition processing on this disturbance image. Then, the evaluation value acquisition unit 17 acquires an evaluation value for the object recognition result in the image recognition processing. Details of this disturbance evaluation process will be described later with reference to the flowchart of FIG. 4.
[0039] In step S40, the evaluation result of the disturbance obtained by the disturbance evaluation process executed in step S30 is output. Here, for example, the evaluation value obtained by the disturbance evaluation process is superimposed on the input image and presented to the user via the GUI 23, thereby visualizing the evaluation result of the disturbance based on the set initial parameters and allowing the user to confirm it.
[0040] In step S50, it is determined whether the disturbance evaluation result output in step S40 is valid. Here, it is determined whether the disturbance evaluation result is valid, for example, based on the content of the input operation performed by the user via the GUI 23 in response to the disturbance evaluation result. As a result, if it is determined that the disturbance evaluation result is valid, the process proceeds to step S60. On the other hand, if it is determined that the disturbance evaluation result is invalid, the process returns to step S20, the initial parameters of the disturbance are reset, and the disturbance evaluation process of step S30 is executed again to determine whether the evaluation result is valid. In this way, the initial parameters of the disturbance are repeatedly set until a valid disturbance evaluation result is obtained for the test image data set, and the values of the initial parameters finally set are used as parameters for disturbance generation in subsequent processes.
[0041] Steps S60 to S90 are processes for adjusting for disturbances. In step S60, an image dataset to be analyzed is acquired. Here, as in step S10, the input image acquisition unit 11 and the training data acquisition unit 12 can acquire a combination of an arbitrary input image containing an object to be recognized in the image recognition process executed by the image recognition unit 16 and a correct answer label as the image dataset to be analyzed. In this case, the combination of the input image and the correct answer label acquired as the image dataset to be analyzed may be the same as that acquired in step S10, or a different combination may be acquired. Furthermore, the image dataset to be analyzed may be larger or smaller than the test image dataset. However, in order to train the image recognition unit 16 using a large amount of image data, it is preferable to acquire a larger image dataset than the test image dataset as the image dataset to be analyzed.
[0042] In step S70, the image data set to be analyzed acquired in step S60 is used to perform a disturbance evaluation process similar to that in step S30.
[0043] In step S80, a disturbance ratio adjustment process is executed based on the disturbance evaluation result obtained by the disturbance evaluation process executed in step S70. Here, the disturbance ratio evaluation unit 21 calculates the disturbance influence degree of the disturbance on the image recognition process by tallying and statistically processing the evaluation values for each disturbance image acquired in the disturbance evaluation process. Then, based on the calculated disturbance influence degree, it is determined whether the disturbance ratio is appropriate, and the determination result is output to the parameter adjustment unit 22. The parameter adjustment unit 22 adjusts the value of the parameter related to the disturbance ratio among the disturbance parameters stored in the disturbance setting value storage unit 13, based on the disturbance ratio determination result input from the disturbance ratio evaluation unit 21. Details of this disturbance ratio adjustment process will be described later with reference to the flowchart of FIG. 5.
[0044] In step S90, it is determined whether the adjustment of the disturbance parameters has been completed based on the determination result of the disturbance ratio obtained by the disturbance ratio adjustment process executed in step S80. Here, if the determination result indicated that the disturbance ratio was appropriate in the disturbance ratio adjustment process executed immediately before, and therefore no adjustment of the parameter value was performed, it is determined that the adjustment of the disturbance parameters has been completed, and the process shown in the flowchart of FIG. 3 is terminated. On the other hand, if the determination result indicated that the disturbance ratio was inappropriate in the disturbance ratio adjustment process executed immediately before, and therefore adjustment of the parameter value was performed, it is determined that the adjustment of the disturbance parameters has not been completed, and the process returns to step S70. In this case, the disturbance evaluation process of step S70 and the disturbance ratio adjustment process of step S80 are executed again using the adjusted parameter value. In this way, the adjustment of the disturbance parameters is repeated until an appropriate disturbance ratio is obtained for the image data set being analyzed.
[0045] FIG. 4 is a flowchart showing the procedure of the disturbance evaluation process executed in steps S30 and S70 of FIG.
[0046] In step S101, the disturbance generator 14 generates a disturbance for the input image. Here, based on the parameters stored in the disturbance setting value storage unit 13, a disturbance is generated that randomly masks the image in each region obtained by dividing the input image into predetermined ranges. Specifically, for example, a disturbance value of "0" or "1" is randomly set for each region obtained by dividing the input image into predetermined ranges. The size of the region and the appearance ratio of "0" and "1" in the disturbance values at this time are determined according to the parameter values stored in the disturbance setting value storage unit 13.
[0047] In step S102, the disturbance superimposing unit 15 combines the disturbance generated in step S101 with the input image to generate a disturbance image. Here, partial images corresponding to each region of the input image are multiplied by the disturbance value of "0" or "1" set in step S101. As a result, in each region where the disturbance value is "0," the partial image corresponding to that portion of the input image is hidden by the disturbance, while in each region where the set disturbance value is "1," the partial image corresponding to that portion of the input image is not hidden by the disturbance, thereby generating a disturbance image in which the disturbance is superimposed on the input image.
[0048] In step S103, the image recognition unit 16 performs inference on the disturbance image generated in step S102. In this inference, an image recognition process using a trained DNN is executed to extract an object to be recognized that is included in the input image before the disturbance is superimposed, and to recognize the type of the object and the range of the area that the object occupies in the image. At this time, the likelihood of the obtained inference result may be calculated. Alternatively, the image recognition unit 16 may be provided outside the analysis device 1, and the analysis device 1 may acquire the inference result executed by this image recognition unit 16 in step S103.
[0049] In step S104, the evaluation value acquisition unit 17 calculates an evaluation value for the inference result of step S103. For example, the object region obtained by inference is compared with the object region indicated by the correct label, and the degree of overlap (IoU: Intersection over Union) between them is calculated as an evaluation value (score) for the inference result. In this case, if the inference result and the correct label completely mismatch, the evaluation value is 0, and if they completely match, the evaluation value is 1. In this case, instead of the correct label, an inference result for an input image without any external disturbance may be acquired and compared with the inference result of step S103 to calculate the evaluation value. The evaluation value may also be calculated taking into account the recognition class (object type) in the inference result. For example, if the recognition class in the inference matches the object type indicated by the correct label, the IoU value is used as the evaluation value, and if they differ, the evaluation value is set to 0 regardless of the IoU value. Alternatively, if the recognition class and the correct label are different but belong to the same category, the IoU value is multiplied by a predetermined penalty rate (e.g., 0.5) to adjust the evaluation value. In addition to this, an evaluation value for the inference result in step S103 can be calculated by any other method.
[0050] In step S105, the attention information generator 18 generates attention information related to the degree of attention in the inference in step S103 based on the evaluation value calculated in step S104. Specifically, for example, the attention information can be generated by calculating the product of the disturbance superimposed on the input image by the disturbance superimposition unit 15 in step S102 and the evaluation value calculated in step S104 for each partial image corresponding to each region of the input image where the disturbance is set. This attention information represents the magnitude of the evaluation value for the inference for each region of the disturbance image where the disturbance value is "1," i.e., for each partial image corresponding to each region not hidden by the disturbance. In other words, the attention information represents regions of the input image for which the same or similar inference results can be obtained by using an image not hidden by the disturbance, regardless of whether or not there is a disturbance, along with the similarity of the inference results. Such information corresponds to the degree of attention in the inference, and it can be said that the greater the value of the attention information (the similarity of the inference results), the higher the degree of attention in the inference. Conversely, image areas with small values of attention information are significantly affected by the parts hidden by the disturbance, and even if the image recognition unit 16 looks at a partial image within that area when performing inference on the disturbance image, it is unable to reproduce the same inference result as the original input image without the disturbance superimposed.
[0051] In step S106, the attention information accumulating unit 19 further accumulates the attention information generated in step S105 with the accumulation result of the attention information obtained in the previous process.
[0052] In step S107, it is determined whether the processing of steps S101 to S106 has been completed a predetermined number of times (for example, N times). If the number of times the processing of steps S101 to S106 has been performed is less than N, the process returns to step S101, and after regenerating the disturbance in step S101, the processing of steps S102 and thereafter is performed to repeatedly accumulate the attention information. On the other hand, if the number of times the processing of steps S101 to S106 has been performed reaches N, the disturbance evaluation processing shown in the flowchart of FIG. 4 is terminated.
[0053] As described above, in the disturbance evaluation process of steps S30 and S70 in Fig. 3, the processes of steps S101 to S107 are repeated N times. As a result, attention information is accumulated each time the process is repeated. Because the disturbance superimposed on the input image is highly random, the accumulation result of attention information after N times has been carried out will have a large attention value for the part corresponding to the attention area, and the value for other parts will be relatively small. Therefore, the attention area can be expressed by the accumulation result of attention information.
[0054] FIG. 5 is a flowchart showing the procedure of the disturbance ratio adjustment process executed in step S80 of FIG.
[0055] In step S201, the tallying unit 211 tally the N evaluation values (scores) calculated by the evaluation value acquisition unit 17 in the disturbance evaluation process in step S70 of Fig. 3. Here, for example, the number of times each evaluation value appears in the N calculation results is counted, and the count values for each obtained evaluation value are plotted as a histogram. This makes it possible to obtain the distribution of evaluation values for the inference results of the image recognition unit 16 for N disturbance images obtained by randomly superimposing disturbances on input images.
[0056] In step S202, the statistical processing unit 212 calculates a statistical value corresponding to the distribution of the evaluation values tallied in step S201. Here, the statistical value corresponding to the distribution of the evaluation values can be calculated by calculating the average value Sa and variance Sv of the tallied results for each evaluation value obtained from N calculation results. Note that the statistical value calculated in step S202 represents the magnitude of the influence that random disturbance superimposed on the input image has on the image recognition process, i.e., the aforementioned disturbance influence degree. In other words, in step S202, statistical information on the distribution of evaluation values for the inference results of the image recognition unit 16 is calculated as the disturbance influence degree.
[0057] In step S203, the determination unit 213 determines the degree of disturbance influence based on the statistical value calculated in step S202. Here, for example, using a determination threshold set in advance by the user via the GUI 23, the determination unit 213 determines whether the degree of disturbance influence is appropriate or whether it is too large or too small, and outputs the obtained determination result to the parameter adjustment unit 22. The method for determining the degree of disturbance influence in step S203 will be described later with reference to FIGS. 6 and 7.
[0058] In step S204, the parameter adjusting unit 22 determines whether the disturbance influence degree is "appropriate," "small influence," or "large influence," based on the determination result of the disturbance influence degree input from the determining unit 213 in step S203. If the disturbance influence degree is determined to be "appropriate," the process proceeds to step S205; if the disturbance influence degree is determined to be "small influence," i.e., the disturbance influence degree is too small, the process proceeds to step S206; if the disturbance influence degree is determined to be "large influence," i.e., the disturbance influence degree is too large, the process proceeds to step S207.
[0059] In step S205, the parameter adjusting unit 22 determines that the parameter adjustment is complete and ends the disturbance ratio adjustment process shown in the flowchart of Fig. 5. In this case, no parameter adjustment is performed, and therefore, in the determination process of step S90 executed following the disturbance ratio adjustment process of step S80 in the flowchart of Fig. 3, it is determined that the adjustment of the disturbance parameters is complete.
[0060] In step S206, the parameter adjustment unit 22 adjusts the parameter values so as to increase the disturbance ratio from the current value for the parameters stored in the disturbance setting value storage unit 13, thereby strengthening the strength of the disturbance. As a result, if the disturbance influence level is too small with the current parameter values and therefore the influence on the inference result is considered to be limited even if the disturbance is superimposed on the input image, the parameter adjustment is performed so as to automatically increase the ratio of the disturbance to be superimposed on the input image by feeding back the determination result of the disturbance influence level obtained in step S203.
[0061] In step S207, the parameter adjustment unit 22 adjusts the parameter values so as to weaken the intensity of the disturbance by lowering the disturbance ratio from the current value for the parameters stored in the disturbance setting value storage unit 13. As a result, if the disturbance influence level is too large with the current parameter values and it becomes difficult to obtain a correct inference result when the disturbance is superimposed on the input image, the parameter adjustment is performed so as to automatically lower the disturbance ratio of the disturbance to be superimposed on the input image by feeding back the determination result of the disturbance influence level obtained in step S203.
[0062] In step S206 or S207, once the parameters in the disturbance set value storage unit 13 have been adjusted so that the disturbance ratio is changed in accordance with the degree of disturbance influence, the disturbance ratio adjustment process shown in the flowchart of FIG. 5 is terminated.
[0063] FIG. 6 is a flowchart showing an example of a method for determining the degree of influence of disturbance, which is performed in step S203 of FIG.
[0064] In step S301, the average value Sa of the statistical values calculated in step S202 in Fig. 5 is compared with a preset upper threshold value Sa_thu for the average value. If the average value Sa is greater than the upper threshold value Sa_thu, the process proceeds to step S304, and if the average value Sa is equal to or less than the upper threshold value Sa_thu, the process proceeds to step S302.
[0065] In step S302, the average value Sa is compared with a preset lower threshold value Sa_thl of the average value. If the average value Sa is less than the lower threshold value Sa_thl, the process proceeds to step S303. If the average value Sa is equal to or greater than the lower threshold value Sa_thl, the disturbance influence degree is determined to be appropriate, and the determination of the disturbance influence degree is terminated.
[0066] In step S303, the variance Sv of the statistical values calculated in step S202 in Fig. 5 is compared with a preset threshold value Sv_th. If the variance Sv is less than the threshold value Sv_th, it is determined that the disturbance influence is too large and the evaluation value is biased toward the low evaluation side, and the determination of the disturbance influence level is terminated. On the other hand, if the variance Sv is equal to or greater than the threshold value Sv_th, it is determined that the disturbance influence level is appropriate, and the determination of the disturbance influence level is terminated.
[0067] In step S304, the variance Sv is compared with a threshold value Sv_th. If the variance Sv is less than the threshold value Sv_th, it is determined that the disturbance influence is too small and the evaluation value is biased toward the high evaluation side, and the determination of the disturbance influence is terminated. On the other hand, if the variance Sv is equal to or greater than the threshold value Sv_th, it is determined that the disturbance influence is appropriate, and the determination of the disturbance influence is terminated.
[0068] FIG. 7 shows an example distribution of evaluation values (scores) for each degree of disturbance influence. Case (1) on the left side of FIG. 7 shows an example distribution of evaluation values (scores) in the attention information obtained by repeating steps S101 to S107 of FIG. 4 N times when the degree of disturbance influence is too large, and an example visualization result of the attention area output in step S40 of FIG. 3. In this case, the disturbance prevents an accurate inference result from being obtained, and therefore the evaluation values in the attention information are concentrated at 0, indicating that the attention area cannot be visualized. Therefore, in the process of FIG. 6, steps S301 and S302 determine that the mean value Sa is less than the lower threshold Sa_thl, and the subsequent step S303 determines that the variance Sv is less than the threshold Sv_th, thereby determining that the degree of disturbance influence is too large. As a result, in step S207 of FIG. 5, the parameter values are adjusted to lower the disturbance ratio.
[0069] In Figure 7, Case (2) shown in the center shows an example of the distribution of evaluation values (scores) in the attention information obtained by repeating steps S101 to S107 of Figure 4 N times when the disturbance influence level is appropriate, and an example of the visualization result of the attention area output in step S40 of Figure 3. In this case, a correct inference result is obtained when a specific attention area (e.g., a person) in the input image is not obscured by the disturbance, whereas an incorrect inference result is not obtained when the attention area is obscured. Therefore, it can be seen that the evaluation values are appropriately distributed in the attention information, and the attention area is visualized in a way that makes it easy to see. Therefore, in the process of Figure 6, in steps S301 and S302, the average value Sa is determined to be equal to or greater than the lower threshold Sa_thl and equal to or less than the upper threshold Sa_thu, and therefore the disturbance influence level is determined to be appropriate. As a result, in step S205 of Figure 5, it is determined that the parameter adjustment is complete, and no parameter value adjustment is performed.
[0070] In Figure 7, Case (3) on the right shows an example distribution of evaluation values (scores) in the attention information obtained by repeating steps S101 to S107 of Figure 4 N times when the disturbance influence is too small, and also shows an example of the visualization result of the attention area output in step S40 of Figure 3. In this case, even if the disturbance is superimposed on the input image, the inference result does not change significantly, making it difficult to determine the location of the attention area within the input image, and the visualization quality of the attention area is degraded. Therefore, in the process of Figure 6, in step S301, the mean value Sa is determined to be greater than the upper threshold Sa_thu, and in the subsequent step S304, the variance Sv is determined to be less than the threshold Sv_th, thereby determining that the disturbance influence is too small. As a result, in step S206 of Figure 5, the parameter values are adjusted to increase the disturbance ratio.
[0071] According to the first embodiment of the present invention described above, the following advantageous effects are achieved.
[0072] (1) The analysis method of image recognition processing using the analysis device 1 is a method for analyzing image recognition processing to recognize a target object reflected in an input image. In this analysis method, a computer processor 101 generates multiple disturbance images by superimposing a disturbance on the input image (step S102), and performs image recognition processing on each of the multiple disturbance images to obtain evaluation values for the object recognition results (step S104). Then, based on the evaluation values obtained for each disturbance image, the degree of disturbance influence on the image recognition processing is calculated (step S202), and the disturbance is adjusted based on the calculated disturbance influence (steps S206 and S207). This method automatically adjusts the disturbance necessary for visualizing a region of interest in inference during image recognition processing, significantly improving the efficiency of analyzing large data sets.
[0073] (2) In step S202, the degree of disturbance influence is calculated based on the distribution of the evaluation values obtained for each disturbance image. Specifically, statistical information (average value Sa, variance Sv) of the distribution of the evaluation values obtained for each disturbance image is calculated, and this statistical information is used as the degree of disturbance influence. In this way, the degree of disturbance influence that the disturbance has on the image recognition process can be calculated appropriately and quantitatively.
[0074] (3) In adjusting the disturbance, if the distribution of the evaluation values is biased toward the high evaluation side (steps S301 and S304: Yes), the intensity of the disturbance is adjusted to be stronger (step S206), and if the distribution of the evaluation values is biased toward the low evaluation side (steps S302 and S303: Yes), the intensity of the disturbance is adjusted to be weaker (step S207). In this way, the intensity of the disturbance can be automatically adjusted to an optimal value according to the distribution of the evaluation values.
[0075] (4) The analysis device 1 is a device for analyzing image recognition processing for recognizing a recognition target object reflected in an input image. The analysis device 1 includes a disturbance superimposition unit 15 that generates multiple disturbance images by superimposing a disturbance on an input image, an evaluation value acquisition unit 17 that acquires an evaluation value for the object recognition results obtained when an image recognition unit 16 executes image recognition processing on each of the multiple disturbance images generated by the disturbance superimposition unit 15, a disturbance ratio evaluation unit 21 that calculates the disturbance influence degree of the disturbance on the image recognition processing based on the evaluation value acquired for each disturbance image by the evaluation value acquisition unit 17, and a parameter adjustment unit 22 that adjusts the disturbance based on the disturbance influence degree calculated by the disturbance ratio evaluation unit 21. As described above, the analysis device 1 can automatically adjust the disturbance necessary for visualizing a region of interest in inference during image recognition processing, thereby significantly improving the efficiency of analyzing large datasets.
[0076] Second Embodiment Next, an analysis method and an analysis device for image recognition processing according to a second embodiment of the present invention will be described. In the first embodiment, an example of adjusting the disturbance ratio among the disturbance parameters was described. In this embodiment, an example of adjusting the disturbance resolution will be described. The hardware configuration of the analysis device 1A according to this embodiment is the same as the hardware configuration of FIG. 1 described in the first embodiment. Therefore, in the following description, the analysis device 1A of this embodiment will be described using the hardware configuration of FIG. 1.
[0077] 8 is a functional block diagram showing an example of the functional configuration of an analysis device 1A according to a second embodiment of the present invention. The analysis device 1A differs from the analysis device 1 described in the first embodiment in that it further includes a disturbance resolution evaluation unit 24.
[0078] The disturbance resolution evaluation unit 24 has a focus area distribution analysis unit 241 and a determination unit 242. The focus area distribution analysis unit 241 analyzes the distribution of focus areas that were focused on in the image recognition process when the image recognition unit 16 recognized an object in the disturbance image, based on the information about the focus areas obtained by the focus information integration unit 19 integrating the focus information. The determination unit 242 determines whether the disturbance resolution is appropriate based on the analysis result of the focus area distribution by the focus area distribution analysis unit 241. The determination result obtained is then notified to the parameter adjustment unit 22 as the evaluation result of the disturbance resolution.
[0079] The parameter adjusting unit 22 adjusts the disturbance resolution in addition to adjusting the disturbance ratio as described in the first embodiment. That is, based on the evaluation result of the disturbance resolution by the disturbance resolution evaluating unit 24, the parameter adjusting unit 22 adjusts the parameters stored in the disturbance setting value storing unit 13 so that the disturbance resolution becomes an appropriate value.
[0080] 9 is a flowchart showing an analysis method of image recognition processing according to a second embodiment of the present invention. In this embodiment, the analysis device 1A can analyze the image recognition results obtained by the image recognition unit 16 by causing the processor 101 to execute the processing shown in the flowchart of FIG.
[0081] In steps S10 to S80, the same processing as in the first embodiment is performed. After the processing in step S80 is performed, in the following step S81, a disturbance resolution adjustment process is performed based on the disturbance evaluation result obtained by the disturbance evaluation process executed in step S70. Here, the disturbance resolution evaluation unit 24 analyzes the distribution of the region of interest using the information of the region of interest generated in the disturbance evaluation process, i.e., the integration result of the information of interest by the information integration unit 19. Then, based on the obtained analysis result, it is determined whether the disturbance resolution is appropriate, and the determination result is output to the parameter adjustment unit 22. The parameter adjustment unit 22 adjusts the value of the parameter related to the disturbance resolution among the disturbance parameters stored in the disturbance setting value storage unit 13 based on the determination result of the disturbance resolution input from the disturbance resolution evaluation unit 24. Details of this disturbance resolution adjustment process will be described later with reference to the flowchart of FIG. 10.
[0082] In step S90, it is determined whether the adjustment of the disturbance parameters has been completed based on the determination result of the disturbance ratio obtained by the disturbance ratio adjustment process executed in step S80 and the determination result of the disturbance resolution obtained by the disturbance resolution adjustment process executed in step S81. Here, if the determination results indicating that the disturbance ratio and disturbance resolution were appropriate were obtained in the disturbance ratio adjustment process and disturbance resolution adjustment process executed immediately before, and therefore no adjustment of the parameter values was performed, it is determined that the adjustment of the disturbance parameters has been completed, and the process shown in the flowchart of FIG. 9 is terminated. On the other hand, if the determination result indicating that the disturbance ratio was inappropriate was obtained in the disturbance ratio adjustment process executed immediately before, or if the determination result indicating that the disturbance resolution was inappropriate was obtained in the disturbance resolution adjustment process executed immediately before, and therefore adjustment of the parameter values was performed, it is determined that the adjustment of the disturbance parameters has not been completed, and the process returns to step S70. In this case, the disturbance evaluation process of step S70, the disturbance ratio adjustment process of step S80, and the disturbance resolution adjustment process of step S81 are executed again using the adjusted parameter values. This allows for iterative adjustment of the disturbance parameters until an appropriate disturbance ratio and disturbance resolution is obtained for the image data set being analyzed.
[0083] FIG. 10 is a flowchart showing the procedure of the disturbance resolution adjustment process executed in step S81 of FIG.
[0084] In step S401, the attention area distribution analysis unit 241 acquires the distribution of the visualization results of the attention area generated by the attention information integration unit 19 integrating N pieces of attention information in the disturbance evaluation process of step S70 in Fig. 9, in the x direction (horizontal direction of the input image) and the y direction (vertical direction of the input image). Note that if a predetermined condition is satisfied, for example, if the size ratio between the x direction and the y direction of the object area indicated by the correct label acquired by the teacher data acquisition unit 12 is equal to or greater than a predetermined value (e.g., 2:1), the distribution of the visualization results of the attention area may be acquired only in one of the directions with the larger size. In this case, the processing from step S402 onward may be performed only for that direction.
[0085] In step S402, the determination unit 242 determines whether the distribution of attention regions acquired in step S401 exceeds the object region indicated by the correct label acquired by the teacher data acquisition unit 12. Here, for example, when the values of the attention regions are normalized in the range of 0 to 1, if a portion of the distribution of attention regions acquired in step S401 having a value of 0.8 or more falls within the object region, it is determined that the distribution of attention regions does not exceed the object region, and the process proceeds to step S403. On the other hand, if that portion protrudes from the object region, it is determined that the distribution of attention regions exceeds the object region, and the process proceeds to step S406.
[0086] In step S403, the attention area distribution analysis unit 241 acquires the spatial frequency of the distribution of the attention areas acquired in step S401. Here, for example, the spatial frequency of the distribution of the attention areas can be acquired by acquiring the number of peaks of the distribution of the attention areas in the object area for each of the x direction and the y direction.
[0087] In step S404, the determination unit 242 determines whether the spatial frequency of the distribution of the region of interest acquired in step S403 is appropriate. Here, for example, as described above, if the number of peaks in the distribution of the region of interest is acquired as the spatial frequency, and the number of peaks is equal to or greater than a predetermined number (e.g., 2), the spatial frequency of the distribution of the region of interest is determined to be appropriate, and the process proceeds to step S405. Note that the spatial frequency of the distribution of the region of interest may be determined to be appropriate if the number of peaks is equal to or greater than a predetermined number in both the x and y directions, or may be determined to be appropriate if the number of peaks is equal to or greater than a predetermined number in at least one of the directions. On the other hand, if the number of peaks is less than the predetermined number, the spatial frequency of the distribution of the region of interest is determined to be insufficient, and the process proceeds to step S406.
[0088] In step S405, the parameter adjusting unit 22 determines that the parameter adjustment is complete and ends the disturbance resolution adjustment process shown in the flowchart of Fig. 10. In this case, the disturbance resolution is deemed appropriate, and no parameter adjustment is performed.
[0089] In step S406, the parameter adjusting unit 22 adjusts the parameter values stored in the disturbance setting value storage unit 13 so as to increase the disturbance resolution from the current value. As a result, the determination results for the disturbance resolution performed in steps S402 and S404 are fed back, and parameter adjustment is performed so as to automatically increase the resolution of the disturbance to be superimposed on the input image.
[0090] In step S406, the parameters in the disturbance set value storage unit 13 are adjusted to increase the insufficient disturbance resolution, and then the disturbance resolution adjustment process shown in the flowchart of FIG. 10 is terminated.
[0091] FIG. 11 shows an example of the visualization results and distribution of the region of interest for each disturbance resolution. In FIG. 11, the upper column shows an example of the visualization of the region of interest obtained when the disturbance resolution is low, and an example of the distribution of this region of interest in the y direction. In this case, it can be seen that the low disturbance resolution results in only one peak in the distribution of the region of interest in the y direction. Therefore, in the processing of FIG. 10, step S404 determines that the spatial frequency of the distribution of the region of interest is insufficient. As a result, in step S406, the parameter values are adjusted to increase the disturbance resolution.
[0092] In Figure 11, the center column shows an example of a visualized region of interest obtained when the disturbance resolution is medium, along with an example of the distribution of this region of interest in the y direction. In this case, it can be seen that there are two peaks in the distribution of the region of interest in the y direction. Therefore, in the process of Figure 10, step S404 determines that the spatial frequency of the distribution of the region of interest is appropriate. As a result, in step S405, it is determined that the parameter adjustment is complete, and no adjustment of the parameter value is performed.
[0093] In Figure 11, the lower section shows an example of a visualized region of interest obtained when the disturbance resolution is high, along with an example of the distribution of this region of interest in the y direction. In this case, it can be seen that there are three peaks in the distribution of the region of interest in the y direction. Therefore, in this case as well, as in the case of a medium disturbance resolution, in the processing of Figure 10, it is determined in step S404 that the spatial frequency of the distribution of the region of interest is appropriate. As a result, it is determined in step S405 that the parameter adjustment is complete, and no adjustment of the parameter value is performed.
[0094] As described above, in the analyzer 1A of this embodiment, a medium to high disturbance resolution is determined to be appropriate, and if the disturbance resolution is low, parameter adjustment is performed to increase the disturbance resolution.
[0095] Note that there is no problem if the disturbance resolution is too fine, so in the analysis device 1A of this embodiment, it is preferable to adjust the disturbance parameters in a direction that increases the insufficient resolution. However, generally, the lower the disturbance resolution, i.e., the coarser the disturbance superimposed on the input image, the easier it is for the evaluation result of the disturbance ratio to converge. Therefore, if the image recognition process is performed with a high disturbance resolution from the beginning, the evaluation result of the disturbance ratio may not converge. Therefore, if the adjustment of the disturbance ratio does not converge for a long time (for example, if it does not converge even after five attempts), the disturbance resolution may be temporarily lowered and the disturbance ratio may be adjusted again.
[0096] According to the second embodiment of the present invention described above, the processor 101, which is a computer, evaluates the resolution of the disturbance based on the distribution of the attention areas in the image recognition process when an object is recognized in the disturbance image (steps S402 and S404). The disturbance is adjusted based on the result of the evaluation of the resolution of the disturbance (step S406). In this way, the resolution of the disturbance can be automatically adjusted to an optimal value according to the distribution of the attention areas.
[0097] Third Embodiment Next, an analysis method and analysis device for image recognition processing according to a third embodiment of the present invention will be described. In the first and second embodiments, an example was described in which only one object to be recognized in the image recognition processing is reflected in the input image. In this embodiment, however, an example is described in which multiple objects to be recognized are reflected in the input image, and parameter adjustment is performed by applying a disturbance to each of the multiple objects. Note that the hardware configuration of the analysis device 1B according to this embodiment is the same as the hardware configuration of FIG. 1 described in the first embodiment. Therefore, in the following description, the analysis device 1B of this embodiment will be described using the hardware configuration of FIG. 1.
[0098] 12 is a functional block diagram showing an example of the functional configuration of an analysis device 1B according to a third embodiment of the present invention. The analysis device 1B differs from the analysis device 1A described in the second embodiment in that it further includes an analysis area determination unit 25 and an analysis area quality determination unit 26.
[0099] The analysis area determination unit 25 determines an analysis area within the input image corresponding to each of a plurality of recognition targets reflected in the input image, based on the correct answer label acquired by the teacher data acquisition unit 12. In this embodiment, for each portion of the input image corresponding to the analysis area set for each recognition target by the analysis area determination unit 25, a disturbance that becomes noise when performing image recognition processing is generated by the disturbance generation unit 14 based on the parameter setting values stored in the disturbance setting value storage unit 13, and the disturbance superposition unit 15 superposes the generated disturbance.
[0100] The analysis area quality determination unit 26 determines whether the analysis area for each recognition object determined by the analysis area determination unit 25 is appropriate or not, based on the disturbance ratio evaluation result by the disturbance ratio evaluation unit 21 and the disturbance resolution evaluation result by the disturbance resolution evaluation unit 24. If it is determined that the analysis area is not appropriate as a result, it instructs the analysis area determination unit 25 to reset the analysis area.
[0101] 13 is a flowchart showing an analysis method of image recognition processing according to the third embodiment of the present invention. In this embodiment, the analysis device 1B can analyze the image recognition results obtained by the image recognition unit 16 by causing the processor 101 to execute the processing shown in the flowchart of FIG.
[0102] In steps S10 to S60, the same processing as in the first embodiment is performed. After the processing in step S60 is performed, in the subsequent step S61, the analysis area determination unit 25 sets analysis areas for each of the multiple objects reflected in the input image included in the image dataset to be analyzed acquired in step S60. Here, for example, for each object area indicated by the correct label included in the image dataset to be analyzed acquired in step S60, the object area is expanded by a predetermined margin rate (e.g., 10%, 15%, 20%, etc.), thereby setting multiple analysis areas with different margin rates for each object.
[0103] In step S62, the overlap between the analysis areas set in step S61 is calculated by the analysis area quality determination unit 26. Here, it is determined whether the analysis areas of each object set for each margin rate overlap on the input image.
[0104] In step S63, the analysis area quality determination unit 26 extracts a combination of analysis areas that do not overlap each other based on the overlap between the analysis areas calculated in step S62. Here, from the combinations of analysis areas for each object set for each margin rate in step S61, a combination of analysis areas that does not overlap each other is extracted, with the highest margin rate. This makes it possible to set the optimal analysis area for each object.
[0105] In step S64, one of the analysis areas is selected from the combination of analysis areas extracted in step S63. In steps S70 to S90, which are performed after step S64, the same processing as that described in the first and second embodiments is performed on the partial image within the analysis area selected in step S64.
[0106] In step S91, it is determined whether all of the analysis areas extracted in step S63 have been selected in step S64. If there are any unselected analysis areas, the process returns to step S64, and after selecting one of the unselected analysis areas in step S64, the processes of steps S70 to S90 are repeated for the partial image within that analysis area. If all analysis areas have been selected, the process shown in the flowchart of FIG. 13 is completed.
[0107] Fig. 14 is a diagram showing an example of an analysis region set for each object. In the example shown in Fig. 14, analysis regions 151 and 152 are set for objects 141 and 142 that appear in the input image by expanding their ranges so that they do not overlap each other.
[0108] In this embodiment, in order to confirm whether the combination of analysis areas extracted in step S63 is appropriate, the processes from step S64 onwards may also be performed on analysis areas with different margin rates. For example, by making the number of times disturbance images are generated less than the above-mentioned predetermined number N and visualizing and presenting the attention areas obtained for each analysis area to the user, the user can confirm that there are no changes in these attention areas.
[0109] According to the third embodiment of the present invention described above, the processor 101, which is a computer, sets an analysis area for each of a plurality of objects reflected in the input image (steps S61 to S63), and calculates the disturbance influence degree for each analysis area to adjust for the disturbance (steps S64 to S81). By doing so, it is possible to shorten the processing time even when a plurality of objects to be recognized are reflected in the input image.
[0110] In the third embodiment of the present invention described above, the analysis device 1B does not need to have the disturbance resolution evaluation unit 24. In this case, the analysis device 1B only adjusts the disturbance ratio, as in the first embodiment, and does not adjust the disturbance resolution.
[0111] Fourth Embodiment Next, a method and apparatus for analyzing image recognition processing according to a fourth embodiment of the present invention will be described. In the third embodiment, an example was described in which, when multiple objects to be recognized in image recognition processing appear in an input image, an analysis region was set for each object and parameter adjustment was performed. However, this parameter adjustment method is not effective when objects are closely spaced. Therefore, in this embodiment, an example is described in which objects appearing in an input image are grouped, images are reconstructed for each object group across the entire data set, and parameter adjustment is performed, thereby efficiently analyzing the entire data set. The hardware configuration of an analysis apparatus 1C according to this embodiment is the same as the hardware configuration shown in FIG. 1 described in the first embodiment. Therefore, in the following description, the analysis apparatus 1C according to this embodiment will be described using the hardware configuration shown in FIG. 1.
[0112] 15 is a functional block diagram showing an example of the functional configuration of an analysis device 1C according to a fourth embodiment of the present invention. The analysis device 1C differs from the analysis device 1B described in the third embodiment in that it further includes an image division and allocation unit 27.
[0113] The image division / allocation unit 27 divides the input image acquired by the input image acquisition unit 11 into units of objects reflected in the input image, thereby acquiring segmented images for each object from the input image. Then, the image to be analyzed is reconstructed by appropriately combining and synthesizing the segmented images acquired for the entire data set. In this embodiment, a disturbance that becomes noise when performing image recognition processing is generated by the disturbance generation unit 14 based on the parameter setting values stored in the disturbance setting value storage unit 13 for the image reconstructed by the image division / allocation unit 27, and the disturbance superposition unit 15 superposes the generated disturbance.
[0114] 16 is a flowchart showing an analysis method of image recognition processing according to a fourth embodiment of the present invention. In this embodiment, the analysis device 1C can analyze the image recognition results obtained by the image recognition unit 16 by having the processor 101 execute the processing shown in the flowchart in FIG.
[0115] In steps S10 to S60, the same processes as those in the first embodiment are performed. After the process in step S60 is performed, in the subsequent step S65, the image division and allocation unit 27 selects one of the image datasets to be analyzed acquired in step S60, and reads the combination of the input image and the correct label in that dataset.
[0116] In step S66, the image division / allocation unit 27 groups each object reflected in the input image included in the data set read in step S61 according to its size. Here, for example, each object is allocated to one of three groups, "large," "medium," or "small," based on the size of the area of each object indicated by the correct label included in the data set. Note that the number of groups to which objects are allocated is not limited to three, and the number of groups to which objects are allocated can be set arbitrarily depending on the specifications of the image recognition process performed by the image recognition unit 16, etc.
[0117] In step S67, it is determined whether grouping has been completed for the objects in the input images included in all the data sets to be analyzed acquired in step S60. If grouping by the process of step S66 has been completed for all the recognition targets reflected in the input images of all the data sets, the process proceeds to step S68. On the other hand, if there are unclassified objects for which the process of step S66 has not been performed, the process returns to step S65, a data set is selected again, and the process of step S66 is performed again to continue grouping the objects.
[0118] In step S68, the image dividing and allocating unit 27 cuts out images of each object grouped in step S66 from the input image and combines them by size to generate a composite image to be analyzed. For example, if the objects are classified into the three groups of "large," "medium," and "small" as mentioned above, partial images within the analysis area set by the analysis area determining unit 25 for the objects in each group are cut out from the input image in which the object is reflected. Then, a different number of partial images are combined for each group to generate a composite image.
[0119] Specifically, for each object classified into the "large" group, for example, one partial image is assigned to each composite image, and the partial image of each object is used as is as the composite image to be analyzed. On the other hand, for each object classified into the "medium" group, for example, two partial images are assigned to each composite image, and two partial images are combined to form one composite image. In this case, two regions of the composite image (e.g., the left and right regions) are assigned to different objects. Also, for each object classified into the "small" group, for example, four partial images are assigned to each composite image, and four partial images are combined to form one composite image. In this case, four regions of the composite image (e.g., the upper left, upper right, lower left, and lower right regions) are assigned to different objects. Note that the number of partial images assigned to each group in the composite image is not limited to this and can be any number.
[0120] In step S69, one of the partial images of each object included in each composite image generated in step S68 is selected from the partial images of each object included in each composite image. In steps S70 to S90 performed after step S69, the same processing as that described in the first and second embodiments is performed on the partial image selected in step S69.
[0121] In step S92, it is determined whether all partial images in all composite images generated in step S68 have been selected in step S69. If there are any unselected composite images or partial images, the process returns to step S69, and after selecting one of them in step S69, the processes of steps S70 to S90 are repeated for that partial image. If all partial images in all composite images have been selected, the process shown in the flowchart of FIG. 16 is completed.
[0122] According to the fourth embodiment of the present invention described above, the processor 101, which is a computer, cuts out partial images corresponding to the analysis area for each object from the input image (steps S65 to S67), generates a composite image by combining the multiple partial images cut out from different input images (step S68), and calculates the disturbance influence degree for each partial image in this composite image to adjust for the disturbance (steps S69 to S81). As a result, when multiple objects to be recognized appear in the input image, the data set can be analyzed more efficiently.
[0123] In the fourth embodiment of the present invention described above, similarly to the third embodiment, the analysis device 1C may not have the disturbance resolution evaluation unit 24. In this case, similarly to the first embodiment, the analysis device 1C only adjusts the disturbance ratio, but does not adjust the disturbance resolution.
[0124] The following modifications may be applied to each of the first to fourth embodiments described above.
[0125] (Variation) In each of the flowcharts of FIGS. 3, 9, 13, and 16, the disturbance may be evaluated for appropriateness during the disturbance evaluation process of FIG. 4, which is executed in steps S30 and S70, to determine whether to continue generating disturbance images. Specifically, for example, in the disturbance evaluation process of FIG. 4, an evaluation value for each disturbance image is acquired in step S104 until the number of generated disturbance images reaches a predetermined upper limit (the number of times N described above). Here, when the number of generated disturbance images is less than the upper limit N, the disturbance ratio evaluation unit 21 aggregates the distribution of the evaluation values for the disturbance images acquired up to that point. If the distribution deviation exceeds a predetermined deviation state (e.g., if the variance Sv is less than the threshold Sv_th described above), the generation of the disturbance image is terminated. Then, after the parameter adjustment unit 22 adjusts the disturbance by changing the disturbance parameters, the disturbance evaluation process is resumed, restarting the generation of the disturbance image from the beginning. This prevents the disturbance evaluation process from being continued due to an inappropriate disturbance, further shortening the processing time.
[0126] The present invention is not limited to the various embodiments and modifications described above, and includes various other modifications. For example, the above-described embodiments have been specifically described to clearly explain the present invention, and are not necessarily limited to those having all of the described configurations. Furthermore, part of the configuration of one embodiment can be replaced with part of the configuration of another embodiment. Furthermore, the configuration of another embodiment can be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment can be deleted, and part of another configuration can be added or replaced with part of another configuration.
[0127] The present invention is not limited to the above-described embodiment, and various modifications are possible without departing from the spirit of the present invention.
[0128] 1, 1A, 1B, 1C: image recognition processing analysis device (analysis device), 11: input image acquisition unit, 12: teacher data acquisition unit, 13: disturbance setting value storage unit, 14: disturbance generation unit, 15: disturbance superposition unit, 16: image recognition unit, 17: evaluation value acquisition unit, 18: attention information generation unit, 19: attention information integration unit, 20: drawing unit, 21: disturbance ratio evaluation unit, 211: aggregation unit, 212: statistical processing unit, 213: judgment unit, 22: parameter adjustment unit, 23: GUI (Graphical User Interface), 24: disturbance resolution evaluation unit, 241: attention area distribution analysis unit, 242: judgment unit, 25: analysis area determination unit, 26: analysis area pass / fail judgment unit, 27: image division / allocation unit
Claims
1. A method for analyzing an image recognition process for recognizing an object to be recognized reflected in an input image, comprising: - generating, by a computer, a plurality of disturbed images obtained by superimposing disturbances on the input image; - obtaining, by the computer, an evaluation value for the recognition result of the object obtained by respectively performing the image recognition process on the plurality of disturbed images; - calculating, by the computer, a disturbance influence degree of the disturbance on the image recognition process based on the evaluation value obtained for each of the disturbed images; - adjusting, by the computer, the disturbance based on the disturbance influence degree. A method for analyzing an image recognition process.
2. The method for analyzing an image recognition process according to claim 1, wherein the disturbance influence degree is calculated based on the distribution of the evaluation values obtained for each of the disturbed images.
3. The method for analyzing an image recognition process according to claim 2, wherein statistical information in the distribution of the evaluation values is calculated and used as the disturbance influence degree.
4. The method for analyzing an image recognition process according to claim 2 or 3, wherein when the distribution of the evaluation values is biased toward the high - evaluation side, the intensity of the disturbance is adjusted to be increased, and when the distribution of the evaluation values is biased toward the low - evaluation side, the intensity of the disturbance is adjusted to be decreased.
5. The method for analyzing an image recognition process according to claim 2 or 3, wherein until the number of generated disturbed images reaches a predetermined upper limit number, an evaluation value of the disturbed image is obtained each time the disturbed image is generated, and when the number of generated disturbed images is less than the upper limit number and the bias of the distribution of the evaluation values exceeds a predetermined bias state, the generation of the disturbed images is terminated, the disturbance is adjusted, and then the generation of the disturbed images is restarted from the beginning.
6. The method for analyzing an image recognition process according to claim 1, wherein: - evaluating, by the computer, the resolution of the disturbance based on the distribution of the attention areas of the image recognition process when the object is recognized in the disturbed image; - adjusting, by the computer, the disturbance based on the evaluation result of the resolution of the disturbance. A method for analyzing an image recognition process.
7. In the method for analyzing an image recognition process according to claim 1, an analysis region is set for each of the plurality of objects reflected in the input image, and the disturbance influence degree is calculated for each analysis region to adjust the disturbance. A method for analyzing an image recognition process.
8. In the method for analyzing an image recognition process according to claim 7, a partial image corresponding to the analysis region is cut out from the input image for each object, a composite image is generated by combining a plurality of the partial images cut out from different input images, and the disturbance influence degree is calculated for each partial image with respect to the composite image to adjust the disturbance. A method for analyzing an image recognition process.
9. An analysis apparatus for analyzing an image recognition process for recognizing an object to be recognized reflected in an input image, comprising: a disturbance superimposing unit that generates a plurality of disturbance images obtained by superimposing a disturbance on the input image; an evaluation value acquisition unit that acquires an evaluation value for the recognition result of the object obtained by respectively executing the image recognition process on the plurality of disturbance images generated by the disturbance superimposing unit; a disturbance ratio evaluation unit that calculates a disturbance influence degree exerted by the disturbance on the image recognition process based on the evaluation value acquired for each disturbance image by the evaluation value acquisition unit; and a parameter adjustment unit that adjusts the disturbance based on the disturbance influence degree calculated by the disturbance ratio evaluation unit.
Citation Information
Patent Citations
Explanatory visualizations for object detection
US20210357644A1
Face shape library construction method and device, equipment and storage medium
CN111581412A
Confrontation sample generation method and device, electronic equipment and storage medium
CN114882323A
Image processing system and image processing method
JP2023001367A