Information processing device, server, and information processing system including the same
The described system addresses the inefficiencies in collecting low-robust data for DNNs by automating the identification and collection process, improving the robustness and reliability of DNNs in real-time applications with reduced human effort and resource usage.
Patent Information
- Application Number
- JP2023216879
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for improving the performance of operational deep neural networks (DNNs) in image recognition tasks face challenges in efficiently collecting and selecting low-robust data, which are necessary for relearning to enhance reliability and stability, due to the reliance on human visual inspection and judgment, leading to high human costs and inefficiencies.
An information processing apparatus and system that includes an image processing unit, inference visualization analysis, robustness determination, and data storage units to automatically identify and collect low-robust data by analyzing the attention areas in DNN inferences, reducing the need for manual visual judgment and optimizing resource usage.
Efficient selection and collection of low-robust data for relearning, enhancing the robustness and reliability of DNNs in real-time applications with reduced computational resources and human intervention.
Smart Images

Figure 2025099897000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus and a server, and more particularly to a technique for improving the effect and efficiency of data collection and data selection for improving the performance of an image recognition program including a machine learning model.
Background Art
[0002] In recent years, image recognition programs using machine learning such as deep neural networks (DNNs) have been put into practical use. For example, image recognition technology has been applied to various devices, services, and solutions such as anomaly inspection systems on production lines, systems for monitoring human intrusion and behavior, and automatic transport devices at logistics sites.
[0003] In order to use DNNs in various applications, it is important to efficiently collect and select learning data. However, it is necessary to collect and select learning data by visual confirmation and qualitative judgment by workers, which increases the human cost. In addition, when the created DNN is continuously operated for a long time, it is necessary to deal with the instability of inferences that were not noticed at the design stage and the performance degradation due to changes in the surrounding environment, and the cost for data collection and selection becomes an issue.
[0004] In order to improve the performance of an in-operation DNN (hereinafter referred to as an operation version DNN), it is necessary to select and collect data effective for performance improvement. For example, in an image recognition task, high performance and high functionality can be achieved by collecting data with recognition errors or data showing a new category (type) that could not be recognized so far and re-learning.
[0005] On the other hand, among the data with correct recognition results, there is also data with low confidence but accidentally correct, data where part of the object to be recognized is hidden by other objects and cannot be recognized, data that cannot be recognized due to the influence of light intensity or noise, etc., that is, data with low recognition stability and robustness (hereinafter referred to as low-robust data). That is, in order to realize an application with high reliability and stability, it is necessary to re-learn not only the data misrecognized by the operational DNN or unknown data, but also the low-robust data for the operational DNN.
[0006] Various existing technologies have been studied as methods for efficiently preparing data effective for the learning of the operational DNN. In Patent Document 1, there is described an information processing apparatus that inputs a plurality of input images into a learned model, causes the learned model to output prediction labels inferred from each of the plurality of input images, and acquires feature regions that are the basis for the prediction of the learned model in each of the plurality of input images, the information processing apparatus comprising: a processing reception means that receives a registration instruction for registering the input image as an image to be processed; a registration means that registers the input image as an image to be processed when the processing reception means receives the registration instruction; an update reception means that receives an update instruction for updating learning data for re-learning the learned model; and an update means that performs a processing operation on at least one of the images to be processed and updates the learning data using the image on which the processing operation has been performed when the update reception means receives the update instruction.
[0007] Also, Patent Document 2 describes a data collection system for collecting image data used by a learning device that recognizes images by machine learning to perform learning. The system includes a requirement definition processing unit that receives a requirement variable representing the requirements of the image data necessary for the learning device to sufficiently learn and requirement definition data specifying the requirement values of the requirement variable. The requirement definition processing unit further receives priority data specifying the priority of the requirement variable, and the requirement definition processing unit further presents the requirement variable and the requirement value in the order of the priority, and presents a requirement value response rate representing the ratio of the requirement variables for which the requirement values are specified among the requirement variables. A data collection system is characterized by providing a requirement confirmation screen.
Prior Art Documents
Patent Documents
[0008]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0009] The invention described in Patent Document 1 efficiently generates learning data for improving the accuracy of image classification AI. However, since the criteria for processing and generating learning data require visual inspection by developers and judgment based on rules of thumb, it is difficult to prepare effective learning data for improving the performance of the operational version of DNN. Further, the invention described in Patent Document 2 appropriately sets the range of image data used as learning data. However, it is the developer who defines the requirements for the collected data, and another technique is required to define low-robust data among the data necessary for improving the performance of the operational version of DNN.
[0010] An object of the present invention is to efficiently collect low-robust data that can improve the inference performance of a machine learning model during operation.
Means for Solving the Problems
[0011] An information processing apparatus is configured by a computer having an arithmetic unit that executes arithmetic processing and a storage unit accessible by the arithmetic unit, and includes an image processing unit that executes an image processing program including a machine learning model and outputs a recognition result of image data, an inference visualization analysis unit that specifies a region of interest in the inference in the image processing unit, a robustness determination unit that determines the robustness of the image data using the specified region of interest, a collected data storage unit that stores image data determined to have a robustness lower than a first predetermined standard in the storage unit, and an interface unit that transmits the image data stored in the collected data storage unit to the outside.
Advantages of the Invention
[0012] According to one aspect of the present invention, data with low robustness for the operation version DNN can be efficiently selected and collected. Problems, configurations, and effects other than those described above will be clarified by the description of the embodiments for carrying out the following invention.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Mode for Carrying Out the Invention
[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each embodiment is an exemplification for explaining the present invention, and is appropriately omitted and simplified for the sake of clarity of explanation. The present invention can be implemented in various other forms. Unless otherwise particularly limited, each component may be singular or plural.
[0015] The positions, sizes, shapes, ranges, etc. of the respective components shown in the drawings may not represent the actual positions, sizes, shapes, ranges, etc. in order to facilitate the understanding of the invention. For this reason, the present invention is not necessarily limited to the positions, sizes, shapes, ranges, etc. disclosed in the drawings. When there are a plurality of components having the same or similar functions, they may be described with the same reference numeral and different subscripts. Also, when it is not necessary to distinguish these plurality of components, the subscripts may be omitted in the description.
[0016] In each embodiment, the processing performed by executing a program may be described. Here, a computer executes a program by a processor (e.g., CPU, GPU), and performs the processing defined by the program while using a storage resource (e.g., memory) and an interface device (e.g., communication port), etc. Therefore, the subject of the processing performed by executing the program may be the processor. Similarly, the subject of the processing performed by executing the program may be a controller, device, system, computer, or node having a processor. The subject of the processing performed by executing the program may be an arithmetic unit or a dedicated circuit for performing a specific processing. Here, the dedicated circuit is, for example, an FPGA (Field Programmable Gate Array), ASIC (Application Specific Integrated Circuit), CPLD (Complex Programmable Logic Device), etc.
[0017] The program may be installed in a computer from a program source. The program source may be, for example, a program distribution server or a storage medium readable by a computer. When the program source is a program distribution server, the program distribution server includes a processor and a storage resource for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. Also, in an embodiment, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0018] In recent years, peripheral recognition applications applying deep neural networks (DNNs), which are a type of machine learning, have become widespread. In programs using machine learning, if there are scenes where recognition errors occur or scenes that need to be learned as new tasks, performance can be improved and reliability and function improvements can be achieved through relearning using such data. On the other hand, in programs using machine learning, compared to rule-based algorithms, the reasons for such inferences are unclear. For example, even if the recognition result is correct, there are cases where misrecognition occurs when there is a little occlusion, truncation, or shooting noise (such as fluctuations in brightness, chroma, noise, etc.) in the image. Such data that is prone to misrecognition is called "low-robust data". Since the recognition result itself of low-robust data is "correct", it is difficult to find it with general evaluation methods that only compare with correct data. On the other hand, by using low-robust data for relearning, the reliability and stability of inferences in the operational DNN can be enhanced.
[0019] As a technique for finding low-robust data, techniques for visualizing the areas of interest in DNN inferences (such as Grad-CAM, RISE, etc.) have been reported. For example, although the entire object to be detected is clearly shown, the inference result in the DNN accurately infers the position of the object, and furthermore, despite the inference confidence score being very high, when continuously inferring the object over time, undetected cases may frequently occur. As a result of visualizing the area of interest for the object, there are cases where only a part of the object is being focused on. In this case, the inference result for the object is likely to be undetected because the area of interest is hidden by other objects or becomes unclear due to the influence of the light amount or noise during shooting. Thus, visualizing the area of interest in inferences is effective for identifying the causes of DNN defects.
[0020] However, collecting and selecting low-robust data based on the results of visualizing the target area requires sensory evaluation and visual judgment by designers and developers. Therefore, when dealing with large-scale datasets in the retraining of machine learning, the man-hours involved become a major issue.
[0021] The present invention aims to improve the efficiency of collecting and selecting low-robust data, which is one of the data necessary for the update and evolution of the operation version of the DNN. In the following embodiments, an example of applying the present invention to an information processing system including an edge device (information processing apparatus) using object detection AI and a server (such as a cloud device or a development environment) connected to the edge device will be described. Different from a development environment such as a server, the edge device has restrictions on the computing resources of the computer and requires real-time processing. Therefore, the information processing system targeted by the present invention is, for example, an anomaly inspection system for a manufacturing line, a system for monitoring human intrusion and behavior, an automatic transport device at a logistics site, a construction machine or an autonomous vehicle at a construction site, etc., which processes images acquired by a camera and notifies the results, or is a device that feeds back to control the device on-site.
[0022] <Example 1> Hereinafter, with reference to FIGS. 1 to 6, the information processing apparatus 100 according to Example 1 of the present invention will be described. As described above, the information processing apparatus 100 is a device with restrictions on computing resources such as an edge device and requires real-time processing.
[0023] FIG. 1 is a block diagram showing the configuration of the information processing apparatus 100 according to Example 1.
[0024] The information processing apparatus 100 according to Example 1 includes a camera 101, an image processing unit 102, other sensors 103 other than the camera, a preprocessing unit 104, an application processing unit 105, an inference visualization analysis unit 106, a robustness determination unit 107, a collected data storage unit 108, and a communication interface 109.
[0025] Camera 101 captures an image to be processed by app processing unit 105. Image processing unit 102 outputs an inference result from the image captured by camera 101 using a machine learning model. Image processing unit 102 is composed of an image processing program including a machine learning model such as a CNN (Convolutional Neural Network) or Transformer for image recognition (hereinafter simply referred to as DNN), and outputs the type and position of the object shown in the input image.
[0026] Other sensor 103 is a sensor other than the camera, and for example, a vibration sensor, a microphone, LiDAR, a radar, an ultrasonic sensor, etc. can be used. The preprocessing unit 104 converts the position, vibration, sound spectrum, and other physical quantities of surrounding objects into data in a format that can be used by the app processing unit 105 through recognition processing such as FFT (Fast Fourier Transform) and point cloud processing, and outputs it.
[0027] The app processing unit 105 realizes the functions provided by the information processing device 100 (for example, object recognition, monitoring of restricted areas, monitoring of human behavior) using the inference result by the image processing unit 102. In addition to the determination result by the image processing unit 102, the app processing unit 105 may realize the functions provided by the information processing device 100 using the data output from the preprocessing unit 104. Various usage methods are conceivable, such as a configuration in which the output of the app is transmitted to the server side of the information processing system through the communication interface (I / F) 109 as shown in the figure, or a configuration in which it is input to the control function of the information processing device itself.
[0028] The inference visualization analysis unit 106 analyzes the inference results by the image processing unit 102, identifies the attention area in the image that contributed to the derivation of the conclusion in the inference by the DNN, and creates a heatmap (see FIGS. 5 and 6) representing the area with high attention. The attention area identified by the inference visualization analysis unit 106 is the area in the image that contributed to the conclusion derived by the DNN. The inference visualization analysis unit 106 may execute a process of calculating the confidence or uncertainty of the inference instead of or in addition to the process of identifying the attention area. As an example of the operation, the robustness determination unit 107 calculates the ratio of the area of the attention area in the image to the area of the object in the image, and evaluates the robustness of the image processing. If the area ratio of the attention area is equal to or less than a predetermined standard (for example, 50%), the image is determined to be a low-robustness image, and the image determined to be the low-robustness image is stored in the collection data storage unit 108 together with the heatmap information. The collection data storage unit 108 stores the image determined to be low-robust according to the determination result of the robustness determination unit 107.
[0029] The communication interface 109 controls communication with other devices. For example, the communication interface 109 transmits the processing result by the application processing unit 105 to an external server.
[0030] FIG. 2 is a block diagram showing the physical configuration of the computer constituting the information processing apparatus 100 according to the first embodiment.
[0031] The information processing apparatus 100 is composed of a computer having a processor (CPU) 11, a memory 12, an auxiliary storage device 13, and a communication interface 14. The information processing apparatus 100 may have an input interface and an output interface.
[0032] Processor 11 is an arithmetic unit that executes programs stored in memory 12. By executing various programs, processor 11 realizes each functional unit of information processing apparatus 100 (for example, image processing unit 102, preprocessing unit 104, application processing unit 105, inference visualization analysis unit 106, robustness determination unit 107, etc.). Note that a part of the processing performed by processor 11 when executing a program may be executed by another arithmetic unit (for example, hardware such as a GPU, ASIC, or FPGA).
[0033] Memory 12 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores invariant programs (for example, BIOS). The RAM is a high-speed and volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by processor 11 and data used during program execution.
[0034] Auxiliary storage device 13 is a large-capacity and non-volatile storage device such as a magnetic storage device (HDD) or a flash memory (SSD). Further, auxiliary storage device 13 stores data accessed by processor 11 during program execution and programs executed by processor 11, and constitutes, for example, collection data storage unit 108. That is, the program is read from auxiliary storage device 13, loaded into memory 12, and executed by processor 11, thereby realizing each function of information processing apparatus 100.
[0035] Communication interface 14 is a network interface device that controls communication with other devices according to a predetermined protocol.
[0036] The programs executed by processor 11 are provided to information processing apparatus 100 via a removable medium (external HDD, optical disk, flash memory, etc.) or a network, and are stored in non-volatile auxiliary storage device 13, which is a non-temporary storage medium. For this reason, information processing apparatus 100 may have an interface for reading data from a removable medium.
[0037] Figure 3 is a flowchart of the low-robustness data determination process executed by the information processing apparatus 100 of Example 1.
[0038] First, the inference visualization analysis unit 106 analyzes the determination result by the image processing unit 102, specifies the attention area in the inference by the DNN from the image, and creates a heat map representing the area with high attention (S301).
[0039] Then, the robustness determination unit 107 calculates the robustness of the specified attention area. As for the robustness, for example, a method of calculating it based on the area ratio of the specified attention area to a reference such as the area of an object on the image can be considered, but it may also be used as an index using, for example, the uniformity of the distribution of the attention area or the total value of the intensity. The area of the reference object may be the area of a rectangular bounding box (bbox) in the case of an object detection task, or the area of the area where the object is drawn (corresponding to a so-called segmentation mask).
[0040] After that, if the area ratio of the specified attention area is equal to or less than a predetermined threshold (for example, 10%), the robustness determination unit 107 determines that the image is low-robust (S303).
[0041] Then, the robustness determination unit 107 stores the images determined to be low-robust in the collected data storage unit 108 (S304).
[0042] As shown in FIGS. 5 and 6, the heat map created by the robustness determination unit 107 is used in the inference in the original image, and the intensity of the white color in the figure is strongly displayed as it contributes more strongly to the derivation of the conclusion. The specification of the heat map is not limited to this, and it may be displayed in a color scale, or only the areas with attention exceeding a predetermined threshold may be colored and displayed.
[0043] For example, in the inference of the pedestrian enclosed by the rectangle in the original image 500 shown in FIG. 4, as shown in FIG. 5, when the attention area spreads over the entire object and the ratio of the area of the attention area to the area of the bounding box 501 is greater than a predetermined threshold, even if a part of the attention area is missing due to occlusion such as being hidden by another object, a correct inference result can be obtained. Therefore, it is determined that the image is not low-robust. On the other hand, in the original image 500 shown in FIG. 4, as shown in FIG. 6, when the attention area is only a part of the object and the ratio of the area of the attention area to the area of the bounding box 501 is smaller than a predetermined threshold, if a part of the attention area is missing due to occlusion or truncation, etc., there is a possibility that a correct inference result such as a detection miss or a type mistake cannot be obtained. Therefore, it is determined that the image is low-robust.
[0044] Also, in step S302, the inference visualization analysis unit 106 may simultaneously execute a process of calculating the confidence or uncertainty of the inference. When the inference visualization analysis unit 106 calculates the confidence or uncertainty of the inference, the robustness determination unit 107 calculates a determination score (for example, taking the product of each other as the robustness score) including the confidence or uncertainty in addition to the ratio of the attention area, determines the data with a low score as low-robust data (S303), and may store the image determined to be low-robust in the collected data storage unit 108 (S304).
[0045] As described above, in the first embodiment of the present invention, among the images correctly recognized by the image processing unit 102, data with low robustness can be efficiently selected and collected, and by applying the misrecognized images to re-learning, the robustness of object recognition by the image processing unit 102 can be improved.
[0046] <Example 2> Hereinafter, with reference to FIGS. 7 to 12, the information processing system according to Embodiment 2 of the present invention will be described. In the information processing system according to Embodiment 2, the server 200 extracts an inference target area from the data collected by the information processing apparatus 100, identifies low-robust data, extracts feature amounts of the identified low-robust data, and generates a data collection criterion. Since the server 200 distributes the generated data collection criterion to the information processing apparatus 100, each information processing apparatus 100 can obtain the effect of always collecting data based on the latest data collection criterion, and there is also an advantage in that the information processing apparatus 100 does not need to perform the process of visualizing the inference target area, thereby reducing the calculation resources. In addition, in Embodiment 2, the differences from Embodiment 1 will be mainly described, and the same components and processes as those in Embodiment 1 will be denoted by the same reference numerals, and the description thereof will be omitted.
[0047] FIG. 7 is a block diagram showing the configuration of the information processing system according to Embodiment 2.
[0048] The information processing system according to Embodiment 2 includes one or more information processing apparatuses 100 and a server 200. The information processing apparatus 100 and the server 200 are communicably connected. In FIG. 7, only the main configuration of the information processing apparatus 100 is shown, and the detailed configuration will be described later with reference to FIG. 9.
[0049] The server 200 includes a data set storage unit 201, an inference target area visualization unit 202, a machine learning model 203, a robustness determination unit 204, a feature amount extraction unit 205, a data collection criterion generation unit 206, and a communication interface 207.
[0050] The dataset storage unit 201 stores the collected data transmitted from the information processing apparatus 100. The inference target area visualization unit 202 outputs an inference result from the data in the collected data stored in the dataset storage unit 201 for which the inference of the machine learning model 203 is correct, using the machine learning model 203. The machine learning model 203 is composed of a DNN (Deep Neural Network) that has learned images labeled with the types of objects shown therein, infers the objects shown in the input image, and outputs a determination result of the objects. Further, the inference target area visualization unit 202 analyzes the images included in the collected data stored in the dataset storage unit 201, and specifies the target area in the inference by the machine learning model 203 from the images. The target area specified by the inference target area visualization unit 202 is the area in the image that contributed to the conclusion derived by the machine learning model. The robustness determination unit 204 calculates the ratio of the area of the target area specified by the inference target area visualization unit 202 to the area of the object in the image, and if the area ratio of the target area is equal to or less than a predetermined threshold value, determines that the image is a low-robustness image, and controls so that an image for which the inference of the machine learning model 203 is correct and the robustness is small is transmitted to the data collection criterion generation unit 206.
[0051] The feature amount extraction unit 205 calculates the feature amounts of the low-robustness images determined by the robustness determination unit 204. The data collection criterion generation unit 206 clusters the image data arranged in the multi-dimensional space (illustrated two-dimensionally in FIG. 11) with the feature amount as the axis, and generates a data collection criterion representing the range of the cluster (see FIG. 11). The communication interface 207 controls communication with other devices. For example, the communication interface 207 transmits the data collection criterion generated by the data collection criterion generation unit 206 to the information processing apparatus 100, and receives the data collected by the information processing apparatus 100.
[0052] FIG. 8 is a block diagram showing the physical configuration of the computer constituting the server 200 of the second embodiment.
[0053] Server 200 is composed of a computer having a processor (CPU) 21, a memory 22, an auxiliary storage device 23, and a communication interface 24. Server 200 may have an input interface 25 and an output interface 26.
[0054] Processor 21 is an arithmetic unit that executes programs stored in memory 22. By the processor 21 executing various programs, each functional unit of the server 200 (for example, the inference target area visualization unit 202, the robustness determination unit 204, the feature amount extraction unit 205, the data collection criterion generation unit 206, etc.) is realized. Note that a part of the processing performed by the processor 21 when executing a program may be executed by another arithmetic unit (for example, hardware such as a GPU, ASIC, FPGA, etc.).
[0055] Memory 22 includes a ROM which is a non-volatile memory element and a RAM which is a volatile memory element. The ROM stores unchanging programs (for example, BIOS), etc. The RAM is a high-speed and volatile memory element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 21 and data used during program execution.
[0056] The auxiliary storage device 23 is a large-capacity and non-volatile storage device such as a magnetic storage device (HDD) or a flash memory (SSD), for example. Also, the auxiliary storage device 23 stores data accessed by the processor 21 during program execution and programs executed by the processor 21, and constitutes, for example, the data set storage unit 201. That is, the program is read from the auxiliary storage device 23, loaded into the memory 22, and executed by the processor 21 to realize each function of the server 200.
[0057] The communication interface 24 is a network interface device that controls communication with other devices according to a predetermined protocol.
[0058] The input interface 25 is an interface to which input devices such as the keyboard 27 and the mouse 28 are connected and which receives inputs from the operator. The output interface 26 is an interface to which output devices such as the display device 29 and a printer (not shown) are connected and which outputs the execution results of the program in a form visible to the user. Note that a terminal (not shown) connected to the server 200 via a network may provide the input device and the output device. In this case, the server 200 may have the function of a web server, and the terminal may access the server 200 using a predetermined protocol (for example, http).
[0059] The program executed by the processor 21 is provided to the server 200 via a removable medium (external HHD, optical disk, flash memory, etc.) or a network and stored in the non-volatile auxiliary storage device 13 which is a non-temporary storage medium. Therefore, the server 200 may preferably have an interface for reading data from the removable medium.
[0060] The server 200 is a computer system configured physically on one computer or on a plurality of computers configured logically or physically, and may operate on a virtual computer constructed on a plurality of physical computer resources. For example, each functional unit may operate on a separate physical or logical computer, or a plurality of them may be combined and operate on one physical or logical computer.
[0061] FIG. 9 is a block diagram showing the configuration of the information processing apparatus 100 according to the second embodiment.
[0062] The information processing apparatus 100 according to the second embodiment includes a camera 101, an image processing unit 102, other sensors 103 other than the camera, a preprocessing unit 104, an application processing unit 105, a feature amount extraction unit 112, a data collection determination unit 110, a collected data storage unit 108, a data collection reference storage unit 111, and a communication interface 109.
[0063] The camera 101, the image processing unit 102, the other sensors 103, the preprocessing unit 104, the application processing unit 105, and the communication interface 109 are the same as the information processing apparatus 100 of the aforementioned Example 1. The image processing unit 102 outputs an inference result from the image captured by the camera 101 using an image recognition program. The image recognition program is composed of a DNN (Deep Neural Network) that has learned images labeled with the types of objects shown, and is a program that infers the position and type of the object shown in the input image and outputs the inference result. Further, the image processing unit 102 outputs the region where the object is determined (for example, a rectangular bounding box).
[0064] The feature extraction unit 112 extracts features using the intermediate layer data of the DNN in the image processing unit 102 and the features of the region used for inference. By using existing techniques such as principal component analysis and t-SNE for feature extraction, multi-dimensional array data such as image data can be extracted as low-dimensional features. The data collection determination unit 110 compares the image features extracted by the feature extraction unit 112 with the image features that are the data collection criteria stored in the data collection criteria storage unit 111. If the similarity of the two features is equal to or greater than a predetermined threshold, it is determined that the image data is to be collected and stored in the collected data storage unit 108. The collected data storage unit 108 stores image data whose image features match the predetermined data collection criteria. The data collection criteria storage unit 111 stores the data collection criteria generated by the data collection criteria generation unit 206 of the server 200. The collection criteria stored in the data collection criteria storage unit 111 are frequently updated by the server 200, enabling the collection of data necessary for improving the performance of the machine learning model 203 during operation.
[0065] FIG. 10 is a flowchart of the low-robust data determination process executed by the information processing apparatus 100 of Example 2.
[0066] First, the image processing unit 102 performs inference on the input image using a DNN (Deep Neural Network) that constitutes the implemented machine learning model. At this time, data of an arbitrary intermediate layer in the DNN is extracted and sent to the feature extraction unit 112 (S311). It is desirable to select a highly sensitive layer that best represents the features of the image as the layer to be extracted. Generally, even if the weight coefficients of the DNN are updated by re-learning, if the algorithm does not change, the highly sensitive layer often does not change. Therefore, the selection of the extraction layer is specified by the developer during algorithm development.
[0067] Next, the feature extraction unit 112 calculates low-dimensional vector data (features) from the data extracted by the image processing unit 102 from the intermediate layer of the DNN (S312). The feature extraction unit 112 converts the multi-dimensional extracted data into low-dimensional features by a feature extraction algorithm such as GAP (Global average pooling) processing, principal component analysis processing, or dimensional compression method (t-SNE (t-Distributed Stochastic Neighbor Embedding)) processing.
[0068] Next, the data collection determination unit 110 determines whether the similarity between the features of the image and the features of the data collection criterion is greater than a predetermined threshold (S313). When the features of the data collection criterion are defined within a range in the multi-dimensional space, the determination of the similarity may be calculated by methods such as the Mahalanobis distance or cosine similarity, and it may be determined whether the calculated similarity is included within the range of the data collection criterion.
[0069] If the similarity between the features of the image and the features of the collection criterion data is greater than a predetermined threshold, the data collection determination unit 110 determines that the features of the image are similar to the features of the collection criterion data, and stores the image as collection data in the collection data storage unit 108 (S314).
[0070] As shown in FIG. 7, in the information processing system of the second embodiment, even if a plurality of information processing apparatuses 100 that perform the same operation are connected to the server 200, as shown in FIG. 12, information processing apparatuses 100 that perform different operations may be connected to the server 200. That is, the machine learning model implemented in the image processing unit 102 may be the same or different depending on the information processing apparatus 100.
[0071] FIG. 12 is a block diagram showing the configuration of the information processing system of the second embodiment when the machine learning model implemented in the image processing unit 102 is different for each information processing apparatus.
[0072] The information processing system shown in FIG. 12 includes a plurality of information processing apparatuses 100 and a server 200. The information processing apparatus 100 and the server 200 are communicably connected. In the information processing system shown in FIG. 12, the same reference numerals are given to the same functions and configurations as those in the information processing system shown in FIG. 7, and the descriptions thereof are omitted.
[0073] In the example shown in FIG. 12, the same machine learning model is implemented in the image processing unit 102 for the information processing apparatus A100 and the information processing apparatus B100, and different machine learning models are implemented in the image processing unit 102 for the information processing apparatus A100, the information processing apparatus C100, and the information processing apparatus D100. The information processing apparatus 100 may change the characteristics of the machine learning model according to its use and installation environment.
[0074] In addition, different dataset storage units 201 are provided corresponding to the machine learning models. The machine learning models 203 used by the inference target area visualization unit 202 are also different for each machine learning model. For example, the data collected from the information processing apparatuses A and B100 in which the machine learning model 1 (203) is implemented is stored in the dataset storage unit 1 (201). The inference target area visualization unit 202 analyzes the images included in the collected data stored in the dataset storage unit 201, and specifies the target area in the inference by the machine learning model 1 (203) from the images. The target area specified by the inference target area visualization unit 202 is the area in the image that contributed to the conclusion derived by the machine learning model. The robustness determination unit 204 determines, using a predetermined threshold, the ratio of the area of the target area specified by the inference target area visualization unit 202 to the area of the object in the image. If the area ratio of the target area is equal to or less than the predetermined threshold, the image is determined to be a low-robustness image, and control is performed so that the image with low robustness is transmitted to the data collection criterion generation unit 206.
[0075] The feature extraction unit 205 calculates the features of the low-robustness images specified by the inference target area visualization unit 202. The data collection criterion generation unit 206 clusters the image data arranged in the multi-dimensional space (illustrated in two dimensions in FIG. 11) with the features as axes, and generates a data collection criterion representing the range of the cluster (see FIG. 11). That is, a data collection criterion for the machine learning model 1 (203) is generated from the data stored in the dataset storage unit 1 and transmitted to the information processing apparatuses A and B100.
[0076] Similarly, data collected from the information processing device C100 in which the machine learning model 2 is implemented is stored in the dataset storage unit 2 (201). The data collection criterion generation unit 206 generates a data collection criterion for the machine learning model 2 (203) from the data stored in the dataset storage unit 2 (201) and transmits it to the information processing device C100. Further, data collected from the information processing device D100 in which the machine learning model 3 (203) is implemented is stored in the dataset storage unit 3 (201). The data collection criterion generation unit 206 generates a data collection criterion for the machine learning model 1 (203) from the data stored in the dataset storage unit 3 (201) and transmits it to the information processing device D100.
[0077] By configuring the information processing system shown in FIG. 12 in this way, the data collection criterion generation unit 206 can generate a data collection criterion for each machine learning model 203, and the data collection criterion distributed to the information processing device 100 can be efficiently updated according to the characteristics of the machine learning model 203, and the information processing device 100 can efficiently collect data.
[0078] As described above, in the second embodiment of the present invention, compared with the first embodiment described above, the amount of calculation required for inference target area visualization in the information processing device 100 can be reduced. Therefore, even if the calculation resources of the information processing device 100 are small, real-time processing can be realized. In addition, since the data collection criterion is distributed from the server 200 to the information processing device 100, data can be collected based on the latest data collection criterion.
[0079] <Example 3> Hereinafter, with reference to FIG. 13, an information processing system according to a third embodiment of the present invention will be described. In the information processing system of the third embodiment, low-robust data is used for re-learning to improve the robustness of the machine learning model. In the third embodiment, differences from the first and second embodiments will be mainly described, and the same components and processes as those in the first and second embodiments will be denoted by the same reference numerals, and their descriptions will be omitted.
[0080] FIG. 13 is a block diagram showing the configuration of the information processing system of the third embodiment.
[0081] The information processing system of Example 3 includes one or more information processing apparatuses 100 and a server 200. The information processing apparatus 100 and the server 200 are communicably connected. In FIG. 13, only the main configuration of the information processing apparatus 100 is illustrated, and the detailed configuration may be the same as that in FIG. 1. The information processing apparatus 100 may have a function of transmitting data collected by the information processing apparatus 100 to the server 200.
[0082] The server 200 includes a dataset storage unit 201, an inference target area visualization unit 202, a machine learning model 203, a robustness determination unit 204, a communication interface 207, a relearning data storage unit 208, a preprocessing unit 209, and a relearning unit 210.
[0083] The processes executed by the dataset storage unit 201, the inference target area visualization unit 202, the machine learning model 203, and the robustness determination unit 204 are the same as those in the aforementioned Example 2. Note that the dataset includes the collected data. That is, the dataset storage unit 201 stores the collected data transmitted from the information processing apparatus 100. The inference target area visualization unit 202 analyzes the images included in the collected data stored in the dataset storage unit 201 using the machine learning model 203 to identify the target areas. The target areas identified by the inference target area visualization unit 202 are the areas in the images that contributed to the conclusions derived by the machine learning model. The machine learning model 203 has a machine learning model composed of a DNN (Deep Neural Network) that has learned images labeled with the types of objects shown therein, and uses this machine learning model to infer the positions and types of objects shown in the input images and outputs the inference results. Also, the inference target area visualization unit 202 analyzes the images included in the collected data stored in the dataset storage unit 201 and identifies the target areas in the images for the inferences by the machine learning model 203. The target areas identified by the inference target area visualization unit 202 are the areas in the images that contributed to the conclusions derived by the machine learning model. The robustness determination unit 204 determines, using a predetermined threshold, the ratio of the area of the target area identified by the inference target area visualization unit 202 to the area of the object in the image. If the area ratio of the target area is equal to or less than the predetermined threshold, the image is determined to be a low-robustness image, and the image with low robustness is stored in the re-learning data storage unit 208.
[0084] The relearning data storage unit 208 stores the collected data determined to have low robustness as relearning data. The preprocessing unit 209 converts the data for relearning stored in the relearning data storage unit 208 into data suitable for relearning by making it difficult to use the target area of the relearning data for recognition. For example, processing such as filling in the target area (filling with other data (e.g., zero)) or blurring the target area is performed. As a result, it becomes possible to relearn to recognize an object by actively using information other than the target area. The relearning unit 210 uses the data preprocessed by the preprocessing unit 209 to relearn the machine learning model 203 and improve the robustness of the machine learning model. Then, the relearning unit 210 transmits the machine learning model with improved robustness through relearning to the information processing apparatus 100 via the communication interface 207. At this time, in order to prevent overfitting with low-robustness data, it is common to relearn including data having a certain degree of robustness. That is, it is preferable to mix low-robustness data (processed by the preprocessing unit 209) with the data used when learning the current DNN for learning.
[0085] As described above, in the third embodiment of the present invention, with respect to the problem that the robustness of the machine learning model does not improve even if the machine learning model is relearned using low-robustness data as it is, by performing preprocessing to make it difficult to recognize the target area of the low-robustness data and using it for relearning, the robustness of the machine learning model can be efficiently improved.
[0086] <Embodiment 4> Hereinafter, with reference to FIG. 14, the information processing system according to the fourth embodiment of the present invention will be described. In the information processing system according to the fourth embodiment, it has both functions of the above-described second and third embodiments, that is, generating a data collection criterion from low-robustness data and using the low-robustness data for relearning to improve the robustness of the machine learning model. In the third embodiment, the differences from the first, second, and third embodiments will be mainly described, and the same components and processes as those in the first, second, and third embodiments are denoted by the same reference numerals, and their descriptions are omitted.
[0087] FIG. 14 is a block diagram showing the configuration of the information processing system according to the fourth embodiment.
[0088] The information processing system according to the third embodiment includes one or more information processing apparatuses 100 and a server 200. The information processing apparatus 100 and the server 200 are communicably connected. In FIG. 14, only the main configuration of the information processing apparatus 100 is illustrated, and the detailed configuration may be the same as that in FIG. 1. The information processing apparatus 100 may have a function of transmitting the data collected by the information processing apparatus 100 to the server 200.
[0089] The server 200 includes a dataset storage unit 201, an inference target area visualization unit 202, a machine learning model 203, a robustness determination unit 204, a feature amount extraction unit 205, a data collection criterion generation unit 206, a communication interface 207, a relearning data storage unit 208, a preprocessing unit 209, and a relearning unit 210.
[0090] The processes executed by the dataset storage unit 201, the inference target area visualization unit 202, the machine learning model 203, the robustness determination unit 204, the feature amount extraction unit 205, the data collection criterion generation unit 206, and the communication interface 207 are the same as those in the second embodiment described above. The processes executed by the dataset storage unit 201, the inference target area visualization unit 202, the machine learning model 203, the robustness determination unit 204, the communication interface 207, the relearning data storage unit 208, the preprocessing unit 209, and the relearning unit 210 are the same as those in the third embodiment described above.
[0091] A series of data and model update processes by adopting the configuration of this embodiment will be described. An operation version DNN is implemented in the image processing unit 102 of the information processing apparatus 100, and performs image processing for a predetermined application in real time. At the same time, in order to determine whether the image acquired by the camera is low-robust for the operation version DNN, feature amounts are extracted from the intermediate data of the DNN, and the data collection determination unit 110 calculates the similarity with the feature amounts stored in the data collection reference storage unit 111 to determine whether data collection is necessary. Low-robust data that needs to be collected is temporarily stored in the collected data storage unit and transmitted to the server 200 at a predetermined timing. In this way, the data acquired by each information processing apparatus 100 at each operation site is stored in the dataset storage unit 201. In order to quantitatively determine the robustness of the data determined by the information processing apparatus 100, the inference target area visualization unit 202 and the robustness determination unit 204 quantify the robustness for each collected data, and if the data seems to be usable for relearning or updating the data collection reference, the image data is input to the subsequent relearning data storage unit 208 and the feature amount extraction unit 205. When a certain amount of data with features different from the currently held feature map (see FIG. 11) is accumulated in the feature amounts extracted by the feature amount extraction unit 205, the reference data in the data collection reference storage unit 111 in each information processing apparatus 100 is updated. On the other hand, relearning may be started when the number of images stored in the relearning data storage unit 208 exceeds a predetermined number, or relearning may be performed when new feature amount data is accumulated in the above-described feature map. Further, in order not to be biased toward data of the same feature, data may be selected so as to evenly cover the feature map and the distance between each feature point is constant. By executing the relearning process shown in Embodiment 3 using the data selected in this way, a DNN with enhanced robustness is generated, and the generated DNN (machine learning model) is distributed to each information processing apparatus 100.
[0092] As described above, in the fourth embodiment of the present invention, since the amount of computation of the information processing apparatus 100 can be reduced, real-time processing can be realized even if the computation resources of the information processing apparatus 100 are scarce. Further, since the data collection criteria are distributed from the server 200 to the information processing apparatus 100, data can be collected according to the latest data collection criteria. Furthermore, in the server 200 of the fourth embodiment, by preprocessing low-robustness data in the preprocessing unit 209 and using it for re-learning, the robustness of the machine learning model can be improved. That is, in the DNN update process, it is expected that data collection and selection based on the rules of thumb of data scientists and developers are no longer necessary, and the man-hours for those tasks can be made more efficient.
[0093] Note that the present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the scope of the appended claims. For example, the above-described embodiments have been described in detail for easy understanding of the present invention, and the present invention is not necessarily limited to those having all the configurations described. Also, a part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Further, the configuration of another embodiment may be added to the configuration of one embodiment. Also, for a part of the configuration of each embodiment, addition, deletion, or replacement with other configurations may be made.
[0094] Also, each of the above-described configurations, functions, processing units, processing means, etc. may be realized in hardware by designing a part or all of them, for example, by an integrated circuit, or may be realized in software by a processor interpreting and executing a program for realizing each function.
[0095] Information such as programs, tables, and files for realizing each function can be stored in a storage device such as a memory, a hard disk, an SSD (Solid State Drive), or a recording medium such as an IC card, an SD card, or a DVD.
[0096] Also, the control lines and information lines show those considered necessary for explanation, and do not necessarily show all the control lines and information lines required for implementation. In reality, it may be considered that almost all components are interconnected.
Explanation of Signs
[0097] 11, 21 Processors 12, 22 Memories 13, 23 Auxiliary Storage Devices 14, 24 Communication Interfaces 25 Input Interface 26 Output Interface 27 Keyboard 28 Mouse 29 Display Device 100 Information Processing Device 101 Camera 102 Image Processing Unit 103 Other Sensors 104 Preprocessing Unit 105 Application Processing Unit 106 Inference Visualization Analysis Unit 107 Robustness Judgment Unit 108 Collected Data Storage Unit 109 Communication Interface 110 Data Collection Judgment Unit 111 Data Collection Criterion Storage Unit 112 Feature Extraction Unit 200 Server 201 Dataset Storage Unit 202 Inference Attention Region Visualization Unit 203 Machine Learning Model 204 Robustness Judgment Unit 205 Feature Extraction Unit 206 Data Collection Criterion Generation Unit 207 Communication Interface 208 Relearning Data Storage Unit 209 Preprocessing Unit 210 Relearning Unit
Claims
1. An information processing apparatus, comprised of a computer having an arithmetic unit that executes arithmetic processing and a storage unit accessible to the arithmetic unit, an image processing unit that executes an image processing program including a machine learning model and outputs a recognition result of image data, an inference visualization analysis unit that specifies a region of interest in the inference in the image processing unit, a robustness determination unit that determines the robustness of the image data using the specified region of interest, a collected data storage unit that stores image data determined to have a robustness lower than a first predetermined criterion in the storage unit, and an interface unit that transmits the image data stored in the collected data storage unit to the outside, wherein the information processing apparatus is characterized by comprising these components.
2. The information processing apparatus according to claim 1, wherein the first predetermined criterion is determined by a ratio of the area of the region of interest to the area of the image.
3. A server that collects image data from an information processing apparatus, comprised of a computer having an arithmetic unit that executes arithmetic processing and a storage unit accessible to the arithmetic unit, a dataset storage unit that determines that it conforms to a data collection criterion and stores the image data collected by the information processing apparatus in the storage unit, an inference region of interest visualization unit that specifies a region of interest in the inference in the image processing for the stored image data, a robustness determination unit that determines the robustness of the image data using the specified region of interest, a feature amount extraction unit that extracts image feature amounts of image data determined to have a robustness lower than a first predetermined criterion, a data collection criterion generation unit that generates a data collection criterion using the extracted image feature amounts, and an interface unit that transmits the generated data collection criterion to the information processing apparatus, wherein the server is characterized by comprising these components.
4. The server communicates with a plurality of information processing apparatuses that execute an image processing program including different machine learning models to recognize image data, the dataset storage unit is configured such that regions for storing collected data are divided for each of the plurality of information processing apparatuses, the data collection criterion generation unit generates a data collection criterion for each of the plurality of information processing apparatuses, and the interface unit transmits the data collection criterion for each of the plurality of information processing apparatuses. The server according to claim 3 is characterized by these features.
5. The server according to claim 3, wherein the first predetermined criterion is determined by a ratio of the area of the attention area to the area of the image.
6. The feature amount extraction unit extracts low-dimensional information from the image feature amounts of the image data determined to have lower robustness than the first predetermined criterion, and The server according to claim 3, wherein the data collection criterion generation unit generates the data collection criterion using a result of clustering a plurality of the image feature amounts to group similar image feature amounts.
7. A relearning data storage unit that stores relearning data for a machine learning model based on the determination result of the robustness, A preprocessing unit that preprocesses the relearning data using the information of the specified attention area, A relearning unit that relearns the machine learning model using the image data processed by the preprocessing unit, and The server according to claim 3, wherein the interface unit transmits the relearned machine learning model to the information processing apparatus.
8. The server according to claim 7, wherein the preprocessing unit processes data for areas of the specified attention area that are equal to or greater than a second predetermined criterion.
9. The server according to claim 8, wherein the second predetermined criterion is determined by a ratio of the area of the attention area to the area of the image.
10. A relearning data storage unit that stores relearning data for a machine learning model based on the determination result of the robustness, A preprocessing unit that preprocesses the relearning data using the information of the specified attention area, A relearning unit that relearns the machine learning model using the image data processed by the preprocessing unit, and The server according to claim 3, wherein the interface unit transmits the relearned machine learning model and the generated data collection criterion to the information processing apparatus.
11. An information processing system including an information processing apparatus and a server, The information processing apparatus, A camera that acquires external information, An image processing unit that performs image processing on the image data acquired by the camera and outputs a result of image recognition, A data collection unit that collects image data conforming to a data collection criterion, A feature amount extraction unit that extracts a first image feature amount of the image data used for the image recognition, A data collection criterion storage unit that stores the data collection criterion, A data collection determination unit that determines whether the extracted first image feature amount conforms to the data collection criterion, A collection data storage unit that stores image data determined to conform to the data collection criteria; A first interface unit for data communication with the outside; and has, The server is, A data set storage unit that stores image data collected by the information processing device, determined to conform to the data collection criteria; An inference attention area analysis unit that specifies an attention area of inference in the image processing for the stored image data; A robustness determination unit that determines the robustness of the image data using the specified attention area; A feature amount extraction unit that extracts a second image feature amount of image data determined to have a robustness lower than a first predetermined standard; A data collection criteria generation unit that generates the data collection criteria using the extracted second image feature amount; A re-learning data storage unit that stores re-learning data of the machine learning model based on the determination result of the robustness; A pre-processing unit that pre-processes the re-learning data using information on the specified attention area; A re-learning unit that re-learns the machine learning model using the image data processed by the pre-processing unit; An information processing system, characterized by comprising: a second interface unit that transmits the generated data collection criteria and the re-learned machine learning model to the information processing device.
Citation Information
Patent Citations
Data collection system
JP2022002058A
Information processing apparatus, information processing method, and program
JP2023084981A