Information processing device, server, and information processing system including these
The information processing apparatus addresses the challenge of collecting low-robust data for improving DNN performance by analyzing attention areas and determining robustness, enabling efficient data collection and enhancing the reliability of DNN inferences.
Patent Information
- Application Number
- PCT/JP2024/027782
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-22
- Filing Date
- 2024-08-02
- Publication Date
- 2025-06-26
AI Technical Summary
Existing technologies face challenges in efficiently collecting and selecting low-robust data for improving the inference performance of operational machine learning models, particularly deep neural networks (DNNs), due to the need for visual inspection and manual judgment.
An information processing apparatus is configured with an image processing unit, an inference visualization analysis unit, a robustness determination unit, and a data storage unit to efficiently identify and collect low-robust data by analyzing the attention areas in DNN inferences and determining data robustness based on predefined standards.
The proposed solution enables the efficient selection and collection of low-robust data, which can be used for re-learning to enhance the robustness and reliability of DNN inferences, thereby improving the performance and stability of image recognition tasks.
Smart Images

Figure JP2024027782_26062025_PF_FP_ABST
Abstract
Description
Information processing device, server, and information processing system including them Incorporation by Reference
[0001] This application claims priority from Japanese Patent Application No. 2023-216879, filed on December 22, 2023, the contents of which are incorporated herein by reference.
[0002] The present invention relates to an information processing device and a server, and more particularly to a technique for improving the effectiveness and efficiency of data collection and data selection for improving the performance of an image recognition program including a machine learning model.
[0003] In recent years, image recognition programs using machine learning such as deep neural networks (DNN) have been put into practical use, and image recognition technology is being applied to a variety of equipment, services, and solutions, such as anomaly inspection systems for manufacturing lines, systems that monitor human intrusion and behavior, and automatic transport equipment at logistics sites.
[0004] In order to use DNNs in various applications, it is important to efficiently collect and select training data, but this requires the collection and selection of training data through visual inspection and qualitative judgment by workers, which increases human costs. Furthermore, if the created DNN is operated over the long term, it becomes necessary to deal with instabilities in inference that were not noticed during the design stage and performance degradation due to changes in the surrounding environment, and the cost of collecting and selecting data becomes an issue.
[0005] In order to improve the performance of a DNN in operation (hereafter referred to as the operational DNN), it is necessary to select and collect data that is effective in improving performance. For example, in an image recognition task, by collecting and re-learning data that has been misrecognized or data that shows a new category (type) that could not be recognized before, it is possible to achieve higher performance and functionality.
[0006] On the other hand, even among data with correct recognition results, there is data with low recognition stability and robustness, such as data with low confidence but correct by chance, data that cannot be recognized when part of the object to be recognized is hidden by another object, and data that cannot be recognized due to the influence of light intensity or noise (hereinafter, these will be referred to as low-robust data). In other words, in order to realize highly reliable and stable applications, it is necessary to re-learn not only data that has been erroneously recognized by the operational DNN or unknown data, but also data that is low-robust for the operational DNN.
[0007] Various existing technologies have been considered as methods for efficiently preparing data that is effective for training an operational DNN. Patent Document 1 (JP 2023-84981 A) describes an information processing device that inputs a plurality of input images to a trained model, causes the trained model to output a predicted label inferred from each of the plurality of input images, and acquires a feature region that serves as the basis for the prediction of the trained model in each of the plurality of input images. The information processing device is characterized by comprising: a processing reception means that receives a registration instruction to register the input image as an image to be processed; a registration means that, when the processing reception means receives the registration instruction, registers the input image as an image to be processed; an update reception means that receives an update instruction to update training data for re-training the trained model; and an update means that, when the update reception means receives the update instruction, performs a processing process on at least one of the images to be processed and updates the training data using the processed image.
[0008] Furthermore, Patent Document 2 (JP 2022-2058 A) describes a data collection system that collects image data used by a learning device that recognizes images through machine learning to perform learning, and that includes a requirements definition processing unit that receives requirement definition data that specifies requirement variables representing the requirements of the image data necessary for the learning device to fully learn and requirement values of the requirement variables, and the requirements definition processing unit further receives priority data that specifies the priority of the requirement variables, and the requirements definition processing unit further presents the requirement variables and the requirement values in order of priority, and also provides a requirement value response rate that indicates the proportion of the requirement variables for which the requirement values are specified.
[0009] The invention described in Patent Document 1 efficiently generates training data for improving the accuracy of image classification AI, but the standards for processing and generating the training data require the developer to visually inspect and determine based on experience, making it difficult to prepare effective training data for improving the performance of the operational DNN. Furthermore, the invention described in Patent Document 2 appropriately sets the range of image data to be used as training data, but the developer defines the requirements for the collected data, and a separate technology is required to define low-robust data among the data necessary to improve the performance of the operational DNN.
[0010] The present invention aims to efficiently collect low-robust data that can improve the inference performance of machine learning models in operation.
[0011] The information processing device is configured by a computer having an arithmetic unit that executes arithmetic processing and a memory unit that is accessible by the arithmetic unit, and is characterized by comprising: an image processing unit that executes an image processing program including a machine learning model and outputs recognition results for image data; an inference visualization analysis unit that identifies an area of interest for inference in the image processing unit; a robustness determination unit that determines the robustness of the image data using the identified area of interest; a collected data storage unit that stores image data whose robustness is determined to be lower than a first predetermined standard in the memory unit; and an interface unit that transmits the image data stored in the collected data storage unit to the outside.
[0012] According to one aspect of the present invention, data that is less robust for an operational DNN can be efficiently selected and collected. Problems, configurations, and effects other than those described above will become apparent from the following description of the preferred embodiments of the present invention.
[0013] 1 is a block diagram showing the configuration of an information processing device according to a first embodiment. FIG. 2 is a block diagram showing the physical configuration of a computer constituting the information processing device according to the first embodiment. FIG. 3 is a flowchart of low robust data determination processing according to the first embodiment. FIG. 4 is a diagram showing an original image used by an inference visualization analysis unit according to the first embodiment. FIG. 5 is a diagram showing an example in which the region of interest according to the first embodiment extends over the entire object. FIG. 6 is a diagram showing an example in which the region of interest according to the first embodiment is only a part of the object. FIG. 7 is a block diagram showing the configuration of an information processing system according to a second embodiment. FIG. 8 is a block diagram showing the physical configuration of a computer constituting a server according to the second embodiment. FIG. 9 is a block diagram showing the configuration of an information processing device according to the second embodiment. FIG. 10 is a flowchart of low robust data determination processing according to the second embodiment. FIG. 11 is a block diagram showing another configuration of an information processing system according to the second embodiment. FIG. 12 is a block diagram showing the configuration of an information processing system according to a third embodiment. FIG. 13 is a block diagram showing the configuration of an information processing system according to a fourth embodiment.
[0014] Hereinafter, embodiments of the present invention will be described with reference to the drawings. Each embodiment is an example for explaining the present invention, and for clarity of explanation, some details have been omitted or simplified as appropriate. The present invention can be implemented in various other forms. Unless otherwise specified, each component may be singular or plural.
[0015] The position, size, shape, range, etc. of each component shown in the drawings may not represent the actual position, size, shape, range, etc. in order to facilitate understanding of the invention. Therefore, the present invention is not necessarily limited to the position, size, shape, range, etc. disclosed in the drawings. When there are multiple components having the same or similar functions, they may be described using the same reference numeral with different subscripts. Furthermore, when it is not necessary to distinguish between these multiple components, the subscripts may be omitted in the description.
[0016] In each embodiment, processing performed by executing a program may be described. Here, a computer executes the program using a processor (e.g., a CPU or a GPU) and performs processing defined by the program using storage resources (e.g., memory) and interface devices (e.g., communication ports). Therefore, the entity performing the processing by executing the program may be the processor. Similarly, the entity performing the processing by executing the program may be a controller, device, system, computer, or node having a processor. The entity performing the processing by executing the program may be a calculation unit or a dedicated circuit that performs specific processing. Here, the dedicated circuit may be, for example, an FPGA (Field Programmable Gate Array), an ASIC (Application Specific Integrated Circuit), or a CPLD (Complex Programmable Logic Device).
[0017] A program may be installed on a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server may include a processor and storage resources for storing the program to be distributed, and the processor of the program distribution server may distribute the program to be distributed to other computers. In addition, in the embodiments, two or more programs may be realized as one program, or one program may be realized as two or more programs.
[0018] In recent years, peripheral recognition applications using deep neural networks (DNNs), a type of machine learning technology, have become widespread. In machine learning-based programs, if there are scenes where recognition errors occurred or scenes that need to be trained as new tasks, retraining using those data can improve performance and achieve improved reliability and functionality. However, compared to rule-based algorithms, machine learning programs often lack clarity regarding the reasoning behind their inferences. For example, even if the recognition result is correct, misrecognition can occur if the image contains slight occlusions, cutoffs, or photographic noise (e.g., fluctuations in brightness, saturation, noise, etc.). Data prone to such misrecognition is called "low-robust data." Because the recognition result itself is "correct," low-robust data is difficult to identify using standard evaluation methods that simply compare it with correct data. On the other hand, using low-robust data for retraining can improve the reliability and stability of inferences in operational DNNs.
[0019] As a technique for finding low-robust data, techniques for visualizing the area of interest in DNN inference (e.g., Grad-CAM and RISE) have been reported. For example, even if the entire object to be detected is clearly captured, the DNN inference result accurately infers the object's location, and the inference reliability score is very high, non-detection may occur frequently when the object is continuously inferred in chronological order. Visualizing the area of interest for the object may result in only a portion of the object being focused on. In this case, the area of interest may be obscured by other objects, or may become unclear due to the amount of light or noise at the time of capture, leading to non-detection of the inference result for the object. Visualizing the area of interest in inference in this way is effective in identifying the cause of DNN malfunctions.
[0020] However, the collection and selection of low-robust data based on the results of visualizing the area of interest requires sensory evaluation and visual judgment by designers and developers, and the labor required for this becomes a major challenge when dealing with large-scale datasets for retraining machine learning.
[0021] The present invention aims to improve the efficiency of collecting and selecting low-robust data, which is one type of data required for updating and evolving operational DNNs. In the following examples, an example is described in which the present invention is applied to an information processing system including an edge device (information processing device) using object detection AI and a server (such as a cloud device or development environment) connected to the edge device. Unlike development environments such as servers, edge devices are subject to constraints on computer computing resources and require real-time processing. Therefore, information processing systems targeted by the present invention include, for example, systems that process images acquired by cameras and notify the results, or devices that provide feedback to on-site equipment control, such as systems for inspecting anomalies on production lines, systems that monitor human intrusions and behavior, automatic transport devices at logistics sites, and construction machinery and autonomous vehicles at construction sites.
[0022] 1 to 6, an information processing device 100 according to a first embodiment of the present invention will be described. As described above, the information processing device 100 is a device that has limited computing resources, such as an edge device, and requires real-time processing.
[0023] FIG. 1 is a block diagram showing the configuration of an information processing apparatus 100 according to the first embodiment.
[0024] The information processing device 100 of Example 1 has a camera 101, an image processing unit 102, other sensors 103 other than a camera, a preprocessing unit 104, an application processing unit 105, an inference visualization analysis unit 106, a robustness determination unit 107, a collected data storage unit 108, and a communication interface 109.
[0025] The camera 101 captures an image to be processed by the application processing unit 105. The image processing unit 102 uses a machine learning model to output an inference result from the image captured by the camera 101. The image processing unit 102 is configured with an image processing program including a machine learning model (hereinafter simply referred to as DNN) such as a convolutional neural network (CNN) or a transformer for image recognition, and outputs the type and position of an object appearing in the input image.
[0026] The other sensor 103 is a sensor other than a camera, and can be, for example, a vibration sensor, a microphone, LiDAR, radar, an ultrasonic sensor, etc. The preprocessing unit 104 converts the positions and vibrations of surrounding objects, sound spectra, and other physical quantities into data in a format usable by the application processing unit 105 and outputs the converted data.
[0027] The application processing unit 105 realizes functions provided by the information processing device 100 (for example, object recognition, monitoring of restricted areas, and monitoring of human behavior) using the inference results from the image processing unit 102. The application processing unit 105 may realize functions provided by the information processing device 100 using data output from the preprocessing unit 104 in addition to the determination results from the image processing unit 102. Various usages are conceivable, such as a configuration in which the output of the application is transmitted to the server side of the information processing system via a communication interface (I / F) 109 as shown in the figure, or a configuration in which the output is input to the control function of the information processing device itself.
[0028] The inference visualization analysis unit 106 analyzes the inference results from the image processing unit 102, identifies areas of interest in the image that contributed to the derivation of a conclusion in the inference by the DNN, and creates a heat map (see FIGS. 5 and 6 ) representing areas with high attention. The areas of interest identified by the inference visualization analysis unit 106 are areas in the image that contributed to the conclusion derived by the DNN. Instead of or in addition to the process of identifying areas of interest, the inference visualization analysis unit 106 may perform a process of calculating the confidence and uncertainty of the inference. As an example of its operation, the robustness determination unit 107 calculates the ratio of the area of the areas of interest in the image to the area of the object in the image to evaluate the robustness of the image processing. If the area ratio of the areas of interest is equal to or less than a predetermined standard (e.g., 50%), the robustness determination unit 107 determines the image to be a low-robust image and stores the image determined to be a low-robust image in the collected data storage unit 108 together with information about the heat map. The collected data storage unit 108 stores the images determined to be low-robust according to the determination result of the robustness determination unit 107.
[0029] The communication interface 109 controls communication with other devices. For example, the communication interface 109 transmits the processing result of the application processing unit 105 to an external server.
[0030] FIG. 2 is a block diagram showing the physical configuration of the computers that constitute the information processing apparatus 100 according to the first embodiment.
[0031] The information processing device 100 is configured as a computer having a processor (CPU) 11, a memory 12, an auxiliary storage device 13, and a communication interface 14. The information processing device 100 may have an input interface and an output interface.
[0032] The processor 11 is a computing device that executes programs stored in the memory 12. The processor 11 executes various programs to realize each functional unit of the information processing device 100 (e.g., the image processing unit 102, the preprocessing unit 104, the application processing unit 105, the inference visualization analysis unit 106, the robustness determination unit 107, etc.). Note that part of the processing performed by the processor 11 by executing the programs may be executed by another computing device (e.g., hardware such as a GPU, an ASIC, or an FPGA).
[0033] The memory 12 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores unchanging programs (e.g., BIOS), etc. The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 11 and data used when the programs are executed.
[0034] The auxiliary storage device 13 is, for example, a large-capacity, non-volatile storage device such as a magnetic storage device (HDD) or a flash memory (SSD). The auxiliary storage device 13 also stores data accessed by the processor 11 when the processor 11 executes a program and the program executed by the processor 11, and constitutes, for example, a collected data storage unit 108. That is, the program is read from the auxiliary storage device 13, loaded into the memory 12, and executed by the processor 11 to realize each function of the information processing device 100.
[0035] The communication interface 14 is a network interface device that controls communication with other devices in accordance with a predetermined protocol.
[0036] The programs executed by the processor 11 are provided to the information processing device 100 via removable media (external HDD, optical disk, flash memory, etc.) or a network, and are stored in a non-volatile auxiliary storage device 13, which is a non-transitory storage medium. For this reason, the information processing device 100 preferably has an interface for reading data from removable media.
[0037] FIG. 3 is a flowchart of the low robustness data determination process executed by the information processing apparatus 100 according to the first embodiment.
[0038] First, the inference visualization analysis unit 106 analyzes the determination result by the image processing unit 102, identifies areas of interest in the DNN inference from within the image, and creates a heat map showing areas with high attention (S301).
[0039] The robustness determination unit 107 then calculates the robustness of the identified region of interest. Robustness can be calculated, for example, by calculating the ratio of the area of the identified region of interest to a reference, such as the area of the object on the image. Alternatively, an index using the uniformity of the distribution of the regions of interest or the total value of their intensities may be used (S302). The reference area of the object may be the area of a rectangular bounding box (bbox) in the case of an object detection task, or the area of the region in which the object is drawn (corresponding to a so-called segmentation mask).
[0040] Thereafter, if the area ratio of the identified region of interest is equal to or less than a predetermined threshold (for example, 10%), the robustness determination unit 107 determines that the image is low robust (S303).
[0041] Then, the robustness determining unit 107 stores the images determined to be low robust in the collected data storage unit 108 (S304).
[0042] The heat map created by the robustness determination unit 107 is used for inference in the original image, and the stronger the contribution to deriving a conclusion, the stronger the intensity of white displayed in the map, as shown in Figures 5 and 6. The specifications of the heat map are not limited to this, and it may be displayed on a color scale, or only areas with a degree of attention exceeding a predetermined threshold may be displayed in color.
[0043] For example, in inferring a pedestrian surrounded by a square in the original image 500 shown in Fig. 4, if the region of interest extends over the entire object as shown in Fig. 5 and the ratio of the area of the region of interest to the area of the bounding box 501 is greater than a predetermined threshold, the inference visualization analysis unit 106 determines that the image is not low robust, because a correct inference result can be obtained even if part of the region of interest is missing due to occlusion, such as being hidden by another object. On the other hand, if the region of interest in the original image 500 shown in Fig. 4 is only a part of the object as shown in Fig. 6 and the ratio of the area of the region of interest to the area of the bounding box 501 is less than a predetermined threshold, the inference visualization analysis unit 106 determines that the image is low robust, because if part of the region of interest is missing due to occlusion, clipping, or the like, a correct inference result may not be obtained, such as missed detection or incorrect type.
[0044] Furthermore, in step S302, the inference visualization analysis unit 106 may simultaneously execute a process of calculating the certainty and uncertainty of the inference. When the inference visualization analysis unit 106 calculates the certainty and uncertainty of the inference, the robustness determination unit 107 calculates a determination score (for example, the product of these values is used as the robustness score) that includes the certainty and uncertainty in addition to the proportion of the attention area, determines data with a low score as low robust data (S303), and stores images determined to be low robust in the collected data storage unit 108 (S304).
[0045] As described above, in the first embodiment of the present invention, it is possible to efficiently select and collect data with low robustness from among images correctly recognized by the image processing unit 102, and to apply incorrectly recognized images to re-learning, thereby improving the robustness of object recognition by the image processing unit 102.
[0046] 7 to 12 , an information processing system according to a second embodiment of the present invention will be described. In the information processing system according to the second embodiment, a server 200 extracts an inference attention region from data collected by an information processing device 100, identifies low-robust data, extracts features from the identified low-robust data, and generates a data collection standard. The server 200 distributes the generated data collection standard to the information processing device 100. This allows each information processing device 100 to always collect data according to the latest data collection standard, and also reduces computational resources by eliminating the need for the information processing device 100 to perform inference attention region visualization processing. In the second embodiment, differences from the first embodiment will be mainly described. The same components and processes as those in the first embodiment are designated by the same reference numerals, and their description will be omitted.
[0047] FIG. 7 is a block diagram showing the configuration of an information processing system according to the second embodiment.
[0048] The information processing system of the second embodiment includes one or more information processing devices 100 and a server 200. The information processing devices 100 and the server 200 are communicatively connected. In Fig. 7, only the main components of the information processing device 100 are shown, and the detailed configuration will be described later with reference to Fig. 9.
[0049] The server 200 includes a dataset storage unit 201 , an inference attention area visualization unit 202 , a machine learning model 203 , a robustness determination unit 204 , a feature extraction unit 205 , a data collection standard generation unit 206 , and a communication interface 207 .
[0050] The dataset storage unit 201 stores collected data transmitted from the information processing device 100. The inference attention region visualization unit 202 uses a machine learning model 203 to output an inference result from data of the collected data stored in the dataset storage unit 201 for which the inference of the machine learning model 203 has been correct. The machine learning model 203 is configured with a deep neural network (DNN) that has learned images labeled with the type of object depicted, infers objects depicted in input images, and outputs a determination result of the object. Furthermore, the inference attention region visualization unit 202 analyzes images included in the collected data stored in the dataset storage unit 201 and identifies an attention region in the image for inference by the machine learning model 203. The attention region identified by the inference attention region visualization unit 202 is an area in the image that contributed to the conclusion derived by the machine learning model. The robustness determination unit 204 calculates the ratio of the area of the attention area identified by the inference attention area visualization unit 202 to the area of the object on the image, and if the area ratio of the attention area is below a predetermined threshold, determines that the image is a low-robust image, and controls so that images in which the inference of the machine learning model 203 is correct and which have low robustness are sent to the data collection standard generation unit 206.
[0051] The feature extraction unit 205 calculates the feature amounts of the low-robust image determined by the robustness determination unit 204. The data collection standard generation unit 206 clusters the image data arranged in a multidimensional space (shown in two dimensions in FIG. 11 ) based on the feature amounts, and generates data collection standards that represent the range of the clusters (see FIG. 11 ). The communication interface 207 controls communication with other devices. For example, the communication interface 207 transmits the data collection standards generated by the data collection standard generation unit 206 to the information processing device 100 and receives data collected by the information processing device 100.
[0052] FIG. 8 is a block diagram showing the physical configuration of the computer that constitutes the server 200 of the second embodiment.
[0053] The server 200 is configured as a computer having a processor (CPU) 21, a memory 22, an auxiliary storage device 23, and a communication interface 24. The server 200 may also have an input interface 25 and an output interface 26.
[0054] The processor 21 is a computing device that executes programs stored in the memory 22. The processor 21 executes various programs to realize each functional unit of the server 200 (e.g., an inference attention area visualization unit 202, a robustness determination unit 204, a feature extraction unit 205, a data collection standard generation unit 206, etc.). Note that part of the processing performed by the processor 21 by executing the programs may be executed by another computing device (e.g., hardware such as a GPU, an ASIC, or an FPGA).
[0055] The memory 22 includes a ROM, which is a non-volatile storage element, and a RAM, which is a volatile storage element. The ROM stores unchanging programs (e.g., BIOS), etc. The RAM is a high-speed, volatile storage element such as a DRAM (Dynamic Random Access Memory), and temporarily stores programs executed by the processor 21 and data used when the programs are executed.
[0056] The auxiliary storage device 23 is, for example, a large-capacity, non-volatile storage device such as a magnetic storage device (HDD) or a flash memory (SSD). The auxiliary storage device 23 also stores data accessed by the processor 21 when the processor 21 executes a program and the program executed by the processor 21, and constitutes, for example, a dataset storage unit 201. That is, the program is read from the auxiliary storage device 23, loaded into the memory 22, and executed by the processor 21 to realize each function of the server 200.
[0057] The communication interface 24 is a network interface device that controls communication with other devices in accordance with a predetermined protocol.
[0058] The input interface 25 is an interface to which input devices such as a keyboard 27 and a mouse 28 are connected and which receives input from an operator. The output interface 26 is an interface to which output devices such as a display device 29 and a printer (not shown) are connected and which outputs the results of program execution in a format that can be viewed by the user. Note that a terminal (not shown) connected to the server 200 via a network may provide the input device and the output device. In this case, the server 200 may have a web server function, and the terminal may access the server 200 using a predetermined protocol (for example, http).
[0059] The programs executed by the processor 21 are provided to the server 200 via removable media (external HDD, optical disk, flash memory, etc.) or a network, and are stored in a non-volatile auxiliary storage device 13, which is a non-transitory storage medium. For this reason, the server 200 should preferably have an interface for reading data from removable media.
[0060] The server 200 is a computer system configured on a single physical computer or on multiple logically or physically configured computers, and may operate on a virtual computer constructed on multiple physical computer resources. For example, each functional unit may operate on a separate physical or logical computer, or multiple functional units may be combined to operate on a single physical or logical computer.
[0061] FIG. 9 is a block diagram showing the configuration of an information processing apparatus 100 according to the second embodiment.
[0062] The information processing device 100 of Example 2 has a camera 101, an image processing unit 102, other sensors 103 other than a camera, a preprocessing unit 104, an application processing unit 105, a feature extraction unit 112, a data collection determination unit 110, a collected data storage unit 108, a data collection standard storage unit 111, and a communication interface 109.
[0063] The camera 101, image processing unit 102, other sensor 103, preprocessing unit 104, application processing unit 105, and communication interface 109 are the same as those of the information processing device 100 in the first embodiment. The image processing unit 102 uses an image recognition program to output an inference result from an image captured by the camera 101. The image recognition program is configured with a deep neural network (DNN) that has learned images labeled with the type of object in the image, and is a program that infers the position and type of an object in an input image and outputs the inference result. The image processing unit 102 also outputs a region (e.g., a rectangular bounding box) in which the object is determined.
[0064] The feature extraction unit 112 extracts features using the intermediate layer data of the DNN in the image processing unit 102 and features of the regions used for inference. Feature extraction can extract multidimensional array data such as image data as low-dimensional features by using existing technologies such as principal component analysis and t-SNE. The data collection determination unit 110 compares the image features extracted by the feature extraction unit 112 with the image features that constitute the data collection criteria stored in the data collection criteria storage unit 111. If the similarity between the two features is equal to or greater than a predetermined threshold, the data collection determination unit 110 determines that the image data should be collected and stores the image data in the collected data storage unit 108. The collected data storage unit 108 stores image data whose image features match the predetermined data collection criteria. The data collection criteria storage unit 111 stores the data collection criteria generated by the data collection criteria generation unit 206 of the server 200. The collection criteria stored in the data collection criteria storage unit 111 are frequently updated by the server 200, enabling the collection of data necessary to improve the performance of the machine learning model 203 in operation.
[0065] FIG. 10 is a flowchart of the low robustness data determination process executed by the information processing apparatus 100 according to the second embodiment.
[0066] First, the image processing unit 102 performs inference on the input image using a deep neural network (DNN) that constitutes the implemented machine learning model. At this time, data from any intermediate layer in the DNN is extracted and sent to the feature extraction unit 112 (S311). It is desirable to select a highly sensitive layer that best expresses the image features as the layer to be extracted. Generally, even if the weight coefficients of the DNN are updated through re-learning, the highly sensitive layer often remains unchanged unless the algorithm is changed. Therefore, the selection of the extraction layer is specified by the developer when developing the algorithm.
[0067] Next, the feature extraction unit 112 calculates low-dimensional vector data (features) from the data extracted from the intermediate layer of the DNN by the image processing unit 102 (S312). The feature extraction unit 112 converts the multidimensional extracted data into low-dimensional features using a feature extraction algorithm such as GAP (Global average pooling), principal component analysis, or dimensionality reduction (t-SNE (t-Distributed Stochastic Neighbor Embedding)).
[0068] Next, the data collection determination unit 110 determines whether the similarity between the feature of the image and the feature of the data collection standard is greater than a predetermined threshold (S313). When the feature of the data collection standard is defined as a range in a multidimensional space, the similarity may be determined by calculating it using a method such as Mahalanobis distance or cosine similarity, and determining whether the calculated similarity is within the range of the data collection standard.
[0069] If the similarity between the image features and the features of the collected reference data is greater than a predetermined threshold, the data collection determination unit 110 determines that the image features are similar to the features of the collected reference data, and stores the image as collected data in the collected data storage unit 108 (S314).
[0070] In the information processing system of the second embodiment, a plurality of information processing devices 100 that perform the same operation may be connected to the server 200 as shown in Fig. 7, or information processing devices 100 that perform different operations may be connected to the server 200 as shown in Fig. 12. That is, the machine learning model implemented in the image processing unit 102 may be the same or different depending on the information processing device 100.
[0071] FIG. 12 is a block diagram showing the configuration of an information processing system according to a second embodiment in which the machine learning model implemented in the image processing unit 102 differs for each information processing apparatus.
[0072] The information processing system shown in Fig. 12 includes a plurality of information processing devices 100 and a server 200. The information processing devices 100 and the server 200 are communicatively connected. In the information processing system shown in Fig. 12, the same functions and configurations as those in the information processing system shown in Fig. 7 are denoted by the same reference numerals, and descriptions thereof will be omitted.
[0073] 12 , the information processing device A100 and the information processing device B100 have the same machine learning model implemented in the image processing unit 102, while the information processing device A100, the information processing device C100, and the information processing device D100 have different machine learning models implemented in the image processing unit 102. The information processing device 100 may change the characteristics of the machine learning model depending on its application and installation environment.
[0074] In addition, different dataset storage units 201 are provided corresponding to the machine learning models. The inference attention region visualization unit 202 also uses different machine learning models 203 for each machine learning model. For example, data collected from the information processing devices A and B 100 in which the machine learning model 1 (203) is implemented is stored in the dataset storage unit 1 (201). The inference attention region visualization unit 202 analyzes images included in the collected data stored in the dataset storage unit 201 and identifies, from the image, an attention region for inference by the machine learning model 1 (203). The attention region identified by the inference attention region visualization unit 202 is an area in the image that contributed to the conclusion derived by the machine learning model. The robustness determination unit 204 determines, using a predetermined threshold, the ratio of the area of the attention region identified by the inference attention region visualization unit 202 to the area of the object in the image. If the area ratio of the attention region is equal to or less than the predetermined threshold, the robustness determination unit 204 determines the image to be a low-robust image and controls so that images with low robustness are sent to the data collection standard generation unit 206.
[0075] The feature extraction unit 205 calculates the feature amounts of the low-robust image identified by the inference attention area visualization unit 202. The data collection standard generation unit 206 clusters the image data arranged in a multidimensional space (illustrated as two-dimensional in FIG. 11 ) with the feature amounts as axes, and generates data collection standards that represent the ranges of the clusters (see FIG. 11 ). That is, data collection standards for the machine learning model 1 (203) are generated from the data stored in the dataset storage unit 1, and are transmitted to the information processing devices A and B 100.
[0076] Similarly, data collected from the information processing device C100 in which machine learning model 2 is implemented is stored in dataset storage unit 2 (201). The data collection criteria generation unit 206 generates data collection criteria for machine learning model 2 (203) from the data stored in dataset storage unit 2 (201), and transmits the generated criteria to the information processing device C100. Furthermore, data collected from the information processing device D100 in which machine learning model 3 (203) is implemented is stored in dataset storage unit 3 (201). The data collection criteria generation unit 206 generates data collection criteria for machine learning model 1 (203) from the data stored in dataset storage unit 3 (201), and transmits the generated criteria to the information processing device D100.
[0077] By configuring the information processing system shown in Figure 12 in this manner, the data collection criteria generation unit 206 can generate data collection criteria for each machine learning model 203, and the data collection criteria delivered to the information processing device 100 can be efficiently updated according to the characteristics of the machine learning model 203, allowing the information processing device 100 to efficiently collect data.
[0078] As described above, in the second embodiment of the present invention, compared to the first embodiment described above, the amount of calculation required for inference attention region visualization in the information processing device 100 can be reduced, and therefore real-time processing can be achieved even with few calculation resources of the information processing device 100. Furthermore, since the data collection standard is distributed from the server 200 to the information processing device 100, data can be collected according to the latest data collection standard.
[0079] <Example 3> An information processing system according to Example 3 of the present invention will be described below with reference to Fig. 13. In the information processing system according to Example 3, low-robust data is used for relearning to improve the robustness of the machine learning model. Note that in Example 3, differences from Examples 1 and 2 will be mainly described, and the same configurations and processes as those in Examples 1 and 2 will be assigned the same reference numerals, and descriptions thereof will be omitted.
[0080] FIG. 13 is a block diagram illustrating the configuration of an information processing system according to the third embodiment.
[0081] The information processing system of Example 3 includes one or more information processing devices 100 and a server 200. The information processing devices 100 and the server 200 are communicatively connected. In Fig. 13, only the main components of the information processing device 100 are shown, and the detailed configuration may be the same as that of Fig. 1, and the information processing device 100 may have a function of transmitting data collected by the information processing device 100 to the server 200.
[0082] The server 200 includes a dataset storage unit 201 , an inference attention area visualization unit 202 , a machine learning model 203 , a robustness determination unit 204 , a communication interface 207 , a re-learning data storage unit 208 , a pre-processing unit 209 , and a re-learning unit 210 .
[0083] The processes performed by the dataset storage unit 201, the inference attention region visualization unit 202, the machine learning model 203, and the robustness determination unit 204 are the same as those in the second embodiment. The dataset includes collected data. That is, the dataset storage unit 201 stores collected data transmitted from the information processing device 100. The inference attention region visualization unit 202 uses the machine learning model 203 to analyze an image included in the collected data stored in the dataset storage unit 201 and identify an attention region. The attention region identified by the inference attention region visualization unit 202 is an area in the image that contributed to a conclusion derived by the machine learning model. The machine learning model 203 has a machine learning model configured by a deep neural network (DNN) that has learned images labeled with the type of object depicted therein. The machine learning model 203 is a program that uses this machine learning model to infer the position and type of an object depicted in an input image and output an inference result. The inference attention region visualization unit 202 analyzes an image included in the collected data stored in the dataset storage unit 201 and identifies an attention region in the image for inference by the machine learning model 203. The attention region identified by the inference attention region visualization unit 202 is a region in the image that contributed to the conclusion derived by the machine learning model. The robustness determination unit 204 determines the ratio of the area of the attention region identified by the inference attention region visualization unit 202 to the area of the object in the image using a predetermined threshold, and if the area ratio of the attention region is equal to or less than the predetermined threshold, determines that the image is a low robust image, and stores the image with low robustness in the re-learning data storage unit 208.
[0084] The relearning data storage unit 208 stores collected data determined to have low robustness as relearning data. The preprocessing unit 209 converts the attention area of the relearning data stored in the relearning data storage unit 208 into data suitable for relearning by making it difficult to use for recognition. For example, the preprocessing unit 209 performs processing such as filling in the attention area (filling in other data (e.g., zeros)) or blurring the attention area. This enables relearning to actively use information other than the attention area to recognize objects. The relearning unit 210 retrains the machine learning model 203 using the data preprocessed by the preprocessing unit 209, improving the robustness of the machine learning model. The relearning unit 210 then transmits the machine learning model, whose robustness has been improved by relearning, to the information processing device 100 via the communication interface 207. At this time, to prevent overlearning with low-robust data, it is common to relearn including data with a certain degree of robustness. That is, it is preferable to mix low-robust data (processed by the pre-processing unit 209) with the data used in training the current DNN for training.
[0085] As described above, in Example 3 of the present invention, in response to the problem that the robustness of a machine learning model does not improve even if the machine learning model is retrained using low-robust data as is, the robustness of the machine learning model can be efficiently improved by preprocessing the low-robust data to change the region of interest to one that is difficult to recognize and using the data for retraining.
[0086] <Example 4> An information processing system according to Example 4 of the present invention will be described below with reference to Fig. 14. The information processing system according to Example 4 combines the functions of both Examples 2 and 3, which generate data collection standards from low-robust data and use the low-robust data for relearning to improve the robustness of the machine learning model. Note that in Example 3, differences from Examples 1, 2, and 3 will be mainly described, and the same configurations and processes as those in Examples 1, 2, and 3 will be assigned the same reference numerals and descriptions thereof will be omitted.
[0087] FIG. 14 is a block diagram showing the configuration of an information processing system according to the fourth embodiment.
[0088] The information processing system of Example 3 includes one or more information processing devices 100 and a server 200. The information processing devices 100 and the server 200 are communicatively connected. In Fig. 14, only the main components of the information processing device 100 are shown, and the detailed configuration may be the same as that in Fig. 1, and the information processing device 100 may have a function of transmitting data collected by the information processing device 100 to the server 200.
[0089] The server 200 includes a dataset storage unit 201, an inference attention area visualization unit 202, a machine learning model 203, a robustness determination unit 204, a feature extraction unit 205, a data collection standard generation unit 206, a communication interface 207, a re-learning data storage unit 208, a pre-processing unit 209, and a re-learning unit 210.
[0090] The processes executed by the dataset storage unit 201, the inference attention region visualization unit 202, the machine learning model 203, the robustness determination unit 204, the feature extraction unit 205, the data collection standard generation unit 206, and the communication interface 207 are the same as those in the above-described Example 2. The processes executed by the dataset storage unit 201, the inference attention region visualization unit 202, the machine learning model 203, the robustness determination unit 204, the communication interface 207, the re-learning data storage unit 208, the pre-processing unit 209, and the re-learning unit 210 are the same as those in the above-described Example 3.
[0091] The flow of a series of data and model updates using the configuration of this embodiment will be described. The image processing unit 102 of the information processing device 100 is equipped with an operational DNN, which performs image processing for a specific application in real time. At the same time, to determine whether an image acquired by a camera is low-robust for the operational DNN, features are extracted from the DNN's intermediate data, and the data collection determination unit 110 calculates the similarity with the features stored in the data collection standard storage unit 111 to determine whether data collection is necessary. Low-robust data that requires collection is temporarily stored in the collected data storage unit and transmitted to the server 200 at a predetermined timing. In this way, data acquired by each information processing device 100 at each operational site is stored in the dataset storage unit 201. To quantitatively determine the robustness of the data determined by the information processing device 100, the inference attention area visualization unit 202 and robustness determination unit 204 quantify the robustness of each collected data. If the data is likely to be useful for re-learning or updating the collected data standard, the image data is input to the downstream re-learning data storage unit 208 and feature extraction unit 205. When a certain amount of data on features extracted by the feature extraction unit 205 differs from the currently stored feature map (see FIG. 11 ) is accumulated, the reference data in the data collection reference storage unit 111 in each information processing device 100 is updated. Alternatively, relearning may be initiated when the number of images stored in the relearning data storage unit 208 exceeds a predetermined number, or when new feature data is accumulated in the feature map. Furthermore, data may be selected so that the distance between feature points is constant to evenly cover the feature map and avoid bias toward data with the same features. By executing the relearning process described in Example 3 using the data selected in this manner, a DNN with improved robustness is generated, and the generated DNN (machine learning model) is distributed to each information processing device 100.
[0092] As described above, in the fourth embodiment of the present invention, the amount of calculation performed by the information processing device 100 can be reduced, thereby enabling real-time processing even with limited calculation resources in the information processing device 100. Furthermore, since the server 200 distributes data collection standards to the information processing device 100, data can be collected according to the latest data collection standards. Furthermore, in the server 200 of the fourth embodiment, the robustness of the machine learning model can be improved by preprocessing low-robust data in the preprocessing unit 209 and using the data for re-learning. In other words, in the DNN update process, data collection and selection based on the empirical rules of data scientists and developers is no longer necessary, and the associated man-hours can be expected to be made more efficient.
[0093] The present invention is not limited to the above-described embodiments, and includes various modifications and equivalent configurations within the spirit and scope of the appended claims. For example, the above-described embodiments have been described in detail to clearly explain the present invention, and the present invention is not necessarily limited to configurations including all of the described configurations. Furthermore, part of the configuration of one embodiment may be replaced with the configuration of another embodiment. Furthermore, the configuration of another embodiment may be added to the configuration of one embodiment. Furthermore, part of the configuration of each embodiment may be added, deleted, or replaced with other configurations.
[0094] Furthermore, the aforementioned configurations, functions, processing units, processing means, etc. may be realized in part or in whole in hardware, for example by designing them as integrated circuits, or may be realized in software by having a processor interpret and execute a program that realizes each function.
[0095] Information such as programs, tables, and files that realize each function can be stored in a storage device such as a memory, hard disk, or SSD (Solid State Drive), or in a recording medium such as an IC card, SD card, or DVD.
[0096] In addition, the control lines and information lines shown are those that are considered necessary for explanation, and do not necessarily represent all the control lines and information lines that are necessary for implementation. In reality, it can be assumed that almost all components are interconnected.
Claims
1. An information processing device, comprising a computer having a calculation unit that executes calculation processing and a memory unit accessible to the calculation unit, and comprising: an image processing unit that executes an image processing program including a machine learning model and outputs recognition results for image data; an inference visualization analysis unit that identifies an area of interest for inference in the image processing unit; a robustness determination unit that determines the robustness of the image data using the identified area of interest; a collected data storage unit that stores image data whose robustness is determined to be lower than a first predetermined standard in the memory unit; and an interface unit that transmits image data stored in the collected data storage unit to the outside.
2. An information processing apparatus according to claim 1, wherein the first predetermined criterion is determined as a ratio of an area of the attention area to an area of the image.
3. A server that collects image data from an information processing device, comprising a computer having a calculation unit that performs calculation processing and a memory unit accessible to the calculation unit, and comprising: a dataset storage unit that stores image data collected by the information processing device that is determined to match a data collection standard and is stored in the memory unit; an inference attention area visualization unit that identifies an attention area for inference in image processing of the stored image data; a robustness determination unit that determines the robustness of the image data using the identified attention area; a feature extraction unit that extracts image features of image data whose robustness is determined to be lower than a first predetermined standard; a data collection standard generation unit that generates data collection standards using the extracted image features; and an interface unit that transmits the generated data collection standards to the information processing device.
4. The server according to claim 3, characterized in that the server communicates with a plurality of information processing devices that recognize image data by executing image processing programs including different machine learning models, the dataset storage unit is configured with separate areas for storing collected data for each of the plurality of information processing devices, the data collection criteria generation unit generates data collection criteria for each of the plurality of information processing devices, and the interface unit transmits the data collection criteria to each of the plurality of information processing devices.
5. The server according to claim 3, wherein the first predetermined criterion is determined as a ratio of an area of the attention area to an area of the image.
6. The server described in claim 3, characterized in that the feature extraction unit extracts low-dimensional information from image features of image data whose robustness is determined to be lower than the first specified standard, and the data collection standard generation unit generates the data collection standard using the results of clustering multiple image features and grouping similar image features.
7. A server as described in claim 3, comprising: a re-learning data storage unit that stores re-learning data for a machine learning model based on the robustness judgment result; a pre-processing unit that pre-processes the re-learning data using information on the identified area of interest; and a re-learning unit that re-learns the machine learning model using image data processed by the pre-processing unit, wherein the interface unit transmits the re-learned machine learning model to the information processing device.
8. The server according to claim 7, characterized in that the preprocessing unit processes data for areas of the identified areas of interest that are equal to or larger than a second predetermined standard.
9. The server according to claim 8, wherein the second predetermined criterion is determined as a ratio of an area of the attention area to an area of the image.
10. A server as described in claim 3, comprising: a re-learning data storage unit that stores re-learning data of a machine learning model based on the robustness judgment result; a pre-processing unit that pre-processes the re-learning data using information of the identified area of interest; and a re-learning unit that re-learns the machine learning model using image data processed by the pre-processing unit, wherein the interface unit transmits the re-learned machine learning model and the generated data collection criteria to the information processing device.
11. An information processing system including an information processing device and a server, wherein the information processing device has: a camera for acquiring external information; an image processing unit for performing image processing on image data acquired by the camera and outputting a result of image recognition; a data collection unit for collecting image data that matches a data collection standard; a feature extraction unit for extracting a first image feature of the image data used for the image recognition; a data collection standard storage unit for storing the data collection standard; a data collection judgment unit for judging whether the extracted first image feature matches the data collection standard; a collected data storage unit for storing image data that is judged to match the data collection standard; and a first interface unit for communicating data with an external device; and the server has: a data set storage unit for storing image data that is judged to match the data collection standard and collected by the information processing device; an inference attention area analysis unit for identifying an attention area of inference in the image processing of the stored image data; a robustness judgment unit for judging the robustness of the image data using the identified attention area; and a feature extraction unit for extracting a second image feature of image data whose robustness is judged to be lower than a first predetermined standard; an information processing system comprising: a data collection criteria generation unit that generates the data collection criteria using the extracted second image feature; a re-learned data storage unit that stores re-learned data for a machine learning model based on a result of the robustness judgment; a pre-processing unit that pre-processes the re-learned data using information on the identified region of interest; a re-learning unit that re-learns the machine learning model using image data processed by the pre-processing unit; and a second interface unit that transmits the generated data collection criteria and the re-learned machine learning model to the information processing device.
Citation Information
Patent Citations
Image processing device, machine tool, and image processing method
JP2021126711A
Learning data collecting apparatus, learning data collecting method, and program
JP2023139099A
Performance indexing device, performance indexing method, and program
WO2023190644A1
Object detecting device, object detecting method, and program
WO2024071347A1