Information processing device, control method, and program
The information processing apparatus addresses overfitting in AI models by acquiring, editing, and grouping data to extract high-quality learning data, enhancing the efficiency and reducing computational costs in structural deformation detection.
Patent Information
- Application Number
- JP2023214845
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-07-02
AI Technical Summary
Existing AI models for detecting structural deformations face issues with overfitting due to the inclusion of low-quality learning data, particularly images with minor edits, leading to inefficient and costly relearning processes.
An information processing apparatus that includes components for acquiring images, detecting deformations, receiving edits, grouping edited data, analyzing overfitting, and extracting high-quality learning data by setting thresholds and performing similarity determinations to reduce overfitting.
The apparatus effectively reduces overfitting by selecting high-quality learning data, improving the efficiency and reducing the computational burden of relearning processes.
Smart Images

Figure 2025098601000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus for detecting deformation of a structure.
Background Art
[0002] In recent years, when inspecting infrastructure structures, it has become possible to detect deformations such as cracks in concrete walls and exposed reinforcing bars by image analysis using an AI model. In order to improve the performance of the AI model, learning data for training is required. At this time, a large amount of computational processing is required for the relearning process, which is costly and time-consuming.
[0003] Therefore, it is required to collect higher-quality learning data excluding learning data that does not contribute to performance improvement. For example, in Patent Document 1, there is a function of not adopting an image with minor editing as learning data among the images edited by the user for the detection result.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] With the technique described in Patent Document 1, it is not always possible to collect high-quality learning data. For example, when an image with many similar images as learning data but not minor editing is collected as learning data, the AI model may overfit to that image.
[0006] In this way, when data edited by the user for the detection result is used as learning data for the purpose of improving the performance of the AI model, there is a possibility that edited data leading to overfitting is used as learning data.
Means for Solving the Problems
[0007] The information processing apparatus includes an acquisition unit that acquires an image of a structure, a detection unit that detects deformation from the image, an editing reception unit that receives editing of the detection result of deformation by the detection unit, a setting unit that sets a method for grouping the edited data edited by the editing reception unit, a grouping unit that groups the edited data by the method set by the setting unit, an analysis unit that analyzes overfitting based on the grouped edited data, an extraction unit that extracts learning data based on the analysis by the analysis unit, and a storage unit that stores the learning data extracted by the extraction unit.
Advantages of the Invention
[0008] According to the present invention, overfitting in learning data regarding deformation of a structure can be reduced.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Mode for Carrying Out the Invention
[0010] [Example 1] The method described in the following embodiments operates in an information processing apparatus 100 having the configuration of the block diagram shown in FIG. 1. The CPU 101 reads the OS and application programs from the storage device 104 or the ROM 102 and loads them into the RAM 103, and various processes proceed by executing them, realizing the functions of the embodiments described later. The application program acquires the input of the user from the input device 105 connected to the computer. Further, information is output to the output device 106 to display the processing result. Further, communication is performed with other computers, servers, devices, etc. connected to the network via the communication device 107. These hardware components are connected to each other by a bus 108 and are configured to be operable from an application program.
[0011] FIG. 2 is an example of the configuration of a system that supports the collection of learning data for an AI model that detects deformations of a structure in Example 1. In order to acquire an image and detect a deformation, it has an image acquisition unit 201 and a deformation detection unit 202. Further, in order to support the collection of learning data for the AI model, it has an editing reception unit 203, an editing data grouping method setting unit 204, an editing data group determination unit 205, an overfitting analysis unit 206, a learning data extraction unit 207, and a learning data storage unit 208.
[0012] Each component of the system that supports the collection of learning data for the AI model is realized by one or more servers and is connected by a LAN 211. The image acquisition unit 201 acquires a plurality of images of an arbitrary part of the structure. The deformation detection unit 202 performs image analysis using an AI model in order to detect a deformation in the acquired image.
[0013] The editing reception unit 203 receives editing for the deformation detection result. The editing data grouping method setting unit 204 sets how to group the editing data. The editing data group determination unit 205 determines to which group the editing data belongs and groups it. The overfitting analysis unit 206 determines whether adopting the editing data as learning data will lead to overfitting based on the grouping information of the editing data. The learning data extraction unit 206 extracts learning data based on the overfitting analysis result. The learning data storage unit 207 stores the extracted learning data.
[0014] FIG. 3 shows an example of the UI screen 300 of the application displayed on the output device 106 to add an image and detect deformation. The image list screen 300 of this system is divided into a folder list panel 301 and an image list panel 302. In order to specify the folder to be registered when saving an image, the user can add a folder by pressing the folder addition button 304. When an arbitrary folder is selected from the folder list 303, the image list 305 registered in the folder is displayed on the image list panel 302. Here, it is possible to switch to the detection result screen 500 described later by operating the image list panel switching tab 306. When the image addition button 307 is pressed, an image can be registered in the folder. Check boxes are provided for each item of the images in the image list 305. When the deformation detection button 308 is pressed, deformation detection using the AI model is executed on the images selected by the check boxes.
[0015] FIG. 4 is a table showing the configuration of deformation data, and shows deformation data 401 which is the result of detecting deformation from an image of a structure. Each row of the deformation data is information regarding a continuous deformation, and is composed of a deformation ID, a deformation type, a pre-editing coordinate list, and a post-editing coordinate list. The deformation ID is a symbol uniquely assigned to identify each deformation.
[0016] The deformation types are for distinguishing deformations such as cracks and rebar exposure. The shape of the deformation is expressed as a polyline or polygon by the coordinate list. The pre-editing coordinate list is automatically detected by an AI model based on the captured image. The post-editing coordinate list is the vertex coordinate list of the deformation after editing described later.
[0017] Figure 5 shows an example of the UI screen 500 of an application that displays the results of detecting deformations on the output device 106. The detection result list 501 displays a list of the results of performing deformation detection. In this example, 502 and 503 are displayed as selectable deformation detection results.
[0018] The deformation detection result display 504 displays information on the deformation detection result selected by selecting the deformation detection result. In this example, the detection result of the selected 502 is displayed. The crack 505 is a drawing of the coordinate information with the deformation ID of D1 in the deformation data 401. Similarly, the rebar exposure 506 corresponds to D2, and the crack 507 corresponds to D3. The edit button 508 accepts editing of the deformation detection result. An editing instruction from the user for misdetection or detection omission is received by the pointer 511, and editing is executed. For example, when the user determines that the crack 505 is a misdetection, the selection of the crack 505 is accepted by the pointer 511, and deletion is accepted by executing a deletion operation with the input device 105. At this time, multiple editing instructions can be received.
[0019] By pressing the save button 509, the detection result displayed in the deformation detection result display 504 becomes edited. The check boxes for cracks or rebar exposure in the detection display 510 can be used to switch the display and non-display of each of the cracks and rebar exposure.
[0020] Figure 6 shows an example of the deformed data after editing. In this example, by setting the post-editing coordinate list with the deformation ID D1 to an empty list, it manages that the crack 505 has been deleted. Also, by updating the post-editing coordinate list with the deformation ID D2 from the pre-editing coordinate list, it manages that the coordinates of the steel bar exposure 506 have been corrected. By not updating the post-editing coordinate list with the deformation ID D3 from the pre-editing coordinate list, it manages that the crack 507 has not been edited. Furthermore, by adding a row with the deformation ID DN, it manages that a crack has been added.
[0021] Figure 7 shows an example of grouping the editing data. Group 1 (bridge type) and Group 2 (deformed area) are columns that can be added by the editing data grouping method setting unit 204. The edited deformed data is classified by group. For example, the data with the deformation ID D1 is classified as bridge C in Group 1 (bridge type) and in the range of 10 mm^2 to in Group 2 (deformed area). At this time, the data with the deformation ID D3 that has not been edited is not classified.
[0022] Figure 8 shows the cumulative number of edited data for each group used in the overfitting analysis unit 206. For the number of edited data 801 of Group 1 (bridge type), the horizontal axis represents Group 1 (bridge type), such as Bridge A, Bridge B, and Bridge C. The vertical axis represents the cumulative number of edited data for each. Also, for the number of edited data 802 of Group 2 (deformed area), the horizontal axis represents Group 2 (deformed area), such as ~1 mm^2, 1 mm^2 to 10 mm^2, 10 mm^2 and above. The vertical axis represents the cumulative number of edited data for each. In the overfitting analysis unit 206, thresholds are manually registered in advance for each group. The thresholds can be set for each attribute of each group. The overfitting analysis unit 206 determines for each group whether the number of edited data of that attribute exceeds the threshold. If none of the groups exceed the threshold, it is determined that there is no overfitting, and the data is adopted as learning data. If a certain group exceeds the threshold, it is determined that there is overfitting and the data is not adopted as learning data. For example, in the edited data grouping example 701, the data with deformation ID D1 is classified as Bridge C in Group 1 (bridge type) and 10 mm^2 and above in Group 2 (deformed area). Since it does not exceed the threshold for the attributes of each group, it is adopted as learning data. On the other hand, the data with deformation ID D1 is classified as Bridge A in Group 1 (bridge type) and 10 mm^2 and above in Group 2 (deformed area). The data with deformation ID D1 does not exceed the threshold for the attribute of Group 2 (deformed area), but exceeds the threshold for the attribute of Group 1 (bridge type), so it is not adopted as learning data.
[0023] Figure 9 is a main flowchart showing the process of extracting learning data in Example 1. The following processes are realized by the CPU 101 controlling each part of the information processing apparatus 100.
[0024] In S901, deformation detection is performed using an AI model. In S902, the result of the deformation detection is displayed. In S903, the user is asked to confirm whether there are any detection omissions or false detections in the deformation detection result, and an editing instruction is received. If there is an editing instruction, the process proceeds to S904 to perform deformation editing. In S905, the edited data is acquired. In S906, the edited data extracted after the editing is grouped. In S907, it is determined whether the number of adopted data for each group in each edited data is within the threshold. If it is within the threshold, it is adopted as learning data in S908. If it is more than the threshold, it is not adopted as learning data in S909.
[0025] [Example 2] In Example 1, the means of manually registering the threshold for whether to adopt each attribute of each group as learning data in advance in the overfitting analysis unit 206 was explained. In Example 2, the means of automatically registering the threshold for whether to adopt each attribute of each group as learning data based on the degree of improvement in the accuracy of the AI model when the edited data is adopted as learning data will be explained.
[0026] Figure 10 shows the degree of improvement in the accuracy of the AI model 1001 when each attribute of the group is adopted as learning data and the threshold 1002 automatically registered based on the degree of improvement in the accuracy. For example, when the degree of improvement in the accuracy of the AI model is large when the edited data classified as Bridge C is adopted as learning data, the threshold for whether to adopt it as learning data for Bridge C is set large. On the other hand, when the degree of improvement in the accuracy of the AI model is small, such as for Bridge B, the threshold is set small.
[0027] As described above, the threshold for whether to adopt each attribute of each group as learning data is automatically registered based on the degree of improvement in the accuracy of the AI model when the edited data is adopted as learning data. Thereby, compared with the means of manually registering the threshold for whether to adopt each attribute of each group as learning data in advance in the overfitting analysis unit 206, the threshold can be registered efficiently.
[0028] [Example 3] In Example 1, the means of manually registering the threshold for whether to adopt each attribute of each group as learning data in advance in the overfitting analysis unit 206 was explained. In Example 3, the means of automatically registering the threshold so as to extract many attributes of a newly added group as learning data later will be explained. For example, when an attribute of bridge C in Group 1 (bridge type) is newly added to the attributes, the threshold of bridge C is set to be larger than the thresholds of bridge A and bridge B that have existed since before bridge C.
[0029] As described above, in this embodiment, the threshold is automatically registered so as to extract many attributes of a newly added group as learning data later. Thereby, compared with the case of manually determining the threshold for adopting each attribute of each group as learning data in the overfitting analysis unit, when there is a possibility that the inspection images of attributes increase and the inspection images of existing attributes decrease, the threshold can be registered efficiently.
[0030] [Example 4] In Example 1, it was explained that for each edited data, it is determined whether the number of adopted data for each group is within the threshold, and if it is within the threshold, it is adopted as learning data, and if it is more than the threshold, it is not adopted as learning data. In Example 4, the means of making a similarity determination with the image of an attribute for which the number of adopted data has not reached the threshold when the number of adopted data is more than the threshold, and determining whether to adopt it as learning data will be explained.
[0031] FIG. 11 is a flowchart in Example 4. In S1101, an image for which it is determined that the number of adopted data is more than the threshold is subjected to a similarity determination with an image of an attribute for which the number of adopted data has not reached the threshold. An image with a high similarity is adopted as learning data, and an image with a low similarity is not adopted as learning data.
[0032] As described above, when the number of adopted data is greater than the threshold, it is determined whether to perform similarity determination with an image of an attribute for which the number of adopted data has not reached the threshold and adopt it as learning data. As a result, it is possible to increase the number of learning data adopted and efficiently extract learning data, rather than simply determining whether the number of adopted data is within the threshold and determining whether to adopt it as learning data.
[0033] [Example 5] In Example 1, an example of analyzing and evaluating overfitting based on a set threshold and extracting learning data based on the result was described. In Example 5, an example is shown in which the user evaluates the editing results and the number of adopted data for the previous and current times, and extracts learning data based on the evaluation result.
[0034] FIG. 12 shows an example of the UI screen 1200 of Example 5 to be displayed on the output device 106. The user can check the differences 1201 and 1202 between the detected result shapes and the editing results at the time of the current inspection and the previous inspection, and the adoption numbers 1203 and 1204 as learning data for each attribute, and analyze and determine whether to adopt it as learning data. When the adoption button 1205 is pressed, the corresponding edited data is adopted as learning data, and the next edited data is displayed. Also, when the non-adoption button 1206 is pressed, it is not adopted as learning data and the next edited data is displayed.
[0035] FIG. 13 is a flowchart in Example 5. In S1301, the editing data and the number of adopted data for the previous and current times are displayed. In S1302, the user evaluates whether there is overfitting from the editing data and the number of adopted data for the previous and current times, and determines whether to adopt the corresponding editing data as learning data.
[0036] As described above, by the user evaluating the editing results and the number of adopted data for the previous and current times, and extracting learning data based on the evaluation result, overfitting can be analyzed based on the improvement in accuracy from the previous time and the increasing trend of the number of adopted data from the previous time, and learning data can be efficiently extracted.
[0037] [Example 6] In Example 1, an example was described in which overfitting was analyzed and evaluated based on a set threshold value, and learning data was extracted based on the result. In Example 6, an example was shown in which the user evaluates the editing result and the number of adopted data at the same location of the same structure, and learning data is extracted based on the evaluation result.
[0038] FIG. 14 shows an example of the UI screen 1400 of Example 6 to be displayed on the output device 106. The user can check the difference 1401 between the detection result shape and the editing result at the time of this inspection, the differences 1402 and 1403 between the detection result shapes and the editing results at a plurality of the same locations of the same structure, and the number of adopted data 1404 as learning data for each attribute, and can analyze and determine whether to adopt it as learning data. When the adopt button 1405 is pressed, the corresponding editing data is adopted as learning data, and the next editing data is displayed. Also, when the non-adopt button 1406 is pressed, it is not adopted as learning data and the next editing data is displayed.
[0039] FIG. 15 is a flowchart in Example 6. In S1501, the editing result and the number of adopted data at the same location of the same structure are displayed. In S1502, the user evaluates whether to overfit from the editing data and the number of adopted data at the same location of the same structure, and determines whether to adopt the corresponding editing data as learning data.
[0040] As described above, by having the user evaluate the editing result and the number of adopted data at the same location of the same structure, and extracting learning data based on the evaluation result, overfitting is analyzed based on the accuracy of the AI model and the number of adopted data, and learning data can be efficiently extracted.
[0041] [Example 7] In Example 1, an example was described in which overfitting was analyzed and evaluated based on a set threshold value, and learning data was extracted based on the result. In Example 7, an example is shown in which the user evaluates the editing result and the number of adopted data in a group of types of changes before and after editing, and learning data is extracted based on the evaluation result.
[0042] FIG. 16 shows an example of the UI screen 1600 of Example 7 to be displayed on the output device 106. The user can check the difference 1601 between the detected result shape and the edited result, and the adoption numbers 1602 and 1603 of each attribute of the group of change types before and after editing as learning data, analyze whether to adopt as learning data, and make a decision. When the adoption button 1604 is pressed, the corresponding edited data is adopted as learning data, and the next edited data is displayed. Also, when the non - adoption button 1605 is pressed, it is not adopted as learning data and the next edited data is displayed.
[0043] FIG. 17 is a flowchart in Example 7. In S1701, the adoption data numbers of the edited data and the group of change types before and after editing are displayed. In S1702, the user evaluates whether there is over - learning and determines whether to adopt the corresponding edited data as learning data.
[0044] As described above, by having the user evaluate the adoption data numbers of the edited result and the group of change types before and after editing, and extracting learning data based on the evaluation results, over - learning can be analyzed based on the accuracy of the AI model and the adoption data numbers, and learning data can be efficiently extracted.
[0045] [Other Embodiments] The present invention can also be realized by supplying a program that realizes one or more functions of the above - described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (for example, ASIC) that realizes one or more functions.
[0046] Note that the present invention is not limited to the above - described embodiments as they are, and at the implementation stage, the components can be modified and embodied without departing from the gist thereof. Also, various inventions can be formed by appropriately combining a plurality of components disclosed in the above - described embodiments. For example, some components can be deleted from all the components shown in the embodiments. Furthermore, components from different embodiments can be appropriately combined.
[0047] In addition, in the above embodiment, at least one of A and B may be only A, only B, or both A and B.
[0048] The disclosure of this embodiment includes the following configurations and methods.
[0049] [Configuration 1] An acquisition means for acquiring an image obtained by photographing a structure; A detection means for detecting deformation from the image; An editing reception means for receiving editing of the detection result of deformation by the detection means; A setting means for setting a method of grouping the edited data edited by the editing reception means; A means for grouping the edited data by the method set by the setting means; An analysis means for analyzing overfitting based on the grouped edited data; An extraction means for extracting learning data based on the analysis by the analysis means; A storage means for storing the learning data extracted by the extraction means; An information processing apparatus characterized by having the above.
[0050] [Configuration 2] The information processing apparatus according to Configuration 1, wherein the analysis means analyzes overfitting based on the degree of improvement in the accuracy of the AI model when the edited data is adopted as learning data.
[0051] [Configuration 3] The information processing apparatus according to Configuration 1 or 2, wherein the extraction means extracts learning data by performing a similarity determination with an image of an attribute that has not reached the threshold when the number of adopted data is more than the threshold.
[0052] [Configuration 4] A display means for displaying the previous and current editing results and the number of adopted data; A means for receiving acceptance or rejection of the adoption of learning data; The information processing apparatus according to any one of Configurations 1 to 3, further comprising
[0053] [Configuration 5] display means for displaying the editing result and the number of adopted data at the same location of the same structure, means for accepting whether learning data can be adopted, The information processing apparatus according to any one of Configurations 1 to 4, further comprising
[0054] [Configuration 6] display means for displaying the editing result and the number of adopted data of the group of types of changes before and after editing, means for accepting whether learning data can be adopted, The information processing apparatus according to any one of Configurations 1 to 5, further comprising
Claims
1. An acquisition means for acquiring an image of a structure, A detection means for detecting deformation from the image, An edit reception means for receiving an edit for the detection result of deformation by the detection means, A setting means for setting a method of grouping the edit data edited by the edit reception means, A means for grouping the edit data by the method set by the setting means, An analysis means for analyzing overfitting based on the grouped edit data, An extraction means for extracting learning data based on the analysis by the analysis means, A storage means for storing the learning data extracted by the extraction means, An information processing apparatus characterized by comprising the above.
2. The information processing apparatus according to claim 1, wherein the analysis means analyzes overfitting based on the degree of accuracy improvement of the AI model when the edit data is adopted as learning data.
3. The information processing apparatus according to claim 1, wherein the extraction means extracts learning data by making a similarity determination with an image of an attribute that does not reach the threshold when the number of adopted data is more than the threshold.
4. A display means for displaying the previous and current edit results and the number of adopted data, A means for receiving acceptance of whether learning data can be adopted, The information processing apparatus according to claim 1, further comprising the above.
5. A display means for displaying the edit results and the number of adopted data at the same location of the same structure, A means for receiving acceptance of whether learning data can be adopted, The information processing apparatus according to claim 1, further comprising the above.
6. A display means for displaying the edit results and the number of adopted data in a group of deformation types before and after editing, A means for receiving acceptance of whether learning data can be adopted, The information processing apparatus according to claim 1, further comprising the above.
7. An acquisition step of acquiring an image of a structure, A detection step of detecting deformation from the image, An edit reception step of receiving an edit for the detection result of deformation by the detection step, A setting step of setting a method of grouping the edit data edited by the edit reception step, A step of grouping the edit data by the method set by the setting step, An analysis step of analyzing overfitting based on the grouped edit data, An extraction step of extracting learning data based on the analysis by the analysis step, A storage step of storing the learning data extracted by the extraction step, A control method characterized by having
8. An acquisition step of acquiring an image obtained by photographing a structure, A detection step of detecting deformation from the image, An editing reception step of receiving editing for the detection result of deformation by the detection step, A setting step of setting a method for grouping the edited data edited by the editing reception step, A step of grouping the edited data by the method set by the setting step, An analysis step of analyzing overfitting based on the grouped edited data, An extraction step of extracting learning data based on the analysis by the analysis step, A storage step of storing the learning data extracted by the extraction step, A program for causing a computer to execute
Citation Information
Patent Citations
Voice analytic synthesizing device
JP1995104799A