Abnormality estimation device, abnormality estimation method, and program

The abnormality estimation device enhances FPCNet by using multiple convolution processes and SWIN transformers to retain spatial information, improving the accuracy of anomaly detection.

JP2025111973APending Publication Date: 2025-07-31PASCO CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024005937
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-18
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Conventional FPCNet models lose spatial information in feature amounts, leading to reduced accuracy in estimating abnormal locations.

Method used

An abnormality estimation device that generates first and second feature amounts through multiple convolution processes, incorporates transposed convolution and attention information synthesis, and uses SWIN transformers to retain spatial information for enhanced estimation accuracy.

Benefits of technology

The device achieves higher accuracy in estimating abnormal locations by maintaining spatial information in attention information, facilitating precise detection of anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025111973000001_ABST
    Figure 2025111973000001_ABST
Patent Text Reader

Abstract

To provide an abnormality estimation device, an abnormality estimation method, and a program that allow more accurate estimation of an abnormal place.SOLUTION: An information processing device as an abnormality estimation device includes a machine learning model that includes an encoding unit, a dilated convolution unit, and a decoding unit. The encoding unit performs convolution processing for a captured image a plurality of times to generate a first feature amount. The dilated convolution unit performs dilated convolution processing for the first feature amount a plurality of times to generate a second feature amount. The decoding unit performs update processing a plurality of times to generate an abnormality estimation image indicating a result of estimation of the degree of possibility of occurrence of an abnormality in each pixel of the captured image. The update processing is a combination of first processing and second processing. In the first processing, an intermediate feature amount is generated by adding the first feature amount to a result of performing transposed convolution for the second feature amount. In the second processing, the second feature amount is updated by generating attention information from the intermediate feature amount and combining the attention information with the intermediate feature amount. Each feature amount and attention information have both a dimension in a channel direction of the captured image and a two-dimensional direction on the image which is different from the channel direction.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an abnormality estimation device, an abnormality estimation method, and a program.

Background Art

[0002] In recent years, the need for inspection and maintenance of infrastructure has been increasing, and the demand for automation and efficiency improvement of these infrastructure inspections has also been growing. For example, in the inspection of port facilities, by inputting an image of a caisson taken by an unmanned aerial vehicle (UAV) or the like into a model trained by machine learning, it is possible to automatically estimate the cracked areas on the caisson, thereby expecting automation and significant efficiency improvement of the work. Conventionally, as a model for automatically estimating cracks from a captured image, Fast Pavement Crack Detection Network (FPCNet) is known (Non-Patent Document 1).

[0003] This FPCNet is an encoder-decoder type model having an encoder that performs convolution processing on an image to extract feature amounts, and a decoder that performs transposed convolution processing on the feature amounts to return them to the original image size. In the decoder of FPCNet, attention information indicating which channel among the feature amounts of a plurality of channels extracted from the image should be focused on is generated, and the feature amounts are emphasized according to the generated attention information. Thereby, FPCNet aims to improve the estimation accuracy of cracks, which are abnormal areas to be detected.

Prior Art Documents

Non-Patent Documents

[0004]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, in the conventional FPCNet, the feature amount is compressed into one-dimensional information having only the dimension of the channel to generate attention information. As a result, there is a problem that spatial information useful for estimating abnormal locations is lost in the feature amount.

[0006] An object of the present invention is to provide an abnormality estimation device, an abnormality estimation method, and a program capable of estimating abnormal locations with higher accuracy.

Means for Solving the Problems

[0007] To achieve the above object, the present invention an encoding unit that generates a first feature amount by performing a plurality of convolution processes on a captured image represented by a set of a plurality of pixels; an extended convolution unit that generates a second feature amount by performing a plurality of extended convolution processes on the first feature amount; a first process of generating an intermediate feature amount by adding the first feature amount to the result of performing a transposed convolution process on the second feature amount, and a second process of generating attention information from the intermediate feature amount and synthesizing the attention information with the intermediate feature amount to update the second feature amount, and a decoding unit that generates an abnormality estimation image indicating an estimation result of the likelihood of an abnormality appearing in each pixel of the captured image by executing a plurality of times a set of the first process and the second process; comprising wherein the first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction, is an abnormality estimation device.

Effects of the Invention

[0008] According to the present invention, there is an effect that abnormal locations can be estimated with higher accuracy.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Mode for Carrying Out the Invention

[0010] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a block diagram showing the functional configuration of an information processing apparatus 1 which is an abnormality estimation apparatus according to the present embodiment.

[0011] The information processing apparatus 1 includes a control unit 11, a storage unit 12, an input / output interface 13 (I / F), a display unit 14, an operation reception unit 15, and the like.

[0012] The control unit 11 comprehensively controls the operation of the information processing apparatus 1. The control unit 11 has a processor that performs arithmetic processing. The processor may be a single general-purpose CPU (Central Processing Unit), or may have a plurality of CPUs, and these may perform arithmetic processing in parallel or independently according to the application or the like. The processor may include those specialized for specific arithmetic processing or image processing. The control unit 11 performs various control processes by reading and executing a program 120 and the like from the storage unit 12.

[0013] The memory unit 12 has a RAM (Random Access Memory) and a non-volatile memory, and stores various data. The RAM provides a working memory space for the control unit 11 and stores temporary data. The non-volatile memory stores and holds programs 120, setting data, etc. The non-volatile memory is, for example, a flash memory, an HDD (Hard Disk Drive), etc., but is not limited thereto. The memory unit 12 may have a ROM (Read Only Memory). An initial control program, etc. may be stored in the ROM. The program 120 includes a control program related to the abnormality estimation process described later. The program 120 includes a machine learning model 121.

[0014] The input / output interface 13 performs input / output of data between the outside of the information processing apparatus 1 (including peripheral devices). The input / output interface 13 has a connection terminal 131 and a communication unit 132. The connection terminal 131 includes, for example, a USB (Universal Serial Bus) terminal, a LAN (Local Area Network) connector, etc. The communication unit 132 controls communication according to a communication protocol (protocol) related to a LAN such as TCP / IP.

[0015] As peripheral devices, a database device 21 which is an auxiliary storage device, and an optical reading device 22 which reads a portable storage medium (optical disk) such as a CDROM, a DVD, and a Blu-ray (registered trademark) may be included. Further, a magnetic tape may be included in the portable storage medium, and a reading device for reading this magnetic tape may be included in the peripheral devices.

[0016] Among the data that the information processing apparatus 1 can acquire from the outside via the input / output interface 13, there is included captured image data 201. The captured image data 201 is data of an image (captured image) obtained by capturing an object for which an abnormality is to be estimated. Here, the object is a caisson in a port facility, and the abnormality is a crack, that is, a crack on the concrete surface of the caisson. The captured image is represented by a set of a plurality of pixels two-dimensionally arrayed in a grid pattern in the vertical and horizontal (xy directions; hereinafter also referred to as the spatial direction). The captured image may be a grayscale monochromatic image or a color image. Note that the captured image data 201 may be stored in the storage unit 12.

[0017] The display unit 14 performs display on the display screen under the control of the control unit 11. The display screen is, for example, a liquid crystal display or an organic EL (Electro-Luminescent) display, etc., but is not limited thereto.

[0018] The operation reception unit 15 receives an input operation from the outside and outputs an operation signal corresponding to the input operation to the control unit 11. The operation reception unit 15 includes, for example, a keyboard and a pointing device, etc. The pointing device may be a mouse. The display unit 14 and / or the operation reception unit 15 may be peripheral devices of the information processing apparatus 1. That is, these may be attached to the main body (computer) of the information processing apparatus 1 including the control unit 11, the storage unit 12, and the input / output interface 13.

[0019] Next, abnormality estimation will be described. The control unit 11 performs anomaly estimation on the captured image using a machine learning model. When estimating cracks as anomalies, targets include, for example, the pavement surfaces of roads, runways, tunnel walls, fences along roads, bridges, dams (including sediment dams, etc.), and weirs, as well as the walls and roofs of various structures. Also, when the target is a bridge, anomalies to be estimated include, for example, peeling, exposed reinforcement, swelling, free lime, water leakage, etc. Note that the machine learning model 121 described later may be separately learned and generated according to these targets and types of anomalies. The image can be captured from above by a UAV or the like, or by a moving vehicle, by a photographer on foot using hand-held shooting, or by fixed shooting using a tripod or the like. Further, the anomaly estimation of the present disclosure may be used for estimating, as anomalies, locations where oak trees have withered from captured images of forests taken by aircraft or satellites.

[0020] Figure 2 is a flowchart showing the control procedure of the anomaly estimation process executed by the information processing apparatus 1. As described above, the control content related to the anomaly estimation process including the anomaly estimation method of the present embodiment is included in the program 120. The control unit 11 acquires the captured image data of the target for anomaly estimation (S1). The control unit 11 inputs the acquired captured image data to the learned machine learning model 121 (S2). The control unit 11 executes the processing content of the machine learning model 121 (S3). The more detailed processing content of the machine learning model 121 will be described later with reference to FIGS. 3, 4, etc.

[0021] The control unit 11 acquires the anomaly estimation image output from the machine learning model 121 (S4). The anomaly estimation image is an image having the same size as the captured image in the spatial direction and is a set of pixels corresponding to each pixel of the captured image. The anomaly estimation image indicates the estimation result of whether an anomaly appears in each pixel of the captured image. For example, the value range of each pixel value of the anomaly estimation image is set to be 0.0 or more and 1.0 or less, and the larger the value of the pixel of the anomaly estimation image corresponding to the pixel where an anomaly is likely to appear in the captured image, and the smaller the value of the pixel of the anomaly estimation image corresponding to the pixel where an anomaly is less likely to appear in the captured image. The control unit 11 outputs a detection result of an anomaly based on the anomaly estimation image (S5). The control unit 11 may, for example, simply output the anomaly estimation image, or may convert the anomaly estimation image into a grayscale image and output it, or may compare a threshold value determined in advance by experiments or the like with each pixel value of the anomaly estimation image and output a binary image indicating the presence or absence of an anomaly. Alternatively, the control unit 11 may obtain the position, size, feature amount, etc. of the portion estimated to be an anomaly and display them in a list. The output may be a display by the display unit 14, or may be performed on the outside or the database device 21, etc. as image data or text data in a predetermined format. Then, the control unit 11 ends the anomaly estimation process.

[0022] Next, the configuration of the machine learning model 121 will be described. The machine learning model 121 performs image recognition as described above. Specifically, an encoder-decoder type including an encoding process (encoding step, encoding means) and a decoding process (decoding step, decoding means) is used for the machine learning model 121. In the encoding process, convolution and pooling are performed to extract feature amounts (Multi-Convolution Feature: MCF, first feature amount) while downsampling in the spatial direction of the image. In the decoding process, the feature amount (Multi-Dilation Feature: MDF, second feature amount) highly related to the abnormality to be detected is upsampled in the spatial direction to update the feature amount and generate an abnormality estimation image. As a method for detecting abnormalities such as cracks by this type of machine learning model, FPCNet is known.

[0023] FIG. 3 is a diagram for explaining the schematic configuration of the machine learning model 121 of the present embodiment. As shown in FIG. 3(a), the machine learning model 121 includes an encoding unit 41, an extended convolution unit 42, and a decoding unit 43. The captured image data is input to the encoding unit 41. The feature amount output from the encoding unit 41 is subjected to extended convolution by the extended convolution unit 42. The feature amount obtained by the extended convolution is input to the decoding unit 43, upsampled to the original spatial size, and output as an abnormality estimation image.

[0024] The parameters in each layer and each function constituting the encoding unit 41, the extended convolution unit 42, and the decoding unit 43 are determined in advance by deep learning. Prior to learning, a large number of learning images obtained by photographing the object and correct data indicating pixels where abnormalities appear and pixels where no abnormalities appear in each learning image are prepared and stored in the storage unit 12 or the database device 21. Also, the parameters of the machine learning model 121 are initialized.

[0025] The correct data can be, for example, an image in which a value of 1.0 is set for pixels with cracks in the learning image of the caisson, and a value of 0.0 is set for pixels without cracks. The correct data is created, for example, by an operator visually checking and manually setting the presence or absence of cracks.

[0026] In deep learning, the control unit 11 sequentially inputs each learning image to the machine learning model 121, and updates the parameters of the machine learning model 121 so that the value of the error function between the abnormal estimation image output by the machine learning model 121 for the learning image and the correct data corresponding to the learning image becomes small. For example, the error can be an L1 error or a cross-entropy error. The parameters are updated by, for example, the error backpropagation method using the gradient descent method. The machine learning model 121 stored in the storage unit 12 through such learning is a learned model that is learned to output an abnormal estimation image representing the likelihood of an abnormality appearing in each pixel of the image when an image of the object is input.

[0027] As shown in FIG. 3(b), the encoding unit 41 has a plurality (N) of convolutional units 411 connected in series (cascaded). Each convolutional unit 411 has a convolutional layer and a pooling layer. The convolutional layer includes, for example, a convolutional process for extracting features and an activation function. The filter for performing the convolutional process can have an appropriate number of channels determined. The number of channels can increase through the processing of the plurality of convolutional units 411. The activation function can be, for example, a Rectified Linear Unit (ReLU).

[0028] The pooling layer performs Max Pooling to extract the maximum value for every predetermined number of pixel ranges, for example, a range of 2×2 pixels. As a result, the sizes in the two-dimensional vertical and horizontal directions (spatial directions) of the image are each halved. That is, the size in the spatial direction becomes 1 / 4. From the input end (upstream side) of the captured image, the nth convolutional unit 411 outputs a three-dimensional feature amount MCF(n) in the two-dimensional vertical and horizontal directions and the channel direction changed as described above, and sends it to the convolutional unit 411 on the downstream side (the (n + 1)th). Also, the feature amounts MCF(n) are each temporarily held and used in the SWIN unit 431(n) of the decoding unit 43. The number of times N (plural) of convolution and the size of the final feature amount MCF(N) in the spatial direction may be determined as appropriate. For example, the number of times of convolution N = 4 may be used, and the size of the feature amount MCF(N) may be 1 / 256 of the size of the original captured image (the length of each side is 1 / 2 4 )). Alternatively, the number of times of convolution etc. may be determined according to the size of the input captured image. In the present embodiment, each of the feature amounts MCF(n) obtained by the nth (1 < n ≦ N) convolutional unit 411 has a configuration having a plurality of components in each of the channel direction, the x direction, and the y direction.

[0029] The extended convolution unit 42 performs extended convolution processing (dilation convolution) on the feature amount MCF(N) obtained by the encoding unit 41 (extended convolution step, extended convolution means). In the extended convolution unit 42, a feature amount MDF(N) (second feature amount) reflecting a plurality of characteristic sizes related to anomalies is generated by multi dilation in which a plurality of results obtained by performing extended convolution processing with a plurality of dilation factors (dilation factor, rate) are combined. The feature amount MDF(N) has a length that is half of the length of each direction in the spatial direction of the feature amount MCF(N) and a number of channels that is twice the number of channels of the feature amount MCF(N). The feature amount MDF(N) is input to the decoding unit 43. The processing configurations of the encoding unit 41 and the extended convolution unit 42 are the same as those of the conventional FPCNet, and detailed descriptions thereof are omitted.

[0030] In the decoding unit 43, the same number (N) of SWIN (Shifted Windows) units 431 as in the encoding unit 41 are connected in series (cascaded). Here, the SWIN units 431 are numbered in ascending order from the output end (downstream side).

[0031] The SWIN unit 431 includes a process of performing transposed convolution on the input feature amount MDF(n) and a process of generating attention information using the SWIN transformer. The SWIN unit 431 updates the feature amount MDF(n) to the feature amount MDF(n - 1) based on the results of these processes and outputs this.

[0032] FIG. 4 is a diagram showing the detailed configuration of the SWIN unit 431. The feature amount MDF(n) is input to the n-th SWIN unit 431(n) from the downstream side (rear). The SWIN unit 431 includes a transposed convolution processing unit 61, a first addition unit 62, a conversion unit 63, SWIN transformers 64, 65, an unfolding unit 66, a data combination unit 67, a convolution unit 68, a second addition unit 69, and the like. Among these, the conversion unit 63, the SWIN transformers 64, 65, and the unfolding unit 66 are included in the attention information generation unit 60 of the present embodiment.

[0033] The transposed convolution processing unit 61 performs transposed convolution on the input feature amount MDF(n), that is, performs the inverse operation of convolution to upsample (expand) the feature amount MDF(n) in the spatial direction. The feature amounts obtained by the transposed convolution processing unit 61 of the n-th (1 < n ≤ N) SWIN unit 431 each have a plurality of components in the x direction and the y direction, respectively, and include spatial information. In the present embodiment, the size of the feature amount MDF(n) is the same as the size of the feature amount MCF(n) output from the n-th convolution unit 411 in the encoding unit 41.

[0034] The first addition unit 62 adds the feature amount obtained by the transposed convolution processing unit 61 and the feature amount MCF(n) obtained by the n-th convolution unit 411 (the feature amount obtained in the n-th update process from the back). The addition here means calculating the sum of the components of the corresponding coordinates in the spatial direction and the channel direction respectively. Hereinafter, the feature amount obtained by the addition in the first addition unit 62 is referred to as an intermediate feature amount. The intermediate feature amount is sent to the attention information generation unit 60, the data combination unit 67, and the second addition unit 69 respectively. Any of the intermediate feature amounts obtained by the n-th (1 < n ≤ N) SWIN unit 431 has a plurality of components in each of the x direction and the y direction and includes spatial information. In the present embodiment, the size of the intermediate feature amount in the n-th SWIN unit 431 is the same as the sizes of the feature amount MDF(n) and the feature amount MCF(n), and the addition by the first addition unit 62 means calculating the sum of the components of the same coordinates. The processing of the transposed convolution processing unit 61 and the processing of the first addition unit 62 correspond to the first processing of the present embodiment.

[0035] In the attention information generation unit 60, in the conversion unit 63, an embedding process is performed in which intermediate feature amounts that are two-dimensional in the spatial direction and one-dimensional in the channel direction are array-converted in the channel direction. This converted data is input to the SWIN Transformer 64. The SWIN Transformers 64 and 65 are a type of Transformer using deep learning, which divides an image into a plurality of windows and independently generates attention information for each of them. In the SWIN Transformer 65, attention information in a window shifted (shifted) from the window division position in the SWIN Transformer 64 is generated. Thereby, the relationship of image features between windows is covered. Since the detailed content of the SWIN Transformers 64 and 65 is well-known, a detailed description thereof is omitted. This attention information obtained for each position in the spatial direction is expanded by the expansion unit 66 into each component in the original channel direction and output. As described above, attention information for each spatial position and channel of the intermediate feature amount is generated. That is, any of the attention information obtained by the attention information generation unit 60 of the n-th (1 < n ≦ N) SWIN unit 431 has a plurality of components in each of the x direction and the y direction and includes spatial information. In the present embodiment, the size of the attention information obtained by the attention information generation unit 60 of the n-th SWIN unit 431 is the same as the feature amounts MDF(n), MCF(n), and the size of the intermediate feature amount in the n-th SWIN unit 431. Each value of the attention information means the weight of the feature related to each spatial position for each channel.

[0036] In the data combining unit 67, the above attention information and intermediate feature amounts are combined by aligning their positions in the spatial direction and stacking them in the channel direction. The combined data obtained by stacking is one-dimensionally convolutionally processed only in the channel direction in the convolutional unit 68, and each feature amount component is weighted while maintaining the size in the spatial direction. In this way, by performing the process of stacking once and then convolving, the influence of weighting (emphasis) by the attention information and the influence of the numerical distribution of the intermediate feature amounts are obtained in an appropriate balance, which is effective for extracting a structure characteristic of cracks.

[0037] In the conventional FPCNet, in the decoding unit 43, the Squeeze-Excitation Up-sampling (SEU) process is performed N times in series. The attention information generation unit in the SEU process does not include the SWIN transformer. The attention information generation unit obtains attention information depending only on the channel and does not include spatial information. As described above, in this embodiment, instead of the SEU process, the SWIN unit 431 is positioned, and attention information is obtained depending on both the channel and the spatial position. Thereby, in the information processing apparatus 1, the loss of spatial information in the generation of attention information is reduced, and the emphasis on abnormalities in the intermediate feature amounts is performed including spatial information.

[0038] The feature amounts convolved in the convolutional unit 68 are further added to the intermediate feature amounts in the second addition unit 69, thereby obtaining the feature amount MDF(n - 1). The processes of the attention information generation unit 60, the data combining unit 67, the convolutional unit 68, and the second addition unit 69 are included in the second process of this embodiment. The combination of the first process and the second process is included in the update process of this embodiment. The update process is executed a number of times (multiple times) corresponding to the number of SWIN units 431. The processes of the data combining unit 67, the convolutional unit 68, and the second addition unit 69 are included in the processes related to the synthesis of this embodiment. In this way, by more directly and strongly retaining the influence of the original intermediate feature amounts, the features of the image included in the intermediate feature amounts are maintained, facilitating the output of the abnormality estimation image by the decoding unit 43.

[0039] The first SWIN unit 431(1) located on the far downstream side of the decoding unit 43 further includes a convolutional unit (not shown) that performs one-dimensional convolution only in the channel direction on the subsequent stage of the second addition unit 69 to convert MDF(n), which is the output of the second addition unit 69 corresponding to each pixel, into a scalar value, and a sigmoid function that converts each of the scalar values into a value between 0.0 and 1.0. An abnormality estimation image is output from the SWIN unit 431(1). The control unit 11 may convert the abnormality estimation image into a grayscale image and output it, or perform noise removal, image enhancement, clarification, etc. by threshold processing such as cutting off pixels whose values in the abnormality estimation image are below the threshold and setting them to zero. Alternatively, as described above, the control unit 11 may output a list display, image data or text data in a predetermined format. In the above manner, an abnormality estimation image is obtained by selectively emphasizing the abnormal part corresponding to the abnormality, particularly cracks, among the feature amounts of the image, and extracting the range of the result estimated as abnormal.

[0040] FIG. 5 is a diagram showing an example of the estimation result of cracks by the abnormality estimation process using the machine learning model 121 of the present embodiment. FIG. 5(a) is a visible light captured image of a caisson in a port facility. FIG. 5(b) is a diagram showing the abnormality estimation result for this visible light captured image. FIG. 5(b) is an example of a binarized image in which the abnormality estimation image is threshold-processed, the pixel values at the positions where cracks are estimated to have occurred are set to 0, and the pixel values at the positions where cracks are not estimated to have occurred are set to 255.

[0041] In FIG. 5(a), it can be seen that the crack C extends mainly left and right while branching within the image. In FIG. 5(b), this crack C is detected and shown as a black area.

[0042] Note that the present invention is not limited to the above-described embodiment, and various modifications are possible. For example, in the above description, each SWIN unit 431 of the decoding unit 43 is assumed to include the SWIN transformer 64, but the transformer is not limited to this. A transformer that can be used for calculating attention information in an image, for example, a vision transformer, etc., may be included in each unit of the decoding unit 43.

[0043] Also, in the above description, an example where the attention information generation unit 60 includes two SWIN transformers 64 is shown, but it is not limited to this. Three or more SWIN transformers 64 may be included in the attention information generation unit 60.

[0044] Also, in the above description, the abnormal estimation image is converted into a binary image and the presence or absence of an abnormality is output in binary, but it is not limited to this. The abnormality may be classified into two or more types, and an output indicating the presence of an abnormality for each type of abnormality and an output indicating that there is no abnormality of any kind may be used to output three or more values. Also, the likelihood of an abnormality appearing in the abnormal estimation image may be represented by a predetermined number of steps. That is, the value representing the likelihood may be discrete values at predetermined intervals. For this purpose, the machine learning model 121 may include, for example, a discretization processing unit that discretizes the output of the decoding unit 43 based on a further predetermined threshold or reference value. The predetermined number of steps may be two steps. In this case, the conversion process from the above-described abnormal estimation image to a binary image, a ternary image, etc. is unnecessary. The discretization processing unit may be added to or made operable in the machine learning model 121 after the learning of the machine learning model 121 is completed.

[0045] Also, the method of synthesizing the attention information and the intermediate feature amount is not limited to the above-described embodiment. It may be synthesized by other methods. For example, instead of the combination and convolution by the data combination unit 67 and the convolution unit 68, the synthesis may be performed by adding the components of the attention information and the intermediate feature amount at the same coordinates. Also, the addition by the second addition unit 69 may be omitted.

[0046] Also, in the above description, it has been described that the number of convolutional units 411 in the encoding unit 41 is equal to the number of SWIN units 431 in the decoding unit 43, but they can also be made different. In this case, for example, when the number of convolutional units 411 is larger than the number of SWIN units 431, the feature amount MCF(n) output by some of the convolutional units 411 may not be input to the SWIN unit 431.

[0047] Also, in the above description, as an example of a computer-readable medium that stores the program 120 related to the control of the abnormality estimation of the present invention, the storage unit 12 composed of a non-volatile memory such as an HDD or a flash memory has been described, but it is not limited to these. As other computer-readable media, other non-volatile memories such as MRAM, and portable recording media such as CD-ROMs and DVD disks can be applied. Also, a carrier wave is also applied to the present invention as a medium for providing the data of the program according to the present invention via a communication line. In addition, the specific configurations, details of the processing operations, procedures, etc. shown in the above embodiments can be appropriately changed without departing from the gist of the present invention. The scope of the present invention includes the scope of the invention described in the claims and its equivalent scope.

[0048] As described above, the information processing apparatus 1 of the present embodiment includes a control unit 11. The control unit 11, as an encoding unit 41, performs convolution processing on a captured image represented by a set of a plurality of pixels a plurality of times to generate a feature amount MCF. The control unit 11, as an extended convolution unit 42, performs extended convolution processing on the feature amount MCF a plurality of times to generate a feature amount MDF. The control unit 11, as a decoding unit 43, performs a first process of generating an intermediate feature amount by adding the feature amount MCF to the result of performing transposed convolution processing on the feature amount MDF, and a second process of generating attention information from the intermediate feature amount and synthesizing the attention information with the intermediate feature amount to update the feature amount MDF. The control unit 11 executes the update process, which is a combination of the first process and the second process, a plurality of times to generate an abnormality estimation image indicating an estimation result of the likelihood that an abnormality appears in each pixel of the captured image. The first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction. In this way, in the information processing apparatus 1, when generating the attention information, the information in the two-dimensional direction on the image is left without being compressed, so that the information in the two-dimensional direction in the specific abnormal range can be used. Therefore, the information processing apparatus 1 can more accurately estimate abnormal locations. Further, the information processing apparatus 1 effectively fuses the information on the feature distribution of the original image included in the feature amount MCF obtained by convolution processing and the information emphasizing the abnormality included in the attention information by generating the intermediate feature amount and synthesizing the attention information and the intermediate feature amount. Also, by this, the information processing apparatus 1 can more accurately estimate the features of abnormal locations while taking advantage of the merits of each piece of information.

[0049] Further, the decoding unit 43 may perform synthesis by performing one-dimensional convolution processing that convolves only in the channel direction by the convolution unit 68 after combining the attention information and the intermediate feature amount in the channel direction by the data combining unit 67. Thereby, the emphasis on the abnormality by the attention information and the feature distribution of the original image included in the intermediate feature amount are synthesized in a well-balanced manner, and abnormal locations are estimated more accurately.

[0050] Further, the decoding unit 43 may perform synthesis by adding the intermediate feature amount again by the second addition unit 69 to the feature amount obtained by performing one-dimensional convolution processing by the convolution unit 68 after combining the attention information and the intermediate feature amount in the channel direction by the data combining unit 67. As a result, it is expected that the influence of the feature distribution in the two-dimensional direction of the image in the feature amount MDF will be maintained without being blurred. Thereby, the information processing apparatus 1 can estimate abnormal locations more accurately.

[0051] Also, the number of times of convolution processing in the encoding unit 41 and the number of times of update processing in the decoding unit 43 may be the same. The decoding unit 43 may use the feature amount MCF(n) generated by the encoding unit 41 in the n-th convolution processing (the n-th convolution unit 411) in the n-th update processing from the back (the n-th SWIN unit 431(n) from the back). In this way, by inputting the feature amount MCF(n) to the SWIN unit 431(n) that obtains intermediate feature amounts of the same size, the feature distribution of the original image is accurately restored. Therefore, the information processing apparatus 1 can estimate abnormal locations more accurately.

[0052] In addition, the abnormality estimation method of this embodiment includes the following three steps: (1) An encoding step of generating a feature amount MCF by performing convolution processing a plurality of times on a captured image represented by a set of a plurality of pixels. (2) An extended convolution step of generating a feature amount MDF by performing extended convolution processing a plurality of times on the feature amount MCF. (3) A first process of generating an intermediate feature amount by adding the feature amount MCF to the result of performing transposed convolution processing on the feature amount MDF, and a second process of generating attention information from the intermediate feature amount and synthesizing the attention information with the intermediate feature amount to update the feature amount MDF. A decoding step of generating an abnormality estimation image showing an estimation result of the likelihood of an abnormality appearing in each pixel of the captured image by executing the update process in a set of the first process and the second process a plurality of times. The first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction. According to such an abnormality estimation method, information about the spatial direction of the image is left in the attention information generated at the time of decoding. Therefore, the abnormal location is estimated with higher accuracy.

[0053] In addition, when a program 120 including the machine learning model 121 of this embodiment is installed and executed on a computer, the abnormal location can be easily and accurately estimated software-wise.

Explanation of Signs

[0054] 1 Information processing apparatus 11 Control unit 12 Storage unit 120 Program 121 Machine learning model 13 Input / output interface 131 Connection terminal 132 Communication unit 14 Display unit 15 Operation reception unit 21 Database apparatus 22 Optical reading apparatus 41 Encoding unit 411 Convolution unit 42 Extended convolution unit 43 Decoding unit 431 SWIN unit 60 Attention information generation unit 61 Transposed convolution processing unit 62 First addition unit 63 Conversion unit 64, 65 SWIN Transformer 66 Expansion unit 67 Data combination unit 68 Convolution unit 69 Second addition unit 201 Photographed image data C Crack

Claims

1. An encoding unit that generates a first feature amount by performing convolution processing a plurality of times on a captured image represented by a set of a plurality of pixels; An extended convolution unit that generates a second feature amount by performing extended convolution processing a plurality of times on the first feature amount; A first process of generating an intermediate feature amount by adding the first feature amount to the result of performing transposed convolution processing on the second feature amount, and a second process of generating attention information from the intermediate feature amount and synthesizing the attention information with the intermediate feature amount to update the second feature amount. A decoding unit that generates an abnormality estimation image indicating an estimation result of the likelihood that an abnormality appears in each pixel of the captured image by executing a plurality of times a combination of update processes; Comprising: The first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction. An abnormality estimation device.

2. The decoding unit performs the synthesis by performing one-dimensional convolution processing of convolving only in the channel direction after combining the attention information and the intermediate feature amount in the channel direction. The abnormality estimation device according to Claim 1.

3. The decoding unit performs the synthesis by adding the intermediate feature amount to the feature amount obtained by performing the one-dimensional convolution processing after combining the attention information and the intermediate feature amount in the channel direction. The abnormality estimation device according to Claim 2.

4. The number of times of the convolution processing in the encoding unit and the number of times of the update processing in the decoding unit are the same. The decoding unit uses the first feature amount generated by the encoding unit in the n-th convolution processing from the back in the n-th update processing from the back. The abnormality estimation device according to Claim 3.

5. An encoding step of generating a first feature amount by performing convolution processing a plurality of times on a captured image represented by a set of a plurality of pixels; An extended convolution step of generating a second feature amount by performing extended convolution processing a plurality of times on the first feature amount; A first process that generates an intermediate feature amount by adding the first feature amount to the result of performing a transposed convolution process on the second feature amount, and a second process that generates attention information from the intermediate feature amount and synthesizes the attention information with the intermediate feature amount to update the second feature amount. A decoding step that repeatedly executes an update process that is a combination of the above to generate an anomaly estimation image indicating an estimation result of the likelihood that an anomaly appears in each pixel of the captured image. including The first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction. Anomaly estimation method.

6. A computer Encoding means for generating a first feature amount by performing a plurality of convolution processes on a captured image represented by a set of a plurality of pixels. Expansion convolution means for generating a second feature amount by performing a plurality of expansion convolution processes on the first feature amount. A first process that generates an intermediate feature amount by adding the first feature amount to the result of performing a transposed convolution process on the second feature amount, and a second process that generates attention information from the intermediate feature amount and synthesizes the attention information with the intermediate feature amount to update the second feature amount. A decoding means that repeatedly executes an update process that is a combination of the above to generate an anomaly estimation image indicating an estimation result of the likelihood that an anomaly appears in each pixel of the captured image. function as The first feature amount, the second feature amount, the intermediate feature amount, and the attention information have a dimension in the channel direction of the captured image and a two-dimensional direction on the image different from the channel direction. Program.