Defect detection method, device and equipment based on event camera and storage medium
The event image data stream is obtained through the event camera and feature extraction and model input, which solves the defect detection accuracy problem of traditional cameras in complex lighting and fast moving target scenarios, achieving more efficient defect recognition.
Patent Information
- Application Number
- CN202510071355.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
In the prior art, traditional cameras are difficult to accurately capture defects in high-reflection and complex lighting scenarios, and cannot effectively handle fast moving targets, resulting in reduced defect detection accuracy.
The event camera is used to obtain the event image data stream, and the event feature map corresponding to the event stream data within the preset time is converted into a high-dimensional event feature map, and input it into the trained defect detection model to obtain the defect detection results.
It improves the accuracy of defect detection, can effectively identify irregular defect features, and enhances the detection capability in complex lighting and fast moving target scenarios.
Smart Images

Figure CN119991603A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of target detection technology, and in particular to a defect detection method, device, equipment and storage medium based on an event camera. Background Art
[0002] With the development and progress of industry, the manufacturing industry is also constantly upgrading its intelligence. However, in industrial manufacturing, defects in workpieces are inevitable. The defects of workpieces not only affect the appearance of the product, but also have a negative impact on its performance. Therefore, defect detection technology plays a vital role in improving production processes and improving product quality.
[0003] In the prior art, defect detection methods are mainly implemented by visual defect detection. Specifically, images are collected by traditional cameras, and defect detection is performed on the collected images using relevant software. However, due to the small dynamic range of traditional cameras, overexposure and other phenomena are prone to occur in complex lighting scenes such as high reflections. In addition, traditional cameras cannot effectively capture fast-moving targets and are prone to dynamic blur. When imaging, the foreground and background of the image are given equal weight, resulting in large information redundancy and reduced computing efficiency, which makes it impossible for traditional cameras to efficiently obtain images of good quality, and therefore it is impossible to accurately obtain defect detection results.
[0004] Based on this, event cameras have begun to attract the attention of researchers in defect detection tasks due to their fast response, high dynamic range, and ability to quickly capture moving targets. However, in the process of acquiring event image data streams using event cameras, the accuracy of defect detection is reduced due to the presence of a large amount of noise and interference. Therefore, how to effectively identify defects in the event image data stream acquired by event cameras is an urgent problem to be solved. Summary of the invention
[0005] Based on this, it is necessary to provide a defect detection method, device, equipment and storage medium based on an event camera to address the above technical problems.
[0006] In a first aspect, an embodiment of the present invention provides a defect detection method based on an event camera, the method comprising:
[0007] Obtain the event feature graph corresponding to the event stream data within a preset time period;
[0008] Converting the event feature graph into a corresponding high-dimensional event feature graph;
[0009] Inputting the high-dimensional event feature graph into a trained defect detection model to obtain a defect detection result;
[0010] Among them, the defect detection model is trained according to the sample training set, and the sample training set includes: sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graphs. The defect detection model includes: an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module and an output module.
[0011] In one embodiment, obtaining an event feature graph corresponding to event stream data within a preset time period includes:
[0012] Acquire multiple event data within a preset time period, wherein each event data is represented by e=(x, y, p, t), where x represents the horizontal coordinate of the event, y represents the vertical coordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event;
[0013] According to the abscissa, the ordinate and the polarity of the plurality of event data, the event feature graph corresponding to the event stream data within a preset time length is obtained.
[0014] In one embodiment, converting the event feature graph into a corresponding high-dimensional event feature graph includes:
[0015] Dividing the event feature graph into a plurality of event sub-feature graphs, and concatenating the plurality of event sub-feature graphs to obtain an initial high-dimensional event feature graph;
[0016] The initial high-dimensional event feature map is input into a cross-stage local network to obtain the high-dimensional event feature map.
[0017] In one embodiment, inputting the high-dimensional event feature graph into a trained defect detection model to obtain a defect detection result includes:
[0018] Inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction module through the input module to extract the target interpolation feature map corresponding to the high-dimensional event feature map;
[0019] Inputting the target interpolation feature map into the deformable pooling downsampling module to obtain a target downsampling feature map corresponding to the target interpolation feature map;
[0020] Inputting the target down-sampling feature map into the spatiotemporal feature extraction module to obtain a target attention feature map corresponding to the target down-sampling feature map;
[0021] The target attention feature map is input into the output module to obtain the defect detection result.
[0022] In one embodiment, the thin plate spline interpolation feature extraction module includes: a thin plate spline interpolation feature extraction submodule, a convolution feature extraction submodule and a fusion submodule, and the high-dimensional event feature map is input into the thin plate spline interpolation feature extraction module through the input module, and the target interpolation feature map corresponding to the high-dimensional event feature map is extracted, including:
[0023] Inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction submodule, and extracting a first initial interpolation feature map based on a thin plate spline interpolation function;
[0024] Inputting the high-dimensional event feature map into the convolution feature extraction submodule for convolution operation to extract a second initial interpolation feature map;
[0025] The first initial interpolation feature map and the second initial interpolation feature map are input into the fusion submodule for fusion operation to obtain the target interpolation feature map.
[0026] In one embodiment, the deformable pooling downsampling module includes: a deformable convolution submodule, a convolution submodule, an average pooling submodule and a weighted fusion module, and the target interpolation feature map is input into the deformable pooling downsampling module to obtain the target downsampling feature map corresponding to the target interpolation feature map, including:
[0027] Inputting the target interpolation feature map into the deformable convolution submodule for convolution operation to extract a first initial down-sampled feature map;
[0028] Inputting the target interpolation feature map into the convolution submodule for convolution operation to extract a second initial down-sampled feature map;
[0029] Inputting the target interpolation feature map into the average pooling submodule for pooling operation to extract a third initial down-sampled feature map;
[0030] The first initial down-sampling feature map, the second initial down-sampling feature map and the third initial down-sampling feature map are input into the weighted fusion module for weighted fusion operation to obtain the target down-sampling feature map.
[0031] In one embodiment, the spatiotemporal feature extraction module includes: a temporal feature extraction submodule and a spatial feature extraction submodule, and the step of inputting the target down-sampling feature map into the spatiotemporal feature extraction module to obtain a target attention feature map corresponding to the target down-sampling feature map includes:
[0032] Inputting the target downsampled feature map into the temporal feature extraction submodule to extract the temporal attention feature map;
[0033] The temporal attention feature map is input into the spatial feature extraction submodule to extract the target attention feature map.
[0034] In a second aspect, an embodiment of the present invention provides a defect detection device based on an event camera, comprising:
[0035] An event feature graph acquisition module is used to acquire an event feature graph corresponding to event stream data within a preset time period;
[0036] A high-dimensional event feature graph acquisition module, used to convert the event feature graph into a corresponding high-dimensional event feature graph;
[0037] A defect detection result acquisition module is used to input the high-dimensional event feature map into a trained defect detection model to obtain a defect detection result;
[0038] Among them, the defect detection model is trained according to the sample training set, and the sample training set includes: sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graphs. The defect detection model includes: an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module and an output module.
[0039] In a third aspect, an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of the event camera-based defect detection method described in the first aspect when executing the computer program.
[0040] In a fourth aspect, a computer-readable storage medium stores a computer program thereon, wherein when the computer program is executed by a processor, the steps of the event camera-based defect detection method described in the first aspect are implemented.
[0041] The technical solution provided by the embodiment of the present invention has the following advantages compared with the prior art:
[0042] An event camera-based defect detection method provided by an embodiment of the present invention adopts this method to obtain an event feature map corresponding to event stream data within a preset time length. The event feature map is converted into a corresponding high-dimensional event feature map. The high-dimensional event feature map is input into a trained defect detection model to obtain a defect detection result. Since the defect detection model is trained based on sample event feature maps corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature map, the defect detection model is composed of a thin plate spline interpolation feature extraction module that can better extract irregular defect features, a deformable pooling downsampling module that reduces the feature map size and retains key spatiotemporal features, and a spatiotemporal feature extraction module that better extracts time dimension features and global space features. Therefore, when performing defect detection, the accuracy of defect detection can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0044] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0045] Figure 1 A schematic flow chart of a defect detection method based on an event camera provided in an embodiment of the present invention;
[0046] Figure 2 A structural schematic diagram of a defect detection model provided by an embodiment of the present invention;
[0047] Figure 3 A schematic diagram of the structure of another defect detection model provided by an embodiment of the present invention;
[0048] Figure 4 A schematic structural diagram of a defect detection device based on an event camera provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0049] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.
[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all of the embodiments.
[0051] In the prior art, defect detection methods are mainly implemented by visual defect detection. Specifically, images are collected by traditional cameras, and defect detection is performed on the collected images using relevant software. However, due to the small dynamic range of traditional cameras, overexposure and other phenomena are prone to occur in complex lighting scenes such as high reflections. In addition, traditional cameras cannot effectively capture fast-moving targets and are prone to dynamic blur. When imaging, the foreground and background of the image are given equal weight, resulting in large information redundancy and reduced computing efficiency, which makes it impossible for traditional cameras to efficiently obtain images of good quality, and therefore it is impossible to accurately obtain defect detection results.
[0052] Based on this, event cameras have begun to attract the attention of researchers in defect detection tasks due to their fast response, high dynamic range, and ability to quickly capture moving targets. However, in the process of acquiring event image data streams using event cameras, the accuracy of defect detection is reduced due to the presence of a large amount of noise and interference. Therefore, how to effectively identify defects in the event image data stream acquired by event cameras is an urgent problem to be solved.
[0053] Therefore, the present invention provides a defect detection method based on an event camera, by obtaining an event feature graph corresponding to event stream data within a preset time length. The event feature graph is converted into a corresponding high-dimensional event feature graph. The high-dimensional event feature graph is input into a trained defect detection model to obtain a defect detection result. Since the defect detection model is trained based on sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graph, the defect detection model is composed of a thin plate spline interpolation feature extraction module that can better extract irregular defect features, a deformable pooling downsampling module that reduces the feature graph size and retains key spatiotemporal features, and a spatiotemporal feature extraction module that better extracts time dimension features and global space features. Therefore, when performing defect detection, the accuracy of defect detection can be improved.
[0054] In one embodiment, Figure 1 As shown, Figure 1 A schematic flow chart of a defect detection method based on an event camera provided by an embodiment of the present invention specifically includes the following steps:
[0055] S10: Obtain an event feature graph corresponding to the event stream data within a preset time length.
[0056] The preset duration refers to the duration set for collecting the corresponding event data stream when acquiring the event feature graph through the event camera, such as 20ms, that is, collecting 20ms of event data stream through the event camera and converting the event data stream into a corresponding one-frame event feature graph. However, the present invention is not limited to this, and those skilled in the art can set it according to actual conditions.
[0057] Specifically, a preset duration is pre-set for the event camera, and when collecting the event data stream, the event data stream of the preset duration is collected, and the event data stream is converted into a corresponding one-frame event feature map.
[0058] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S10 may be:
[0059] S101: Acquire multiple event data within a preset time period.
[0060] Each event data is represented by e=(x, y, p, t), where x represents the horizontal coordinate of the event, y represents the vertical coordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event.
[0061] Multiple event data refers to multiple events included in a preset time length. Following the above embodiment, the preset time length is set to 20ms, and then 20ms contains multiple events, that is, it can be understood that event data of multiple events can be collected within 20ms. However, it is not limited to this, and the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0062] It should be noted that the polarity of the event reflects the trend of light intensity change at the current moment. When the light intensity increases, the event polarity is determined to be 1, indicating that the event is an ON event (i.e., a positive polarity event), and the polarity corresponding to the event is taken as 1. When the light intensity decreases, the event polarity is determined to be -1, indicating that the event is an OFF event (i.e., a negative polarity event). It can be understood that when there is a change in light intensity and exceeds a certain threshold, that is, an event occurs, the pixel value of the event is determined to be 1. Based on this, since the event camera only responds to changes in light intensity, compared with traditional cameras, it does not require energy integration process and AD conversion process, and has the characteristics of fast imaging speed and large dynamic range.
[0063] S102: Obtain an event feature graph corresponding to the event stream data within a preset time length according to the horizontal coordinates, vertical coordinates and polarities of the plurality of event data.
[0064] Specifically, multiple event data within a preset time length are collected by an event camera, and further, an event feature graph corresponding to the event stream data within the preset time length is obtained according to the horizontal coordinates, vertical coordinates and polarities of the collected multiple event data.
[0065] Exemplarily, following the above embodiment, the preset duration is set to 20ms, and multiple event data are collected within 20ms. Each event data has a corresponding horizontal coordinate and vertical coordinate of each pixel. Based on this, when the light intensity changes, an event occurs. According to the horizontal coordinate and vertical coordinate of the event, the pixel value of the polarity of the event is filled into the pixel matrix to obtain the event feature map. However, it is not limited to this, and the present invention is not specifically limited. Those skilled in the art can set it according to actual conditions.
[0066] In this way, this embodiment can obtain multiple event data within a preset time length, and obtain an event feature graph according to the horizontal coordinate, vertical coordinate, polarity of the event, and timestamp of the event respectively contained in the multiple event data within the preset time length, rather than obtaining the event feature graph based on the event data at a single moment, thereby improving the quality of obtaining the event feature graph.
[0067] S11: Convert the event feature map into a corresponding high-dimensional event feature map.
[0068] Specifically, after obtaining the event feature graph corresponding to the event stream data within a preset time length, the event feature graph is converted into a corresponding high-dimensional event feature graph.
[0069] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S11 may be:
[0070] S111: Divide the event feature graph into multiple event sub-feature graphs, and concatenate the multiple event sub-feature graphs to obtain an initial high-dimensional event feature graph.
[0071] Among them, multiple event sub-feature maps refer to multiple event sub-feature maps obtained by slicing the event feature map according to the image pixel position according to a preset slicing module such as the Focus module. Each pixel contained in the multiple event sub-feature maps comes from a different position of the event feature map, thereby forming a compression of spatial information. In this way, the spatial resolution of the event feature map is reduced, thereby reducing the complexity of the subsequent calculation process.
[0072] Specifically, the acquired event feature graph is sliced and divided to obtain a plurality of event sub-feature graphs, and the plurality of event sub-feature graphs are spliced to obtain an initial high-dimensional event feature graph.
[0073] Exemplarily, the input event feature map is segmented according to the pixel position by the Focus module, and divided into four event sub-feature maps, which are the upper left corner, upper right corner, lower left corner, and lower right corner of the event feature map, respectively, and the four event sub-feature maps are spliced to obtain an initial high-dimensional event feature map. However, this is not limited to this, and the present invention is not specifically limited, and those skilled in the art can set it according to actual conditions.
[0074] S112: Input the initial high-dimensional event feature map into the cross-stage local network to obtain a high-dimensional event feature map.
[0075] Among them, the cross-stage local network is a network that can obtain richer gradient fusion information while reducing the amount of calculation. It mainly divides the initial high-dimensional event feature map into two parts, and then fuses these two parts through a cross-stage layer. In this way, by separating the gradient flow, the gradient flow can be propagated on different network paths, which can reduce the amount of calculation and improve the reasoning speed and accuracy.
[0076] Specifically, the initial high-dimensional event feature map is input into the cross-stage local network, and the cross-stage local network is used to further extract and integrate features of the initial high-dimensional event feature map to obtain the high-dimensional event feature map.
[0077] In this way, this embodiment can convert the event feature map into a high-dimensional event feature map, thereby reducing the spatial resolution of the event feature map, thereby reducing the complexity of subsequent calculation processes and improving the reasoning speed and accuracy.
[0078] S12: Input the high-dimensional event feature graph into the trained defect detection model to obtain the defect detection result.
[0079] Among them, the defect detection model is trained based on the sample training set, and the sample training set includes: sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graphs.
[0080] The preset category defects refer to defects of different categories in the sample event feature map, such as spots, scratches, and stains, but are not limited thereto. The present invention is not specifically limited thereto, and those skilled in the art can set them according to actual conditions. The label data refers to the location information and category information of each preset category defect in the sample event feature map.
[0081] For further reference, Figure 2 As shown, the defect detection model includes: an input module 20, a thin plate spline interpolation feature extraction module 21, a deformable pooling downsampling module 22, a spatiotemporal feature extraction module 23 and an output module 24.
[0082] Among them, the thin plate spline interpolation feature extraction module 21 is based on the thin plate spline interpolation algorithm, which can adapt to various data types and irregular shapes, has high accuracy and stability, and achieves accurate fitting of data by minimizing the bending ability of the difference function, effectively handles irregular data distribution, and suppresses noise interference. Based on this, the present invention uses the thin plate spline interpolation feature extraction module based on the thin plate spline interpolation algorithm to extract feature data of high-dimensional event feature maps, which can better obtain irregular defect features, thereby improving the accuracy of defect detection.
[0083] The deformable pooling downsampling module 22 is used to perform sampling processing on the data of the feature map.
[0084] The spatiotemporal feature extraction module 23 is used to obtain the features of the feature graph in the time dimension and the global spatial dimension, so as to capture the temporal correlation of defect features and better obtain feature information in the global space, thereby improving the accuracy of defect detection.
[0085] Specifically, the high-dimensional event feature graph is input into a defect detection model trained according to a sample training set including sample event feature graphs corresponding to sample event stream data within multiple preset time lengths and label data of each preset category of defects contained in the sample event feature graphs, to obtain the defect detection result output by the defect detection model, wherein the defect detection model includes an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module and an output module.
[0086] In this way, the defect detection method based on the event camera provided in this embodiment obtains the event feature graph corresponding to the event stream data within a preset time length. The event feature graph is converted into the corresponding high-dimensional event feature graph. The high-dimensional event feature graph is input into the trained defect detection model to obtain the defect detection result. Since the defect detection model is trained based on the sample event feature graphs corresponding to the sample event stream data within multiple preset time lengths, and the label data of each preset category defect contained in the sample event feature graph, the defect detection model is composed of a thin plate spline interpolation feature extraction module that can better extract irregular defect features, a deformable pooling downsampling module that reduces the feature map size and retains key spatiotemporal features, and a spatiotemporal feature extraction module that better extracts time dimension features and global space features. Therefore, when performing defect detection, the accuracy of defect detection can be improved.
[0087] Optionally, based on the above embodiment, in some embodiments of the present invention, an implementation manner of S12 may be:
[0088] S121: Inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction module through the input module, and extracting the target interpolation feature map corresponding to the high-dimensional event feature map.
[0089] Optionally, based on the above embodiments, in some embodiments of the present invention, reference is made to Figure 3 As shown, the thin plate spline interpolation feature extraction module 21 includes a thin plate spline interpolation feature extraction submodule 211, a convolution feature extraction submodule 212 and a fusion submodule 213. An implementation of S121 may be:
[0090] S1211: Input the high-dimensional event feature map to the thin plate spline interpolation feature extraction submodule, and extract a first initial interpolation feature map based on the thin plate spline interpolation function.
[0091] Among them, the thin plate spline interpolation function is used to interpolate the high-dimensional event feature map.
[0092] Optionally, in this embodiment, the thin plate spline interpolation function may be defined by the following expression:
[0093]
[0094] Among them, r represents the Euclidean distance between the input feature map and the control point, and ε represents a regularization term to avoid the value of the logarithmic function being 0.
[0095] Specifically, the obtained high-dimensional event feature map is input into the thin plate spline interpolation feature extraction submodule, and the high-dimensional event feature map is interpolated by the thin plate spline interpolation function in the thin plate spline interpolation feature extraction submodule to obtain an interpolation base feature map, and further feature extraction is performed on the interpolation base feature map to obtain a first initial interpolation feature map.
[0096] S1212: Input the high-dimensional event feature map into the convolution feature extraction submodule for convolution operation to extract a second initial interpolation feature map.
[0097] Specifically, the obtained high-dimensional event feature map is input into a convolution feature extraction submodule, and a convolution operation is performed on the high-dimensional event feature map by the convolution feature extraction submodule to extract a second initial interpolation feature map.
[0098] S1213: Input the first initial interpolation feature map and the second initial interpolation feature map into a fusion submodule for fusion operation to obtain a target interpolation feature map.
[0099] Specifically, the first initial interpolation feature map and the second initial interpolation feature map are input into a fusion submodule, and the fusion submodule performs a fusion operation on the first initial interpolation feature map and the second initial interpolation feature map to obtain a target interpolation feature map.
[0100] In this way, the present invention utilizes a thin plate spline interpolation feature extraction module based on a thin plate spline interpolation algorithm to extract feature data of a high-dimensional event feature graph, and can better obtain irregular defect features, thereby improving the accuracy of defect detection results.
[0101] S122: Input the target interpolation feature map into the deformable pooling downsampling module to obtain a target downsampling feature map corresponding to the target interpolation feature map.
[0102] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3 As shown, the deformable pooling downsampling module 22 includes: a deformable convolution submodule 221, a convolution submodule 222, an average pooling submodule 223 and a weighted fusion module 224. An implementation of S122 may be:
[0103] S1221: Input the target interpolation feature map into the deformable convolution submodule for convolution operation to extract the first initial down-sampling feature map.
[0104] Among them, deformable convolution is a sampling convolution for feature learning and extraction of irregular features, which can achieve implicit alignment between frames and retain important inter-frame temporal feature information while downsampling. Based on this, the deformable convolution submodule can retain key spatiotemporal features during the convolution downsampling operation of the target interpolation feature map, and better obtain the defect features with irregular shapes.
[0105] Specifically, the obtained target interpolation feature map is input into a deformable convolution submodule, and a deformable convolution downsampling operation is performed on the target interpolation feature map through the deformable convolution submodule to obtain a first initial downsampling feature map.
[0106] S1222: Input the target interpolation feature map into the convolution submodule for convolution operation to extract the second initial down-sampling feature map.
[0107] Specifically, the obtained target interpolation feature map is input into a convolution submodule, and a convolution downsampling operation is performed on the target interpolation feature map through the convolution submodule to obtain a second initial downsampling feature map.
[0108] S1223: Input the target interpolation feature map into the average pooling submodule for pooling operation to extract the third initial down-sampled feature map.
[0109] Specifically, the obtained target interpolation feature map is input into the average pooling submodule, and the target interpolation feature map is pooled by the average pooling submodule to obtain the third initial down-sampling feature map.
[0110] S1224: Input the first initial down-sampling feature map, the second initial down-sampling feature map, and the third initial down-sampling feature map into a weighted fusion module for performing a weighted fusion operation to obtain a target down-sampling feature map.
[0111] Specifically, the first initial down-sampling feature map, the second initial down-sampling feature map and the third initial down-sampling feature map are input into the weighted fusion module, and the weighted fusion operation is performed on the first initial down-sampling feature map, the second initial down-sampling feature map and the third initial down-sampling feature map through the weighted fusion sub-module to obtain the target down-sampling feature map.
[0112] Optionally, based on the above embodiment, in some embodiments of the present invention, the first initial down-sampling feature map, the second initial down-sampling feature map and the third initial down-sampling feature map are input into a weighted fusion module for weighted fusion operation, and the target down-sampling feature map is obtained, which can be defined by the following expression:
[0113] X=FuseConv(DeConv(x)*W1+Conv(x)*W2+AvgConv(x)*W3)
[0114] Among them, FuseConv represents a 1x1 fused convolution, DeConv(x) represents the first initial down-sampling feature map obtained by the deformable convolution submodule, W1 represents the weight corresponding to the first initial down-sampling feature map, Conv(x) represents the second initial down-sampling feature map obtained by the convolution submodule, W2 represents the weight corresponding to the second initial down-sampling feature map, AvgConv(x) represents the third initial down-sampling feature map obtained by the average pooling submodule, and W3 represents the weight corresponding to the third initial down-sampling feature map.
[0115] It should be noted that one implementation method for determining the weight values corresponding to W1, W2, and W3 may be to initialize the weight values of W1, W2, and W3 with random values, wherein the weight values corresponding to W1, W2, and W3 are added to a preset value, which is, for example, 3. During the model training process, the weight values corresponding to W1, W2, and W3 are continuously learned and adjusted until the model converges to determine the target weight values corresponding to W1, W2, and W3.
[0116] S123: Input the target down-sampling feature map into the spatiotemporal feature extraction module to obtain the target attention feature map corresponding to the target down-sampling feature map.
[0117] Optionally, based on the above embodiments, in some embodiments of the present invention, continue to refer to Figure 3As shown, the spatiotemporal feature extraction module 23 includes: a temporal feature extraction submodule 231 and a spatial feature extraction submodule 232. An implementation of S123 may be:
[0118] S1231: Input the target downsampled feature map into the temporal feature extraction submodule to extract the temporal attention feature map.
[0119] Among them, the time feature extraction submodule is used to obtain the information of the feature map in the time dimension to ensure that the time continuity information of the target downsampled feature map can be extracted.
[0120] Specifically, the target down-sampling feature map is input into the time feature extraction submodule, and the time information feature is extracted from the target down-sampling feature map through the time feature extraction submodule to obtain the time attention feature map.
[0121] S1232: Input the temporal attention feature map into the spatial feature extraction submodule to extract the target attention feature map.
[0122] Among them, the spatial feature extraction submodule is used to obtain the information of the feature map in the spatial dimension to ensure that the spatial feature information of the feature map can be extracted.
[0123] Specifically, the temporal attention feature map is input into the spatial feature extraction submodule, and the spatial information feature is extracted from the temporal attention feature map through the spatial feature extraction submodule to obtain the target attention feature map.
[0124] It should be noted that since the complexity of obtaining spatial feature information is lower than that of obtaining temporal feature information, when the temporal attention feature map is output to the spatial feature extraction submodule through the fully connected layer for the first time, the relevant parameters of the fully connected layer are set to 0 to ensure that the spatial information of the feature map is obtained first, thereby improving the model convergence efficiency.
[0125] S124: Input the target attention feature map into the output module to obtain the defect detection result.
[0126] Specifically, the target attention feature map is input into the output module, and the defect detection result is output through the output module.
[0127] In this way, the defect detection method based on event camera provided by the present invention, by inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction module through the input module, extracting the target interpolation feature map corresponding to the high-dimensional event feature map; inputting the target interpolation feature map into the deformable pooling downsampling module, obtaining the target downsampling feature map corresponding to the target interpolation feature map; inputting the target downsampling feature map into the spatiotemporal feature extraction module, obtaining the target attention feature map corresponding to the target downsampling feature map; inputting the target attention feature map into the output module, obtaining the defect detection result. Since the thin plate spline interpolation feature extraction module can better extract irregular defect features, the deformable pooling downsampling module can reduce the feature map size and retain key spatiotemporal features, and the spatiotemporal feature extraction module can better extract time dimension features and global space features, therefore, when performing defect detection, the accuracy of defect detection can be improved.
[0128] Optionally, based on the above embodiments, in some embodiments of the present invention, in order to verify that the defect detection method based on the event camera of the present invention can improve the accuracy of obtaining defect detection results, single-frame target detection methods such as Faster-RCNN, YoloV5, YoloV7, and video target detection methods such as RDN, MEGA, and YOLOV are selected for simulation experiments.
[0129] The following table shows the simulation experiment comparison results of the present invention with multiple single-frame target detection methods and multiple video target detection methods.
[0130] method mAP@0.4 AP@Point Mark AP@Scratch AP@Stain Faster-RCNN 0.210 0.000 0.536 0.095 YoloV5 0.569 0.393 0.756 0.559 YoloV7 0.543 0.471 0.644 0.514 RDN 0.512 0.553 0.476 0.507 MEGA 0.401 0.356 0.509 0.349 YOLOV 0.537 0.112 0.628 0.670 TIFF-EDD 0.617 0.512 0.701 0.639 The present invention 0.668 0.521 0.830 0.651
[0131] Among them, the average precision (AP) in the above table is used to measure the detection accuracy of the defect detection model for defects, which is obtained by calculating the average value of the precision under different recall rates. The above mean average precision (mAP) is used to evaluate the performance of the model on multiple categories. The AP of each category is calculated, and then the average value of multiple APs is taken to obtain a comprehensive indicator to evaluate the average performance of the model on all categories. The "@0.4" in mAP@0.4 refers to a specific IoU (Intersection over Union) threshold. IoU is a metric in object detection that measures the degree of overlap between the predicted bounding box and the true bounding box. By setting a preset IoU threshold, it is determined whether the predicted box is a true positive. Specifically, when the IoU value of the predicted box and the true box is higher than the preset IoU threshold, the prediction is considered to be correct. The preset IoU threshold I can be 0.4, 0.5, or 0.7. However, it is not limited to this, and the present invention is not specifically limited. People in this field can set it according to actual conditions.
[0132] Based on this, it can be seen from the above table that compared with other multiple single-frame target detection methods and multiple video target detection methods, the present invention has higher improvements in mAP@0.4, AP@spots, AP@scratches and AP@stains.
[0133] It should be understood that although Figures 1 to 3 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figures 1 to 3 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0134] In one embodiment, Figure 4 As shown, a defect detection device based on an event camera is provided, comprising: an event feature map acquisition module 10, a high-dimensional event feature map acquisition module 11, and a defect detection result acquisition module 12.
[0135] Among them, the event feature graph acquisition module 10 is used to obtain the event feature graph corresponding to the event stream data within a preset time length.
[0136] The high-dimensional event feature graph acquisition module 11 is used to convert the event feature graph into a corresponding high-dimensional event feature graph.
[0137] The defect detection result acquisition module 12 is used to input the high-dimensional event feature map into the trained defect detection model to obtain the defect detection result. The defect detection model is trained based on a sample training set, and the sample training set includes: sample event feature maps corresponding to sample event stream data within a plurality of preset time lengths, and label data of each preset category defect contained in the sample event feature map. The defect detection model includes: an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module, and an output module.
[0138] In the above embodiment, the event feature graph corresponding to the event stream data within a preset time length is obtained through the event feature graph acquisition module. The event feature graph is converted into a corresponding high-dimensional event feature graph through the high-dimensional event feature graph acquisition module. The high-dimensional event feature graph is input into the trained defect detection model through the defect detection result acquisition module to obtain the defect detection result. Since the defect detection model is trained based on the sample event feature graphs corresponding to the sample event stream data within multiple preset time lengths, and the label data of each preset category defect contained in the sample event feature graph, the defect detection model is composed of a thin plate spline interpolation feature extraction module that can better extract irregular defect features, a deformable pooling downsampling module that reduces the feature graph size and retains key spatiotemporal features, and a spatiotemporal feature extraction module that better extracts time dimension features and global space features. Therefore, when performing defect detection, the accuracy of defect detection can be improved.
[0139] For the specific definition of the defect detection device based on the event camera, please refer to the definition of the defect detection method based on the event camera above, which will not be repeated here. Each module in the above-mentioned server can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0140] An embodiment of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the defect detection method based on the event camera provided in the embodiment of the present invention can be implemented. For example, when the processor executes the computer program, the defect detection method based on the event camera provided in the embodiment of the present invention can be implemented. Figures 1 to 3 The technical solution of any of the method embodiments shown has similar implementation principles and technical effects, which will not be repeated here.
[0141] The embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. The computer program is executed by a processor to implement the defect detection method based on an event camera provided by the embodiment of the present invention. For example, when the computer program is executed by the processor, the computer program can be implemented. Figures 1 to 3 The technical solution of any of the method embodiments shown has similar implementation principles and technical effects, which will not be repeated here.
[0142] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static random access memory (SRAM) and dynamic random access memory (DRAM).
[0143] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0144] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.
Claims
1. A defect detection method based on an event camera, characterized in that: include: Obtain the event feature graph corresponding to the event stream data within a preset time period; Converting the event feature graph into a corresponding high-dimensional event feature graph; Inputting the high-dimensional event feature graph into a trained defect detection model to obtain a defect detection result; Among them, the defect detection model is trained according to the sample training set, and the sample training set includes: sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graphs. The defect detection model includes: an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module and an output module.
2. The method according to claim 1, characterized in that The obtaining of the event feature graph corresponding to the event stream data within a preset time period includes: Acquire multiple event data within a preset time period, wherein each event data is represented by e=(x, y, p, t), where x represents the horizontal coordinate of the event, y represents the vertical coordinate of the event, p represents the polarity of the event, and t represents the timestamp of the event; According to the abscissa, the ordinate and the polarity of the plurality of event data, the event feature graph corresponding to the event stream data within a preset time length is obtained.
3. The method according to claim 1, characterized in that The converting the event feature graph into a corresponding high-dimensional event feature graph comprises: Dividing the event feature graph into a plurality of event sub-feature graphs, and concatenating the plurality of event sub-feature graphs to obtain an initial high-dimensional event feature graph; The initial high-dimensional event feature map is input into a cross-stage local network to obtain the high-dimensional event feature map.
4. The method according to claim 1, characterized in that Inputting the high-dimensional event feature graph into the trained defect detection model to obtain defect detection results includes: Inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction module through the input module to extract the target interpolation feature map corresponding to the high-dimensional event feature map; Inputting the target interpolation feature map into the deformable pooling downsampling module to obtain a target downsampling feature map corresponding to the target interpolation feature map; Inputting the target down-sampling feature map into the spatiotemporal feature extraction module to obtain a target attention feature map corresponding to the target down-sampling feature map; The target attention feature map is input into the output module to obtain the defect detection result.
5. The method according to claim 4, characterized in that The thin plate spline interpolation feature extraction module includes: a thin plate spline interpolation feature extraction submodule, a convolution feature extraction submodule and a fusion submodule. The high-dimensional event feature map is input into the thin plate spline interpolation feature extraction module through the input module, and a target interpolation feature map corresponding to the high-dimensional event feature map is extracted, including: Inputting the high-dimensional event feature map into the thin plate spline interpolation feature extraction submodule, and extracting a first initial interpolation feature map based on a thin plate spline interpolation function; Inputting the high-dimensional event feature map into the convolution feature extraction submodule for convolution operation to extract a second initial interpolation feature map; The first initial interpolation feature map and the second initial interpolation feature map are input into the fusion submodule for fusion operation to obtain the target interpolation feature map.
6. The method according to claim 4, characterized in that The deformable pooling downsampling module includes: a deformable convolution submodule, a convolution submodule, an average pooling submodule and a weighted fusion module. The target interpolation feature map is input into the deformable pooling downsampling module to obtain a target downsampling feature map corresponding to the target interpolation feature map, including: Inputting the target interpolation feature map into the deformable convolution submodule for convolution operation to extract a first initial down-sampled feature map; Inputting the target interpolation feature map into the convolution submodule for convolution operation to extract a second initial down-sampled feature map; Inputting the target interpolation feature map into the average pooling submodule for pooling operation to extract a third initial down-sampled feature map; The first initial down-sampling feature map, the second initial down-sampling feature map and the third initial down-sampling feature map are input into the weighted fusion module for weighted fusion operation to obtain the target down-sampling feature map.
7. The method according to claim 4, characterized in that The spatiotemporal feature extraction module includes: a temporal feature extraction submodule and a spatial feature extraction submodule. The step of inputting the target down-sampling feature map into the spatiotemporal feature extraction module to obtain a target attention feature map corresponding to the target down-sampling feature map includes: Inputting the target downsampled feature map into the temporal feature extraction submodule to extract the temporal attention feature map; The temporal attention feature map is input into the spatial feature extraction submodule to extract the target attention feature map.
8. A defect detection device based on an event camera, characterized in that: include: An event feature graph acquisition module is used to acquire an event feature graph corresponding to event stream data within a preset time period; A high-dimensional event feature graph acquisition module, used to convert the event feature graph into a corresponding high-dimensional event feature graph; A defect detection result acquisition module is used to input the high-dimensional event feature map into a trained defect detection model to obtain a defect detection result; Among them, the defect detection model is trained according to the sample training set, and the sample training set includes: sample event feature graphs corresponding to sample event stream data within multiple preset time lengths, and label data of each preset category of defects contained in the sample event feature graphs. The defect detection model includes: an input module, a thin plate spline interpolation feature extraction module, a deformable pooling downsampling module, a spatiotemporal feature extraction module and an output module.
9. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the event camera-based defect detection method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the event camera-based defect detection method according to any one of claims 1 to 7 are implemented.