Defect detection method based on industrial endoscope, terminal device and storage medium
By using binocular cameras and target detection networks in industrial endoscopes, combined with light source enhanced illumination, and generating three-dimensional point cloud data, the efficiency and accuracy issues of defect detection in low-light environments are solved, achieving efficient defect detection.
Patent Information
- Application Number
- CN202411969211.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2044-12-30
AI Technical Summary
When an industrial endoscope is used in a low-light environment, the user's attention is distracted and the monocular camera's measurement algorithm has low accuracy, resulting in low defect detection efficiency.
An industrial endoscope with multiple detection modules, equipped with a binocular camera and a light source, performs defect detection in combination with a target detection network, including the backbone structure, text encoder, neck structure, and head structure. The target detection frame is generated by fusion of image and text features, and brightening operations are performed under enhanced lighting to construct three-dimensional point cloud data.
It improves the accuracy and efficiency of defect detection, reduces the workload of operators, is suitable for deployment in independent environments, and improves the convenience and safety of detection.
Smart Images

Figure CN119399185B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application belong to the technical field of computer vision, and in particular relate to a defect detection method, terminal device, and storage medium based on an industrial endoscope. Background Art
[0002] The industrial endoscope is a commonly used visual inspection instrument. Without disassembling or destroying the industrial equipment, the industrial endoscope is inserted into the interior of the industrial equipment to perform non-destructive testing on the industrial equipment.
[0003] Currently, when users use industrial endoscopes, they usually freeze the image when the monocular camera observes defects inside the industrial equipment, and use the monocular camera's internal and external parameters and other information to measure the geometric information of the defect.
[0004] However, the internal environment of industrial equipment has low brightness. When users operate the industrial endoscope while observing the image, their attention is easily distracted, resulting in the omission of certain defects. Moreover, the measurement algorithm of the monocular camera has low accuracy, resulting in low efficiency in detecting defects in industrial equipment. Summary of the Invention
[0005] In view of this, an embodiment of the present application provides a defect detection method, terminal device and storage medium based on an industrial endoscope, so as to improve the efficiency of defect detection of industrial equipment using an industrial endoscope.
[0006] A first aspect of an embodiment of the present application provides a defect detection method based on an industrial endoscope, wherein the industrial endoscope has multiple detection modules and a target detection network; each detection module has a binocular camera and a light source, and the target detection network includes a backbone structure, a text encoder, a neck structure, and a head structure; the method includes:
[0007] Determining target detection information; the target detection information includes target text information indicating a defect, and raw image data collected by the binocular camera inside the industrial equipment under the illumination of the light source;
[0008] Inputting the original image data into the backbone structure to extract original image features at multiple scales;
[0009] Inputting the target text information into the text encoder to extract original text features at multiple scales;
[0010] Inputting a plurality of the original image features and a plurality of the original text features into the neck structure for interaction to generate a plurality of image-text fusion features;
[0011] Inputting the plurality of image-text fusion features into the header structure to detect a target detection frame representing the defect; the target detection frame has a confidence level;
[0012] If the confidence level is greater than or equal to a preset threshold, a brightening operation is performed on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data is image data collected by the binocular camera inside the industrial equipment under the condition of enhanced illumination by the light source;
[0013] Constructing three-dimensional point cloud data based on the target detection frame in the target image data;
[0014] Geometric information is detected for the defect in the point cloud data.
[0015] A second aspect of an embodiment of the present application provides a defect detection device based on an industrial endoscope, wherein the industrial endoscope has multiple detection modules and a target detection network; each detection module has a binocular camera and a light source; the target detection network includes a backbone structure, a text encoder, a neck structure, and a head structure; the device includes:
[0016] a target detection information determination module, configured to determine target detection information; the target detection information includes target text information indicating a defect, and raw image data captured by the binocular camera inside the industrial equipment under the illumination of the light source;
[0017] An original image feature extraction module, configured to input the original image data into the backbone structure to extract original image features of multiple scales;
[0018] An original text feature extraction module, configured to input the target text information into the text encoder to extract original text features of multiple scales;
[0019] An image-text fusion feature generation module, configured to input a plurality of the original image features and a plurality of the original text features into the neck structure for interaction to generate a plurality of image-text fusion features;
[0020] a target detection frame generation module, configured to input the plurality of image-text fusion features into the header structure to detect a target detection frame representing the defect; the target detection frame has a confidence level;
[0021] a target image data generation module configured to, if the confidence level is greater than or equal to a preset threshold, perform a brightening operation on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data being image data captured by the binocular camera inside the industrial equipment under the condition of enhanced illumination by the light source;
[0022] A point cloud data construction module, configured to construct three-dimensional point cloud data based on the target detection frame in the target image data;
[0023] A geometric information detection module is used to detect geometric information of the defect in the point cloud data.
[0024] The third aspect of an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the computer program, it implements the defect detection method based on the industrial endoscope as described in the first aspect above.
[0025] A fourth aspect of an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the defect detection method based on the industrial endoscope as described in the first aspect above.
[0026] A fifth aspect of the embodiments of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the industrial endoscope-based defect detection method described in the first aspect.
[0027] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0028] In an embodiment of the present application, target detection information is determined; the target detection information includes target text information representing a defect, and original image data collected by a binocular camera inside an industrial device under conditions of light source illumination; the original image data is input into a backbone structure to extract original image features of multiple scales; the target text information is input into a text encoder to extract original text features of multiple scales; multiple original image features and multiple original text features are input into a neck structure for interaction to generate multiple image-text fusion features; the multiple image-text fusion features are input into a head structure to detect a target detection frame representing a defect; the target detection frame has a confidence level; if the confidence level is greater than or equal to a preset threshold, a brightening operation is performed on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data is image data collected by a binocular camera inside an industrial device under conditions of enhanced light source illumination; three-dimensional point cloud data is constructed based on the target detection frame in the target image data; and geometric information of defect detection is detected in the point cloud data. The embodiment of the present application has a built-in target detection network in the industrial endoscope, detects defects in the original image data based on open target text information, and realizes adaptive defect tracking, allowing users to focus on operating the industrial endoscope and effectively prevent missing defects. In addition, the information collected by the binocular camera is richer, and three-dimensional modeling based on the binocular camera can improve the accuracy of the geometric information of the detected defects, thereby improving the efficiency of defect detection in industrial equipment.
[0029] In addition, the target detection network has a clear structure, the detection process is a single forward transmission, it occupies few resources, and has a fast detection speed. It is suitable for deployment in industrial endoscopes to achieve local detection, and does not rely on cloud servers or other high-computing power equipment, thereby improving the convenience and safety of industrial endoscopes in independent environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the embodiments or descriptions of the prior art. Obviously, the drawings described below are only some embodiments of the present application, and those skilled in the art can derive other drawings based on these drawings without inventive effort.
[0031] Figure 1 is a schematic diagram of a defect detection method based on an industrial endoscope provided in an embodiment of the present application;
[0032] Figure 2 This is a schematic structural diagram of an industrial endoscope provided in an embodiment of the present application;
[0033] Figure 3 This is a schematic structural diagram of a probe provided in an embodiment of the present application;
[0034] Figure 4 This is a schematic diagram of the structure of a target detection network provided by an embodiment of the present application;
[0035] Figure 5 is a schematic diagram of an activation function provided in an embodiment of the present application;
[0036] Figure 6 is a schematic diagram of a defect detection device based on an industrial endoscope provided in an embodiment of the present application;
[0037] Figure 7 This is a schematic diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0038] In the following description, specific details such as specific system structures and technologies are provided for the purpose of illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present application with unnecessary details.
[0039] The technical solution of this application is described below through specific embodiments.
[0040] Reference Figure 1 , which shows a schematic diagram of a defect detection method based on an industrial endoscope provided in an embodiment of the present application, which may specifically include the following steps:
[0041] Step 101: Determine target detection information.
[0042] like Figure 2 As shown, an industrial endoscope generally includes components such as a probe 201, a display screen 202, and a control console 203, and can be used to detect defects inside industrial equipment such as pipelines.
[0043] like Figure 3 As shown, the probe of the industrial endoscope has multiple sets of detection modules with different angles, and each set of detection modules has a binocular camera 2011 and a light source 2022.
[0044] For example, the probe of the industrial endoscope has two sets of detection modules, one set of detection modules is located at the end of the probe and is used to collect image data toward the front, and the other set of detection modules is located on the side of the probe and is used to collect image data toward the side.
[0045] Users can control the switching between different detection modules on the console to detect defects under different visual conditions such as direct view or side view, observe the image data under different visual conditions such as direct view and side view on the display screen, and compare the geometric information of defect measurements under different visual conditions such as direct view and side view.
[0046] The detection module at the end of the probe captures image data by extracting the pipe ring and performing distortion correction on it. The corrected ring is then expanded into a panoramic image, providing foundational support for subsequent processing. The panoramic image created by the direct-viewing detection module helps the side-viewing detection module locate the coordinates of defects within industrial equipment.
[0047] The inner wall ring image expansion step is roughly divided into the following two steps:
[0048] 1. Tangential expansion: the difference between the inner and outer circle radii in the circular image is used as the width of the expanded rectangle, and the circumference of the outer circle is used as the length of the rectangle, so as to expand the annular image into a rectangular image.
[0049] 2. Radial expansion: We propose radial stretching based on the geometric relationship between the pipe inner wall and the imaging surface. The coordinates of the original point p(x,y) after tangential transformation are p′(x′,y′), while the coordinates of the desired expanded image are p″(x″,y″). The relationship between the two is x″ - x′, y″ = f(y′).
[0050] In this way, a panoramic expansion image under the direct viewing angle can be obtained. By adopting scale-invariant feature transformation for feature extraction, the robust RANSAC algorithm is used to remove the mismatched points after matching the feature points. The pixel gray value weighted (fading in and out) fusion algorithm is used to fuse the matching result images obtained above, and finally the complete reproduction of the planar flaw detection area under the direct viewing angle is achieved.
[0051] The industrial endoscope in this embodiment combines the three-dimensional measurement functions of direct view and side view. During the actual inspection process, there is no need to take out the industrial endoscope and re-switch the lens to re-locate the defect, thereby significantly reducing the operator's workload and manual errors.
[0052] In this embodiment, the industrial endoscope is configured with a high-performance computing module, such as the NPU (Neural Processing Unit) of RK3588, and the target detection network constructed and trained offline is transplanted into the computing module. That is, the industrial endoscope has a target detection network, which is used to detect defects in industrial equipment according to user needs (represented in text information).
[0053] Furthermore, if Figure 4 As shown in the figure, the target detection network includes the backbone structure Backbone, the text encoder TextEncoder, the neck structure Neck and the head structure Head.
[0054] Among them, the backbone network is responsible for extracting local and global features from the image data collected by the probe.
[0055] The text encoder TextEncoder is responsible for extracting local or global features from the text information input by the user.
[0056] The neck structure Neck is responsible for fusing the features extracted by the backbone network Backbone with the features extracted by the text encoder TextEncoder to enhance the expressiveness of the features.
[0057] The head structure is responsible for performing specific detection tasks, mapping the features extracted by the neck structure to the final output space, and generating a target detection frame representing the defect.
[0058] For this target detection network, the training samples mainly consist of image data of various defects inside industrial equipment and text information describing the defects. The loss function consists of two parts: one is the contrastive loss from image to text, and the other is the contrastive loss from text to image.
[0059] When the industrial endoscope is running, the target detection network can be loaded and target detection information can be received.
[0060] The target detection information includes target text information indicating defects input by the user, and raw image data collected by the binocular camera inside the industrial equipment under the condition of illumination by a light source with a preset intensity Dbase.
[0061] When the target detection network is easy to update online, the target text information input by the user can be a fully open vocabulary. At this time, the target detection network is trained with richer samples and can understand vocabulary with more complex semantics.
[0062] When there are certain difficulties in updating the target detection network, the target text information input by the user can be semi-open vocabulary. At this time, the target detection network is trained with a fixed amount of samples and can understand vocabulary with a certain degree of semantic complexity.
[0063] The so-called semi-open means that the target text information entered by the user is as close as possible to the samples used during training, which makes it easier for the target detection network to understand the semantics and improve the accuracy of defect detection.
[0064] In a specific implementation, the industrial endoscope has built-in categories and a restrictive vocabulary library representing defects. The restrictive vocabulary library contains multiple restrictive phrases. Both the categories and the restrictive vocabulary library are used to train the target detection network.
[0065] Among them, the backbone structure Backbone, neck structure Neck and head structure Head can form a basic network that can realize target detection. The function of this basic network is to detect defects in various categories. These categories are usually fixed, and industrial endoscopes may be used by users to detect defects inside different types of industrial equipment, which increases new detection needs. If the category is modified each time to match the new detection needs, the basic network is modified and retrained. Users also reduce the update frequency due to the long time required to fully update the basic network each time. The performance growth of the target detection network may stagnate, which will bring huge incremental costs to the training and maintenance of the target detection network.
[0066] Therefore, in this embodiment, a text encoder TextEncoder can be added to the basic network, and the neck structure Neck can be improved to form a target detection network. Without destroying the ability to detect defects under various categories, the ability to learn open words composed of the restricted vocabulary is added to detect defects that meet the categories and / or restricted vocabulary, effectively expanding the capabilities of the target detection network. The categories of the target detection network are set to some mainstream and general categories, and the restricted vocabulary is used to expand the categories, that is, the priority of the category is higher than the priority of the restricted vocabulary. When matching new detection requirements, there is no need to modify the category, only to modify the restricted vocabulary and fine-tune the target detection network. The user can update the parameters and restricted vocabulary of the target detection network online when appropriate. The update amount is small and the operation is convenient, so that the performance of the target detection network is continuously enhanced, which brings convenience to the training and maintenance of the target detection network.
[0067] At this time, the original text information input by the user may be received, and word segmentation processing may be performed on the original text information using a dictionary or other means to obtain a plurality of original phrases.
[0068] The similarity between each piece of original text information and the category representing the defect is calculated using methods such as cosine similarity, Jaccard distance, and bag-of-words model.
[0069] If the similarity is within a preset first similarity interval, indicating that the original text information and the category representing the defect have a high similarity, the category can be directly set as the target text information.
[0070] If the similarity is within the preset second similarity interval, it means that the original text information and the category representing the defect have a medium similarity, and it can be considered that the original text information matches the category representing the defect. Then, an original phrase with a matching category is searched as the first candidate phrase.
[0071] Dependency parsing (NLP) is performed on each original phrase in the original text information to filter out some interference such as stop words and expected words, and identify other original phrases that have a dependency relationship with the first candidate phrase as the second candidate phrase.
[0072] Use cosine similarity, Jaccard distance, bag-of-words model, etc. to search for a qualifying phrase that matches the second candidate word in a preset qualifying vocabulary.
[0073] The category matching the first candidate word group and the qualified word group matching the second candidate word group are set as target text information.
[0074] If the similarity is within the preset third similarity interval, indicating that the similarity between the original text information and the category representing the defect is medium to low, a preset restrictive word library is searched for restrictive phrases matching each original phrase using methods such as cosine similarity, Jaccard distance, and bag-of-words model.
[0075] A qualified phrase that matches the original phrase is set as the target text information.
[0076] The value in the first similarity interval is greater than the value in the second similarity interval, and the value in the second similarity interval is greater than the value in the third similarity interval.
[0077] Step 102: Input the original image data into the backbone structure to extract original image features at multiple scales.
[0078] like Figure 4 As shown, the original image data Image is input into the backbone structure Backbone, and the backbone structure Backbone extracts original image features of multiple scales from the original image data Image in sequence.
[0079] In a specific implementation, the original image features include first image level features, second image level features and third image level features; wherein the scale of the first image level features is larger than the scale of the second image level features, and the scale of the second image level features is larger than the scale of the third image level features.
[0080] In a structure such as Figure 4 As shown in the figure, the backbone structure Backbone includes 5 convolution modules ConvModule, 4 C2f modules CSPLayer_2Conv and a fast spatial pyramid pooling module (Spatial Pyramid Pooling Fast, SPPF).
[0081] In this structure, the original image data Image is sequentially input into the convolution module ConvModule, the convolution module ConvModule, the C2f module CSPLayer_2Conv, the convolution module ConvModule, and the C2f module CSPLayer_2Conv for processing to obtain the first image level feature P1.
[0082] The first image-level feature P1 is sequentially input into the convolution module ConvModule and the C2f module CSPLayer_2Conv for processing to obtain the second image-level feature P2.
[0083] The second image level feature P2 is sequentially input into the convolution module ConvModule, the C2f module CSPLayer_2Conv, and the fast spatial pyramid pooling module SPPF for processing to obtain the third image level feature P3.
[0084] Step 103: Input the target text information into a text encoder to extract original text features at multiple scales.
[0085] like Figure 4 As shown, the target text information Text is input into the text encoder TextEncoder, and TextEncoder extracts original text features of multiple scales from the target text information Text in turn.
[0086] In a structure such as Figure 4 As shown in the figure, the text encoder TextEncoder includes a bidirectional encoder representations from transformers (BERT), a first convolutional layer Conv_1, a second convolutional layer Conv_2, a third convolutional layer Conv_3 and a pyramid pooling module (PPM). Accordingly, the original text features include the first text level feature T1, the second text level feature T2, the third text level feature T3 and the fourth text level feature T4.
[0087] Then, in this structure, the target text information Text is input into the bidirectional encoder BERT to extract one-dimensional text sequence features.
[0088] Perform dot multiplication of the transpose of the text sequence feature and the text sequence feature, expressed as , where x is the text sequence feature and T is the transpose, thereby converting the text sequence feature into a two-dimensional text matrix feature, which can locally amplify the features in the text. When semi-open words are used as input, the difference between local features can be increased, which facilitates the subsequent use of the attention mechanism and improves the accuracy of target detection.
[0089] The text matrix features are input into the first convolutional layer Conv_1 to perform a convolution operation to extract the first text level features T1.
[0090] The first text-level feature T1 is input into the second convolutional layer Conv_2 to perform a convolution operation to extract the second text-level feature T2.
[0091] The second text-level feature T2 is input into the third convolutional layer Conv_3 to perform a convolution operation to extract the third text-level feature T3.
[0092] The first text-level feature T1, the second text-level feature T2, and the third text-level feature T3 are input into the pyramid pooling module PPM for interactive fusion to extract the fourth text-level feature T4. The pyramid pooling module PPM has different receptive fields and can enrich the information content of the fourth text-level feature T4.
[0093] In one design, Figure 4 As shown, the pyramid pooling module PPM includes the first self-attention layer Self-Attention_1, the second self-attention layer Self-Attention_2, the third self-attention layer Self-Attention_3 and the fourth convolutional layer Conv_4.
[0094] The first text-level feature T1 is input into the first self-attention layer Self-Attention_1, and the first text-level feature T1 is converted into the first text-attention feature under self-attention (i.e., the first text-level feature T1 is used as the Q (query) matrix, K (key) matrix and V (value) matrix).
[0095] The second text-level feature T2 is input into the second self-attention layer Self-Attention_2, and the second text-level feature T2 is converted into the second text attention feature under self-attention (i.e., the second text-level feature T2 is used as the Q matrix, K matrix and V matrix).
[0096] The third text-level feature is input into the third self-attention layer Self-Attention_3, and the third text-level feature T3 is converted into the third text-attention feature under self-attention (i.e., the third text-level feature T3 is used as the Q matrix, K matrix and V matrix).
[0097] Use functions such as Concat to concatenate the first text attention feature, the second text attention feature, and the third text attention feature into the fourth text attention feature.
[0098] The fourth text attention feature is input into the fourth convolutional layer Conv_4 to perform a convolution operation to extract the fourth text level feature T4.
[0099] Step 104: Input multiple original image features and multiple original text features into the neck structure for interaction to generate multiple image-text fusion features.
[0100] like Figure 4 As shown in FIG, multiple original image features and multiple original text features are input into the neck structure Neck for interaction at different levels, thereby generating multiple image-text fusion features.
[0101] In a structure such as Figure 4As shown, the neck structure Neck includes the first attention layer Attention_1, the second attention layer Attention_2, the third attention layer Attention_3, the fourth attention layer Attention_4 and the fifth attention layer Attention_5, as well as 4 C2f modules CSPLayer_2Conv and 2 convolution modules ConvModule.
[0102] Correspondingly, the image-text fusion feature includes a first image-text level feature C1, a second image-text level feature C2 and a third image-text level feature C3.
[0103] Then, in this structure, the first image-level feature P1 and the first text-level feature T1 are input into the first attention layer Attention_1, and the first image-level feature P1 is used as the Q matrix, and the first text-level feature T1 is used as the K matrix and the V matrix for operation. According to the attention of the first text-level feature T1 to the first image-level feature P1, the first text-level feature T1 is fused into the first image-level feature P1 to obtain the first image-text candidate feature L1.
[0104] The second image-level feature P2 and the second text-level feature T2 are input into the second attention layer Attention_2. The second image-level feature P2 is used as the Q matrix, and the second text-level feature T2 is used as the K matrix and the V matrix for operation. According to the attention of the second text-level feature T2 to the second image-level feature P2, the second text-level feature T2 is fused into the second image-level feature P2 to obtain the second image-text candidate feature L2.
[0105] The third image-level feature P3 and the third text-level feature T3 are input into the third attention layer Attention_3, and the third image-level feature P3 is used as the Q matrix, and the third text-level feature T3 is used as the K matrix and the V matrix for operation. According to the attention of the third text-level feature T3 to the third image-level feature P3, the third text-level feature T3 is fused into the third image-level feature P3 to obtain the third image-text candidate feature L3.
[0106] A first upsampling operation Upsample is performed on the third image-text candidate feature L3.
[0107] If the first upsampling operation Upsample is completed, the second image-text candidate feature L2 and the third image-text candidate feature L3 are concatenated to form a fourth image-text candidate feature L4.
[0108] The fourth image-text candidate feature L4 is input into the C2f module CSPLayer_2Conv to extract the fifth image-text candidate feature L5.
[0109] A second upsampling operation Upsample is performed on the fifth image-text candidate feature L5.
[0110] If the second upsampling operation Upsample is completed, the first image-text candidate feature L1 and the fifth image-text candidate feature L5 are concatenated to form a sixth image-text candidate feature L6.
[0111] The sixth image-text candidate feature L6 is input into the C2f module CSPLayer_2Conv to extract the first image-text level feature C1.
[0112] The first image-text level feature C1 is input into the convolution module ConvModule to extract the seventh image-text candidate feature L7.
[0113] The fifth image-text candidate feature L5 and the fourth text-level feature T4 are input into the fourth attention layer Attention_4. The fifth image-text candidate feature L5 is used as the Q matrix, and the fourth text-level feature T4 is used as the K matrix and the V matrix for operation. According to the attention of the fourth text-level feature T4 to the fifth image-text candidate feature L5, the fourth text-level feature T4 is fused into the fifth image-text candidate feature L5 to obtain the eighth image-text candidate feature L8.
[0114] The seventh image and text candidate feature L7 and the eighth image and text candidate feature L8 are concatenated to form the ninth image and text candidate feature L9.
[0115] The ninth image-text candidate feature L9 is input into the C2f module CSPLayer_2Conv to extract the second image-text level feature C2.
[0116] The second image-text level feature C2 is input into the convolution module ConvModule to extract the tenth image-text candidate feature L10.
[0117] The third image-text candidate feature L3 and the tenth image-text candidate feature L10 are concatenated to form the eleventh image-text candidate feature L11.
[0118] The eleventh image-text candidate feature L11 is input into the C2f module CSPLayer_2Conv to extract the third image-text level feature C3.
[0119] Step 105: Input multiple image-text fusion features into the header structure to detect the target detection frame representing the defect.
[0120] like Figure 4 As shown, multiple image-text fusion features are input into the head structure Head, and the multiple image-text fusion features of the head structure Head are mapped into a target detection frame for representing defects, wherein the content in the target detection frame conforms to the target text information, and the target detection frame has confidence.
[0121] In one structure, the head structure Head has three decoupling modules Decoupled, and the first image-text level feature C1, the second image-text level feature C2 and the third image-text level feature C3 are respectively input into each decoupling module Decoupled to generate three candidate detection boxes.
[0122] Among them, each decoupling module Decoupled includes two branch structures, both of which contain a convolution module ConvModule and a convolution layer Conv2d. One branch structure is responsible for using features (such as the first image-text level feature C1, the second image-text level feature C2, and the third image-text level feature C3) to generate a box Bbox, and the other branch structure is responsible for using features (such as the first image-text level feature C1, the second image-text level feature C2, and the third image-text level feature C3) to generate a type Cls (i.e., target text information). The two are jointly represented as a candidate detection box.
[0123] The three candidate detection frames are processed through IOU (intersection-over-union) and NMS (non-maximum suppression) to generate the final target detection frame.
[0124] Step 106 : If the confidence level is greater than or equal to a preset threshold, a brightening operation is performed on the original image data based on the multiple frames of candidate image data to obtain target image data.
[0125] In practical applications, the confidence level can be compared with a preset threshold. If the confidence level is greater than or equal to the preset threshold, it means that the confidence level of the target detection box is high.
[0126] At this time, considering the low brightness of the environment inside the industrial equipment, if the original image data is directly used for 3D modeling, the modeling loss will be high and the accuracy of the point cloud data will be low, which will affect the accuracy of the geometric information for defect detection.
[0127] In this regard, this embodiment controls the light source to enhance (i.e., increase the intensity) the illumination at multiple time steps, i.e., equivalent to the intensity Dbase serving as a benchmark, controls the light source to enhance the intensity by ΔD (i.e., Dbase+ΔD) at time step T1 and collects a frame of candidate image data, controls the light source to continue enhancing the intensity by ΔD (i.e., Dbase+2ΔD) at time step T2 and collects a frame of candidate image data, controls the light source to continue enhancing the intensity by ΔD (i.e., Dbase+3ΔD) at time step T3 and collects a frame of candidate image data, and so on, thereby collecting multiple frames of candidate image data. Then, the candidate image data is the image data collected by the binocular camera inside the industrial equipment under the condition of enhanced (i.e., increased intensity) illumination by the light source.
[0128] While keeping the content of the original image data unchanged, a brightening operation is performed on the original image data based on multiple frames of candidate image data to obtain target image data. Using the target image data for three-dimensional modeling can reduce modeling losses and increase the accuracy of point cloud data, thereby improving the accuracy of geometric information for defect detection.
[0129] In one embodiment of the present application, step 106 may include the following steps:
[0130] Step 1061 : Extract candidate brightness image data from each frame of candidate image data.
[0131] In this embodiment, the pixels under the brightness channel may be separated from each frame of candidate image data to obtain candidate brightness image data.
[0132] Among them, the brightness channel is the Y channel in the YUV format, the L channel in the HSL format, and so on.
[0133] Step 1062 : Fusing multiple frames of candidate brightness image data into target brightness image data based on the time step of enhanced illumination of the light source.
[0134] In this embodiment, multiple frames of candidate brightness image data may be synchronously fused into target brightness image data with reference to a time step of enhanced illumination by a light source.
[0135] In a specific implementation, a target point for light source calibration is located in multiple frames of candidate brightness image data. The target point is mostly located near the center point of the candidate brightness image data, indicating a point with the highest brightness value.
[0136] An activation function is set for each frame of candidate brightness image data according to the time step of enhanced illumination of the light source; wherein the activation function is a single-peak curve. Initially, the horizontal coordinate of the peak value of the activation function is 0, and the horizontal coordinate of the peak value of the activation function shifts toward the positive direction of the horizontal axis as the time step of enhanced illumination of the light source increases.
[0137] For example, Figure 5 As shown, the first curve 501 is the activation function at the initial time step T1, the second curve 502 is the activation function at the initial time step T2, the third curve 503 is the activation function at the initial time step T3, and so on.
[0138] In the candidate brightness image data, the distance between each pixel and the target point is calculated, such as the Euclidean distance.
[0139] Substitute the distance as the horizontal coordinate into the corresponding activation function and map it to the weight, that is, the weight is the vertical coordinate corresponding to the horizontal coordinate in the activation function.
[0140] For pixel points at the same position in multiple frames of candidate brightness image data, a normalization operation is performed on multiple weights so that the sum of the multiple weights is 1.
[0141] If the normalization operation is completed, the pixel points at the same position in the multiple frames of candidate brightness image data are weightedly summed using multiple weights to obtain the target brightness image data.
[0142] Then, the target brightness image data can be expressed as:
[0143]
[0144] in, is the brightness value of pixel i in the target brightness image data, S is the set of time steps, At time step T, the weight of pixel i in the candidate brightness image data is, is the brightness value of pixel i in the candidate brightness image data at time step T.
[0145] Step 1063: Under the framework of the image enhancement model Retinex, the target brightness image data is fused into the original image data to obtain the target image data.
[0146] The theory behind the Retinex image enhancement model states that the raw image data seen by the user is determined by the light emitted by the light source and the reflection of the incident light from the surface of the industrial equipment. Specifically, light emitted by the light source is reflected by the industrial equipment and enters the binocular camera, forming the raw image data. In this process, the reflectivity is determined by the industrial equipment itself and is unaffected by the light emitted by the light source. This can be expressed as: L = I·T, where L is the raw image data, I is the illumination component of the ambient light, and T is the reflected component that carries target detail information.
[0147] If I is regarded as illumination image data and T is regarded as reflection image data, then the original image data can be modeled as the superposition of illumination image data I and reflection image data T.
[0148] Within the framework of the Retinex image enhancement model, the illumination image data I determines the dynamic range that each pixel in the original image data can achieve, and the reflection image data T determines the intrinsic properties of the original image data. The properties of the illumination image data I can be discarded from the original image data L, thereby separating the original appearance of the industrial equipment (i.e., the reflection image data T). The reflection image data T is then fused with the target brightness image data with more uniform illumination to form the target image data, eliminating the impact of uneven illumination and improving the quality of the target image data. This can be expressed as: L'=I'·T, where L' is the target image data, I' is the target brightness image data, and T is the reflection image data.
[0149] Step 107: Construct three-dimensional point cloud data based on the target detection frame in the target image data.
[0150] In practical applications, methods such as block matching and semi-global matching (SGM) can be used to compare the blocks in the target image data (especially the target detection frame) corresponding to the binocular camera to determine the disparity value between each pixel point, and then perform three-dimensional modeling based on the disparity value, calculate the three-dimensional coordinates of each pixel point, and thus construct three-dimensional point cloud data.
[0151] The three-dimensional coordinates can be expressed as: Z = (f × B) / d, where Z is the distance from the pixel to the binocular camera, f is the focal length of the binocular camera, B is the baseline distance between the binocular cameras themselves, and d is the parallax.
[0152] Step 108: Detect geometric information of defects in the point cloud data.
[0153] In this embodiment, various measurements are performed on the point cloud data of the defect model according to business requirements to obtain geometric information of the internal defects of the industrial equipment, such as length, area, minimum bounding rectangle, etc.
[0154] There is a certain matching cost when building 3D modeling based on binocular cameras, which affects the accuracy.
[0155] To this end, the binocular matching cost value can be used to determine the confidence of the defect edge, and the nearby point with the highest confidence can be selected when measuring geometric information.
[0156] In addition, considering that defects such as cracks, rust, and defects will change in physical form, the brightness difference in the object's form will be magnified under the condition of enhanced lighting. Therefore, the brightness of point cloud data can be combined to improve the accuracy of defect detection.
[0157] In a specific implementation, the point cloud data is vertically projected onto a horizontal plane.
[0158] On the horizontal plane, a window of size n×n (n is an odd number, n>1) is added to the current point cloud data; wherein the current point cloud data is located at the center of the window.
[0159] The average brightness value of other point cloud data located at non-center positions of the window is calculated to obtain the average brightness.
[0160] The absolute value of the difference between the brightness value of the current point cloud data and the average brightness of other point cloud data is taken as the brightness deviation value.
[0161] For each point cloud data, the brightness deviation value is mapped to the confidence level of the defect edge; the confidence level is positively correlated with the brightness deviation value, that is, the larger the brightness deviation value, the higher the confidence level of the point cloud data belonging to the defect edge; conversely, the smaller the brightness deviation value, the lower the confidence level of the point cloud data belonging to the defect edge.
[0162] Multiple point cloud data with the highest confidence levels are selected to construct the edge of the defect, thereby generating geometric information about the defect.
[0163] In an embodiment of the present application, target detection information is determined; the target detection information includes target text information representing a defect, and original image data collected by a binocular camera inside an industrial device under conditions of light source illumination; the original image data is input into a backbone structure to extract original image features of multiple scales; the target text information is input into a text encoder to extract original text features of multiple scales; multiple original image features and multiple original text features are input into a neck structure for interaction to generate multiple image-text fusion features; the multiple image-text fusion features are input into a head structure to detect a target detection frame representing a defect; the target detection frame has a confidence level; if the confidence level is greater than or equal to a preset threshold, a brightening operation is performed on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data is image data collected by a binocular camera inside an industrial device under conditions of enhanced light source illumination; three-dimensional point cloud data is constructed based on the target detection frame in the target image data; and geometric information of defect detection is detected in the point cloud data. The embodiment of the present application has a built-in target detection network in the industrial endoscope, detects defects in the original image data based on open target text information, and realizes adaptive defect tracking, allowing users to focus on operating the industrial endoscope and effectively prevent missing defects. In addition, the information collected by the binocular camera is richer, and three-dimensional modeling based on the binocular camera can improve the accuracy of the geometric information of the detected defects, thereby improving the efficiency of defect detection in industrial equipment.
[0164] In addition, the target detection network has a clear structure, the detection process is a single forward transmission, it occupies few resources, and has a fast detection speed. It is suitable for deployment in industrial endoscopes to achieve local detection, and does not rely on cloud servers or other high-computing power equipment, thereby improving the convenience and safety of industrial endoscopes in independent environments.
[0165] It should be noted that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0166] Reference Figure 6, shows a schematic diagram of a defect detection device based on an industrial endoscope provided in an embodiment of the present application. The industrial endoscope has multiple detection modules and a target detection network; each detection module has a binocular camera and a light source, and the target detection network includes a backbone structure, a text encoder, a neck structure, and a head structure; the device may specifically include the following modules:
[0167] The target detection information determination module 601 is configured to determine target detection information, wherein the target detection information includes target text information indicating a defect and raw image data collected by the binocular camera inside the industrial equipment under the illumination of the light source.
[0168] The original image feature extraction module 602 is used to input the original image data into the backbone structure to extract original image features of multiple scales;
[0169] The original text feature extraction module 603 is used to input the target text information into the text encoder to extract original text features of multiple scales;
[0170] An image-text fusion feature generation module 604 is configured to input a plurality of the original image features and a plurality of the original text features into the neck structure for interaction to generate a plurality of image-text fusion features;
[0171] The target detection frame generation module 605 is used to input the plurality of the image-text fusion features into the header structure to detect a target detection frame representing the defect; the target detection frame has a confidence level;
[0172] a target image data generating module 606 configured to, if the confidence level is greater than or equal to a preset threshold, perform a brightening operation on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data being image data captured by the binocular camera inside the industrial equipment under the condition of enhanced illumination by the light source;
[0173] A point cloud data construction module 607 is configured to construct three-dimensional point cloud data based on the target detection frame in the target image data;
[0174] The geometric information detection module 608 is used to detect geometric information of the defect in the point cloud data.
[0175] In one embodiment of the present application, the target detection information determination module 601 is further configured to:
[0176] receiving original text information input by a user, performing word segmentation processing on the original text information to obtain a plurality of original phrases;
[0177] Calculating the similarity between the original text information and the category representing the defect;
[0178] If the similarity is within a preset first similarity interval, setting the category as target text information;
[0179] If the similarity is within a preset second similarity interval, searching for the original phrase that matches the category as a first candidate phrase;
[0180] identifying other original phrases that have a dependency relationship with the first candidate phrase as second candidate phrases;
[0181] Searching for a qualifier phrase matching the second candidate word in a preset qualifier word library;
[0182] setting the category matching the first candidate word group and the qualified word group matching the second candidate word group as target text information;
[0183] If the similarity is within a preset third similarity interval, searching a preset restrictive word library for a restrictive word group that matches the original word group;
[0184] Setting the restricted phrase that matches the original phrase as target text information;
[0185] The value in the first similarity interval is greater than the value in the second similarity interval, and the value in the second similarity interval is greater than the value in the third similarity interval;
[0186] The categories and the restricted vocabulary are both used to train the object detection network.
[0187] In one embodiment of the present application, the text encoder includes a bidirectional encoder BERT, a first convolutional layer, a second convolutional layer, a third convolutional layer and a pyramid pooling module; the original text features include first text-level features, second text-level features, third text-level features and fourth text-level features;
[0188] The original text feature extraction module 603 is further used to:
[0189] Input the target text information into the bidirectional encoder BERT to extract one-dimensional text sequence features;
[0190] Converting the text sequence features into two-dimensional text matrix features;
[0191] Inputting the text matrix features into the first convolutional layer to extract first text-level features;
[0192] Inputting the first text-level features into the second convolutional layer to extract second text-level features;
[0193] Inputting the second text-level features into the third convolutional layer to extract third text-level features;
[0194] The first text-level feature, the second text-level feature, and the third text-level feature are input into the pyramid pooling module to extract a fourth text-level feature.
[0195] In one embodiment of the present application, the pyramid pooling module includes a first self-attention layer, a second self-attention layer, a third self-attention layer and a fourth convolutional layer;
[0196] The original text feature extraction module 603 is further used to:
[0197] Inputting the first text-level feature into the first self-attention layer and converting it into a first text-attention feature under self-attention;
[0198] Inputting the second text-level features into the second self-attention layer and converting them into second text-attention features under self-attention;
[0199] Inputting the third text-level feature into the third self-attention layer and converting it into a third text-attention feature under self-attention;
[0200] splicing the first text attention feature, the second text attention feature and the third text attention feature into a fourth text attention feature;
[0201] The fourth text attention feature is input into the fourth convolutional layer to extract the fourth text level feature.
[0202] In one embodiment of the present application, the original image features include first image-level features, second image-level features, and third image-level features; the scale of the first image-level features is larger than the scale of the second image-level features, and the scale of the second image-level features is larger than the scale of the third image-level features;
[0203] The neck structure includes the first attention layer, the second attention layer, the third attention layer, the fourth attention layer and the fifth attention layer, as well as 4 C2f modules and 2 convolution modules;
[0204] The image-text fusion features include first image-text level features, second image-text level features and third image-text level features;
[0205] The image-text fusion feature generation module 604 is further used to:
[0206] Inputting the first image-level feature and the first text-level feature into the first attention layer, and fusing the first text-level feature into the first image-level feature based on the attention of the first text-level feature to the first image-level feature, to obtain a first image-text candidate feature;
[0207] Inputting the second image-level feature and the second text-level feature into the second attention layer, and fusing the second text-level feature into the second image-level feature based on the attention of the second text-level feature to the second image-level feature, to obtain a second image-text candidate feature;
[0208] Inputting the third image-level feature and the third text-level feature into the third attention layer, and fusing the third text-level feature into the third image-level feature based on the attention of the third text-level feature to the third image-level feature, to obtain a third image-text candidate feature;
[0209] Performing a first upsampling operation on the third image-text candidate feature;
[0210] If the first upsampling operation is completed, the second image-text candidate feature and the third image-text candidate feature are concatenated into a fourth image-text candidate feature;
[0211] Inputting the fourth image-text candidate feature into the C2f module to extract the fifth image-text candidate feature;
[0212] Performing a second upsampling operation on the fifth image-text candidate feature;
[0213] If the second upsampling operation is completed, the first image-text candidate feature and the fifth image-text candidate feature are concatenated into a sixth image-text candidate feature;
[0214] Inputting the sixth image-text candidate feature into the C2f module to extract the first image-text level feature;
[0215] Inputting the first image-text level feature into the convolution module to extract the seventh image-text candidate feature;
[0216] Inputting the fifth image-text candidate feature and the fourth text-level feature into the fourth attention layer, and fusing the fourth text-level feature into the fifth image-text candidate feature based on the attention of the fourth text-level feature to the fifth image-text candidate feature, to obtain an eighth image-text candidate feature;
[0217] splicing the seventh image-text candidate feature and the eighth image-text candidate feature into a ninth image-text candidate feature;
[0218] Inputting the ninth image-text candidate feature into the C2f module to extract the second image-text level feature;
[0219] Inputting the second image-text level feature into the convolution module to extract the tenth image-text candidate feature;
[0220] splicing the third image-text candidate feature and the tenth image-text candidate feature into an eleventh image-text candidate feature;
[0221] The eleventh image-text candidate feature is input into the C2f module to extract the third image-text level feature.
[0222] In one embodiment of the present application, the target image data generating module 606 is further configured to:
[0223] extracting candidate brightness image data from the candidate image data of each frame respectively;
[0224] fusing multiple frames of candidate brightness image data into target brightness image data according to a time step of enhanced illumination of the light source;
[0225] Under the framework of the image enhancement model Retinex, the target brightness image data is fused into the original image data to obtain target image data.
[0226] In one embodiment of the present application, the target image data generating module 606 is further configured to:
[0227] Locating a target point for calibrating the light source in multiple frames of candidate brightness image data;
[0228] An activation function is set for each frame of the candidate brightness image data according to the time step of the enhanced illumination of the light source; the activation function is a single-peak curve, and at an initial time step, the abscissa of the peak value of the activation function is 0, and the abscissa of the peak value of the activation function shifts toward the positive direction of the abscissa as the time step of the enhanced illumination of the light source increases;
[0229] In the candidate brightness image data, calculating the distance between each pixel and the target point;
[0230] Substituting the distance as the horizontal coordinate into the activation function and mapping it into a weight;
[0231] performing a normalization operation on the plurality of weights for the pixel points located at the same position in the plurality of frames of the candidate brightness image data;
[0232] If the normalization operation is completed, a plurality of weights are used to perform weighted summation on the pixel points at the same position in the multiple frames of the candidate brightness image data to obtain the target brightness image data.
[0233] In one embodiment of the present application, the geometric information detection module 608 is further configured to:
[0234] Projecting the point cloud data onto a horizontal plane;
[0235] On a horizontal plane, a window is added to the current point cloud data; the current point cloud data is located at the center of the window;
[0236] Calculating an average brightness value of other point cloud data located at non-central positions of the window to obtain an average brightness;
[0237] Taking the absolute value of the difference between the brightness value of the current point cloud data and the average brightness of the other point cloud data as the brightness deviation value;
[0238] For each of the point cloud data, mapping the brightness deviation value to a confidence level of the defect edge; the confidence level is positively correlated with the brightness deviation value;
[0239] The plurality of point cloud data with the highest confidence levels are selected to generate geometric information for the defect.
[0240] An embodiment of the present application provides a defect detection device based on an industrial endoscope. By using this device, each step in the aforementioned method embodiments can be implemented.
[0241] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment part.
[0242] Reference Figure 7 , shows a schematic diagram of a terminal device provided by an embodiment of the present application. Figure 7 As shown, the terminal device 400 in the embodiment of the present application includes: a processor 410, a memory 420, and a computer program 421 stored in the memory 420 and executable on the processor 410. When the processor 410 executes the computer program 421, the steps of each embodiment of the above-mentioned industrial endoscope-based defect detection method are implemented. Alternatively, when the processor 410 executes the computer program 421, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0243] Exemplarily, the computer program 421 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 420 and executed by the processor 410 to complete the present application. The one or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments can be used to describe the execution process of the computer program 421 in the terminal device 400. For example, the computer program 421 can be divided into a target detection information determination module, an original image feature extraction module, an original text feature extraction module, a graphic and text fusion feature generation module, a target detection frame generation module, a target image data generation module, a point cloud data construction module, and a geometric information detection module. The specific functions of each module are as follows:
[0244] a target detection information determination module, configured to determine target detection information; the target detection information includes target text information indicating a defect, and raw image data captured by the binocular camera inside the industrial equipment under the illumination of the light source;
[0245] An original image feature extraction module, configured to input the original image data into the backbone structure to extract original image features of multiple scales;
[0246] An original text feature extraction module, configured to input the target text information into the text encoder to extract original text features of multiple scales;
[0247] An image-text fusion feature generation module, configured to input a plurality of the original image features and a plurality of the original text features into the neck structure for interaction to generate a plurality of image-text fusion features;
[0248] a target detection frame generation module, configured to input the plurality of image-text fusion features into the header structure to detect a target detection frame representing the defect; the target detection frame has a confidence level;
[0249] a target image data generation module configured to, if the confidence level is greater than or equal to a preset threshold, perform a brightening operation on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data being image data captured by the binocular camera inside the industrial equipment under the condition of enhanced illumination by the light source;
[0250] A point cloud data construction module, configured to construct three-dimensional point cloud data based on the target detection frame in the target image data;
[0251] A geometric information detection module is used to detect geometric information of the defect in the point cloud data.
[0252] The terminal device 400 may be the industrial endoscope in the aforementioned embodiments, which may be connected to a desktop computer, a cloud server or other computing device via wired or wireless communication. These computing devices provide computing power support for the industrial endoscope. The terminal device 400 may include, but is not limited to, a processor 410 and a memory 420. Those skilled in the art will appreciate that Figure 7 It is only an example of the terminal device 400 and does not constitute a limitation of the terminal device 400. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the terminal device 400 may also include input and output devices, network access devices, buses, etc.
[0253] The processor 410 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0254] The memory 420 may be an internal storage unit of the terminal device 400, such as a hard disk or memory of the terminal device 400. The memory 420 may also be an external storage device of the terminal device 400, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped on the terminal device 400. Furthermore, the memory 420 may include both an internal storage unit of the terminal device 400 and an external storage device. The memory 420 is used to store the computer program 421 and other programs and data required by the terminal device 400. The memory 420 may also be used to temporarily store data that has been output or is to be output.
[0255] An embodiment of the present application also discloses a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the defect detection method based on the industrial endoscope as described in the aforementioned embodiments is implemented.
[0256] An embodiment of the present application further discloses a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the industrial endoscope-based defect detection method as described in the aforementioned embodiments is implemented.
[0257] An embodiment of the present application further discloses a computer program product. When the computer program product is run on a computer, the computer is enabled to execute the industrial endoscope-based defect detection method described in the aforementioned embodiments.
[0258] The above embodiments are intended only to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they may still modify the technical solutions described in the above embodiments or replace some of the technical features therein with equivalents; and such modifications or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present application and should be included within the scope of protection of the present application.
Claims
1. A defect detection method based on an industrial endoscope, characterized in that: The industrial endoscope comprises multiple detection modules and a target detection network, categories representing defects and a definite word library, wherein the definite word library comprises multiple definite word groups, and the categories and the definite word library are both used to train the target detection network; each detection module comprises a binocular camera and a light source, and the target detection network comprises a backbone structure, a text encoder, a neck structure, and a head structure; the method comprises: Determining target detection information; the target detection information includes target text information indicating a defect, and raw image data collected by the binocular camera inside the industrial equipment under the illumination of the light source; Inputting the original image data into the backbone structure to extract original image features at multiple scales; Inputting the target text information into the text encoder to extract original text features at multiple scales; Inputting a plurality of the original image features and a plurality of the original text features into the neck structure for interaction to generate a plurality of image-text fusion features; Inputting the plurality of image-text fusion features into the header structure to detect a target detection frame representing the defect; the target detection frame has a confidence level; If the confidence level is greater than or equal to a preset threshold, a brightening operation is performed on the original image data based on multiple frames of candidate image data to obtain target image data; the candidate image data is image data collected by the binocular camera inside the industrial equipment under the condition of controlling the light source to enhance illumination at multiple time steps; Constructing three-dimensional point cloud data based on the target detection frame in the target image data; detecting geometric information of the defect in the point cloud data; The brightening operation is performed on the original image data according to the multiple frames of candidate image data to obtain the target image data, including: extracting candidate brightness image data from the candidate image data of each frame respectively; Locating a target point for calibrating the light source in multiple frames of candidate brightness image data; An activation function is set for each frame of the candidate brightness image data according to the time step of the enhanced illumination of the light source; the activation function is a single-peak curve, and at an initial time step, the abscissa of the peak value of the activation function is 0, and the abscissa of the peak value of the activation function shifts toward the positive direction of the abscissa as the time step of the enhanced illumination of the light source increases; In the candidate brightness image data, calculating the distance between each pixel and the target point; Substituting the distance as the horizontal coordinate into the activation function and mapping it into a weight; performing a normalization operation on the plurality of weights for the pixel points located at the same position in the plurality of frames of the candidate brightness image data; If the normalization operation is completed, then for the pixel points located at the same position in the multiple frames of the candidate brightness image data, a weighted sum is performed on the pixel points using multiple weights to obtain the target brightness image data; Under the framework of the image enhancement model Retinex, the target brightness image data is fused into the original image data to obtain target image data.
2. The method according to claim 1, characterized in that The determining of target detection information includes: receiving original text information input by a user, performing word segmentation processing on the original text information to obtain a plurality of original phrases; Calculating the similarity between the original text information and the category representing the defect; If the similarity is within a preset first similarity interval, setting the category as target text information; If the similarity is within a preset second similarity interval, searching for the original phrase that matches the category as a first candidate phrase; identifying other original phrases that have a dependency relationship with the first candidate phrase as second candidate phrases; Searching for a qualifier phrase matching the second candidate word in a preset qualifier word library; setting the category matching the first candidate word group and the qualified word group matching the second candidate word group as target text information; If the similarity is within a preset third similarity interval, searching a preset restrictive word library for a restrictive word group that matches the original word group; Setting the restricted phrase that matches the original phrase as target text information; The value in the first similarity interval is greater than the value in the second similarity interval, and the value in the second similarity interval is greater than the value in the third similarity interval; The categories and the restricted vocabulary are both used to train the object detection network.
3. The method according to claim 1, characterized in that The text encoder includes a bidirectional encoder BERT, a first convolutional layer, a second convolutional layer, a third convolutional layer and a pyramid pooling module; the original text features include first text-level features, second text-level features, third text-level features and fourth text-level features; The step of inputting the target text information into the text encoder to extract original text features at multiple scales includes: Input the target text information into the bidirectional encoder BERT to extract one-dimensional text sequence features; Converting the text sequence features into two-dimensional text matrix features; Inputting the text matrix features into the first convolutional layer to extract first text-level features; Inputting the first text-level features into the second convolutional layer to extract second text-level features; Inputting the second text-level features into the third convolutional layer to extract third text-level features; The first text-level feature, the second text-level feature, and the third text-level feature are input into the pyramid pooling module to extract a fourth text-level feature.
4. The method according to claim 3, characterized in that The pyramid pooling module includes a first self-attention layer, a second self-attention layer, a third self-attention layer and a fourth convolutional layer; The step of inputting the first text-level feature, the second text-level feature, and the third text-level feature into the pyramid pooling module to extract the fourth text-level feature includes: Inputting the first text-level feature into the first self-attention layer and converting it into a first text-attention feature under self-attention; Inputting the second text-level features into the second self-attention layer and converting them into second text-attention features under self-attention; Inputting the third text-level feature into the third self-attention layer and converting it into a third text-attention feature under self-attention; splicing the first text attention feature, the second text attention feature and the third text attention feature into a fourth text attention feature; The fourth text attention feature is input into the fourth convolutional layer to extract the fourth text level feature.
5. The method according to claim 4, characterized in that The original image features include first image level features, second image level features and third image level features; The scale of the first image-level feature is greater than the scale of the second image-level feature, and the scale of the second image-level feature is greater than the scale of the third image-level feature; The neck structure includes the first attention layer, the second attention layer, the third attention layer, the fourth attention layer and the fifth attention layer, as well as 4 C2f modules and 2 convolution modules; The image-text fusion features include first image-text level features, second image-text level features and third image-text level features; The step of inputting the plurality of original image features and the plurality of original text features into the neck structure for interaction to generate a plurality of image-text fusion features includes: Inputting the first image-level feature and the first text-level feature into the first attention layer, and fusing the first text-level feature into the first image-level feature based on the attention of the first text-level feature to the first image-level feature, to obtain a first image-text candidate feature; Inputting the second image-level feature and the second text-level feature into the second attention layer, and fusing the second text-level feature into the second image-level feature based on the attention of the second text-level feature to the second image-level feature, to obtain a second image-text candidate feature; Inputting the third image-level feature and the third text-level feature into the third attention layer, and fusing the third text-level feature into the third image-level feature based on the attention of the third text-level feature to the third image-level feature, to obtain a third image-text candidate feature; Performing a first upsampling operation on the third image-text candidate feature; If the first upsampling operation is completed, the second image-text candidate feature and the third image-text candidate feature are concatenated into a fourth image-text candidate feature; Inputting the fourth image-text candidate feature into the C2f module to extract the fifth image-text candidate feature; Performing a second upsampling operation on the fifth image-text candidate feature; If the second upsampling operation is completed, the first image-text candidate feature and the fifth image-text candidate feature are concatenated into a sixth image-text candidate feature; Inputting the sixth image-text candidate feature into the C2f module to extract the first image-text level feature; Inputting the first image-text level feature into the convolution module to extract the seventh image-text candidate feature; Inputting the fifth image-text candidate feature and the fourth text-level feature into the fourth attention layer, and fusing the fourth text-level feature into the fifth image-text candidate feature based on the attention of the fourth text-level feature to the fifth image-text candidate feature, to obtain an eighth image-text candidate feature; splicing the seventh image-text candidate feature and the eighth image-text candidate feature into a ninth image-text candidate feature; Inputting the ninth image-text candidate feature into the C2f module to extract the second image-text level feature; Inputting the second image-text level feature into the convolution module to extract the tenth image-text candidate feature; splicing the third image-text candidate feature and the tenth image-text candidate feature into an eleventh image-text candidate feature; The eleventh image-text candidate feature is input into the C2f module to extract the third image-text level feature.
6. The method according to claim 1, characterized in that The detecting geometric information of the defect in the point cloud data includes: Projecting the point cloud data onto a horizontal plane; On a horizontal plane, a window is added to the current point cloud data; the current point cloud data is located at the center of the window; Calculating an average brightness value of other point cloud data located at non-central positions of the window to obtain an average brightness; Taking the absolute value of the difference between the brightness value of the current point cloud data and the average brightness of the other point cloud data as the brightness deviation value; For each of the point cloud data, mapping the brightness deviation value to a confidence level of the defect edge; wherein the confidence level is positively correlated with the brightness deviation value; The plurality of point cloud data with the highest confidence levels are selected to generate geometric information for the defect.
7. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the defect detection method based on the industrial endoscope is implemented as described in any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the defect detection method based on an industrial endoscope according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Natural language relation extraction method and device, electronic equipment and storage medium
CN114064852A
Pipeline defect identification and positioning method based on target detection and binocular vision
CN114067197A
Open vocabulary target detection method, system and device and medium
CN118673465A