Underground cable defect identification method and system based on multi-modal data

Through multimodal data fusion and deep learning technology, the problem that a single modal image is difficult to identify underground cable defects is solved, and higher recognition accuracy and robustness are achieved, reducing the maintenance cost of cable failures.

CN119992160APending Publication Date: 2025-05-13海南电力产业发展有限责任公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411963727.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art relies on single-modal image data in underground cable defect recognition, which is difficult to fully reflect the characteristics of cable defects, and lacks effective feature mapping and classification strategies, resulting in insufficient recognition accuracy.

Method used

The underground cable defect recognition method based on multimodal data is adopted. By collecting X-ray images, ultrasonic images and RGB color image data, pre-processing and manual annotation, a network model is constructed to extract the features of the multimodal image, and the model is trained through feature mapping and optimization loss function to achieve accurate identification of defect categories.

Benefits of technology

Through the fusion of multimodal data and deep learning technology, the accuracy and robustness of underground cable defect identification are significantly improved, the resistance to noise and complex environments is enhanced, and the losses and maintenance costs caused by cable failures are reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992160A_ABST
    Figure CN119992160A_ABST
Patent Text Reader

Abstract

The invention discloses an underground cable defect identification method and system based on multi-modal data, and relates to the technical field of underground cable defect detection and identification, and the method comprises the steps: collecting the multi-modal image data of an underground cable, carrying out the manual marking, determining the defect type of an image, forming a training data set, and constructing a network model through the processed data, feature extraction and feature coding are conducted on the multi-modal image through feature coding, a coding feature matrix set is obtained, the coding feature matrix set is mapped to the same projection space, mapping features are obtained, the mapping features are processed, the defect category is predicted, and a network model is trained by optimizing a loss function, so that the defect classification is predicted. And inputting a to-be-identified underground cable multi-modal image into the trained network model, and outputting a defect category identification result through the network model. The cable defect identification method adapts to the defect identification requirements of cables of different types and specifications, and brings remarkable economic benefits to power enterprises by reducing the loss of cable faults and the maintenance cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of underground cable defect detection and identification, and in particular to an underground cable defect identification method and system based on multimodal data. Background Art

[0002] With the acceleration of urbanization and the continuous improvement of infrastructure construction, underground cables, as an important part of power transmission, are crucial to the stability and reliability of urban power supply. However, due to the complexity and unpredictability of the underground environment, underground cables will inevitably have various defects and failures during use, such as aging, damage, and moisture of the insulation layer. These defects not only affect the transmission performance of the cable, but may also cause safety accidents, posing a serious threat to people's lives and property.

[0003] Traditional underground cable detection methods mainly rely on manual inspections and simple electrical tests, which are not only inefficient but also difficult to accurately identify tiny defects inside the cable. In recent years, with the rapid development of image processing technology and deep learning algorithms, image-based cable defect recognition methods have gradually become a research hotspot.

[0004] However, the existing related technologies still have many shortcomings in underground cable defect identification. First, most existing methods rely only on single-modal image data, such as using only X-ray images or ultrasonic images, which limits the accuracy and robustness of recognition. Due to the variety of underground cable defect types, single-modal images often cannot provide enough information to accurately identify all types of defects. Secondly, the existing image preprocessing and feature extraction methods are often relatively simple and fail to effectively extract the essential features that can characterize cable defects, resulting in poor recognition results. In addition, the existing technology does not perform in-depth fusion processing of multimodal data, lacks effective feature mapping and classification strategies, and the accuracy of defect recognition needs to be improved. Summary of the invention

[0005] In view of the above-mentioned existing problems, the present invention provides an underground cable defect identification method and system based on multimodal data to solve the problem that a single-modal image is difficult to fully reflect the defect characteristics of the cable and lacks effective feature mapping and classification strategies.

[0006] In order to solve the above technical problems, a method for underground cable defect identification based on multimodal data is proposed, including:

[0007] Collect multimodal image data of underground cables, manually annotate the collected image data, determine the defect category of the image, and form a training data set; use the processed data to build a network model, use feature coding to extract and encode features of the multimodal image, and obtain a coding feature matrix group; map the coding feature matrix group to the same projection space to obtain mapping features, and process the mapping features to predict the defect category, and train the network model by optimizing the loss function; input the multimodal image of the underground cable to be identified into the trained network model, and output the defect category identification result through the network model.

[0008] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the collecting of multimodal image data of underground cables includes collecting X-ray image data, ultrasonic image data and RGB color image data, and preprocessing the collected data;

[0009] The preprocessing includes removing blur and overexposed images, enhancing image contrast, denoising and sharpening, and performing image cropping and scaling.

[0010] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the forming of the training data set includes manually annotating the pre-processed image, determining the defect category of the image, and forming the training data set;

[0011] The manual annotation includes selecting an image annotation tool, defining defects, formulating annotation methods and category classification standards, assigning category labels to defects, recording the size, shape and location of the defects, and saving annotation data; the determination of the defect category of the image includes determining the defect category of the image according to the defined defect category classification standard, integrating the annotated image and annotation information into a data set, and dividing the data set into a training set, a validation set and a test set.

[0012] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the construction of the network model includes selecting a network architecture according to task requirements, designing a multimodal fusion layer, defining a loss function and determining an optimizer;

[0013] The feature extraction includes initializing an input X-ray image of an underground cable, the image sequentially passes through three feature extraction submodules, the first submodule uses a convolution layer of size 13×13 to extract large-scale features of the image and obtain an output feature map, the second submodule uses a convolution layer of size 9×9 to extract medium-scale features and obtain an output feature map, the third submodule uses a convolution layer of size 5×5 to extract small-scale features and obtain an output feature map, and the output feature maps of the three submodules are cascaded to merge into a cascaded feature map, and further convolution operations are performed to output coded features, and all modal images are subjected to feature extraction to obtain a coded feature matrix group;

[0014] The formula contents of the three feature extraction submodules are:

[0015]

[0016] in, is the feature map output by the first feature extraction submodule, Cat(·) is the cascade operation, I X is the input image, C 13 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 13×13, C 11 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 11×11, C 7 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 11×11. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 3×3. It is the feature map output by the second feature extraction submodule. The feature map output by the third feature extraction submodule;

[0017] The further convolution formula is:

[0018]

[0019] Among them, F X It is the feature map finally output by the feature extraction module. It is the feature map output by the first feature extraction submodule. It is the feature map output by the second feature extraction submodule. is the feature map output by the third feature extraction submodule, C 1 The convolution layer with a convolution kernel size of 1×1 is used for feature extraction, C 3 Feature extraction is performed for the convolution layer with a convolution kernel size of 3×3.

[0020] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the obtaining of mapping features includes three fully connected layers and one activation function layer, and performing feature projection on the encoding feature matrix group, mapping it to the same feature space, obtaining mapping features, and performing feature screening through the activation function layer;

[0021] The three fully connected layer formulas are expressed as:

[0022] F y =R(W 1 F z +W 2 F C +W 3 F RGB )

[0023] Among them, F y is the feature vector output by the feature mapping block, R(·) is the activation function, and W 1 , W 2 and W 3 is the weight matrix of the fully connected layer, F z 、F C and F RGB are the encoding features of the three types of image data.

[0024] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the prediction of defect categories includes performing feature processing on mapping features to obtain defect classification identification information, using a network model to predict multimodal image samples to obtain the predicted probability that each sample belongs to each defect type;

[0025] The training network model includes using a cross entropy function to constrain the prediction results of the network model, calculating the gradient value of the loss function with respect to the model parameters, updating the parameters using a gradient descent method, adjusting the weights and biases of the network model according to the calculated gradient, and continuously repeating the loss calculation and parameter updating process until a preset number of training rounds is reached, and calculating the accuracy of the trained network model using a validation set;

[0026] The cross entropy function includes

[0027]

[0028] Where M is the total number of defect types, y ij is the true value of the jth defect of the i-th image sample, is the model prediction result, L(θ) is the loss function, N is the total number of image samples, θ is the parameter of the loss function with respect to the model, and i and j are variable indices.

[0029] As a preferred solution of the underground cable defect identification method based on multimodal data described in the present invention, wherein: the output defect category identification result includes inputting the multimodal image of the underground cable to be identified into the trained network model, and outputting the defect category identification result through the convolution layer, activation function layer, BN normalization layer and full connection layer of the network model;

[0030] The output defect category recognition result also includes a convolution layer receiving input data and performing convolution, passing the convolved data to an activation function layer, activating the data using a ReLU activation function, performing batch normalization on the activated data in a BN normalization layer, and outputting the defect category recognition result through a fully connected layer.

[0031] Another object of the present invention is to provide an underground cable defect identification system based on multimodal data. The present invention timely discovers and handles potential fault hazards by learning the correlation and complementarity between different modal images; the system of the present invention can reveal the characteristics of cable defects from different angles and depths by fusing multimodal image information of X-ray images, ultrasonic images and RGB color images, significantly improving the recognition accuracy, and utilizing multimodal data fusion to enhance the system's resistance to noise, occlusion and incomplete information, thereby improving robustness and ensuring stable defect identification in complex environments. It adapts to the defect identification needs of cables of different types and specifications, and reduces the losses and maintenance costs caused by cable failures, thereby bringing significant economic benefits to power companies.

[0032] As a preferred solution of the underground cable defect identification system based on multimodal data described in the present invention, it is characterized by including a data collection and preprocessing module, a training data set construction module, a network model construction and feature extraction module, a feature mapping and defect prediction module, and a model training and result output module.

[0033] The data collection and preprocessing module is used to collect multimodal image data of underground cables and perform preprocessing.

[0034] The training data set construction module is used to manually annotate the preprocessed images, determine the defect categories of the images, and divide the data set into a training set, a verification set, and a test set.

[0035] The network model construction and feature extraction module is used to design the network model, perform feature extraction and encoding, select the network architecture, define the loss function and determine the optimizer, extract the large, medium and small scale features of the image through convolution layers of different sizes, cascade and further convolute the extracted feature maps, and output the encoded features.

[0036] The feature mapping and defect prediction module is used to map the encoding feature matrix group to the same projection space, process the mapping features, and predict the defect category.

[0037] The model training and result output module is used to use the cross entropy function to constrain the prediction results of the network model, calculate the gradient value of the loss function with respect to the model parameters, and use the gradient descent method to update the parameters until the model converges. The multimodal image of the underground cable to be identified is input into the trained network model, and the defect category identification results are output through each layer of the network model.

[0038] A computer device includes a memory and a processor, wherein the memory stores a computer program, and is characterized in that when the processor executes the computer program, the steps of a method for identifying underground cable defects based on multimodal data are implemented.

[0039] A computer-readable storage medium stores a computer program thereon, characterized in that when the computer program is executed by a processor, the steps of a method for identifying underground cable defects based on multimodal data are implemented.

[0040] Beneficial effects of the present invention: The present invention defines the image defect categories and constructs a training data set by collecting and manually annotating multimodal image data of underground cables, thereby providing high-quality annotated data for network model training, ensuring that the model learns accurate defect features and improving the accuracy and reliability of defect recognition; preprocessing improves data quality, reduces the impact of noise and irrelevant information on model training, and enhances the generalization ability and robustness of the model; fine-grained defect feature division and standardized data sets provide clear training goals for the network model, which helps the model learn defect features more effectively; a network model is constructed and feature extraction and encoding are performed, and convolutional layers and cascades of different scales are used to extract and encode defect features. The multimodal image features are deeply mined and integrated to provide rich feature representation for defect recognition. By mapping the encoded feature matrix group to the same projection space, the fusion and dimensionality reduction of multimodal features are realized, thereby improving the accuracy of defect recognition. By predicting the defect category and optimizing the loss function to train the network model, the cross entropy function and gradient descent method are used to accurately adjust the model parameters to achieve high-accuracy defect classification. The multimodal image to be identified is input into the trained network model, and through end-to-end processing of convolutional layer, ReLU activation function layer, BN normalization layer and fully connected layer, the nonlinear processing capability and training stability of the network are improved, thereby achieving efficient and accurate defect recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work, among which:

[0042] Figure 1 An overall flow chart of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0043] Figure 2 A schematic diagram of the structure of a network model of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0044] Figure 3 A schematic diagram of the structure of a feature encoding module of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0045] Figure 4 A schematic diagram of the structure of a feature extraction submodule of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0046] Figure 5A schematic diagram of the structure of a feature alignment mapping module of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0047] Figure 6 A schematic structural diagram of a defect classification module of an underground cable defect identification method based on multimodal data provided by an embodiment of the present invention.

[0048] Figure 7 A system solution flow chart of an underground cable defect identification system based on multimodal data provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the above-mentioned purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work should fall within the scope of protection of the present invention.

[0050] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0051] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it an embodiment that is mutually exclusive with other embodiments, either individually or selectively.

[0052] The present invention is described in detail with reference to schematic diagrams. When describing the embodiments of the present invention, for the sake of convenience, the cross-sectional diagrams showing the device structure will not be partially enlarged according to the general scale, and the schematic diagrams are only examples, which should not limit the scope of protection of the present invention. In addition, in actual production, the three-dimensional dimensions of length, width and depth should be included.

[0053] At the same time, in the description of the present invention, it should be noted that the directions or positional relationships indicated by the terms "upper, lower, inner and outer" are based on the directions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific direction, be constructed and operated in a specific direction, and therefore cannot be understood as limiting the present invention. In addition, the terms "first, second or third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0054] In the present invention, unless otherwise clearly specified and limited, the terms "install, connect, connect" should be understood in a broad sense, for example: it can be a fixed connection, a detachable connection or an integral connection; it can also be a mechanical connection, an electrical connection or a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0055] Example 1, reference Figure 1-Figure 6 , which is the first embodiment of the present invention, provides an underground cable defect identification method based on multimodal data, comprising:

[0056] S1: Collect multimodal image data of underground cables, manually annotate the collected image data, determine the defect category of the image, and form a training data set.

[0057] The collecting of multimodal image data of underground cables includes collecting X-ray image data, ultrasonic image data and RGB color image data, and preprocessing the collected data;

[0058] The preprocessing includes removing blur and overexposed images, enhancing image contrast, denoising and sharpening, and performing image cropping and scaling.

[0059] It should be noted that the forming of the training data set includes manually annotating the pre-processed images, determining the defect categories of the images, and forming the training data set;

[0060] The manual annotation includes selecting an image annotation tool, defining defects, formulating annotation methods and category classification standards, assigning category labels to defects, recording the size, shape and location of the defects, and saving annotation data; the determination of the defect category of the image includes determining the defect category of the image according to the defined defect category classification standard, integrating the annotated image and annotation information into a data set, and dividing the data set into a training set, a validation set and a test set.

[0061] S2: Use the processed data to build a network model, use feature coding to extract and encode features of the multimodal image, and obtain a coding feature matrix group.

[0062] Furthermore, the construction of the network model includes selecting a network architecture according to task requirements, designing a multimodal fusion layer, defining a loss function and determining an optimizer;

[0063] like Figure 2 ,The constructed network model includes three feature encoding modules, a feature alignment mapping module, and a defect classification module; Figure 3 and Figure 4, the structure of each feature encoding module is consistent, except for the input content; at the same time, the structure of each feature extraction submodule is similar, while the convolution kernel size of the convolution layer is different. For example, convolution layer C13 represents a convolution layer with a convolution kernel size of 13*13, and so on. The feature extraction includes initializing the input X-ray image of the underground cable, and the image passes through three feature extraction submodules in sequence. The first submodule uses a convolution layer of size 13×13 to extract large-scale features of the image and obtain an output feature map. The second submodule uses a 9×9 convolution layer to extract medium-scale features and obtain an output feature map. The third submodule uses a 5×5 convolution layer to extract small-scale features and obtain an output feature map. The output feature maps of the three submodules are cascaded and merged into a cascaded feature map, and further convolution operations are performed to output coding features, and all modal images are feature extracted to obtain a coding feature matrix group;

[0064] The formula contents of the three feature extraction submodules are:

[0065]

[0066] in, is the feature map output by the first feature extraction submodule, Cat(·) is the cascade operation, I X is the input image, C 13 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 13×13, C 11 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 11×11, C 7 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 11×11. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 3×3. It is the feature map output by the second feature extraction submodule. The feature map output by the third feature extraction submodule;

[0067] The further convolution formula is:

[0068]

[0069] Among them, F X It is the feature map finally output by the feature extraction module. It is the feature map output by the first feature extraction submodule. It is the feature map output by the second feature extraction submodule. is the feature map output by the third feature extraction submodule, C 1 The convolution layer with a convolution kernel size of 1×1 is used for feature extraction, C 3 Feature extraction is performed for the convolution layer with a convolution kernel size of 3×3.

[0070] S3: Map the encoding feature matrix group to the same projection space, obtain the mapping features, process the mapping features, predict the defect category, and train the network model by optimizing the loss function.

[0071] Further, such as Figure 5 ,The feature alignment mapping module includes ,three fully connected layers and an activation function layer, and ,feature projection is performed on the encoded feature matrix group ,mapped to the same feature space, the mapping features are obtained, and ,feature screening is performed through the activation function layer;

[0072] The three fully connected layer formulas are expressed as:

[0073] F y =R(W 1 F z +W 2 F C +W 3 F RGB )

[0074] Among them, F y is the feature vector output by the feature mapping block, R(·) is the activation function, and W 1 , W 2 and W 3 is the weight matrix of the fully connected layer, F z 、F C and F RGB are the encoding features of the three types of image data.

[0075] Furthermore, the predicting defect category includes performing feature processing on the mapping features to obtain defect classification identification information, and using the network model to predict the multimodal image samples to obtain the predicted probability that each sample belongs to each defect type;

[0076] The training network model includes using a cross entropy function to constrain the prediction results of the network model, calculating the gradient value of the loss function with respect to the model parameters, updating the parameters using a gradient descent method, adjusting the weights and biases of the network model according to the calculated gradient, and continuously repeating the loss calculation and parameter updating process until a preset number of training rounds is reached, and calculating the accuracy of the trained network model using a validation set;

[0077] The cross entropy function includes

[0078]

[0079] Where M is the total number of defect types, y ij is the true value of the jth defect of the i-th image sample, is the model prediction result, L(θ) is the loss function, N is the total number of image samples, θ is the parameter of the loss function with respect to the model, and i and j are variable indices;

[0080] The gradient descent method formula is:

[0081]

[0082] Among them, α is the learning rate, θ is the parameter of the current loss function about the model, and θ new is the updated loss function about the model parameters, is the gradient of the loss function with respect to the model parameters;

[0083] The formula for calculating accuracy is:

[0084]

[0085] Among them, Accuracy is the accuracy, N is the total number of image samples, y ij is the true value of the jth defect of the i-th image sample, is the model prediction result, l(·) is the indicator function, and i and j are variable indices.

[0086] S4: Input the multimodal image of the underground cable to be identified into the trained network model, and output the defect category identification result through the network model.

[0087] It should be noted that if Figure 6 ,The defect classification module includes inputting the multimodal image of the underground cable to be identified into the ,trained network model, and outputting the defect category identification ,result through the convolution layer, activation function layer, BN normalization layer and ,fully connected layer of the network model;

[0088] The output defect category recognition result also includes a convolution layer receiving input data and performing convolution, passing the convolved data to an activation function layer, activating the data using a ReLU activation function, performing batch normalization on the activated data in a BN normalization layer, and outputting the defect category recognition result through a fully connected layer;

[0089] The ReLU activation function is expressed as:

[0090] F relu =max(0,F conv )

[0091] Among them, F relu is the output of the activation function, F conv is the output of the convolutional layer.

[0092] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0093] Example 2, reference Figure 7 , which is the second embodiment of the present invention, provides an underground cable defect identification system based on multimodal data, including a data collection and preprocessing module 100, a training data set construction module 200, a network model construction and feature extraction module 300, a feature mapping and defect prediction module 400 and a model training and result output module 500.

[0094] The data collection and preprocessing module 100 is used to collect multimodal image data of underground cables and perform preprocessing.

[0095] The training data set construction module 200 is used to manually annotate the preprocessed images, determine the defect categories of the images, and divide the data set into a training set, a validation set, and a test set.

[0096] The network model construction and feature extraction module 300 is used to design the network model, perform feature extraction and encoding, select the network architecture, define the loss function and determine the optimizer, extract the large, medium and small scale features of the image through convolution layers of different sizes, cascade and further convolute the extracted feature maps, and output the encoded features.

[0097] The feature mapping and defect prediction module 400 is used to map the encoding feature matrix group to the same projection space, process the mapping features, and predict the defect category.

[0098] The model training and result output module 500 is used to use the cross entropy function to constrain the prediction results of the network model, calculate the gradient value of the loss function with respect to the model parameters, and use the gradient descent method to update the parameters until the model converges. The multimodal image of the underground cable to be identified is input into the trained network model, and the defect category identification results are output through each layer of the network model.

[0099] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

[0100] Embodiment 3, the third embodiment of the present invention, is different from the first two embodiments in that:

[0101] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program codes.

[0102] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0103] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0104] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

Claims

1. A method for identifying underground cable defects based on multimodal data, characterized in that: include, Collect multimodal image data of underground cables, manually annotate the collected image data, determine the defect categories of the images, and form a training data set; The processed data is used to construct a network model, and feature extraction and feature encoding are performed on the multimodal image using feature encoding to obtain a coding feature matrix group; Map the encoding feature matrix group to the same projection space, obtain the mapping features, process the mapping features, predict the defect category, and train the network model by optimizing the loss function; The multimodal image of the underground cable to be identified is input into the trained network model, and the defect category identification result is output through the network model.

2. The underground cable defect identification method based on multimodal data according to claim 1, characterized in that: The collecting of multimodal image data of underground cables includes collecting X-ray image data, ultrasonic image data and RGB color image data, and preprocessing the collected data; The preprocessing includes removing blur and overexposed images, enhancing image contrast, denoising and sharpening, and performing image cropping and scaling.

3. The underground cable defect identification method based on multimodal data according to claim 2, characterized in that: The forming of the training data set includes manually labeling the preprocessed images, determining the defect categories of the images, and forming the training data set; The manual annotation includes selecting an image annotation tool, defining defects, formulating annotation methods and category classification standards, assigning category labels to defects, recording the size, shape and location of the defects, and saving annotation data; the determination of the defect category of the image includes determining the defect category of the image according to the defined defect category classification standard, integrating the annotated image and annotation information into a data set, and dividing the data set into a training set, a validation set and a test set.

4. The underground cable defect identification method based on multimodal data according to claim 3, characterized in that: The construction of the network model includes selecting a network architecture according to task requirements, designing a multimodal fusion layer, defining a loss function and determining an optimizer; The feature extraction includes initializing an input X-ray image of an underground cable, the image sequentially passes through three feature extraction submodules, the first submodule uses a convolution layer of size 13×13 to extract large-scale features of the image and obtain an output feature map, the second submodule uses a convolution layer of size 9×9 to extract medium-scale features and obtain an output feature map, the third submodule uses a convolution layer of size 5×5 to extract small-scale features and obtain an output feature map, and the output feature maps of the three submodules are cascaded to merge into a cascaded feature map, and further convolution operations are performed to output coded features, and all modal images are subjected to feature extraction to obtain a coded feature matrix group; The formula contents of the three feature extraction submodules are: in, is the feature map output by the first feature extraction submodule, Cat(·) is the cascade operation, I X is the input image, C 13 (I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 13×13, C 11 (I X ) is the feature map obtained by the convolution layer with a convolution kernel size of 11×11, C7(I X ) is the feature map obtained through the convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 11×11. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the first feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 7×7. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 5×5. The feature map output by the second feature extraction submodule is obtained by passing through a convolution layer with a convolution kernel size of 3×3. It is the feature map output by the second feature extraction submodule. The feature map output by the third feature extraction submodule; The further convolution formula is: Among them, FX is the feature map finally output by the feature extraction module. It is the feature map output by the first feature extraction submodule. It is the feature map output by the second feature extraction submodule. It is the feature map output by the third feature extraction submodule. C1 is the convolution layer with a convolution kernel size of 1×1 for feature extraction, and C3 is the convolution layer with a convolution kernel size of 3×3 for feature extraction.

5. The underground cable defect identification method based on multimodal data according to claim 4, characterized in that: The mapping feature acquisition includes three fully connected layers and one activation function layer, and feature projection is performed on the encoding feature matrix group, mapped to the same feature space, the mapping feature is acquired, and feature screening is performed through the activation function layer; The three fully connected layer formulas are expressed as: F y =R(W1F z +W2F C +W3F RGB ) Among them, F y is the feature vector output by the feature map block, R(·) is the activation function, W1, W2 and W3 are the weight matrices of the fully connected layer, and F z 、F C and F RGB are the encoding features of the three types of image data.

6. The underground cable defect identification method based on multimodal data according to claim 5, characterized in that: The predicting defect category includes performing feature processing on the mapping features to obtain defect classification identification information, and using the network model to predict the multimodal image samples to obtain the predicted probability that each sample belongs to each defect type; The training network model includes using a cross entropy function to constrain the prediction results of the network model, calculating the gradient value of the loss function with respect to the model parameters, updating the parameters using a gradient descent method, adjusting the weights and biases of the network model according to the calculated gradient, and continuously repeating the loss calculation and parameter updating process until a preset number of training rounds is reached, and calculating the accuracy of the trained network model using a validation set; The cross entropy function includes Where M is the total number of defect types, y ij is the true value of the jth defect of the i-th image sample, is the model prediction result, L ( θ ) is the loss function, N is the total number of image samples, θ is the parameter of the loss function with respect to the model, and i and j are variable indices.

7. The underground cable defect identification method based on multimodal data according to claim 6, characterized in that: The outputting of the defect category recognition result comprises inputting the multimodal image of the underground cable to be identified into the trained network model, and outputting the defect category recognition result through the convolution layer, activation function layer, BN normalization layer and full connection layer of the network model; The output defect category recognition result also includes a convolution layer receiving input data and performing convolution, passing the convolved data to an activation function layer, activating the data using a ReLU activation function, performing batch normalization on the activated data in a BN normalization layer, and outputting the defect category recognition result through a fully connected layer.

8. A system using the underground cable defect identification method based on multimodal data as claimed in any one of claims 1 to 7, characterized in that: It includes data collection and preprocessing module, training data set construction module, network model construction and feature extraction module, feature mapping and defect prediction module, and model training and result output module; The data collection and preprocessing module is used to collect multimodal image data of underground cables and perform preprocessing; The training data set construction module is used to manually annotate the preprocessed images, determine the defect categories of the images, and divide the data set into a training set, a validation set, and a test set; The network model construction and feature extraction module is used to design the network model, perform feature extraction and encoding, select the network architecture, define the loss function and determine the optimizer, extract the large, medium and small scale features of the image through convolutional layers of different sizes, cascade and further convolve the extracted feature maps, and output the encoded features; The feature mapping and defect prediction module is used to map the encoding feature matrix group to the same projection space, process the mapping features, and predict the defect category; The model training and result output module is used to use the cross entropy function to constrain the prediction results of the network model, calculate the gradient value of the loss function with respect to the model parameters, and use the gradient descent method to update the parameters until the model converges. The multimodal image of the underground cable to be identified is input into the trained network model, and the defect category identification results are output through each layer of the network model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a method for identifying underground cable defects based on multimodal data described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a method for identifying underground cable defects based on multimodal data described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Robot vision detection method and system based on image enhancement

    CN120563475A

  • Cable surface defect accurate identification method and device for cable cutting machine

    CN120890974A