Visual detection method for suspension state defects of high-speed rail overhead line system in open environment
By combining pre-trained defect detection models, improved GAN and regional attention mechanisms, as well as CLIP models and language generation models, the problems of scarcity of data and diversity of defect types in suspension state defect detection in high-speed rail contact network are solved, and efficient and accurate defect detection and visual marking are achieved.
Patent Information
- Application Number
- CN202411913110.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
There are problems in the detection of suspension status defects of high-speed rail contact networks, such as scarce data, diverse and complex defect types, and insufficient modal alignment, resulting in low detection accuracy and low efficiency.
The pre-trained defect detection model is used to combine an improved generative adversarial network (GAN) and regional attention mechanism to detect images to be detected of the high-speed rail contact network suspension equipment. Through the CLIP model and language generation model, image features are aligned with text features to achieve accurate identification and position marking of defect types.
It improves the accuracy and efficiency of defect detection of high-speed rail contact network suspension equipment, can effectively detect in zero samples or sparse samples, and improves the interpretability and practicality of the detection through visual marking.
Smart Images

Figure CN119941640A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for visually detecting suspension state defects of a high-speed railway contact network in an open environment, and belongs to the technical field of target detection. Background Art
[0002] As an important part of the modern transportation network, high-speed rail undertakes a large number of passenger and freight tasks, and its operation safety is directly related to the safety of passengers' lives and property. The contact network is one of the key facilities of high-speed electrified railways, responsible for providing a stable power supply for electric locomotives. The suspension state of the contact network and its abnormal changes may lead to unstable power transmission, thus affecting the normal operation of the train, and may even cause electrical failures and traffic accidents. Therefore, health monitoring and fault diagnosis of the contact network are of great significance to the safe operation of high-speed rail.
[0003] At present, the detection methods of contact network mainly rely on manual inspection, vehicle-mounted inspection and drone photography. Although manual inspection can make an intuitive judgment on the condition of the contact network, it is inefficient and easily affected by factors such as weather and environment. It cannot meet the actual needs of long high-speed rail lines, intensive operations, and frequent maintenance. Vehicle-mounted inspection monitors the status of the contact network through vehicle-mounted equipment, but its cost is high and the accuracy of the test results is limited by the performance of the equipment and the environment. Although drone photography can improve inspection efficiency, its application in large-scale, high-speed, and high-frequency inspections still has certain limitations, especially in complex environments and severe weather conditions. The clarity and accuracy of the image may be affected, which in turn affects the recognition effect of defects.
[0004] With the rapid development of artificial intelligence (AI) technology, especially the widespread application of computer vision and deep learning, automated defect detection technology has gradually become an important direction for high-speed rail contact network monitoring. Image processing methods based on computer vision can quickly and efficiently detect the suspension state of the contact network, thereby timely discovering anomalies and defects and avoiding the shortcomings of manual inspections. However, existing image processing methods still face some challenges in practical applications, which are mainly reflected in the following aspects:
[0005] Data scarcity: Image datasets of high-speed rail contact network suspension defects are usually scarce, especially some special types of defect samples are extremely rare. Due to the lack of sufficient labeled data, it is difficult to train deep learning models, resulting in insufficient generalization ability of the model, affecting the accuracy and reliability of defect detection.
[0006] Defects are diverse and complex: The abnormal suspension state of the high-speed rail contact network may include various types of defects, such as split pin defects, broken insulators, loose pipe caps, etc. Different types of defect images may have similar features. How to effectively distinguish these subtle defect differences is one of the challenges faced by existing technologies.
[0007] Modal alignment problem: In multimodal learning, the fusion and alignment of image and text descriptions is an important issue. Existing methods often process image and text information separately and lack an effective modal alignment mechanism, which results in the inability to fully combine the features of images and text when processing complex scenes, affecting the accurate detection and description of defects. Summary of the invention
[0008] The purpose of the present invention is to provide a visual detection method for high-speed railway contact network suspension state defects in an open environment. Aiming at the situations of zero samples or sparse samples and low detection accuracy in the detection of high-speed railway contact network suspension state defects, artificial intelligence and deep learning technologies are used to solve the problems of data scarcity, defect diversity, poor data enhancement effect and insufficient alignment of image and text modalities in the detection of high-speed railway contact network suspension state defects.
[0009] In order to solve the above technical problems, the present invention is implemented by adopting the following technical solutions.
[0010] The present invention provides a method for visually detecting suspension state defects of a high-speed railway contact network in an open environment, comprising:
[0011] According to the acquired image to be inspected of the high-speed railway overhead line suspension equipment, based on the pre-trained defect detection model, the inspection is performed to obtain the inspection result, including the image to be inspected with the defect position marked;
[0012] Wherein, the training method of the defect detection model includes:
[0013] Feature extraction is performed on the acquired real image of the defects of the high-speed railway overhead line suspension equipment to obtain sub-images of different defect types;
[0014] Improve the generative adversarial network (GAN) and introduce regional attention to expand the sub-graphs of different defect types to obtain the generated images of defects in the high-speed rail contact network suspension equipment;
[0015] According to the language generation model and the natural language description of the types of suspension state defects of the high-speed railway overhead line suspension equipment, a text description library is constructed to obtain the text description of the defects of the high-speed railway overhead line suspension equipment;
[0016] Based on the CLIP model, the real image and the generated image are fed into an image encoder, and the text description is fed into a text encoder to extract image features and text features;
[0017] Based on the attention mechanism, the image features and the text features are aligned, the similarity between the image features and the text features is calculated, the defect type is determined according to the similarity, and the defect type is mapped to the real image and the defect position is marked.
[0018] Furthermore, the acquired real images of defects of the high-speed railway contact network suspension equipment include real images of defects of the high-speed railway contact network suspension equipment in normal state and abnormal state.
[0019] Furthermore, feature extraction is performed on the acquired real image of the defects of the high-speed railway overhead line suspension equipment to obtain sub-images of different defect types, including:
[0020] The acquired real image of the defects of the high-speed railway overhead line suspension equipment is sent to the ROI region extraction network to perform feature extraction at different levels to obtain defect region candidate frames of different sizes;
[0021] Based on defect area candidate boxes of different sizes, the non-maximum suppression method is used to save the candidate boxes whose confidence exceeds the threshold, and the corresponding area of the real image is intercepted and saved according to the coordinates of the candidate boxes exceeding the threshold to obtain sub-images of different defect types.
[0022] Furthermore, the generative adversarial network (GAN) is improved and regional attention is introduced to expand the sub-graphs of different defect types, and the generated images of defects of high-speed rail contact network suspension equipment are obtained, including:
[0023] The regional attention module and regional attention mechanism are introduced into the generator and discriminator of the generative adversarial network (GAN) respectively to obtain the improved generative adversarial network (GAN).
[0024] Furthermore, the loss function of the improved generative adversarial network GAN is expressed as:
[0025] ;
[0026] In the formula, Represents the loss function of the improved generative adversarial network GAN, represents the loss function of the GAN before improvement, represents the hyperparameter that controls the diversity regularization weight, represents the diversity regularization loss function;
[0027] Among them, the loss function of the improved generative adversarial network GAN is , expressed as:
[0028] L GAN = <m> min < / m> <m> max < / m> E x ~ pdata [ logD ( x ) ] + E z ~ pz [ log(1-D ( G ( z ))) ] ;
[0029] In the formula, Denotes the discriminator D training generator Minimize the loss function when Denotes the generator G training discriminator When maximizing the loss function, From real data The actual data distribution The expected value of the real data sampled from represents the logarithmic function, Represents real data is the probability of a real image, Representation Discriminator The logarithmic expected value of the probability that the generated sample is false, where is the noise distribution, Represents the first noise sampled in the noise distribution The generated image by generator G, Represents the probability that the image generated by the generator G is classified as true by the discriminator D;
[0030] Among them, the diversity regularization loss function , expressed as
[0031] L div =− E z 1 , z 2 ~ pz [ dist ( G ( z 1 ), G ( z 2 )) ] ;
[0032] In the formula, represents the expected value between the output results of two random noises generated by the generator G, where are the second and third noise in the noise distribution, respectively. Represents the measure of the second noise Generated image after generator G and the third noise Generated image after generator G The function of the difference between , represents the dot product operation, Represents the modulus length between the two.
[0033] Furthermore, it also includes classifying and saving the generated images of defects of the high-speed railway contact network suspension equipment according to the defect types and labels corresponding to the defect types.
[0034] Furthermore, according to the language generation model and the natural language description of the types of suspension state defects of the high-speed railway contact network suspension equipment, a text description library is constructed to obtain the text description of the defects of the high-speed railway contact network suspension equipment, including:
[0035] According to the types of suspension state defects of high-speed rail overhead contact network suspension equipment, write a natural language description of each defect type;
[0036] Based on the natural language description of each defect type, a language generation model is used to generate diverse sentences;
[0037] Building a text description library based on the diversified sentences and labels corresponding to the defect types, and obtaining a text description of the defects of the high-speed railway overhead line suspension equipment;
[0038] The types of suspension state defects of the high-speed railway overhead line suspension equipment include cotter pin defects, nut defects and loose pipe caps;
[0039] The natural language description of the suspension state defect type of the high-speed railway contact network suspension equipment includes physical characteristics and physical forms.
[0040] Furthermore, the image features and the text features are aligned based on the attention mechanism, the similarity between the image features and the text features is calculated, the defect type is determined according to the similarity, and the defect type is mapped to the real image and the defect position is marked, including:
[0041] (1) Embed image features and text features into the same space and convert them into query vectors, key vectors, and value vectors as output features to be passed to the attention mechanism. Use multiple independent attention heads to process the query vectors, key vectors, and value vectors of each pair of image features and text features in parallel to obtain the outputs of all attention heads.
[0042] (2) Calculate the dot product between the query vector and the key vector of the image feature and the text feature to obtain the similarity score and normalize it to obtain the cross-modal feature association weight of the image feature and the text feature;
[0043] (3) Based on the cross-modal feature association weights of image features and text features, the outputs of all attention heads are concatenated or aggregated using the weighted average method to align image features and text features;
[0044] (4) Calculate the similarity between each pair of image features and text features through the contrast loss function, determine the text feature that best matches the image feature based on the similarity, determine the defect type based on the text description corresponding to the text feature, map it to the real image and mark the defect location.
[0045] Furthermore, the calculation expression of the dot product between the query vector and the key vector of the image feature and the text feature is:
[0046] ;
[0047] In the formula, represents the similarity score, represents the query vector, Key Vector The transposed vector of .
[0048] Furthermore, the contrast loss function is expressed as:
[0049] ;
[0050] In the formula, represents the contrast loss function, represents the logarithmic function, represents the exponential function with base e, It is For image features and text features corresponding to real images, Represents the total number of pairs of image features and text features corresponding to the real image, represents the temperature parameter, represents all text features corresponding to real images, Represents cosine similarity.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] 1. The present invention detects the images to be detected of the high-speed railway contact network suspension equipment by using a pre-trained defect detection model. The training process of the defect detection model not only extracts and expands features based on the real images of the defects of the high-speed railway contact network suspension equipment, but also introduces the generative adversarial network GAN improved by regional attention to enhance the diversity and accuracy of the defect images. At the same time, combined with the language generation model and the CLIP model, the image features are aligned with the text features, and the accurate identification and position marking of the defect types of the high-speed railway contact network suspension equipment are realized, which not only improves the accuracy and efficiency of defect detection, but also can automatically generate detection images containing defect position marks, providing intuitive and convenient information support for subsequent maintenance and management, and significantly improving the safety and reliability of the high-speed railway contact network suspension equipment.
[0053] 2. By optimizing the loss function of GAN and introducing the regional attention mechanism, the present invention can expand the subgraphs of different defect types under limited defect data, obtain the generated images of defects of high-speed rail contact network suspension equipment, significantly expand the sample size and distribution, and can effectively deal with zero sample or sparse sample scenarios;
[0054] 3. The present invention also sends the acquired real image of the high-speed railway overhead line suspension equipment defects into the ROI region extraction network to perform feature extraction at different levels, capture the details of different target scales, accurately extract defect region candidate frames of different sizes and effectively intercept the corresponding area of the real image, thereby realizing comprehensive detection from global to local;
[0055] 4. The present invention also aligns image features with text features based on the CLIP model and attention mechanism, solves the similarity problem between categories through the attention mechanism, realizes efficient defect detection under zero-sample conditions, and improves the interpretability and practicality of detection through visual marking. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 The figure is a flow chart of a method for visually detecting suspension state defects of a high-speed railway contact network in an open environment provided by an embodiment of the present invention;
[0057] Figure 2 The figure is a schematic diagram of a process of defect detection on a sample to be tested provided by an embodiment of the present invention;
[0058] Figure 3 The figure is a schematic diagram of a process of extracting hierarchical features from a ROI region provided by an embodiment of the present invention;
[0059] Figure 4 The figure shows a schematic diagram of the overall architecture of the improved generative adversarial network GAN provided by an embodiment of the present invention;
[0060] Figure 5 Shown is a schematic diagram of an improved generative adversarial network GAN incorporating regional attention provided by an embodiment of the present invention;
[0061] Figure 6 FIG. 1 is a schematic diagram of alignment of image features and text features provided by an embodiment of the present invention; DETAILED DESCRIPTION
[0062] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. The embodiments of the present invention and the technical features in the embodiments may be combined with each other unless there is a conflict.
[0063] The term "and / or" is only a description of the association relationship between related objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " generally indicates that the related objects are in an "or" relationship.
[0064] Example 1
[0065] This embodiment introduces a method for visually detecting defects in the suspension state of a high-speed railway contact network in an open environment, including:
[0066] like Figure 2As shown, according to the acquired image to be inspected of the high-speed railway contact network suspension equipment, based on the pre-trained defect detection model, detection is performed to obtain the detection result, including the image to be inspected with the defect position marked.
[0067] Wherein, the training method of the defect detection model is as follows: Figure 1 As shown, the following steps are included:
[0068] Step 1: Extract features from the acquired real image of defects in the high-speed rail contact network suspension equipment to obtain sub-images of different defect types.
[0069] The purpose of the present invention is to extract key information that can represent different defect types from the real image of the high-speed railway overhead line suspension equipment defect obtained. The extracted features are the local area, shape, color, texture, etc. in the real image of the high-speed railway overhead line suspension equipment defect obtained. The accuracy and efficiency of defect detection can be improved through feature extraction.
[0070] Step 2: Improve the generative adversarial network (GAN) and introduce regional attention to expand the sub-graphs of different defect types to obtain the generated images of defects in the high-speed railway contact network suspension equipment.
[0071] The present invention improves the generative adversarial network (GAN) so that the improved generative adversarial network (GAN) can generate more realistic and diverse defect images of high-speed rail contact network suspension equipment. At the same time, the present invention introduces a regional attention mechanism so that the improved generative adversarial network (GAN) can focus on the key areas in the image when generating images, thereby generating more refined defect images, which can enhance the generalization ability of the defect detection model, because the generated images can be used as additional training data to enhance the performance of the model, and the smoothness also improves the generation quality of the defect image, making the generated image closer to the real defect image.
[0072] Step 3: Based on the language generation model and the natural language description of the types of suspension state defects of the high-speed railway contact network suspension equipment, a text description library is constructed to obtain a text description of the defects of the high-speed railway contact network suspension equipment.
[0073] The present invention constructs a text description library based on a language generation model and a natural language description of the types of suspension state defects of high-speed railway contact network suspension equipment. The text description provides detailed and accurate information about the defect types and matches them with image features, thereby providing additional semantic information for defect detection. This enables the defect detection model to understand the meaning of the defect type and enhances the interpretability of the defect detection model because the detection results can be explained by text descriptions.
[0074] Step 4: Based on the CLIP model, the real image and the generated image are fed into an image encoder, and the text description is fed into a text encoder to extract image features and text features.
[0075] The present invention uses the CLIP model to send real images and generated images into an image encoder, and sends text descriptions into a text encoder, and extracts image features and text features through the image encoder and the text encoder respectively, thereby improving the accuracy and efficiency of feature extraction because the CLIP model has been pre-trained on a large amount of data.
[0076] Step 5: Align the image features with the text features based on the attention mechanism, calculate the similarity between the image features and the text features, determine the defect type based on the similarity, map it to the real image and mark the defect location.
[0077] The present invention aligns image features with text features based on the attention mechanism, calculates the similarity between image features and text features, determines the defect type according to the similarity, and maps the defect position to the real image for marking, thereby achieving efficient defect detection under zero-sample conditions. Since the defect detection model can use the pre-trained CLIP model and attention mechanism to identify new defect types, the accuracy and practicality of defect detection are also improved, because the detection results can be directly visualized and marked in the real image.
[0078] Example 2
[0079] Based on the same inventive concept as Example 1, this example introduces a method for visually detecting defects in the suspension state of a high-speed railway contact network in an open environment, comprising:
[0080] According to the acquired image to be inspected of the high-speed railway contact network suspension equipment, detection is performed based on a pre-trained defect detection model to obtain a detection result, including the image to be inspected with the defect position marked.
[0081] Wherein, the training method of the defect detection model includes:
[0082] Step 1: Extract features from the acquired real image of the defects of the high-speed rail contact network suspension equipment to obtain sub-images of different defect types.
[0083] First, it is necessary to collect high-definition image data of defects in high-speed rail contact network suspension equipment for preprocessing, including image enhancement, to improve imaging quality, increase the number of defect samples, and improve the quality and analyzability of data. Features of different levels in defects in high-speed rail contact network suspension equipment are extracted from the preprocessed images. By training the feature pyramid network, high, medium and low-level features in the preprocessed images are automatically learned and extracted, so as to achieve accurate identification and classification of different defect types and abnormal states, and obtain sub-images of defect areas.
[0084] This embodiment uses a contact network suspension status detection and monitoring device, such as a 4C device, installed on a contact network operating vehicle or a special vehicle. At a certain operating speed, high-speed imaging is performed on the components of the contact network suspension system to form image data of the contact network.
[0085] In this embodiment, the real images of the defects of the high-speed railway contact network suspension equipment obtained include real images of the defects of the high-speed railway contact network suspension equipment in normal state and abnormal state, so the detection results also include the state type of the high-speed railway contact network suspension equipment, that is, normal state or abnormal state.
[0086] In this embodiment, feature extraction is performed on the acquired real image of the defects of the high-speed railway overhead line suspension equipment to obtain sub-images of different defect types, including:
[0087] The acquired real image of the high-speed rail contact network suspension equipment defect is sent to the ROI region extraction network to perform feature extraction at different levels to obtain defect region candidate frames of different sizes, such as Figure 3 As shown;
[0088] Based on defect area candidate boxes of different sizes, the non-maximum suppression method is used to save the candidate boxes whose confidence exceeds the threshold, and the corresponding area of the real image is intercepted and saved according to the coordinates of the candidate boxes exceeding the threshold to obtain sub-images of different defect types.
[0089] Step 2: Improve the generative adversarial network (GAN) and introduce regional attention to expand the sub-graphs of different defect types to obtain the generated images of the defects of the high-speed rail contact network suspension equipment, including:
[0090] The regional attention module and regional attention mechanism are introduced into the generator and discriminator of the generative adversarial network (GAN) respectively, and an improved generative adversarial network (GAN) is obtained. It uses a limited number of sub-images of different defect types to generate diverse and realistic defect images, and expands the scale and distribution of the training sample set.
[0091] In this embodiment, the loss function of the improved generative adversarial network GAN is expressed as:
[0092] ;
[0093] In the formula, Represents the loss function of the improved generative adversarial network GAN, represents the loss function of the GAN before improvement, represents the hyperparameter that controls the diversity regularization weight, represents the diversity regularization loss function, which is used to generate more defect subgraphs from different perspectives.
[0094] Among them, the loss function of the improved generative adversarial network GAN is , expressed as:
[0095] L GAN = <m> min < / m> <m> max < / m> E x ~ pdata [ logD ( x ) ] + E z ~ pz [ log(1-D ( G ( z ))) ] ;
[0096] In the formula, Denotes the discriminator D training generator Minimize the loss function when Denotes the generator G training discriminator Maximize the loss function when , which is used to correctly distinguish between real images and generated images; From real data The actual data distribution The expected value of the real data sampled from represents the logarithmic function, Represents real data is the probability of a real image, Representation Discriminator The logarithmic expected value of the probability that the generated sample is false, where is the noise distribution, Represents the first noise sampled in the noise distribution The generated image by generator G, Represents the probability that the image generated by the generator G is classified as true by the discriminator D;
[0097] Among them, the diversity regularization loss function , expressed as
[0098] L div =− E z 1 , z 2 ~ pz [ dist ( G ( z 1 ), G ( z 2 )) ] ;
[0099] In the formula, Represents the expected value between the output results generated by two random noises of generator G, which is used to constrain the diversity between generated images and measure whether generator G can generate diverse images. are the second and third noise in the noise distribution, respectively. Represents the measure of the second noise Generated image after generator G and the third noise Generated image after generator G The function of the difference between , represents the dot product operation, Represents the modulus length between the two.
[0100] In this embodiment, the method also includes classifying and saving the generated images of defects of the high-speed railway contact network suspension equipment according to the defect types and labels corresponding to the defect types.
[0101] Second noise Generated image after generator G and the third noise Generated image after generator G It is the core data for evaluating the generator's generation ability and diversity. By comparing the difference between the two, it can be detected whether the generator generates diverse samples and control the hyperparameters of the diversity regularization weight. It is used to control the weight of the diversity regularization term in the overall loss function, balancing the diversity and authenticity of the generated images. When is larger, the generator will pay more attention to the differences between generated samples, encouraging the generation of diverse generated images with defects, but will sacrifice some image quality; When it is smaller, the generator will focus more on producing high-quality but potentially repetitive images. , which can achieve the best balance between diversity and authenticity and prevent the mode collapse problem.
[0102] Introducing the regional attention module in the generator can enable the defect detection model to focus on the detail generation of the defect area and enhance the authenticity of the local structure. At the same time, adding the regional attention mechanism in the discriminator can refine the ability to distinguish the defect area from the global background, thereby ensuring that the generated image is more realistic in detail and overall. The overall architecture of the improved generative adversarial network GAN is shown in the figure below. Figure 4 As shown in Figure 2, the improved GAN incorporating regional attention is shown in Figure 2. Figure 5 shown.
[0103] Step 3: Based on the language generation model and the natural language description of the types of suspension state defects of the high-speed railway contact network suspension equipment, a text description library is constructed to obtain a text description of the defects of the high-speed railway contact network suspension equipment, including:
[0104] According to the types of suspension state defects of high-speed rail overhead contact network suspension equipment, write a natural language description of each defect type;
[0105] Based on the natural language description of each defect type, a language generation model is used to generate diverse sentences;
[0106] A text description library is constructed according to the diversified sentences and labels corresponding to the defect types to obtain a text description of the defects of the high-speed railway contact network suspension equipment.
[0107] In this embodiment, the types of suspension state defects of the high-speed railway overhead line suspension equipment include cotter pin defects, nut defects and loose pipe caps;
[0108] In this embodiment, the natural language description of the suspension state defect type of the high-speed railway contact network suspension equipment includes physical characteristics and physical forms.
[0109] The text description library constructed in this embodiment is shown in Table 1:
[0110] Table 1 Text description library
[0111] Defect Type describe Example scenario Keywords Split pin defect Broken, missing or bent cotter pins may cause the connection parts to loosen and affect the stability of the structure. The split pin breaks and causes the parts to fall off; the split pin bends and affects the normal connection Split pin, broken, missing, bent, loose connection Nut Defect Loosening, rusting or falling off of the nuts may reduce the reliability of the fixed structure. The rust on the nuts makes it impossible to tighten the bolts; the loose nuts cause the equipment to shake Nut, loose, rusted, fallen off, fixed structure Loose pipe cap Improper installation or loosening of the pipe cap may cause reduced sealing performance or component damage. Loose caps lead to fluid leaks; missing caps expose equipment Cap, missing, loose, sealing performance, equipment damage ...... ...... ...... ......
[0112] Step 4: Based on the CLIP model, the real image and the generated image are fed into an image encoder, and the text description is fed into a text encoder to extract image features and text features.
[0113] Inputting the real image and the generated image into a CLIP-based image encoder to obtain a vector of fixed dimension to represent image features;
[0114] The text descriptions in the text description library are converted into corresponding feature vectors to represent text features.
[0115] Step 5: Align the image features with the text features based on the attention mechanism, calculate the similarity between the image features and the text features, determine the defect type according to the similarity, map it to the real image and mark the defect location, including:
[0116] (1) Embed image features and text features into the same space and convert them into query vectors, key vectors, and value vectors as output features to be passed to the attention mechanism. Use multiple independent attention heads to process the query vectors, key vectors, and value vectors of each pair of image features and text features in parallel to obtain the outputs of all attention heads.
[0117] (2) Calculate the dot product between the query vector and the key vector of the image feature and the text feature to obtain the similarity score and normalize it to obtain the cross-modal feature association weight of the image feature and the text feature;
[0118] In this embodiment, the calculation expression of the dot product between the query vector and the key vector of the image feature and the text feature is:
[0119] ;
[0120] In the formula, represents the similarity score, represents the query vector, Key Vector The transposed vector of .
[0121] (3) Based on the cross-modal feature association weights of image features and text features, the outputs of all attention heads are concatenated or aggregated using the weighted average method to align image features and text features, such as Figure 6 As shown;
[0122] (4) Calculate the similarity between each pair of image features and text features through the contrast loss function, determine the text feature that best matches the image feature based on the similarity, determine the defect type based on the text description corresponding to the text feature, map it to the real image and mark the defect location.
[0123] In this embodiment, the contrast loss function is expressed as:
[0124] ;
[0125] In the formula, represents the contrast loss function, represents the logarithmic function, represents the exponential function with base e, It is For image features and text features corresponding to the real image, i.e., the correct text, Represents the total number of pairs of image features and text features corresponding to the real image, represents the temperature parameter, represents all text features corresponding to real images, Represents cosine similarity.
[0126] In this embodiment, the contrast loss function can maximize the similarity numerator of the positive sample pair, that is, maximize the similarity numerator of the correct pair, and minimize the similarity denominator of the negative sample pair, that is, minimize the similarity denominator of the incorrect pair, through the similarity between each pair of image features and text features. The numerator in the contrast loss function is the similarity of the positive sample pair, that is, The similarity between image features and text features, and the denominator is the sum of the similarities of all negative sample pairs, that is, all pairs of image features and text features. By maximizing the numerator and minimizing the denominator, the defect detection model can learn the alignment of image features and text features in the shared feature space, and by optimizing the contrast loss function, it can better achieve effective alignment of text features and image features.
[0127] Finally, the text feature that best matches the image feature is determined based on the similarity, the defect type is determined based on the text description corresponding to the text feature, and is mapped to the real image and the defect position is marked.
[0128] Example 3
[0129] Based on the same inventive concept as other embodiments, this embodiment introduces a computer-readable storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, the steps of the method in the above-mentioned embodiment 1 or 2 are implemented.
[0130] Example 4
[0131] Based on the same inventive concept as other embodiments, the present invention further provides a computer program product, including computer instructions, which implement the steps of the method in the above-mentioned embodiment 1 or 2 when executed by a processor.
[0132] In summary, the present invention solves the problem of sparse or zero samples in the abnormal state of the high-speed rail contact network suspension. It can be used in multiple scenarios such as daily operation and maintenance, safety monitoring, accident prevention, line maintenance, construction and planning of high-speed rail. The present invention detects the image to be detected of the high-speed rail contact network suspension equipment by using a pre-trained defect detection model. The training process of the defect detection model not only extracts and expands features based on the real image of the defect of the high-speed rail contact network suspension equipment, but also introduces the generative adversarial network GAN with improved regional attention to enhance the diversity and accuracy of the defect image. At the same time, combined with the language generation model and the CLIP model, the image features are aligned with the text features, and the accurate identification and position marking of the defect type of the high-speed rail contact network suspension equipment are realized, which not only improves the accuracy and efficiency of defect detection, but also can automatically generate a detection image containing defect position marks, which provides intuitive and convenient information support for subsequent maintenance and management, and significantly improves the safety and reliability of the high-speed rail contact network suspension equipment.
[0133] By optimizing the loss function of GAN and introducing the regional attention mechanism, the present invention can expand the subgraphs of different defect types with limited defect data, and obtain the generated images of defects of high-speed railway contact network suspension equipment, which significantly expands the sample size and distribution and can effectively cope with zero-sample or sparse-sample scenarios.
[0134] The present invention also sends the acquired real image of the defects of the high-speed railway contact network suspension equipment into the ROI region extraction network to perform feature extraction at different levels, capture the details of different target scales, and can accurately extract defect area candidate frames of different sizes and effectively intercept the corresponding areas of the real image, thereby realizing comprehensive detection from global to local.
[0135] The present invention also aligns image features with text features based on the CLIP model and attention mechanism, solves the similarity problem between categories through the attention mechanism, realizes efficient defect detection under zero-sample conditions, and improves the interpretability and practicality of detection through visual marking.
[0136] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0137] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0138] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0140] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.
Claims
1. A method for visually detecting the suspension state defects of a high-speed railway contact network in an open environment, characterized in that: include: According to the acquired image to be inspected of the high-speed railway overhead line suspension equipment, based on the pre-trained defect detection model, the inspection is performed to obtain the inspection result, including the image to be inspected with the defect position marked; Wherein, the training method of the defect detection model includes: Feature extraction is performed on the acquired real image of the defects of the high-speed railway overhead line suspension equipment to obtain sub-images of different defect types; Improve the generative adversarial network (GAN) and introduce regional attention to expand the sub-graphs of different defect types to obtain the generated images of defects in the high-speed rail contact network suspension equipment; According to the language generation model and the natural language description of the types of suspension state defects of the high-speed railway overhead line suspension equipment, a text description library is constructed to obtain the text description of the defects of the high-speed railway overhead line suspension equipment; Based on the CLIP model, the real image and the generated image are fed into an image encoder, and the text description is fed into a text encoder to extract image features and text features; Based on the attention mechanism, the image features and the text features are aligned, the similarity between the image features and the text features is calculated, the defect type is determined according to the similarity, and the defect type is mapped to the real image and the defect position is marked.
2. The method for visually detecting the suspension state defects of high-speed railway contact network in an open environment according to claim 1 is characterized in that: The acquired real images of defects in the high-speed railway contact network suspension equipment include real images of defects in the high-speed railway contact network suspension equipment in normal states and abnormal states.
3. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 1 is characterized in that: Feature extraction is performed on the acquired real image of the defects of the high-speed railway overhead line suspension equipment to obtain sub-images of different defect types, including: The acquired real image of the defects of the high-speed railway overhead line suspension equipment is sent to the ROI region extraction network to perform feature extraction at different levels to obtain defect region candidate frames of different sizes; Based on defect area candidate boxes of different sizes, the non-maximum suppression method is used to save the candidate boxes whose confidence exceeds the threshold, and the corresponding area of the real image is intercepted and saved according to the coordinates of the candidate boxes exceeding the threshold to obtain sub-images of different defect types.
4. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 1 is characterized in that: By improving the generative adversarial network (GAN) and introducing regional attention, the sub-graphs of different defect types are expanded to obtain the generated images of defects in the high-speed rail contact network suspension equipment, including: The regional attention module and regional attention mechanism are introduced into the generator and discriminator of the generative adversarial network (GAN) respectively to obtain the improved generative adversarial network (GAN).
5. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 4 is characterized in that: The loss function of the improved generative adversarial network GAN is expressed as: ; In the formula, Represents the loss function of the improved generative adversarial network GAN, represents the loss function of the GAN before improvement, represents the hyperparameter that controls the diversity regularization weight, represents the diversity regularization loss function; Among them, the loss function of the improved generative adversarial network GAN is , expressed as: ; In the formula, Denotes the discriminator D training generator Minimize the loss function when Denotes the generator G training discriminator When maximizing the loss function, From real data The actual data distribution The expected value of the real data sampled from represents the logarithmic function, Represents real data is the probability of a real image, Representation Discriminator The logarithmic expected value of the probability that the generated sample is false, where is the noise distribution, Represents the first noise sampled in the noise distribution The generated image by generator G, Represents the probability that the image generated by the generator G is classified as true by the discriminator D; Among them, the diversity regularization loss function , expressed as ; In the formula, represents the expected value between the output results of two random noises generated by the generator G, where are the second and third noise in the noise distribution, respectively. Represents the measure of the second noise Generated image after generator G and the third noise Generated image after generator G The function of the difference between , represents the dot product operation, Represents the modulus length between the two.
6. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 1 is characterized in that: It also includes classifying and saving the generated images of defects in the high-speed railway contact network suspension equipment according to the defect types and labels corresponding to the defect types.
7. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 6 is characterized in that: According to the language generation model and the natural language description of the types of suspension state defects of the high-speed railway contact network suspension equipment, a text description library is constructed to obtain the text description of the defects of the high-speed railway contact network suspension equipment, including: According to the types of suspension state defects of high-speed rail overhead contact network suspension equipment, write a natural language description of each defect type; Based on the natural language description of each defect type, a language generation model is used to generate diverse sentences; Building a text description library based on the diversified sentences and labels corresponding to the defect types, and obtaining a text description of the defects of the high-speed railway overhead line suspension equipment; The types of suspension state defects of the high-speed railway overhead line suspension equipment include cotter pin defects, nut defects and loose pipe caps; The natural language description of the suspension state defect type of the high-speed railway contact network suspension equipment includes physical characteristics and physical forms.
8. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 7 is characterized in that: Based on the attention mechanism, the image features and the text features are aligned, the similarity between the image features and the text features is calculated, the defect type is determined according to the similarity, and the defect type is mapped to the real image and the defect location is marked, including: (1) Embed image features and text features into the same space and convert them into query vectors, key vectors, and value vectors as output features to be passed to the attention mechanism. Use multiple independent attention heads to process the query vectors, key vectors, and value vectors of each pair of image features and text features in parallel to obtain the outputs of all attention heads. (2) Calculate the dot product between the query vector and the key vector of the image feature and the text feature to obtain the similarity score and normalize it to obtain the cross-modal feature association weight of the image feature and the text feature; (3) Based on the cross-modal feature association weights of image features and text features, the outputs of all attention heads are concatenated or aggregated using the weighted average method to align image features and text features; (4) Calculate the similarity between each pair of image features and text features through the contrast loss function, determine the text feature that best matches the image feature based on the similarity, determine the defect type based on the text description corresponding to the text feature, map it to the real image and mark the defect location.
9. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 8, characterized in that: The calculation expression of the dot product between the query vector and the key vector of the image feature and text feature is: ; In the formula, represents the similarity score, represents the query vector, Key Vector The transposed vector of .
10. The method for visually detecting the hanging state defects of high-speed railway contact network in an open environment according to claim 8, characterized in that: The contrast loss function is expressed as: ; In the formula, represents the contrast loss function, represents the logarithmic function, represents the exponential function with base e, It is For image features and text features corresponding to real images, Represents the total number of pairs of image features and text features corresponding to the real image, represents the temperature parameter, represents all text features corresponding to real images, Represents cosine similarity.
Citation Information
Cited By
Defect detection method and equipment for cutterhead and storage medium
CN120598866A
Outdoor power transmission channel inspection intelligent detection method, system and equipment
CN120635692A
An outdoor power transmission channel inspection intelligent detection method, system and device
CN120635692B