Machine vision AI large model-oriented track flaw detection method and device
Through the generative large model, the data samples are expanded and multi-scale feature fusion model is constructed, the problem of scarce data and insufficient correlation between multimodal features in track damage detection is solved, efficient identification and real-time detection of complex track damage is achieved, and detection accuracy and reliability of maintenance decisions are improved.
Patent Information
- Application Number
- CN202510432365.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-25
AI Technical Summary
The existing orbital damage detection methods have poor generalization capabilities under scarce samples, and cannot efficiently identify hidden cracks and rare damages. The complexity of multi-source data and multi-modal feature correlation modeling is insufficient, and the semantic extraction capability is limited. It is difficult to comprehensively analyze the damage patterns in complex orbital environments.
Generative large-scale models are used to expand images, and local feature extraction modules and global feature extraction modules are built. Multi-scale fusion is carried out through feature fusion modules. Combined with classifiers and description generators, loss function training track flaw detection model is built, adversarial training is used for adversarial training of the generator and discriminator, attention generators and multi-stage generation networks are introduced to capture the features of the damaged area, and multi-scale feature extraction and semantic description are realized.
It improves the detection coverage and accuracy of rare and complex track damage types, optimizes maintenance decision efficiency and reliability, and realizes refined modeling and real-time detection of multiple types of damage such as track cracks and wear.
Smart Images

Figure CN120375046A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of machine vision, artificial intelligence, and railway transportation safety. Specifically, it relates to an intelligent detection method, device, electronic device, computer-readable storage medium, and computer program product for track damage based on a machine vision-driven ultrasonic full-focus phased array acquisition device and a generative large model, which are used to efficiently and accurately identify track cracks and other types of structural damage. Background Art
[0002] Track flaw detection technology occupies a core position in global railway traffic safety and is an important link to ensure the efficient operation of railway transportation. With the rapid expansion of the global railway network, the total mileage has exceeded 1.5 million kilometers, posing higher requirements for the real-time performance and accuracy of damage detection. Currently, commonly used detection technologies such as magnetic particle detection and ultrasonic detection are mature, but they have limitations in terms of detection speed, accuracy, and intelligence level. The global non-destructive testing market is developing towards the intelligent direction, and the development of generative AI large models provides a new idea for track flaw detection, such as expanding the dataset, optimizing real-time analysis capabilities, and improving the automation level. In the future, the wide application of these technologies in high-speed railway lines, heavy-haul railways, and urban rail transit will promote the industry efficiency to increase by more than 20%, while reducing the accident rate.
[0003] The difficulties in track flaw detection include the technical bottlenecks of data and algorithms. Existing track damage detection needs to handle complex damage forms, such as micro-cracks (width less than 0.1 mm), spalling (area less than 50 mm 2 ) and corrosion, which are difficult to accurately identify through traditional detection methods. In addition, the data volume during high-speed inspection of tracks (more than 60 km / h) is huge, and traditional algorithms cannot process it in real time. More importantly, due to the rarity of certain types of track damage (such as deep cracks), the existing dataset is insufficient, and the damage types covered by the public dataset (such as RAIL-Defect) are only 60% of those in actual applications. Generative AI large models can generate scarce data through data augmentation techniques and analyze complex features using the semantic extraction ability of machine vision, providing new possibilities for track flaw detection.
[0004] At present, large AI models are providing a new breakthrough for rail flaw detection technology. Their applications are mainly reflected in data augmentation, semantic understanding, and real-time analysis, fundamentally solving the bottleneck problems of traditional detection technologies. In terms of data augmentation, large models generate representative virtual data by simulating real rail damages, including rare types of cracks, spalls, and corrosion, making up for the deficiencies of real datasets. Research shows that with the dataset expanded by generative AI, the coverage rate of damage data has been greatly improved, optimizing the adaptability of algorithms to minor and rare damages. In addition, through powerful machine vision capabilities, generative large models can not only analyze single-modal features but also perform deep semantic association modeling on multi-source data, providing more accurate pattern descriptions for rail damage detection. For example, large models can automatically extract the geometric shape features of rail spalls and their potential relationships with environmental factors (such as temperature and pressure changes), thus improving the detection ability for concealed damages (such as internal cracks with a depth of up to 5 cm). In terms of real-time analysis, by combining large models with edge computing technology and through model parameter optimization and lightweight inference algorithms, real-time processing of rail damages can be achieved. Currently, the SNCF (French National Railways) in France has adopted a system that combines full-focus phased array ultrasound and AI algorithms, with its processing efficiency increased by more than 50% and the false detection rate reduced to 3%. DB Netz AG (German Rail Network) in Germany has improved the accuracy of complex damage detection to the millimeter level through multi-modal data fusion technology and significantly optimized the detection stability. The Chinese railway department is testing a real-time detection system driven by generative large models. For example, the TFDS freight car fault detection system optimized by the Huawei Cloud Pangu large model is adopted by the Zhengzhou North Car Depot, and its application is mainly concentrated in the field of freight trains. It is still in the experimental stage in terms of large-scale application. There is still a need to further improve the multi-modal data processing ability and the adaptability to industrial-grade hardware to meet higher economic and performance requirements.
[0005] Currently, the existing methods have the following three problems:
[0006] 1. Scarce data samples affect the generalization ability of the model: High-precision rail flaw detection depends on rich labeled data. However, due to the low incidence of rail damage events, samples of some key damages (such as deep cracks) are extremely scarce and even difficult to obtain. Although the samples automatically generated by large models can alleviate this problem to a certain extent, they often fail to capture the complex features and environmental diversity of real damages. Therefore, when dealing with these rare damage types, the accuracy of the detection system is still low, and the detection effect is unstable, restricting the universality and promotion ability of the technology;
[0007] 2. Insufficient matching between data features and the model: The data in existing rail flaw detection systems is multi-source and complex, including rail surface images, stress wave signals, etc. The data distribution is uneven and the noise interference is significant, making it difficult to directly match the training requirements of generative large models. In addition, the number of rare damage samples (such as hidden cracks) is limited, resulting in the lack of generalization ability of the model during training, and it is prone to overfitting to edge features or missed detection. This problem of data-model adaptability restricts the further improvement of detection accuracy;
[0008] 3. Limited semantic extraction ability and lack of in-depth correlation analysis: Currently, the detection methods mainly rely on single or a small number of modal feature extractions, making it difficult to comprehensively analyze the damage patterns in complex rail environments. There are limitations in the existing systems when dealing with the in-depth correlation modeling of multi-source data, and they cannot fully extract the potential correlations between damages and environmental factors (such as pressure and temperature changes), resulting in insufficient prediction ability and missing potential hidden dangers. Summary of the Invention
[0009] The existing rail damage detection method has poor model generalization ability under scarce sample conditions and cannot efficiently identify hidden cracks and rare damages. Aiming at the complexity of multi-source data and the lack of multi-modal feature correlation modeling, the present invention proposes a rail flaw detection method based on a generative large model to improve detection accuracy and real-time performance.
[0010] Aiming at the deficiencies of the existing technology, as Figure 6 shown, the present invention proposes a rail flaw detection method for a machine vision AI large model, including:
[0011] Image expansion step: Obtain the original images with labeled rail damage labels, use the generative model to expand the samples of the original images labeled as scarce damages to obtain expanded images, and gather all the original images and expanded images as the training image set; construct a rail flaw detection model including the local feature extraction module, the global feature extraction module, the feature fusion module, the classifier, and the description generator;
[0012] Feature extraction step: Respectively extract the local damage features and global damage features of the images in the training image set through the local feature extraction module and the global feature extraction module to obtain multi-dimensional information such as the dynamic change features of cracks, the damage edge features, and the spatial frequency distribution;
[0013] Model training step: Multiscale fuse the local damage features and the global damage features through the feature fusion module to obtain fused features, obtain the predicted damage type and generate semantic descriptions through the classifier and the description generator, and construct a loss function according to the predicted damage type and the damage label to train the rail flaw detection model;
[0014] Model flaw detection steps: Input the track image to be flaw-detected into the trained track flaw detection model to obtain its damage type and semantic description.
[0015] The described track flaw detection method for the machine vision AI large model, including that the generative model is an adversarial network including a generator and a discriminator;
[0016] The generator G takes the random noise vector z as input and learns to generate a generated image similar to the real track damage image x through learning; the discriminator D is used to distinguish whether the input image is a real track damage image and promotes the optimization of the generator G through feedback; the generator and the discriminator interact with each other in the adversarial training, and the objective function is:
[0017]
[0018] where, L adv is the adversarial loss, x is the real sample, G(z) is the generated image, and D(x) and D(G(z)) are the discrimination results of the discriminator for the real image and the generated image respectively;
[0019] The perceptual loss L per is used to measure the difference between the generated image and the real track damage image in the high-level feature space;
[0020]
[0021] where, represents the deep features extracted by the pre-trained network; the adversarial network is trained through the perceptual loss and the adversarial loss.
[0022] The described track flaw detection method for the machine vision AI large model, including that the generator includes an attention generator to ensure that the noise vector z only exists in the damaged area of the generated image, and the structure of the attention generator is as follows:
[0023] A(x) = σ(Conv s (x)) · σ(Conv c (x))
[0024] where, σ is the sigmoid function, Conv s and Conv c are the convolution operations of the spatial and channel attention modules respectively.
[0025] The described track flaw detection method for the machine vision AI large model, including that the generative network is a multi-stage generative network, and uses the focus mask pyramid {m0, m1,..., m n}Guide different stages of the generation process; the input for each stage includes the product of the random noise and the focus mask, where only the damaged area contains noise and other areas are set to zero; the generation process can be expressed as:
[0026]
[0027] where y n is the generated image generated in the nth stage, z n is the random noise, m n is the focus mask in the nth stage, is the result of upsampling the image generated in the previous stage.
[0028] The described rail flaw detection method for the machine vision AI large model includes that the local feature extraction module is used to extract the detailed features of rail damage; extract the basic features of the image data input to the local feature extraction module, and then perform batch normalization:
[0029]
[0030] where is the normalized image feature, μ and σ 2 are the mean and variance respectively, and ∈ is the stabilization term; the normalized feature is processed by the smoothing and non-linear enhancement ability of the Swish activation function:
[0031] Swish(x) = x · σ(x)
[0032] The local feature extraction module constructs a feature extraction block through a combination of dilated convolution, depthwise separable convolution, and pointwise convolution. The output of the feature extraction block is:
[0033] y i = f(x i *(k d + k p )) + b
[0034] where k d is the dilated convolution kernel, k p is the pointwise convolution kernel, * represents the convolution operation, and f is the Swish activation function; the feature output by the local feature extraction module passes through the global pooling layer and the fully connected layer to generate the local feature output F l ;
[0035] The global feature extraction network receives the local feature, performs batch normalization on it, and then processes it through the ReLU activation function; captures global features at different scales through multi-dilation rate convolution:
[0036]
[0037] Among them, k di is a convolution kernel with a dilation rate of d i . Through parallel operations with multiple dilation rates, it can capture features at different scales; through skip connections, multi-scale features are fused to achieve the joint expression of global and local features; the features after multi-scale fusion are further dimension-reduced through a multi-scale pooling layer to obtain the global feature output F g ;
[0038] The described rail flaw detection method for a large machine vision AI model includes the local feature output F l and the global feature output F g to achieve semantic interaction through attention weighting:
[0039] F c = α·F l +(1 - α)·F g
[0040] α = σ(MLP([F l ; F g ))
[0041] where [F l ; F g represents the concatenation of local and global features, σ is the Sigmoid activation function. MLP is a multi-layer perceptron for dynamic interaction modeling of local and global features.
[0042] As shown in Figure 7 , the present invention also proposes a rail flaw detection device for a large machine vision AI model, including:
[0043] An image expansion module, which acquires the original image with the labeled rail damage label, uses a generative model to expand the samples of the original image labeled as scarce damage to obtain an expanded image, and combines all the original images and the expanded images as a training image set; constructs a rail flaw detection model including the local feature extraction module, the global feature extraction module, a feature fusion module, a classifier, and a description generator;
[0044] A feature extraction module, which respectively extracts the local damage features and the global damage features of the images in the training image set through the local feature extraction module and the global feature extraction module to obtain multi-dimensional information such as the dynamic change features of cracks, the damage edge features, and the spatial frequency distribution;
[0045] A model training module, which multi-scale fuses the local damage features and the global damage features through the feature fusion module to obtain fused features, and through the classifier and the description generator, obtains the predicted damage type and generates a semantic description, and constructs a loss function according to the predicted damage type and the damage label to train the rail flaw detection model;
[0046] The model flaw detection module inputs the track image to be flaw detected into the trained track flaw detection model to obtain its damage type and semantic description.
[0047] The present invention also provides an electronic device, including the track flaw detection device for the machine vision AI large model described in claim 7. The electronic device is either connected to an information display device, and the information display device is used to display the damage type with the display parameters, attributes set by the user or through an artificial intelligence model.
[0048] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the track flaw detection method for the machine vision AI large model described in any one of claims 1-6 are implemented.
[0049] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the track flaw detection method for the machine vision AI large model described in any one of claims 1-6 are implemented.
[0050] As can be seen from the above solutions, the advantages of the present invention are as follows:
[0051] 1. By expanding data samples through a generative large model, the detection coverage rate and accuracy for rare and complex track damage types are improved.
[0052] 2. The intelligent reasoning module based on semantic extraction provides visual and structured analysis results, optimizing the efficiency and reliability of maintenance decision-making.
[0053] 3. Dynamic modeling and multi-scale feature fusion technology enable refined modeling and real-time detection of various types of damage such as track cracks and wear. Brief Description of the Drawings
[0054] Figure 1 It is the overall flowchart of track flaw detection based on the AI large model;
[0055] Figure 2 It is the key technology module diagram of track flaw detection based on the AI large model;
[0056] Figure 3 It is the damage sample expansion technology diagram based on the generative large model;
[0057] Figure 4 It is the local feature extraction network diagram;
[0058] Figure 5 It is the global feature extraction network diagram;
[0059] Figure 6 It is the method flowchart of the present invention;
[0060] Figure 7 This is the block diagram of the device of the present invention;
[0061] Figure 8 This is the schematic structural diagram of the first electronic device of the present invention;
[0062] Figure 9 This is the schematic structural diagram of the application environment of the first electronic device of the present invention;
[0063] Figure 10 This is the schematic structural diagram of the second electronic device of the present invention.
[0064] Reference numerals:
[0065] A - The first electronic device;
[0066] B - The rail flaw detection device for machine vision AI large model;
[0067] C - Data acquisition device;
[0068] D - Information display device;
[0069] 1000 - The second electronic device;
[0070] Ⅰ - Computing unit;
[0071] Ⅱ - ROM;
[0072] Ⅲ - RAM;
[0073] Ⅳ - Bus;
[0074] Ⅴ - Interface;
[0075] Ⅵ - Input unit;
[0076] Ⅶ - Output unit;
[0077] Ⅷ - Storage medium;
[0078] Ⅸ - Communication unit. Detailed implementation manners
[0079] It should be noted that in this application, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0080] Without further limitations, an element limited by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the said element.
[0081] The processor described in the present invention is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or it can be an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0082] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and by calling data stored in the memory.
[0083] In a specific implementation, as an embodiment, the processor can include one or more CPUs. Each of these processors can be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smart phones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0084] The memory is used to store the software program for implementing the solution of the present invention and is controlled by the processor for execution. The specific implementation method can refer to the above method embodiments and will not be elaborated here.
[0085] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device can include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements.
[0086] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, or a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0087] It should also be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context before and after.
[0088] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0089] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0090] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections between each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0093] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0094] In the technical difficulties of ultrasonic image track flaw detection of the present invention, the generative large model is introduced to solve the core problems such as data scarcity, damage feature modeling, and semantic analysis. For the common complex damage types during the operation of the track, including transverse damage, longitudinal damage, damage caused by improper use or handling, etc., intelligent detection and damage description are realized. The detection data is collected by an ultrasonic full-focus phased array device, which can obtain two-dimensional images of internal damage of the track with high resolution, and form a fully covered focus area through multi-channel transmission and reception to ensure the accuracy and reliability of the detection results.
[0095] In terms of data augmentation, the present invention uses a generative network to solve the problem of scarce track flaw detection samples. Based on limited ultrasonic image data, the generative network introduces controllable damage feature perturbations in the latent space through layer-by-layer conditional modeling, mapping the complex distribution into a high-dimensional Gaussian distribution. Subsequently, the signal is gradually reconstructed through the reverse generation process to generate synthetic samples highly consistent with the characteristics of actual ultrasonic images. These samples cover cracks and wear features of different depths and morphologies, which can improve the model's recognition ability for rare and complex damage scenarios, as well as the generalization ability of the detection method in few-shot scenarios.
[0096] In terms of damage feature modeling, the present invention deeply mines the spatial and frequency domain features of ultrasonic images through a neural network, and establishes a multi-scale modeling method that correlates global and local features. It can accurately capture the dynamic characteristics of crack propagation, such as the non-linear propagation mode of crack boundaries and the changes in stress concentration areas. In addition, for bolt hole cracks, multi-dimensional features are extracted to reconstruct the complex morphology of cracks in different directions; for rolling contact wear, the focus is on analyzing the material wear law and periodic deformation characteristics of the contact area. This feature modeling method elevates low-level image features to high-level semantic representations, providing a technical basis for the unified detection of multiple types of damage.
[0097] In terms of semantic analysis, the present invention adopts FCdDN (Fully Convolutional Dense Dilated Network) as a generative AI large model, and converts complex ultrasonic signal patterns into semantic descriptions that are easy for engineers to interpret through context correlation analysis. The model has a structure of sparse convolution and dense connection, and effectively captures multi-scale damage information through dilated convolution technology, and can automatically label the damage type (such as fatigue crack, corrosion pit) and its geometric parameters (such as position, length, depth, direction). For the corrugation wear of rail, the semantic description generated by the model includes information such as corrugation period and amplitude. This analysis ability bridges the gap between machine vision and manual interpretation, providing reliable technical support for track maintenance. In summary, the present invention proposes the following key technical points:
[0098] Key Point 1: Damage sample augmentation technology based on a generative large model. Different types and depths of track damage data are generated through a generative large model, including scarce micro-crack and concealed crack samples, to improve the model's recognition ability for rare damage. The augmented data covers multiple scenarios, improving the detection generalization and robustness of the model in complex environments.
[0099] Key Point 2: Intelligent Detection and Semantic Reasoning Technology for Track Damage. Through the collaborative work of the local feature extraction module and the global feature extraction module, the multi-scale features of track damage in the spatial and frequency domains are accurately captured. The local feature extraction module first extracts local detail information such as the edges and cracks of the damage from the ultrasonic image, providing key local clues for global modeling; at the same time, the global feature extraction module supplements the detailed semantics from a global perspective, capturing the macroscopic distribution of the damage, the expansion pattern, and the potential associations of environmental factors. Through the interaction of these two modules, the system can generate semantic descriptions of the accurate damage type and geometric parameters (such as position, depth, length, direction, etc.). Especially in the detection of concealed damages such as deep cracks and micro-cracks, this technology can significantly improve the interpretability and recognition accuracy, ensuring a comprehensive analysis of complex damages.
[0100] Key Point 3: Multi-scale Feature Fusion and Dynamic Modeling Method. Based on the multi-scale feature fusion technology, the global semantic features and local edge features of track damage are extracted to ensure accurate modeling of complex forms such as cracks and wear. Dynamic modeling can capture the non-linear pattern of crack propagation, adapt to the changing characteristics of different damage areas, and achieve high-precision detection.
[0101] To make the above features and effects of the present invention more clearly understandable, specific embodiments are hereinafter given and described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The scope of protection of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0102] The overall process of track flaw detection based on the AI large model is as Figure 1 shown, and the internal key technical modules are as Figure 2As shown below. First, high-resolution ultrasonic images are collected by a phased array ultrasonic device, and a generative large model is used to expand damage samples to effectively solve the problem of data scarcity. This step uses the generative large model to generate new training data for scarce track damage samples (such as deep cracks, micro-cracks, etc.) to ensure the diversity and integrity of the damage dataset. Then, the system extracts local damage features and global damage features in the image through a local feature extraction module and a global feature extraction module respectively, obtaining multi-dimensional information such as the dynamic changes of cracks, damage edge features, and spatial frequency distribution, providing a comprehensive basis for subsequent damage analysis. Next, the local and global features are multi-scale fused through a feature fusion module to further enhance the ability to represent complex damage information. On this basis, the system uses a classifier and a description generator to accurately classify damage types and generate semantic descriptions. The semantic descriptions are trained through supervised learning. During the training process, ultrasonic images and corresponding damage types and geometric parameters (such as position, depth, length, direction, etc.) are used as labeled data. The generative AI large model maps the damage patterns of ultrasonic images into easily understandable semantic descriptions through context correlation analysis. The training process requires labeled data for supervised learning to train a model that can accurately identify and generate damage descriptions. Finally, after multi-level information processing and intelligent reasoning, the system outputs detailed damage features and geometric parameters (such as position, depth, length, direction, etc.), providing a scientific basis for the subsequent maintenance and safety assessment of the track. The output damage features include the type of damage (such as fatigue cracks, corrosion pits, corrugated wear, etc.) and its geometric parameters, such as position, depth, length, and direction. In addition, specific damage types may also include other detailed information related to them, such as the period and amplitude of the corrugations, the crack propagation mode, etc.
[0103] 1. Damage Sample Expansion Technology Based on Generative Large Model
[0104] The damage sample expansion technology based on the generative large model proposed in this invention aims to solve the problem of scarce track damage data, especially in the case of insufficient samples of types such as rapid device damage and bolt breakage. To address this challenge, a generative adversarial network is used to generate high-quality track damage images. Compared with traditional data augmentation methods, the generative adversarial network can generate diverse and highly accurate damage samples through the adversarial training of the generator and discriminator, while maintaining the overall structure of the image and the details of the damage area, as Figure 3 shown.
[0105] In a generative adversarial network, the generator G of module 101 and the discriminator D of module 102 are the core components. The generator takes a random noise vector z as input and learns to generate fake images similar to the real track damage images x, while learning how to generate fake images similar to the real track damage images during the training process; the discriminator is responsible for distinguishing whether the input image is a real image and promotes the optimization of the generator through feedback. The two interact with each other in adversarial training, and the objective function is:
[0106]
[0107] where L adv is the adversarial loss, x is the real sample, G(z) is the image generated by the generator, and D(x) and D(G(z)) are the discrimination results of the discriminator for the real image and the generated image, respectively.
[0108] To further improve the quality of the generated images, the present invention introduces a perceptual loss and a reconstruction loss. The perceptual loss is used to measure the difference between the generated image and the real image in the high-level feature space, and the formula is:
[0109]
[0110] where represents the deep features of the image, which are the features extracted by the generator of module 101. This loss ensures the consistency of the generated image and the real image in high-level semantics. The reconstruction loss and the adversarial loss are used together to train the adversarial generation network. The adversarial loss ensures that the generator can generate realistic images, while the reconstruction loss helps the generator maintain the consistency of the details and structure of the image.
[0111] Aiming at the characteristic that the damaged area in the track damage image is usually limited to a local area, the present invention introduces a focusing mechanism and an attention generator, so that the generator can focus on the damaged area, enhance the detailed performance of the damaged area, and at the same time maintain the structural integrity of the non-damaged area. The focusing mechanism ensures that the noise in the generated image only exists in the damaged area, while the background area remains consistent. A convolutional-based attention module is adopted, which combines channel attention and spatial attention mechanisms and focuses on the most important feature content and key positions in the image respectively. The structure of the attention module in the generator is as follows:
[0112] A(x) = σ(Conv s (x)) · σ(Conv c (x))
[0113] where, σ is the sigmoid function, Conv s and Conv c are the convolution operations of the spatial and channel attention modules respectively. Through the attention mechanism, the generator can capture the details of the damaged area more accurately.
[0114] To improve the diversity and quality of sample generation, the present invention proposes to use a multi-stage generation network. At each stage, the focus mask pyramid {m0, m1,..., m n} is used to guide different stages of the generation process. The input of each stage includes the product of random noise and the focus mask. Only the damaged area contains noise, and other areas are set to zero. This method ensures that the generated images only vary in the damaged area, and the background area remains consistent. The generation process can be expressed as:
[0115]
[0116] where, y n is the image generated in the nth stage, z n is the random noise, m n is the focus mask in the nth stage, is the result of upsampling the image generated in the previous stage. Through this multi-stage generation network, the features of the damaged area can be accurately generated while retaining the overall structure of the image.
[0117] Through the above technologies, the present invention can effectively expand the track damage samples, generate diverse and high-quality damaged images, and thus provide sufficient training data for track damage detection. This not only solves the problem of sample scarcity but also greatly improves the model's recognition ability for complex damage types, promoting the development of track damage intelligent detection technology.
[0118] 2. Intelligent Detection and Semantic Reasoning Technology for Track Damage
[0119] The intelligent detection and semantic reasoning technology for track damage takes the local feature extraction network and the global feature extraction network as the core, and constructs an efficient multi-scale semantic modeling method for accurate detection and classification of track damage.
[0120] The local feature extraction network, as shown in Figure 4 , focuses on extracting the detailed features of track damage, especially the local patterns of micro-defects such as cracks and pits. The input ultrasonic data generates basic features through the initial feature extraction module, and then batch normalization is performed. The formula is:
[0121]
[0122] where, is the normalized feature, μ and σ 2are the mean and variance respectively, and ∈ is the stability term. The normalized features are processed by the smoothing and non-linear enhancement ability of the Swish activation function:
[0123] Swish(x) = x·σ(x)
[0124] Subsequently, the local feature extraction module 201 constructs a feature extraction block through a combination of dilated convolution, depthwise separable convolution, and pointwise convolution. The output of each block is:
[0125] y i = f(x i *(k d + k p )) + b
[0126] where k d is the dilated convolution kernel, k p is the pointwise convolution kernel, * represents the convolution operation, and f is the Swish activation function. This module can expand the receptive field to enhance the local pattern modeling ability while maintaining the detailed features. After passing through 6 groups of local feature extraction modules, the features extract important information through the global pooling layer and generate the local feature output through the fully connected layer.
[0127] The global feature extraction network, as Figure 5 , aims to capture the global semantic patterns of track damage, especially outstanding in the case where multiple damage patterns coexist in ultrasonic data. The input ultrasonic data also undergoes initial feature extraction and batch normalization processing, and then is processed by the ReLU activation function. The core module 202 captures global features at different scales through multi-dilation rate convolution. The formula is as follows:
[0128]
[0129] where is the convolution kernel with a dilation rate of d i . Through parallel operations with multiple dilation rates, it can capture features at different scales. The multi-scale features are fused through the skip connection module 203 to achieve the joint expression of global and local features. The features after multi-scale fusion are further dimension-reduced through the multi-scale pooling layer, reducing redundancy while retaining the key semantics. The design of the global feature extraction network enables the global features to have stronger semantic expression ability and can effectively handle the complex scenarios of track damage.
[0130] 3. Multi-scale Feature Fusion and Dynamic Modeling Method
[0131] In the detection of track damage, multi-scale feature fusion and dynamic modeling are one of the key methods to improve the detection accuracy and adaptability to complex scenarios. For this purpose, this section first starts from the interactive fusion of local and global features and adopts a multi-scale feature fusion strategy. To achieve the assistance of local features in global modeling and the supplementation of global features to local details, a bidirectional feature interaction module is introduced. The local feature output F l and the global feature output F g realize dynamic semantic interaction through attention weighting, and the process is as follows:
[0132] F c = α·F l +(1 - α)·F g
[0133] α = σ(MLP([F l ; F g ))
[0134] where F c is the finally fused feature map, that is, the feature fusion result in Figure 1 , [F l ; F g represents the concatenation of local and global features, σ is the Sigmoid activation function. MLP is a multi-layer perceptron for dynamically interacting and modeling local and global features. Among them, the dynamic modeling strategy adjusts their contribution ratio by learning the dynamic weight coefficient α between local and global features. The model learns these weight coefficients through the attention mechanism and dynamically balances the fusion of local features (such as cracks, pits, etc.) and global features (such as the overall shape of the track). In complex track damage scenarios, this strategy ensures a reasonable trade-off between local details and global semantics, so that the final fused feature map can not only accurately detect minor damages but also effectively capture the overall damage pattern, thereby improving the detection accuracy.
[0135] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment, and in order to reduce repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied in the above embodiment.
[0136] As Figure 7 shown, the present invention also proposes a track flaw detection device for a machine vision AI large model, including:
[0137] An image expansion module obtains the original images with marked track damage labels, uses a generative model to expand the samples of the original images marked as scarce damage to obtain expanded images, and combines all the original images and the expanded images as a training image set; constructs an orbital flaw detection model including the local feature extraction module, the global feature extraction module, a feature fusion module, a classifier, and a description generator;
[0138] A feature extraction module respectively extracts the local damage features and global damage features of the images in the training image set through the local feature extraction module and the global feature extraction module to obtain multi-dimensional information such as the dynamic change features of cracks, damage edge features, and spatial frequency distribution;
[0139] A model training module performs multi-scale fusion of the local damage features and the global damage features through the feature fusion module to obtain fusion features, and through the classifier and the description generator, obtains the predicted damage type and generates a semantic description. According to the predicted damage type and the damage label, constructs a loss function to train the orbital flaw detection model;
[0140] A model flaw detection module inputs the track image to be flaw-detected into the trained orbital flaw detection model to obtain its damage type and semantic description.
[0141] As Figure 8 shown, in another embodiment of the present invention, a first electronic device A is also proposed, including the orbital flaw detection device for the machine vision AI large model described above.
[0142] As Figure 9 shown, the first electronic device A can also be connected to a data acquisition device C and an information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to acquire the track images to be flaw-detected, such as the orbital flaw detection ultrasonic images described in the embodiments of the present invention, and the information display device D is used to display the damage type and semantic description analyzed by the present invention.
[0143] Among them, the information display device D can process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be preset manually. For example, the data output by the first electronic device A is visually displayed, and it can display according to the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll and play, etc. The key information specified by the user is presented to the user. For example, news trends, system information, etc. The user can understand this information more timely without having to access the secondary page or scroll the page, saving the user's operations. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key attention information according to the user's previous usage habits, such as viewing duration, click times, editing times, etc., and then automatically present rich and necessary key information to the user.
[0144] The present invention also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a readable storage medium. When the computer program is executed by a processor, the computer can execute the rail flaw detection method for the machine vision AI large model provided by the above-mentioned various methods.
[0145] In another embodiment, the present invention further provides a storage medium VIII for storing a computer program for executing the rail flaw detection method for the machine vision AI large model. It should be understood that the storage medium in the embodiments of the present invention may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchlink dynamic random access memory (SLDRAM), and direct rambus random access memory (DR RAM).
[0146] Figure 10 FIG. shows a schematic block diagram of a second electronic device 1000 that can be used to implement embodiments of the present invention. The second electronic device 1000 is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 may also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present invention described and / or claimed herein. The second electronic device 1000 may be the same as or different from the first electronic device A.
[0147] The second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory II (ROM) or a computer program loaded from a storage medium VIII into a random access memory (RAM) III. In the RAM III, various programs and data required for the operation of the device 1000 can also be stored. The computing unit I, the ROM II, and the RAM III are connected to each other through a bus IV. An input / output (I / O) interface V is also connected to the bus IV.
[0148] A plurality of components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard, a mouse, etc.; an output unit VII, such as various types of displays, speakers, etc.; a storage medium VIII, such as a magnetic disk, an optical disc, etc.; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0149] The computing unit I can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit I include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various special artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit I executes the various methods and processes described above, such as method steps S1 - S4. For example, in some embodiments, the method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 1000 via the ROM II and / or the communication unit IX. When the computer program is loaded into the RAM III and executed by the computing unit I, one or more steps of the method described above can be executed. Alternatively, in other embodiments, the computing unit I can be configured to execute the method by any other appropriate means (e.g., by means of firmware).
[0150] Although the embodiments of the present invention have been disclosed as above, they are not limited to only the applications listed in the specification and the embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to specific details and the examples shown and described herein.
Claims
1. An orbital flaw detection method for machine vision AI large models, characterized in that, Including: An image expansion step, obtaining the original image with marked track damage labels, using a generative model to perform sample expansion on the original images marked as scarce damage to obtain expanded images, and aggregating all the original images and the expanded images as a training image set; constructing a track flaw detection model including the local feature extraction module, the global feature extraction module, a feature fusion module, a classifier, and a description generator; A feature extraction step, respectively extracting the local damage features and the global damage features of the images in the training image set through the local feature extraction module and the global feature extraction module to obtain multi-dimensional information such as the dynamic change features of cracks, the damage edge features, and the spatial frequency distribution; A model training step, performing multi-scale fusion of the local damage features and the global damage features through the feature fusion module to obtain fused features, obtaining the predicted damage type and generating a semantic description through the classifier and the description generator, and constructing a loss function according to the predicted damage type and the damage labels to train the track flaw detection model; A model flaw detection step, inputting the track image to be flaw detected into the trained track flaw detection model to obtain its damage type and semantic description.
2. The rail flaw detection method for the machine vision AI large model according to claim 1, characterized in that The generative model is a generative adversarial network including a generator and a discriminator; The generator G takes a random noise vector z as input and generates a generated image similar to the real track damage image x through learning; the discriminator D is used to distinguish whether the input image is a real track damage image and promotes the optimization of the generator G through feedback; the generator and the discriminator interact with each other in the adversarial training, and the objective function is: Among them, L adv is the adversarial loss, x is the real sample, G(z) is the generated image, and D(x) and D(G(z)) are the discrimination results of the discriminator for the real image and the generated image, respectively; Perceptual loss L per It is used to measure the difference between the generated image and the real track damage image in the high-level feature space; Among them, represents the deep features extracted by the pre-trained network; the adversarial network is trained by this perceptual loss and this adversarial loss.
3. The rail flaw detection method for machine vision AI large model according to claim 2, wherein, The generator includes an attention generator to ensure that the noise vector z only exists in the damaged area of the generated image, and the structure of the attention generator is as follows: A(x) = σ(Conv s (x)) · σ(Conv c (x)) Among them, σ is the sigmoid function, Conv s and Conv c are the convolution operations of the spatial and channel attention modules respectively.
4. The rail flaw detection method for the machine vision AI large model according to claim 2 or 3, characterized in that, The generation network is a multi-stage generation network that uses a focus mask pyramid {m0, m1,..., m n} to guide different stages of the generation process; the input for each stage includes the product of the random noise and the focus mask, where only the damaged area contains noise and other areas are set to zero; the generation process can be expressed as: where y n is the generated image generated in the nth stage, z n is the random noise, m n is the focus mask in the nth stage, is the result of upsampling the image generated in the previous stage.
5. The rail flaw detection method for machine vision AI large models according to claim 1 or 4, characterized in that, The local feature extraction module is used to extract the detailed features of track damage; extracting the basic features of the image data input to the local feature extraction module, and then performing batch normalization: Among them, is the normalized image feature, and μ and σ 2 are the mean and variance respectively, and ∈ is the stabilizing term; the normalized feature is processed by the smoothing and non-linear enhancement ability of the Swish activation function: Swish(x) = x·σ(x) The local feature extraction module constructs a feature extraction block through a combination of dilated convolution, depthwise separable convolution, and pointwise convolution, and the output of the feature extraction block is: y i = f(x i *(k d + k p )) + b Among them, k d is the extended convolution kernel, and k p is the pointwise convolution kernel. * represents the convolution operation, f is the Swish activation function; the features output by the local feature extraction module pass through the global pooling layer and the fully connected layer to generate the local feature output F l ; The global feature extraction network receives the local features and performs batch normalization on them, and then processes them through a ReLU activation function; capturing global features at different scales through multi-dilation rate convolution: Among them, is a convolution kernel with a dilation rate of d i , which can capture features of different scales through parallel operations with multiple dilation rates; fuse multi-scale features through skip connections to achieve joint expression of global and local features; the features after multi-scale fusion are further dimension-reduced through a multi-scale pooling layer to obtain the global feature output F g .
6. The rail flaw detection method for the machine vision AI large model according to claim 5, characterized in that, Local feature output F l and global feature output F g Semantic interaction is achieved through attention weighting: f c = α·F l + (1 - α)·F g α = σ(MLP([F l ; F g )) where [F l ; F g represents the concatenation of local and global features, and σ is the Sigmoid activation function. MLP is a multi-layer perceptron used for dynamic interaction modeling of local and global features.
7. An orbital flaw detection device for machine vision AI large models, characterized in that, Including: An image expansion module, obtaining the original image with marked track damage labels, using a generative model to perform sample expansion on the original images marked as scarce damage to obtain expanded images, and aggregating all the original images and the expanded images as a training image set; constructing a track flaw detection model including the local feature extraction module, the global feature extraction module, a feature fusion module, a classifier, and a description generator; A feature extraction module, respectively extracting the local damage features and the global damage features of the images in the training image set through the local feature extraction module and the global feature extraction module to obtain multi-dimensional information such as the dynamic change features of cracks, the damage edge features, and the spatial frequency distribution; A model training module, which performs multi-scale fusion on the local damage feature and the global damage feature through the feature fusion module to obtain a fused feature, and obtains a predicted damage type and generates a semantic description through the classifier and the description generator. A loss function is constructed based on the predicted damage type and the damage label to train the rail flaw detection model; A model flaw detection module, which inputs a rail image to be flaw detected into the trained rail flaw detection model to obtain its damage type and semantic description.
8. An electronic device, characterized in that, It includes the rail flaw detection device for the machine vision AI large model described in claim 7. The electronic device is connected to an information display device, and the information display device is used to display the damage type with display parameters, attributes set by the user or through an artificial intelligence model.
9. A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the rail flaw detection method for the machine vision AI large model described in any one of claims 1-6 are implemented.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the rail flaw detection method for the machine vision AI large model described in any one of claims 1-6 are implemented.