A method for locating faults in railway vehicle images

By obtaining marked and unmarked railway vehicle image samples, preprocessing and feature fusion, and building a deep learning model, the problems of low manual inspection efficiency and environmental impact in railway vehicle fault detection are solved, and the rapid and accurate positioning and efficient detection of faults are achieved.

CN118761961BActive Publication Date: 2025-05-27INST OF COMPUTING TECH CHINA ACAD OF RAILWAY SCI +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410756099.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-13
Publication Date
2025-05-27
Estimated Expiration
2044-06-13

AI Technical Summary

Technical Problem

The existing railway vehicle fault detection methods rely on manual inspection, which are inefficient and error-prone. Image recognition technology is difficult to accurately locate in changing environments, light and climate change affect clarity, and rely heavily on manual labeling.

Method used

By obtaining marked and unmarked railway vehicle image samples, preprocessing and feature fusion, building a model training set, using deep learning technology to train the failed component positioning model, reducing the influence of human factors, and improving detection accuracy and efficiency.

Benefits of technology

It realizes rapid and accurate positioning of railway vehicle failures, reduces human errors, improves the reliability and accuracy of inspection results, improves detection efficiency, and provides solid guarantees for safe railway transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118761961B_ABST
    Figure CN118761961B_ABST
Patent Text Reader

Abstract

The present application provides a method for locating railway vehicle image faults, including: obtaining marked railway vehicle image samples and unmarked TFDS image samples; performing preprocessing and fusion processing to obtain pre-trained image samples and fused image samples; generating marked vehicle component image samples according to vehicle component location labels, performing preprocessing to obtain training image samples, and constructing a training image set; semi-automatically annotating unmarked vehicle component image samples with the pre-trained image samples to construct a pre-trained image set, and training a large model to obtain a fault component location pre-trained model, which is trained with the training image set to obtain a fault component location model; inputting the TFDS image to be processed into the fault component location model to locate the positions of vehicle fault components on the TFDS image to be processed. Thereby, the influence of human factors on the detection results is reduced, the reliability and accuracy of the detection results are improved, the efficiency and accuracy of railway vehicle fault detection are greatly enhanced, and a solid guarantee is provided for railway safety transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and particularly relates to a method for locating faults in railway vehicle images. Background Art

[0002] The rapid expansion of the railway transportation industry and the increase in train speed have put forward higher requirements for the safety, reliability and operation efficiency of vehicles. Vehicle faults are not only a common cause of train delays and suspensions, but also pose a major hidden danger to passenger safety. Therefore, achieving rapid and accurate fault location has become the key to maintaining the safe and smooth operation of railway transportation. Traditional railway vehicle fault detection methods mainly rely on manual visual inspection, facing the dilemmas of large workload and low efficiency, and are prone to errors due to human factors, affecting the detection quality. Although the railway has adopted the Truck Running Fault Detection System (TFDS) on the trackside, an advanced system integrating high-speed imaging, big data processing, precise positioning and automatic control, which has strengthened the safety monitoring ability, this system still highly relies on manual image review, accompanied by high labor costs and subjective risks. For example, the fatigue and experience differences of inspectors may both lead to misjudgments.

[0003] In recent years, the leap of computer vision and deep learning technologies has opened up a new path for the recognition and detection of railway vehicle fault images. However, there are still bottlenecks in the precise positioning of fault components in existing image recognition technologies. Especially in the changeable railway operation environment, the fluctuations of external conditions such as light and climate seriously affect the image clarity and fault recognition accuracy. In addition, the high dependence on manual annotation is also a major challenge restricting the development of the technology. Summary of the Invention

[0004] To solve the above technical problems, the present application provides a method for locating railway vehicle image faults, including: obtaining labeled railway vehicle image samples and unlabeled TFDS image samples; preprocessing the labeled railway vehicle image samples with the unlabeled TFDS image samples and forming image samples for model pre-training with the unlabeled TFDS image samples; obtaining TFDS image samples and performing image feature fusion on the TFDS image samples to obtain TFDS fused image samples; generating labeled vehicle component image samples according to vehicle component location labels; preprocessing the labeled vehicle component image samples and forming image samples for model training with the TFDS fused image samples; semi-automatically annotating some unlabeled vehicle component image samples according to the image samples for model pre-training to construct an image set for model pre-training; constructing an image set for model training based on the image samples for model training and the vehicle component location labels; training a target large model based on the image set for model pre-training to obtain a pre-trained model for fault component location; training the pre-trained model for fault component location based on the image set for model training to obtain a model for fault component location; obtaining a TFDS image to be processed and inputting it into the model for fault component location to locate the position of the vehicle fault component on the TFDS image to be processed. In the present application, by providing a method for locating railway vehicle image faults, the influence of human factors on the detection results is reduced, the reliability and accuracy of the detection results are improved, and the efficiency and accuracy of railway vehicle fault detection are greatly enhanced, providing a solid guarantee for railway safety transportation. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1 is a flowchart of a method for locating railway vehicle image faults according to an embodiment of the present application. DETAILED DESCRIPTION

[0006] Figure 1 is a flowchart of a method for locating railway vehicle image faults according to an embodiment of the present application. As Figure 1 shown, it includes:

[0007] S101. Obtain labeled railway vehicle image samples and unlabeled TFDS image samples; S102. Preprocess the labeled railway vehicle image samples with the unlabeled TFDS image samples and form image samples for model pre-training with the unlabeled TFDS image samples; S103. Obtain TFDS image samples and perform image feature fusion on the TFDS image samples to obtain TFDS fused image samples;

[0008] S104. Generate a marked vehicle component image sample according to the vehicle component positioning label; S105. Preprocess the marked vehicle component image sample and form an image sample for model training by fusing it with the TFDS fusion image sample; S106. Semi-automatically annotate some unmarked vehicle component image samples according to the image sample for model pre-training to construct an image set for model pre-training; S107. Construct an image set for model training based on the image sample for model training and the vehicle component positioning label; S108. Train a target large model based on the image set for model pre-training to obtain a pre-trained model for fault component positioning; S109. Train the pre-trained model for fault component positioning based on the image set for model training to obtain a fault component positioning model; S110. Obtain a TFDS image to be processed and input it into the fault component positioning model to locate the position of the vehicle fault component on the TFDS image to be processed.

[0009] In this application, by providing a method for positioning railway vehicle image faults, a marked railway vehicle image sample and an unmarked TFDS image sample are obtained; the unmarked TFDS image sample is used to preprocess the marked railway vehicle image sample and form an image sample for model pre-training by fusing it with the unmarked TFDS image sample; a TFDS image sample is obtained, and image feature fusion is performed on the TFDS image sample to obtain a TFDS fusion image sample; a marked vehicle component image sample is generated according to the vehicle component positioning label; the marked vehicle component image sample is preprocessed and form an image sample for model training by fusing it with the TFDS fusion image sample; some unmarked vehicle component image samples are semi-automatically annotated according to the image sample for model pre-training to construct an image set for model pre-training; an image set for model training is constructed based on the image sample for model training and the vehicle component positioning label; a target large model is trained based on the image set for model pre-training to obtain a pre-trained model for fault component positioning; the pre-trained model for fault component positioning is trained based on the image set for model training to obtain a fault component positioning model; a TFDS image to be processed is obtained and input into the fault component positioning model to locate the position of the vehicle fault component on the TFDS image to be processed. Thereby, the influence of human factors on the detection result is reduced, the reliability and accuracy of the detection result are improved, the efficiency and accuracy of railway vehicle fault detection are greatly improved, and a solid guarantee is provided for railway safe transportation.

[0010] Optionally, the positioning method further includes: digitally collecting a railway vehicle to obtain a digital image; constructing a structural model of each vehicle component based on the digital image, and based on the structural model of each vehicle component, semi-automatically annotating some unlabeled vehicle component image samples according to the image samples for model pre-training, so as to construct an image set for model pre-training.

[0011] In this embodiment, through the digital collection of railway vehicles, detailed digital images are obtained, and further based on these images, the structural models of each vehicle component are accurately constructed, laying a solid foundation for subsequent model pre-training. Secondly, through the structural models of each vehicle component, some unlabeled vehicle component image samples are semi-automatically annotated more efficiently, greatly reducing the workload of manual annotation and improving the accuracy and efficiency of annotation. Finally, these semi-automatically annotated image samples are integrated into an image set for model pre-training, providing rich and high-quality data support for the training of the subsequent fault component positioning model, thereby significantly improving the accuracy and reliability of fault component positioning.

[0012] An exemplary code is provided in an embodiment of the present application for implementing the positioning method of railway vehicle components. The method includes digitally collecting a railway vehicle to obtain a digital image, constructing a structural model of each vehicle component based on these images, and then based on these structural models, semi-automatically annotating some unlabeled vehicle component image samples to construct an image set for model pre-training.

[0013] import cv2

[0014] import numpy as np

[0015] from sklearn.cluster import KMeans

[0016] from PIL import Image

[0017] import os

[0018] # Digitally collect a railway vehicle to obtain a digital image

[0019] # Image file path

[0020] def digitize_railway_vehicle(vehicle_id, image_folder):

[0021] digitized_images = []

[0022] for filename in os.listdir(image_folder):

[0023] if filename.startswith(f'vehicle_{vehicle_id}_'):

[0024] image_path = os.path.join(image_folder, filename)

[0025] image = cv2.imread(image_path)

[0026] if image is not None:

[0027] digitized_images.append(image)

[0028] return digitized_images

[0029] # According to the digitized images, use KMeans for color segmentation to simulate the structural model

[0030] def build_component_models(digitized_images, n_clusters):

[0031] component_models = {}

[0032] for i, image in enumerate(digitized_images):

[0033] # Convert to floating point and reshape to a two-dimensional array

[0034] pixels = image.reshape((-1, 3)).astype(np.float32)

[0035] # Use KMeans for color segmentation

[0036] kmeans = KMeans(n_clusters=n_clusters, random_state=0).fit(pixels)

[0037] # Replace each pixel with its nearest cluster center

[0038] segmented_image = kmeans.cluster_centers_[kmeans.labels_].astype(np.uint8)

[0039] # Reshape back to the original image shape

[0040] segmented_image = segmented_image.reshape(image.shape)

[0041] # Each segmented part represents a vehicle component

[0042] component_models[f'part_{i + 1}'] = segmented_image

[0043] return component_models

[0044] # Semi - automatic annotation

[0045] def semi_automatic_annotation(component_models, unlabeled_images_folder):

[0046] annotated_images = []

[0047] for filename in os.listdir(unlabeled_images_folder):

[0048] image_path = os.path.join(unlabeled_images_folder, filename)

[0049] image = cv2.imread(image_path)

[0050] # Integrate annotation tools for annotation

[0051] # Add the image to the annotation list

[0052] annotated_images.append(image)

[0053] # Annotation is complete

[0054] return annotated_images

[0055] # Build the image set for model pre-training

[0056] def build_pretraining_dataset(annotated_images, output_folder):

[0057] if not os.path.exists(output_folder):

[0058] os.makedirs(output_folder)

[0059] for i, image in enumerate(annotated_images):

[0060] # Save the original image and perform preprocessing steps such as data augmentation and normalization

[0061] output_path = os.path.join(output_folder, f'annotated_image_{i}.jpg')

[0062] cv2.imwrite(output_path, image)

[0063] print(f"Pretraining dataset saved to {output_folder}")

[0064] In the above code, the digitize_railway_vehicle function converts a physical vehicle into a digital image that can be processed by a computer, laying the foundation for subsequent modeling and analysis. Then, the build_component_models function uses these digital images to construct the structural models of each vehicle component through image segmentation technology. These models are the key to identifying and locating vehicle components. Subsequently, the semi_automatic_annotation function simulates the semi-automatic annotation process, annotates unlabeled images, and integrates image annotation tools to support more accurate and efficient annotation work. The build_pretraining_dataset function integrates the annotated images into a pre-training dataset. After preprocessing and augmentation, these datasets are used to train localization or recognition models to improve the performance of the models.

[0065] Optionally, the digital acquisition of the railway vehicle to obtain a digital image includes: based on a set acquisition strategy, performing multi-batch digital acquisitions on different railway vehicles of the same model to obtain digital images corresponding to the railway vehicles of the same model. The acquisition strategy is configured based on the used acquisition device and acquisition method. The acquisition device includes a line array camera, a area array camera, and a laser scanner. The acquisition methods include line array camera shooting, area array camera shooting, and laser scanner scanning.

[0066] In this embodiment, by setting a clear acquisition strategy, selecting a suitable acquisition device and acquisition method according to actual needs, such as a line array camera, an area array camera, or a laser scanner, and corresponding acquisition methods such as line array camera shooting, area array camera shooting, or laser scanner scanning, multi-batch digital acquisitions are performed on different railway vehicles of the same model. At the same time, this acquisition method ensures the comprehensiveness and accuracy of the data, and also greatly improves the efficiency and flexibility of data acquisition. Through the rich digital images collected, a more refined and accurate vehicle component structure model is constructed, providing strong data support for subsequent identification and positioning work, thereby significantly enhancing the intelligent level and operation efficiency of railway vehicle management.

[0067] An exemplary code is provided in an embodiment of the present application to show how to perform multi-batch digital acquisitions on different railway vehicles of the same model based on a set acquisition strategy.

[0068] import os

[0069] import random

[0070] from abc import ABC, abstractmethod

[0071] # Acquisition device class (abstract base class)

[0072] class CaptureDevice(ABC):

[0073] @abstractmethod

[0074] def capture_image(self, vehicle_id):

[0075] pass

[0076] # Line array camera class

[0077] class LineScanCamera(CaptureDevice):

[0078] def capture_image(self, vehicle_id):

[0079] # Call the SDK function to actually capture the image

[0080] # Return the image file name

[0081] return f"line_scan_image_{vehicle_id}.jpg"

[0082] # Area scan camera class

[0083] class AreaScanCamera(CaptureDevice):

[0084] def capture_image(self, vehicle_id):

[0085] # Call the SDK function to capture the image

[0086] # Return a simulated image file name

[0087] return f"area_scan_image_{vehicle_id}.jpg"

[0088] # Simulated laser scanner class

[0089] class LaserScanner(CaptureDevice):

[0090] def capture_image(self, vehicle_id):

[0091] # Laser scanners usually generate point cloud data, return an image

[0092] return f"laser_scan_image_{vehicle_id}.jpg"

[0093] # Acquisition strategy configuration

[0094] def configure_capture_strategy(device_type):

[0095] if device_type == 'line_scan':

[0096] return LineScanCamera()

[0097] elif device_type == 'area_scan':

[0098] return AreaScanCamera()

[0099] elif device_type == 'laser_scan':

[0100] return LaserScanner()

[0101] else:

[0102] raise ValueError("Invalid device type")

[0103] # Multi-batch digitization acquisition function

[0104] def digitize_railway_vehicles(vehicle_model, n_batches, device_type,image_folder):

[0105] if not os.path.exists(image_folder):

[0106] os.makedirs(image_folder)

[0107] =capture_device = configure_capture_strategy(device_type)

[0108] for batch in range(n_batches):

[0109] for vehicle_id in range(1, 101):# Capture 100 vehicles per batch

[0110] image_filename = capture_device.capture_image(vehicle_id)

[0111] # Record the file name and save the processed image data

[0112] print(f"Batch {batch+1}: Captured image for vehicle {vehicle_id} of model {vehicle_model}, saved as {image_filename}")

[0113] # Implement the saving logic to save the image to the folder

[0114] # save_image(os.path.join(image_folder, image_filename))

[0115] In this exemplary code, the abstract base class CaptureDevice represents a capture device, and three simulated subclasses are created for it: LineScanCamera, AreaScanCamera, and LaserScanner. Each subclass has a capture_image method that calls the corresponding hardware SDK to capture images. The configure_capture_strategy function is used to configure the capture strategy according to the device type and return the corresponding device instance. The digitize_railway_vehicles function is a process for multi-batch digitized capture of different railway vehicles of the same model, which is used to receive the vehicle model, the number of batches, the device type, and the image storage folder as parameters, and perform simulated image capture based on these parameters.

[0116] Optionally, constructing the structural models of vehicle components based on the digitized images includes:

[0117] Performing fusion processing on the digitized images corresponding to different railway vehicles of the same model to construct the structural models of vehicle components.

[0118] In this embodiment, the digitized images corresponding to different railway vehicles of the same model are subjected to fusion processing, making full use of the image data collected in multiple batches and from multiple angles to eliminate the possible information loss or perspective limitations of a single image. Moreover, the fusion processing can integrate the useful information in multiple images, thereby constructing a more complete, accurate, and detailed structural model of vehicle components, which not only helps with subsequent vehicle component identification and positioning work but also improves the intelligent level of the entire railway vehicle management and provides more reliable technical support for railway operations.

[0119] This application embodiment provides an exemplary code showing how to perform fusion processing on the digitized images of different railway vehicles of the same model.

[0120] import cv2

[0121] import numpy as np

[0122] import os

[0123] from PIL import Image

[0124] # Obtain the directory containing different digital images of railway vehicles of the same model

[0125] images_dir = 'path_to_images_directory'

[0126] # Load all images of vehicles of the same model

[0127] images = []

[0128] for filename in os.listdir(images_dir):

[0129] if filename.endswith('.jpg') or filename.endswith('.png'): # Only process image files

[0130] img_path = os.path.join(images_dir, filename)

[0131] img = cv2.imread(img_path)

[0132] if img is not None: # Ensure the image is loaded successfully

[0133] images.append(img)

[0134] # Ensure all images have the same dimensions

[0135] if images:

[0136] base_shape = images[0].shape

[0137] for img in images:

[0138] if img.shape != base_shape:

[0139] print(f"Warning: Image {filename} has a different shape and will beresized.")

[0140] img = cv2.resize(img, (base_shape[1], base_shape[0]))

[0141] # Image fusion, using the average fusion algorithm

[0142] fused_image = np.zeros(base_shape, dtype=np.float32)

[0143] for img in images:

[0144] fused_image = np.add(fused_image, img)

[0145] fused_image = np.divide(fused_image, len(images)) # Calculate the average value

[0146] # Convert the fused image to an 8-bit unsigned integer and save it

[0147] fused_image_uint8 = np.uint8(np.clip(fused_image, 0, 255))

[0148] fused_image_path = 'path_to_save_fused_image / fused_vehicle_model.jpg'

[0149] cv2.imwrite(fused_image_path, fused_image_uint8)

[0150] # Based on the fused image, use image processing techniques

[0151] # Build the structural model of each vehicle component

[0152] print(f"Fusion of {len(images)} images completed. Saved to {fused_image_path}")

[0153] In this exemplary code, all the digital images of railway vehicles of the same model are loaded and adjusted to have the same size. Then, these images are fused and the fused image is saved.

[0154] Optionally, based on the set acquisition strategy, multiple digital acquisitions are performed on different railway vehicles of the same model to obtain the digital images corresponding to the railway vehicles of the same model, including: performing multiple digital acquisitions on different railway vehicles of the same model respectively by using the same acquisition method and different acquisition devices to obtain the digital images corresponding to the railway vehicles of the same model; fusing the digital images corresponding to the railway vehicles of the same model to obtain the single-time digital image corresponding to the single sampling method; fusing multiple single-time digital images corresponding to the single sampling method to obtain the digital image corresponding to the railway vehicles of the same model under the same sampling method; and obtaining the digital image corresponding to the railway vehicles of the same model according to the digital images corresponding to the railway vehicles of the same model under different sampling methods.

[0155] In this embodiment, based on the set acquisition strategy, by using the same acquisition method but different acquisition devices, multiple digital acquisitions are performed on different railway vehicles of the same model respectively, ensuring the diversity and reliability of the data and avoiding the limitations that may be brought by a single device or acquisition method. Then, the digital images corresponding to the railway vehicles of the same model are initially fused to obtain the single-time digital image corresponding to the single sampling method. This fusion technology can integrate the information of multiple images and improve the clarity and detail expressiveness of the images. Further, multiple single-time digital images corresponding to the single sampling method are fused to obtain a more comprehensive and accurate digital image of the railway vehicles of the same model under the same sampling method. By integrating the digital images corresponding to the railway vehicles of the same model under different sampling methods, a comprehensive, detailed and multi-angle digital image of the railway vehicles of this model is constructed.

[0156] The embodiment of this application provides an exemplary code to show the logic of performing multiple digital acquisitions on different railway vehicles of the same model to obtain their corresponding digital images.

[0157] # The function performs image acquisition and returns image data

[0158] def collect_image(vehicle_id, device_id, sampling_method):

[0159] # Perform image acquisition using the specified device and sampling method

[0160] # Return image data

[0161] return f"Image_{vehicle_id}_{device_id}_{sampling_method}.jpg" # Return the image file name

[0162] # The function is used for image fusion, which is simplified to merging a list of file names here

[0163] def fuse_images(image_filenames):

[0164] # Use an image processing library to fuse images

[0165] # Merge the file names

[0166] return "Fused_" + "_".join(image_filenames)

[0167] # Set the acquisition strategy

[0168] sampling_methods = ['method1','method2'] # List of sampling methods

[0169] device_ids = ['deviceA', 'deviceB'] # List of acquisition devices

[0170] vehicle_type = 'RailwayVehicleTypeX' # Railway vehicles of the same type

[0171] # Initialize the dictionary to store the collected images

[0172] collected_images = {}

[0173] for sampling_method in sampling_methods:

[0174] collected_images[sampling_method] = {}

[0175] for device_id in device_ids:

[0176] # Perform multiple digital acquisitions on different railway vehicles of the same type vehicle_id = f"{vehicle_type}_Vehicle{ord(device_id[0])}" # Vehicle ID

[0177] image_filename = collect_image(vehicle_id, device_id, sampling_method)

[0178] if sampling_method not in collected_images:

[0179] collected_images[sampling_method] = {}

[0180] if device_id not in collected_images[sampling_method]:

[0181] collected_images[sampling_method][device_id]= []

[0182] collected_images[sampling_method][device_id].append(image_filename)

[0183] # Fuse the digital images corresponding to railway vehicles of the same model

[0184] fused_single_sampling_images = {}

[0185] for sampling_method, device_images in collected_images.items():

[0186] for device_id, image_filenames in device_images.items():

[0187] # Fuse all the images of the same device under the same sampling method

[0188] fused_image_filename = fuse_images(image_filenames)

[0189] if sampling_method not in fused_single_sampling_images:

[0190] fused_single_sampling_images[sampling_method] = []

[0191] fused_single_sampling_images[sampling_method].append(fused_image_filename)

[0192] # Fuse multiple single-digitized images corresponding to a single sampling method

[0193] final_fused_images = {}

[0194] for sampling_method, fused_image_filenames in fused_single_sampling_images.items():

[0195] # Fuse all fused images under the same sampling method

[0196] final_fused_image_filename = fuse_images(fused_image_filenames)

[0197] final_fused_images[sampling_method] = final_fused_image_filename

[0198] # final_fused_images contains the digitized images corresponding to different sampling methods for railway vehicles of the same model

[0199] print(f"Final fused images for {vehicle_type}: {final_fused_images}")

[0200] In this exemplary code, the main task of the collect_image function is to perform image acquisition. According to the preset acquisition strategy, different acquisition devices and sampling methods are used to perform multiple digital acquisitions on different railway vehicles of the same model, ensuring that the acquired image data is diverse and accurate, thus obtaining rich image data to provide a basis for subsequent analysis and processing. Then, the fuse_images function is responsible for fusing the acquired image data. It receives multiple image files as input and then uses image processing techniques to register, align, and fuse these images, integrating the useful information in multiple images to generate a more comprehensive and accurate digital image.

[0201] Optionally, obtaining the digital image corresponding to the railway vehicle of the same model according to the digital images corresponding to the railway vehicle of the same model under different sampling methods includes: obtaining the vehicle structure digital model and the maintenance part digital model; fusing the digital images corresponding to the railway vehicle of the same model under different sampling methods, the vehicle structure digital model, and the maintenance part digital model to obtain the digital image corresponding to the railway vehicle of the same model.

[0202] In this embodiment, by comprehensively applying the digital images under different sampling methods and combining the vehicle structure digital model and the maintenance part digital model, it is possible to accurately obtain the comprehensive digital image of the railway vehicle of the same model. It not only includes the diversity of image acquisition but also combines the image information under different sampling methods with the accurate models of the vehicle structure and maintenance parts through image fusion technology, thereby constructing a detailed and accurate digital representation of the vehicle, improving the integrity of data collection, and ensuring the accuracy of the analysis results.

[0203] The embodiment of the present application provides an exemplary code to demonstrate the logic of obtaining the corresponding digital image according to the digital images of the railway vehicle of the same model under different sampling methods.

[0204] # Image file names and 3D model file paths

[0205] image_filenames_by_sampling_method = {

[0206] 'method1': ['image1_method1.jpg', 'image2_method1.jpg'],

[0207] 'method2': ['image1_method2.jpg', 'image2_method2.jpg']

[0208] }

[0209] vehicle_structure_model_path = 'vehicle_structure.3dm'# 3D model file format

[0210] repair_parts_model_path = 'repair_parts.3dm'

[0211] # Image fusion function

[0212] def fuse_images(image_filenames):

[0213] # Return the filename representing the fused image

[0214] return "fused_image.jpg"

[0215] # Simulated 3D model fusion function

[0216] def fuse_3d_models(structure_model_path, parts_model_path):

[0217] # Here, return a filename representing the fused 3D model

[0218] return "fused_3d_model.3dm"

[0219] # Image fusion process

[0220] fused_images_by_method = {}

[0221] for method, filenames in image_filenames_by_sampling_method.items():

[0222] fused_images_by_method[method] = fuse_images(filenames)

[0223] # Obtain the fused 2D image

[0224] # Obtain the fused 3D model

[0225] fused_2d_image = "fused_all_methods_2d.jpg"# Assume this is the 2D image fused by all methods

[0226] fused_3d_model = fuse_3d_models(vehicle_structure_model_path, repair_parts_model_path)

[0227] # File names for combining 2D images and 3D models,

[0228] # indicating their integration into the digital representation of the same type of railway vehicle

[0229] # Image mapping to the 3D model

[0230] combined_representation = {

[0231] '2d_image': fused_2d_image,

[0232] '3d_model': fused_3d_model

[0233] }

[0234] print(f"Combined representation for the railway vehicle model:{combined_representation}")

[0235] In this exemplary code, the collect_images function is responsible for collecting digital images of the same type of railway vehicle under multiple sampling methods. By calling different image acquisition devices or algorithms, detailed image data of the vehicle under various conditions is obtained. These image data not only cover the appearance and structure of the vehicle, but also include close-up images of internal components and key areas. Then, the fuse_digital_representations function receives the digital image data from the collect_images function, combines the digital models of the vehicle structure and repair parts, performs high-level data fusion, and combines the 2D images with the 3D model using image processing techniques and 3D modeling techniques to generate a comprehensive, accurate, and easy-to-understand digital representation.

[0236] Optionally, the set of acquisition devices composed of different acquisition methods and different acquisition devices is denoted as:

[0237]

[0238] Perform multi-batch digital acquisitions on different railway vehicles of the same model to obtain the digital images corresponding to the railway vehicles of the same model, denoted as:

[0239]

[0240] In this embodiment, by constructing a set that includes different acquisition methods and multiple acquisition devices under the same acquisition method, multi-batch digital acquisitions are realized for different railway vehicles of the same model. Moreover, the diversified set of acquisition methods ensures the diversity and accuracy of the data. Different acquisition methods and devices may capture different features and details of the vehicle. Further, the digital images contain rich information, which provides a comprehensive understanding of the vehicle state.

[0241] The embodiment of this application provides an exemplary code that represents functions and classes to represent the devices of different acquisition methods and the digital acquisition process.

[0242] class CollectionDevice:

[0243] def __init__(self, method_id, device_id):

[0244] self.method_id = method_id

[0245] self.device_id = device_id

[0246] def collect_image(self, vehicle_id):

[0247] # Image acquisition process, calling the acquisition logic

[0248] image_data = f"Simulated image data for vehicle {vehicle_id}collected by {self.device_id}"

[0249] return image_data

[0250] class RailwayVehicleDigitization:

[0251] def __init__(self, vehicle_id):

[0252] self.vehicle_id = vehicle_id

[0253] self.images = {} # Used to store image data collected multiple times

[0254] def collect_images(self, collection_devices, batch_number=1):

[0255] # batch_number differentiates different collection batches

[0256] images_batch = []

[0257] for device in collection_devices:

[0258] image_data = device.collect_image(self.vehicle_id)

[0259] images_batch.append(image_data)

[0260] # Simplified processing to record the data of the last collection

[0261] self.images[f"Batch {batch_number}"] = images_batch

[0262] print(f"Collected images for vehicle {self.vehicle_id} in batch {batch_number}")

[0263] def get_images(self, batch_number):

[0264] # Returns the image data of the specified batch

[0265] return self.images.get(f"Batch {batch_number}", [])

[0266] In this exemplary code, by defining the CollectionDevice class, the behaviors of different acquisition methods and devices are simulated. The RailwayVehicleDigitization class is responsible for coordinating the acquisition processes of these devices. Through the collect_images method, multi-batch digitization acquisition of different railway vehicles of the same model is achieved. The get_images method obtains the image data of the specified batch for subsequent analysis and processing.

[0267] Optionally, based on the following formula, the digitized images corresponding to railway vehicles of the same model under different sampling methods, the vehicle structure digital model, and the maintenance part digital model are fused to obtain the digitized images corresponding to railway vehicles of the same model:

[0268] Based on the following formula, for railway vehicles of the same model, the digitized images obtained by one-time acquisition using one acquisition method and different acquisition devices are fused and calculated to obtain the single-time digitized images of the railway vehicles of this model obtained using different acquisition devices based on this acquisition method, denoted as :

[0269] , where is the affine transformation matrix, is the acquisition weight value assigned to different acquisition devices, and the weight value is ;

[0270] Based on the following formula, the digitized images obtained from multiple acquisitions are fused to obtain the single-batch digitized images of the railway vehicles of the same model under the same acquisition method, denoted as :

[0271]

[0272] where is the affine transformation matrix, β is the weight value assigned to different batch acquisitions, and the weight value β is ;

[0273] Based on the following formula, the digitized images obtained from different acquisition methods for railway vehicles of the same model are fused to obtain the digitized images of railway vehicles of the same model, denoted as :

[0274]

[0275] where is the fusion constant, is the weight value assigned to different acquisition methods, is ;

[0276] Fuse the digital images corresponding to the same type of railway vehicle under different sampling methods, the digital model of the vehicle structure, and the digital model of the maintenance parts to obtain the digital image corresponding to the same type of railway vehicle, denoted as :

[0277] , where 、 、 are the digital images corresponding to the same type of railway vehicle as weight values respectively. The weight values can take , is the digital model of the vehicle structure, is the digital model of the maintenance parts.

[0278] In this embodiment, by assigning weights to different acquisition devices, the images acquired by different devices under the same acquisition method are effectively fused to obtain an optimized image under a single acquisition. Subsequently, using the data acquired multiple times, image fusion is performed batch by batch to ensure the consistency and stability of the data, thereby obtaining the single-batch digital image S1 of the railway vehicle of this type under a certain acquisition method. Further, the digital images S1 obtained under different acquisition methods are fused again to eliminate potential differences and form a unified digital image S2. By fusing this unified digital image with the digital model of the vehicle structure and the digital model of the maintenance parts, a complete digital representation S integrating multi-faceted information is obtained. This representation not only includes the overall appearance and internal structure of the vehicle, but also incorporates detailed information of the maintenance parts, which not only improves the comprehensiveness and accuracy of the data, but also provides strong data support for subsequent vehicle maintenance, fault diagnosis, and performance optimization, thus significantly enhancing the efficiency and effectiveness of railway vehicle management and maintenance.

[0279] This application embodiment provides an exemplary code to demonstrate the implementation method of forming a digital representation of a vehicle by combining the digital images of the same type of railway vehicle, the digital model of the vehicle structure, and the digital model of the maintenance parts under different sampling methods using a technology fusion method.

[0280] # Function for loading images, models, and performing affine transformation and image fusion

[0281] def apply_affine_transform_and_weight(image, transform_matrix,weight):

[0282] # Apply affine transformation and weight

[0283] # Perform image processing logic

[0284] def fuse_images_by_device(images, transform_matrix, weights):

[0285] # Fuse images according to the device and weights

[0286] fused_image = 0

[0287] for image, weight in zip(images, weights):

[0288] fused_image += apply_affine_transform_and_weight(image, transform_matrix, weight)

[0289] return fused_image

[0290] def fuse_images_by_batch(batches_of_images, transform_matrix):

[0291] # Fuse images according to the batch

[0292] fused_image = 0

[0293] total_batches = len(batches_of_images)

[0294] for batch in batches_of_images:

[0295] fused_image += sum(batch) / len(batch)# batch is a list of images from different devices in the same batch

[0296] fused_image *= transform_matrix# transform_matrix is the fusion constant

[0297] return fused_image / total_batches# Divide by the number of batches to get the average fused image

[0298] def fuse_images_by_method(fused_images_by_method, weights):

[0299] # Fuse images according to the acquisition method

[0300] total_methods = len(fused_images_by_method)

[0301] fused_image = 0

[0302] for image, weight in zip(fused_images_by_method, weights):

[0303] fused_image += image * weight

[0304] return fused_image

[0305] def fuse_with_models(fused_image, vehicle_structure_model, repair_parts_model, a, b, c):

[0306] # Fuse with the vehicle structure model and repair parts model

[0307] return a * fused_image + b * vehicle_structure_model + c * repair_parts_model

[0308] In this exemplary code, the fuse_images_by_device function fuses the images collected by different devices under the same acquisition method according to the acquisition weight values of the devices. Then, the fuse_images_by_batch function further fuses the images obtained from multiple acquisitions to obtain a single batch of digital images under this acquisition method. Then, the fuse_images_by_method function fuses the images obtained from different acquisition methods to obtain the overall digital image of this type of railway vehicle. The fuse_with_models function fuses the digital image with the vehicle structure digital model and repair parts digital model to obtain a more comprehensive and accurate digital representation of the railway vehicle.

[0309] Optionally, the positioning method further includes: establishing a physical structure model of the railway vehicle according to the structure diagram of the railway vehicle; fusing the TFDS historical train operation images with the physical structure model of the railway vehicle to generate a vehicle structure digital model.

[0310] In this embodiment, by using the structure diagram of the railway vehicle, a physical structure model of the vehicle is accurately established, which lays a solid foundation for subsequent digital processing. Then, the historical running images of TFDS are fused with the established physical structure model to achieve the transformation from image data to the digital model of the vehicle structure. At the same time, this fusion process not only effectively combines the actual images with the theoretical model, but also visually presents the detailed information of the vehicle structure through digital means, quickly identifies potential fault points of the vehicle, improves the pertinence and efficiency of maintenance work, and thus ensures the safe operation of railway vehicles. This digital model is also convenient for storage and sharing, providing strong support for the informatization and intelligentization of railway vehicle management and maintenance.

[0311] An exemplary code is provided in the embodiment of this application to demonstrate the process of establishing the physical structure model of the railway vehicle and fusing the historical running images of TFDS to generate the digital model of the vehicle structure.

[0312] import numpy as np

[0313] import cv2 # Use OpenCV for image processing

[0314] from some_3d_modelling_library import create_3d_model # 3D modeling library

[0315] # Create a physical structure model according to the structure diagram

[0316] def create_physical_model(structure_diagram):

[0317] # structure_diagram is a data structure containing information such as vehicle dimensions and component positions

[0318] # Return the physical structure model representation, including dictionaries of dimensions and positions

[0319] model_params = {

[0320] 'length': structure_diagram['length'],

[0321] 'width': structure_diagram['width'],

[0322] 'height': structure_diagram['height'],

[0323] # The list 'components' contains component information

[0324] 'components': structure_diagram['components']

[0325] }

[0326] return model_params

[0327] # Fuse historical driving images with the physical structure model to generate a digital model

[0328] def fuse_images_with_model(tfds_images, physical_model):

[0329] # The list 'tfds_images' contains multiple images

[0330] # The 'physical_model' is the physical structure model created by the above function

[0331] # Process each image and annotate the position and size of the physical structure model on the image

[0332] # Initialize an empty 3D scene for the digital model

[0333] digital_model = create_3d_model()

[0334] # Iterate through each image and add the image and annotation to the corresponding position in the digital model

[0335] for image in tfds_images:

[0336] # Process the image to match the physical structure model

[0337] processed_image = cv2.resize(image, (int(physical_model['width']),int(physical_model['height'])))

[0338] # Add the processed image and annotation to the digital model

[0339] # Render the image into the 3D scene

[0340] print(f"Adding processed image to digital model at position({physical_model['length']}, {physical_model['width']}, {physical_model['height']})")

[0341] # Return the generated digital model of the vehicle structure

[0342] return digital_model

[0343] In this exemplary code, a physical structure model is created based on the structure diagram of the railway vehicle. This model contains the key dimensions and component information of the vehicle. Then, the physical structure model is fused with the TFDS historical train operation images. Through image processing technology and 3D modeling technology, a digital model of the vehicle structure is generated. This digital model not only contains the physical dimensions and component information of the vehicle, but also incorporates the image data during the actual train operation process.

[0344] Optionally, the fusing of the TFDS historical train operation images with the physical structure model of the railway vehicle to generate a digital model of the vehicle structure includes: fusing the TFDS historical train operation images with the physical structure model of the railway vehicle to obtain the initial physical structure model of the railway vehicle; based on the component fine-tuning knowledge description, adjusting the initial physical structure model of the railway vehicle to generate a digital model of the vehicle structure.

[0345] In this embodiment, advanced image processing and 3D modeling technologies are adopted to effectively fuse the TFDS historical train operation images with the physical structure model of the railway vehicle, thereby generating the initial physical structure model of the vehicle. The visual information in the historical train operation images is fully utilized, and the precise dimensions and component positions of the physical structure model are also combined. Secondly, based on the component fine-tuning knowledge description, this initial model is finely adjusted to ensure the accuracy and integrity of each component. Through this fusion and adjustment process, a high-precision and high-reliability digital model of the vehicle structure is finally obtained. This model not only provides an intuitive and accurate reference for vehicle maintenance and fault detection, but also provides strong support for vehicle performance analysis and optimization.

[0346] An exemplary code is provided in an embodiment of this application:

[0347] import numpy as np

[0348] from some_image_processing_library import process_image # Image processing library

[0349] from some_3d_modelling_library import create_3d_model, adjust_model # 3D modeling library

[0350] # Fuse TFDS images and physical structure model to obtain the initial physical structure model

[0351] def fuse_images_with_model(tfds_images, physical_model):

[0352] # Assume tfds_images contains a list of multiple images

[0353] # physical_model contains a dictionary or object with vehicle physical structure information

[0354] # Initialize an empty 3D model

[0355] initial_3d_model = create_3d_model(physical_model)

[0356] # The process_image function can process the image and extract key information

[0357] # Add the processed result of each image to the 3D model

[0358] for image in tfds_images:

[0359] processed_image_data = process_image(image) # The processed data can be directly used for the 3D model

[0360] # Integrate processed_image_data into initial_3d_model

[0361] return initial_3d_model

[0362] # Adjust the model based on the component fine-tuning knowledge description

[0363] Def adjust_model_with_knowledge(initial_model, fine_tuning_knowledge):

[0364] # fine_tuning_knowledge contains a dictionary or object with fine-tuning information such as part positions and dimensions

[0365] adjusted_model = adjust_model(initial_model, fine_tuning_knowledge)

[0366] return adjusted_model

[0367] In this exemplary code, the fuse_images_with_model function is used to fuse TFDS historical train images with the physical structure model of a railway vehicle, combining the visual information in the images with the exact dimensions and part positions of the physical model, thus generating an initial physical structure model of the railway vehicle. Then, the adjust_model_with_knowledge function is used to adjust the initial physical structure model based on the part fine-tuning knowledge description, ensuring that each part in the model meets the actual dimension and position requirements. Finally, a high-precision and high-reliability digital model of the vehicle structure is generated.

[0368] Optionally, based on the structural models of the various parts of the vehicle, semi-automatically annotate some unlabeled vehicle part image samples according to the image samples for model pre-training to construct an image set for model pre-training, including: locating the vehicle parts on the image samples for model pre-training based on the structural models of the various parts of the vehicle; obtaining the annotation points of the vehicle parts on the image samples for model pre-training;

[0369] Perform inference on some unlabeled vehicle part image samples according to the positioning of the vehicle parts and the annotation points of the vehicle parts to label the vehicle parts on the some unlabeled vehicle part image samples and construct an image set for model pre-training.

[0370] In this embodiment, by adopting advanced image processing and machine learning technologies and based on the structural models of various vehicle components, the vehicle components on the image samples for model pre-training are effectively and accurately located. This not only improves the accuracy of annotation but also reduces the workload of manual annotation. Secondly, the annotation points of the vehicle components on these image samples are automatically obtained to ensure the accuracy of the annotation information. Based on the location and annotation point information of the vehicle components, intelligent inference is performed on some unlabeled vehicle component image samples to achieve automatic or semi-automatic marking of the vehicle components on these samples. This not only improves the efficiency and accuracy of image annotation but also helps to construct a richer and higher-quality image set for model pre-training, laying a solid foundation for subsequent model training and performance improvement.

[0371] The embodiment of this application provides an exemplary code to show how to achieve semi-automatic annotation of vehicle component images by combining image processing and deep learning technologies, so as to construct an image set for model pre-training.

[0372] import cv2 # Use OpenCV for image processing

[0373] import numpy as np

[0374] from tensorflow.keras.models import load_model # Use the Keras API of TensorFlow to load the pre-trained model

[0375] # Locate vehicle components based on the vehicle component structure model

[0376] def locate_vehicle_parts(image, part_models):

[0377] # part_models is a dictionary where the keys are part names and the values are pre-trained models for location

[0378] located_parts = {}

[0379] for part_name, part_model in part_models.items():

[0380] # Use a prediction function to locate the parts

[0381] # part_model.predict returns the location of the part

[0382] part_location = part_model.predict(image) located_parts[part_name] = part_location

[0383] return located_parts

[0384] # Get the annotation points of vehicle parts

[0385] def get_annotation_points(image, part_locations):

[0386] # part_locations is a dictionary containing the location information of vehicle parts

[0387] annotation_points = {}

[0388] for part_name, location in part_locations.items():

[0389] # Each part location contains the information required for annotation points annotation_points[part_name] = location

[0390] return annotation_points

[0391] # Infer and annotate unlabeled image samples

[0392] def infer_and_annotate_unlabeled_images(unlabeled_images, labeled_images, part_models):

[0393] # labeled_images contains the labeled image samples for assisting in inference

[0394] annotated_images = []

[0395] for image in unlabeled_images:

[0396] located_parts = locate_vehicle_parts(image, part_models)

[0397] # Use the labeled images to optimize the annotation results

[0398] annotation_points = get_annotation_points(image, located_parts)

[0399] # annotation_points now contains enough information to generate the annotated image

[0400] # Add annotation_points to the result list. In practice, it may be necessary to generate the annotated image file

[0401] annotated_images.append((image, annotation_points))

[0402] # Construct the image set for model pre-training

[0403] return annotated_images

[0404] In this exemplary code, some functions are defined to simulate the processes of locating vehicle parts, obtaining annotation points, and inferring and annotating unlabeled images. Specifically, the locate_vehicle_parts function uses a pre-trained vehicle part location model to locate vehicle parts in the image. Secondly, the get_annotation_points function obtains the annotation points of vehicle parts based on the location results. Furthermore, the infer_and_annotate_unlabeled_images function infers unlabeled image samples and performs semi-automatic annotation based on the annotated image samples and the pre-trained location model.

[0405] Optionally, training the target large model based on the image set for model pre-training to obtain a pre-trained model for fault part location includes: performing unsupervised training on the target large model based on the unlabeled TFDS image samples in the image set for model pre-training to obtain an unsupervised large model; performing enhanced training on the unsupervised large model based on the digital images of railway vehicles in the image set for model pre-training to obtain an original large model; performing supervised training on the original large model based on the labeled railway vehicle image samples in the image set for model pre-training to obtain an initial large model; and performing supervised training on the initial large model based on the vehicle part image samples after semi-automatic annotation processing to obtain a pre-trained model for fault part location.

[0406] In this embodiment, through unsupervised training, the unlabeled TFDS image samples in the image set are used to perform initial unsupervised learning on the target large model, thereby obtaining an unsupervised large model with basic feature extraction capabilities. Secondly, in order to further enhance the accuracy and generalization ability of the model, the rich digital images of railway vehicles in the image set are used to perform enhanced training on the unsupervised large model, resulting in a more refined original large model. Subsequently, in order to enable the model to accurately identify and locate the faulty components of railway vehicles, the original large model is trained with supervision using the labeled railway vehicle image samples in the image set, obtaining an initial large model that can initially identify vehicle components. The vehicle component image samples after semi-automatic annotation processing are used to perform further supervised training on the initial large model. Through this step, a pre-trained model that can accurately locate the faulty components of railway vehicles is obtained, improving the training efficiency and accuracy of the model and greatly reducing the workload of manual annotation.

[0407] An exemplary code is provided in the embodiment of this application, showing that the entire training process starts from unsupervised learning, transitions to supervised learning, and finally achieves accurate positioning of faulty components.

[0408] import torch.nn as nn

[0409] import torch.optim as optim

[0410] from torch.utils.data import DataLoader, Dataset

[0411] from torchvision import transforms

[0412] # A Dataset class for loading the image set used for model pre-training

[0413] class PretrainingDataset(Dataset):

[0414] def __init__(self, image_paths, labels=None, transform=None):

[0415] # Initialize, load image paths, labels (if any), and preprocessing transforms

[0416] self.image_paths = image_paths

[0417] self.labels = labels

[0418] self.transform = transform

[0419] def __len__(self):

[0420] return len(self.image_paths)

[0421] def __getitem__(self, idx):

[0422] # Get the image and label (if available) based on the index

[0423] image_path = self.image_paths[idx]

[0424] image =...# Code to load the image

[0425] if self.labels is not None:

[0426] label = self.labels[idx]

[0427] return image, label

[0428] else:

[0429] return image

[0430] # Load the large model class

[0431] class LargeModel(nn.Module):

[0432] def __init__(self):

[0433] super(LargeModel, self).__init__()

[0434] # Define the model architecture

[0435] def forward(self, x):

[0436] # Define the forward pass

[0437] return x

[0438] # Initialize the large model, optimizer, and loss function

[0439] model = LargeModel()

[0440] criterion = nn.CrossEntropyLoss() # Or other loss functions suitable for the problem

[0441] optimizer = optim.Adam(model.parameters(), lr = 0.001) # Use the Adam optimizer

[0442] # Perform unsupervised training

[0443] def unsupervised_training(model, unlabeled_dataloader):

[0444] # Training loop

[0445] model.train()

[0446] for epoch in range(num_epochs):

[0447] for images in unlabeled_dataloader:

[0448] # Assume that unsupervised training does not require labels

[0449] outputs = model(images)

[0450] # Calculate the loss

[0451] loss =...

[0452] # Backpropagation and optimization

[0453] optimizer.zero_grad()

[0454] loss.backward()

[0455] optimizer.step()

[0456] # Perform supervised training

[0457] def supervised_training(model, labeled_dataloader, criterion, optimizer):

[0458] model.train()

[0459] for epoch in range(num_epochs):

[0460] for images, labels in labeled_dataloader:

[0461] outputs = model(images)

[0462] loss = criterion(outputs, labels)

[0463] optimizer.zero_grad()

[0464] loss.backward()

[0465] optimizer.step()

[0466] # Dataset

[0467] unlabeled_tfds_paths = [...] # List of paths to unlabeled TFDS image samples

[0468] labeled_railway_paths = [...] # List of paths to labeled railway vehicle image samples, including labels

[0469] semi_annotated_paths = [...] # List of paths to vehicle part image samples after semi-automatic annotation processing, including labels

[0470] # Preprocessing transforms

[0471] transform = transforms.Compose([...])

[0472] # Create data loaders

[0473] unlabeled_tfds_dataset = PretrainingDataset(unlabeled_tfds_paths,transform=transform)

[0474] unlabeled_tfds_dataloader = DataLoader(unlabeled_tfds_dataset, batch_size=32, shuffle=True)

[0475] labeled_railway_dataset = PretrainingDataset(labeled_railway_paths, labels=..., transform=transform)

[0476] labeled_railway_dataloader = DataLoader(labeled_railway_dataset, batch_size=32, shuffle=True)

[0477] semi_annotated_dataset = PretrainingDataset(semi_annotated_paths, labels=..., transform=transform)

[0478] semi_annotated_dataloader = DataLoader(semi_annotated_dataset, batch_size=32, shuffle=True)

[0479] # Conduct unsupervised training

[0480] print("Conducting unsupervised training...")

[0481] unsupervised_training(model, unlabeled_tfds_dataloader)

[0482] print("Unsupervised large model training completed.")

[0483] # Load the model weights after unsupervised training (or directly use the above model)

[0484] # model.load_state_dict(...)

[0485] # Conduct supervised enhanced training based on digital images of railway vehicles

[0486] print("Conducting enhanced training based on railway vehicles...")

[0487] supervised_training(model, labeled_railway_dataloader, criterion, optimizer)

[0488] print("The original large model training is completed.")

[0489] Load the weights of the original large model (or directly use the model above)

[0490] model.load_state_dict(...)

[0491] Perform initial supervised training based on the labeled railway vehicle image samples

[0492] print("Performing initial supervised training...")

[0493] supervised_training(model, labeled_railway_dataloader, criterion, optimizer)

[0494] print("The original large model training is completed.")

[0495] Load the weights of the initial large model (or directly use the model above)

[0496] model.load_state_dict(...)

[0497] Perform final supervised training based on the vehicle component image samples after semi-automatic annotation processing

[0498] print("Performing pre-training for fault component localization...")

[0499] supervised_training(model, semi_annotated_dataloader, criterion, optimizer)

[0500] print("The pre-training model for fault component localization is completed.")

[0501] In this exemplary code, the PretrainingDataset class and the LargeModel class for loading the image dataset are defined as the target large model. Then, the model, optimizer, and loss function are initialized. Next, the model is trained unsupervised by calling the unsupervised_training function, using the unlabeled TFDS image sample dataset to obtain an unsupervised large model. After that, the unsupervised large model is enhanced-trained using the labeled railway vehicle image sample dataset, and the original large model is obtained by calling the supervised_training function. Then, the original large model is continuously subjected to initial supervised training using the labeled railway vehicle images to obtain an initial large model. The initial large model is finally supervised-trained using the vehicle component image samples processed by semi-automatic annotation to obtain a pre-trained model for fault component localization.

[0502] Optionally, the image feature fusion of the TFDS image samples to obtain TFDS fusion image samples includes: representing each TFDS image sample as a pixel value matrix; based on a set feature extraction function, extracting feature points corresponding to the TFDS image sample based on the pixel value matrix; determining the corresponding feature points on different TFDS image samples; determining the geometric transformation function of different TFDS image samples, and aligning the corresponding feature points on different TFDS image samples based on the geometric transformation function.

[0503] Fusing the TFDS image samples with aligned feature points to obtain TFDS fusion image samples;

[0504] Based on the following formula, the image quality of the TFDS fusion image samples is evaluated to obtain an image quality evaluation value, so as to determine whether the TFDS fusion image samples meet the set quality standard value, and when they do not meet, the above fusion processing steps are re-executed:

[0505] , where E is the evaluation function, represents the image quality evaluation value, represents the TFDS fusion image sample, represents the reference image, SSIM represents the structural similarity calculation value, and PSNR represents the peak signal-to-noise ratio calculation value.

[0506] In this embodiment, each image sample is represented as a pixel value matrix, and the image data is converted into a numerical form that can be processed by a computer. Then, based on a set feature extraction function, feature points corresponding to the image sample are extracted from the pixel value matrix, and these feature points represent the key information in the image. Next, the corresponding feature points on different image samples are determined to ensure that this key information can be accurately corresponding during the fusion process. After determining the corresponding feature points, the geometric transformation function between different image samples is further determined, and based on this function, the feature points on different images are aligned to ensure that they can match together during fusion. After that, the aligned image samples are further fused to generate TFDS fused image samples. This process integrates the information of multiple images into one image, improving the richness and accuracy of the information. To ensure the quality of the fused image, an image quality evaluation formula is adopted. This formula combines two metrics, structural similarity (SSIM) and peak signal-to-noise ratio (PSNR), and calculates an image quality evaluation value through calculation, which can intuitively reflect the similarity and clarity between the fused image and the reference image. Among them, when the evaluation value does not meet the set quality standard value, the above fusion processing steps are re-executed until a TFDS fused image sample that meets the requirements is obtained. This process not only ensures the accuracy of image fusion but also guarantees the high quality of the final fused image through cyclic optimization, thus providing users with more accurate and reliable image information.

[0507] An exemplary code is provided in an embodiment of this application.

[0508] import numpy as np

[0509] from skimage.measure import compare_ssim

[0510] from skimage.metrics import peak_signal_noise_ratio

[0511] # images is a list containing TFDS image samples, and each image is a numpy array

[0512] images = [...] # Load images

[0513] def extract_keypoints(image):

[0514] # Use SIFT as the feature extraction function

[0515] sift = cv2.SIFT_create()

[0516] kp, des = sift.detectAndCompute(image, None)

[0517] return kp, des

[0518] def align_images(images, keypoints, descriptors):

[0519] # Usually use RANSAC, FLANN or other methods for image registration and alignment

[0520] # Transformation matrix transforms

[0521] aligned_images = []

[0522] for img, kp, des, transform in zip(images, keypoints, descriptors, transforms):

[0523] # Apply the transformation matrix

[0524] h, w = img.shape[:2]

[0525] aligned_img = cv2.warpAffine(img, transform, (w, h))

[0526] aligned_images.append(aligned_img)

[0527] return aligned_images

[0528] def blend_images(aligned_images):

[0529] # Average blending, more complex blending methods may be required in practice

[0530] fusion_img = np.mean(aligned_images, axis = 0).astype(np.uint8)

[0531] return fusion_img

[0532] def evaluate_image_quality(fusion_img, reference_img):

[0533] # Calculate SSIM and PSNR

[0534] ssim_value = compare_ssim(fusion_img, reference_img, multichannel=True)

[0535] psnr_value = peak_signal_noise_ratio(fusion_img, reference_img)

[0536] # a is the weight value, set to 0.5 here

[0537] a = 0.5

[0538] E = a * ssim_value + (1 - a) * psnr_value

[0539] return E

[0540] # Initialize variables

[0541] keypoints = []

[0542] descriptors = []

[0543] transforms = [] # This should contain the calculated geometric transformation matrices

[0544] # Extract keypoints

[0545] for img in images:

[0546] kp, des = extract_keypoints(img)

[0547] keypoints.append(kp)

[0548] descriptors.append(des)

[0549] # Align images

[0550] aligned_images = align_images(images, keypoints, descriptors,transforms)

[0551] # Fuse images

[0552] fusion_img = blend_images(aligned_images)

[0553] # Code to load the reference image

[0554] reference_img =...

[0555] # Evaluate image quality

[0556] quality_score = evaluate_image_quality(fusion_img, reference_img)

[0557] # If the quality standard is not met, re - execute the fusion processing steps

[0558] if quality_score < SOME_THRESHOLD:

[0559] print("The image quality does not meet the requirements, re - executing the fusion process...")

[0560] # Re - call the align_images and blend_images functions, recalculate the transforms

[0561] else:

[0562] print("The TFDS - fused image sample meets the quality standard, and the image quality evaluation value is:", quality_score)

[0563] In this exemplary code, feature points are extracted for each image sample, and these points are used to align different images. Then, the aligned images are fused to generate a TFDS - fused image sample. An evaluation function that combines structural similarity (SSIM) and peak signal - to - noise ratio (PSNR) is used to evaluate the quality of the fused image. If the image quality does not meet the set quality standard, the process of image alignment and fusion will be re - executed until a fused image that meets the requirements is obtained.

[0564] Optionally, training the pre - trained model for fault component localization based on the model training image set to obtain a fault component localization model includes: determining candidate regions of vehicle components on the model training images in the model training image set based on the pre - trained model for fault component localization;

[0565] Based on the attention mechanism module, identify the area of the predetermined vehicle component in the candidate areas of the vehicle components; based on the logical processing module, infer the areas of other vehicle components according to the area of the predetermined component; based on the region of interest extraction module, extract the area of the predetermined component and the areas of other vehicle components from the images for model training; adjust the model parameters of the pre-trained model for fault component localization according to the area of the predetermined component, the areas of other vehicle components, and the markings of the vehicle components, so as to train the pre-trained model for fault component localization to obtain a fault component localization model.

[0566] In this embodiment, based on the pre-trained model for fault component localization, the candidate areas of vehicle components in the image set for model training are initially determined, which provides a basis for subsequent fine-grained recognition. Subsequently, by introducing the attention mechanism module, the recognition focuses on the area of the predetermined vehicle component, reducing the computational amount and improving the recognition accuracy. Immediately afterwards, the logical processing module infers the positions of other vehicle components through logical reasoning based on the identified area of the predetermined component, which further expands the detection range and enhances the generalization ability of the system. Then, using the region of interest extraction module, the areas of the predetermined component and other key components are extracted from the complex image background, making subsequent processing more efficient. According to these accurately extracted areas of vehicle components and their corresponding markings, the parameters of the pre-trained model for fault component localization are adjusted to continuously optimize the model performance until a precise and reliable fault component localization model is obtained. The entire process combines attention mechanism, logical reasoning, and region extraction technologies to achieve precise localization of fault components, providing strong support for subsequent fault analysis and maintenance.

[0567] This application embodiment provides an exemplary code.

[0568] import torchvision.transforms as transforms

[0569] from torchvision.models.detection import FasterRCNN # Use Faster R-CNN as the base model

[0570] from torch.utils.data import DataLoader, Dataset

[0571] # Custom Dataset class for loading and preprocessing image data

[0572] class VehicleDataset(Dataset):

[0573] def __init__(self, image_paths, annotations):

[0574] self.image_paths = image_paths

[0575] self.annotations = annotations

[0576] self.transform = transforms.Compose([transforms.ToTensor()])# Preprocessing

[0577] def __len__(self):

[0578] return len(self.image_paths)

[0579] def __getitem__(self, idx):

[0580] image_path = self.image_paths[idx]

[0581] image =...# Logic to load the image

[0582] annotation = self.annotations[idx]

[0583] image = self.transform(image)

[0584] return image, annotation

[0585] # Pre-trained Faster R-CNN model

[0586] pretrained_model = FasterRCNN(...)

[0587] # Attention mechanism module def attention_mechanism(features):

[0588] # features is the feature map extracted from a certain layer of the model

[0589] attention_weights =...# Logic to calculate attention weights

[0590] attended_features = features * attention_weights # Apply attention weights

[0591] return attended_features

[0592] # Logical processing module, infer other parts based on predefined parts

[0593] def logical_processing(predefined_parts, detected_boxes):

[0594] # predefined_parts is the location of predefined parts, detected_boxes is the location of parts detected by the model

[0595] other_parts_locations =... # Logic to infer the locations of other parts

[0596] return other_parts_locations

[0597] # Region of interest extraction module (extract from the image based on region locations)

[0598] def extract_rois(image, rois):

[0599] # rois are the coordinates of the regions of interest

[0600] extracted_rois =... # Logic to extract regions from the image

[0601] return extracted_rois

[0602] # Training loop

[0603] def train_model(model, dataset, num_epochs, optimizer, criterion):

[0604] dataloader = DataLoader(dataset, batch_size=..., shuffle=True)

[0605] for epoch in range(num_epochs):

[0606] for images, annotations in dataloader:

[0607] # Forward propagation: Determine candidate regions

[0608] with torch.no_grad():

[0609] predictions = model(images)

[0610] # Apply attention mechanism

[0611] attended_features = attention_mechanism(model.features)# Obtain feature maps from model.features

[0612] # Logical processing: Infer other components

[0613] other_parts = logical_processing(predictions['boxes'][0], predictions['boxes'])# Only process the first batch of data as an example

[0614] # Extract regions of interest

[0615] rois = predictions['boxes'].tolist() + other_parts# Add the inferred components to rois

[0616] extracted_rois = extract_rois(images, rois)

[0617] # Update model parameters according to the loss

[0618] optimizer.zero_grad()

[0619] loss.backward()

[0620] optimizer.step()

[0621] return model

[0622] In this exemplary code, a dataset containing vehicle images and corresponding annotations is loaded. Then, using a pre-trained Faster R-CNN model as a starting point, the model is fine-tuned through a training loop. In each iteration, candidate regions of vehicle parts on the image are determined through forward propagation. Then, a predefined vehicle part in the candidate region is identified using an attention mechanism module, and the recognition results are filtered and sorted by a logic processing module according to predefined thresholds and confidence levels to finally determine the exact positions of the vehicle parts.

[0623] Optionally, before performing image feature fusion on the TFDS image samples to obtain TFDS fused image samples, it includes: performing spatial transformation processing on the TFDS image samples according to the acquisition system error of the TFDS image samples to obtain spatially transformed image samples, where the acquisition system error at least includes: imaging angle, perspective relationship, and light intensity; enhancing the spatially transformed image samples to generate enhanced image samples with differences in color, brightness, and contrast; performing affine graphic transformation at the pixel level on the enhanced image samples to normalize the enhanced image samples to obtain normalized image samples; performing multi-scale rotation, scaling, and stretching operations on the normalized image samples to obtain unified scale image samples with different resolutions; performing image data domain enhancement and amplification processing on the unified scale image samples to obtain amplified image samples; performing grayscale processing on the amplified image samples to obtain grayscale image samples for use when performing image feature fusion on the TFDS image samples.

[0624] In this embodiment, according to the possible systematic errors in the image acquisition process, such as imaging angle, perspective relationship, and light intensity, spatial transformation processing is performed on the original image samples to correct these errors and obtain more accurate spatially transformed image samples. Next, in order to enhance the diversity of the image samples and the robustness of model training, the spatially transformed image samples are enhanced to generate enhanced image samples with differences in color, brightness, and contrast. Subsequently, through pixel-level affine graphic transformation, the enhanced image samples are normalized to ensure the consistency of the image samples in the color space, thereby obtaining normalized image samples. Secondly, in order to capture the features of the image at different scales, multi-scale rotation, scaling, and stretching operations are performed on the normalized image samples to obtain unified scale image samples with different resolutions. Again, in order to further increase the diversity and richness of the image samples, enhanced amplification processing in the image data domain is performed on the unified scale image samples to generate more amplified image samples. In order to reduce the complexity of image processing and highlight the key information in the image, the amplified image samples are grayscaled to obtain grayscale image samples, which will be used in the subsequent image feature fusion process to improve the effect of feature fusion and the performance of the model.

[0625] An exemplary code is provided in an embodiment of this application to show how to perform a series of preprocessing steps on TFDS image samples and finally obtain grayscale image samples for image feature fusion.

[0626] import tensorflow as tf

[0627] import tensorflow_datasets as tfds

[0628] import cv2

[0629] import numpy as np

[0630] # Loaded TFDS dataset

[0631] # dataset, info = tfds.load('your_dataset_name', with_info=True,...)

[0632] # Preprocessing function

[0633] def spatial_transform(image, transformation_params):

[0634] rows, cols, _ = image.shape

[0635] M = cv2.getRotationMatrix2D((cols / / 2, rows / / 2), transformation_params['angle'], 1)

[0636] transformed_image = cv2.warpAffine(image, M, (cols, rows))

[0637] return transformed_image

[0638] def enhance_image(image):

[0639] # Color, brightness, and contrast enhancement

[0640] hsv = cv2.cvtColor(image, cv2.COLOR_BGR2HSV)

[0641] hsv[:,:,2] = cv2.equalizeHist(hsv[:,:,2])

[0642] enhanced_image = cv2.cvtColor(hsv, cv2.COLOR_HSV2BGR)

[0643] return enhanced_image

[0644] def normalize_image(image):

[0645] # Normalize to a unified color space

[0646] # Example: Normalize the image to the range [0, 1]

[0647] normalized_image = image / 255.0

[0648] return normalized_image

[0649] def multi_scale_augmentation(image, scales):

[0650] # Multi-scale rotation, scaling, stretching operations

[0651] multi_scale_images = []

[0652] for scale in scales:

[0653] # Resize the image

[0654] scaled_image = cv2.resize(image, None, fx=scale, fy=scale, interpolation=cv2.INTER_LINEAR)

[0655] # Add other operations such as rotation

[0656] multi_scale_images.append(scaled_image)

[0657] return multi_scale_images

[0658] def domain_augmentation(images):

[0659] # Image data domain augmentation processing

[0660] # Add Gaussian noise

[0661] augmented_images = []

[0662] for image in images:

[0663] mean = 0.0

[0664] var = 0.01

[0665] sigma = var**0.5

[0666] gauss = np.random.normal(mean, sigma, image.shape)

[0667] gauss = gauss.astype('uint8')

[0668] augmented_image = cv2.add(image, gauss)

[0669] augmented_images.append(augmented_image)

[0670] return augmented_images

[0671] def grayscale_conversion(images):

[0672] # Grayscale processing

[0673] gray_images = [cv2.cvtColor(image, cv2.COLOR_BGR2GRAY) for image in images]

[0674] return gray_images

[0675] # Function to load and preprocess TFDS image samples

[0676] def preprocess_tfds_image(dataset, transformation_params, scales):

[0677] def _preprocess(example):

[0678] # example['image'] is the image data

[0679] image = example['image']

[0680] # Spatial transformation processing

[0681] spatial_transformed_image = spatial_transform(image.numpy(), transformation_params)

[0682] # Color, brightness, and contrast enhancement

[0683] enhanced_image = enhance_image(spatial_transformed_image)

[0684] # Normalization processing

[0685] normalized_image = normalize_image(enhanced_image)

[0686] # Multi-scale rotation, scaling, and stretching operations

[0687] multi_scale_images = multi_scale_augmentation(normalized_image,scales)

[0688] # Image data domain augmentation processing

[0689] augmented_images = domain_augmentation(multi_scale_images)

[0690] # Grayscale processing

[0691] gray_images = grayscale_conversion(augmented_images)

[0692] # Keep a grayscale image sample for feature fusion

[0693] gray_image_sample = gray_images[0]

[0694] # Add the grayscale image sample as a new feature to the example

[0695] example['gray_image'] = tf.convert_to_tensor(gray_image_sample, dtype=tf.uint8)

[0696] return example

[0697] # Use the map function of tf.data to preprocess the dataset

[0698] return dataset.map(_preprocess, num_parallelcalls=tf.data.AUTOTUNE)

[0699] In this exemplary code, the spatial_transform function performs spatial transformation processing on the image according to the given transformation parameters to correct the distortion introduced by the acquisition system error, and obtains spatial transformation image samples; the enhance_image function enhances the color, brightness, and contrast of the spatially transformed image samples to generate differentially enhanced image samples, making the key information in the image more prominent; the normalize_image function normalizes the enhanced image samples, unifying their color space to a specified range so that the model can adopt a unified scale when processing different images; the multi_scale_augmentation function performs multi-scale rotation, scaling, and stretching operations on the normalized image samples to generate unified scale image samples with different resolutions to improve the model's recognition ability for targets of different sizes and angles; the domain_augmentation function performs image data domain enhancement and augmentation processing on the multi-scale image samples, such as adding noise and random cropping, to further increase the robustness of the model; the grayscale_conversion function converts the enhanced and augmented image samples into grayscale images because grayscale images can extract image features more effectively in some cases; the _preprocess function is a function that integrates all the above steps. It takes a sample from a TFDS dataset as input, and sequentially performs spatial transformation, color enhancement, normalization, multi-scale transformation, data domain enhancement, and grayscaling, etc. Then, the processed grayscale image is added as a new feature to the sample and the updated sample is returned; the preprocess_tfds_image function uses the tf.data.Dataset.map method to apply the _preprocess function to the entire dataset to batch process all image samples. Here, num_parallel_calls = tf.data.AUTOTUNE means using automatic parallelization to accelerate the data preprocessing process.

[0700] Optionally, the normalizing the enhanced image samples by performing pixel-level linear scaling on the enhanced image samples to obtain normalized image samples includes: performing affine image transformation on the enhanced image samples of different sizes to obtain enhanced image samples of a fixed size; separating the color channels of all pixel points on the enhanced image samples of a fixed size, and calculating the maximum pixel value and the minimum pixel value for each separated color channel; for each color channel of each pixel point, calculating the difference between its corresponding pixel value and the minimum pixel value, and dividing it by the difference between the maximum pixel value and the minimum pixel value to obtain the normalized channel pixel value; for all pixel points, combining the normalized channel pixel values corresponding to different color channels to obtain the normalized image samples.

[0701] In this embodiment, the advantage of this technology is that it adopts an efficient and accurate normalization method, which can significantly improve the accuracy and stability of image analysis. By performing affine image transformation on enhanced image samples of different sizes, it can ensure that all image samples have a unified size, which provides a standardized basis for subsequent processing. Subsequently, color channel separation is performed on the image samples of fixed size, enabling each color channel to be processed independently, which helps to more finely adjust the color information of the image. Then, for each color channel, the maximum pixel value and the minimum pixel value are calculated, and based on this, the normalized channel pixel value of each pixel point is calculated. This step can eliminate the scale difference of pixel values in the image, making the color distributions of different image samples more consistent. Merging and processing the normalized channel pixel values of different color channels to obtain normalized image samples. This process not only retains the original information of the image but also enables the image samples to have a unified color space globally, providing strong support for subsequent image analysis. The entire processing flow is logically rigorous and easy to operate, greatly improving the efficiency and accuracy of image processing.

[0702] The embodiment of this application provides an exemplary code, which shows how to perform normalization processing on enhanced image samples.

[0703] import cv2

[0704] import numpy as np

[0705] def normalize_image(image, target_size=(256, 256)):

[0706] # Perform affine image transformation on enhanced image samples of different sizes to obtain enhanced image samples of fixed size

[0707] resized_image = cv2.resize(image, target_size, interpolation=cv2.INTER_LINEAR)

[0708] # Perform color channel separation on all pixel points on the enhanced image samples of fixed size

[0709] # Process images in BGR format

[0710] channels = cv2.split(resized_image)

[0711] # For each separated color channel, calculate the maximum pixel value and the minimum pixel value

[0712] min_vals = [np.min(chan) for chan in channels]

[0713] max_vals = [np.max(chan) for chan in channels]

[0714] # Initialize the array for normalized channel pixel values

[0715] normalized_channels = []

[0716] # For each color channel at each pixel, calculate the difference between its corresponding pixel value and the minimum pixel value,

[0717] # and divide by the difference between the maximum pixel value and the minimum pixel value to obtain the normalized channel pixel value

[0718] for chan, min_val, max_val in zip(channels, min_vals, max_vals):

[0719] # Avoid division by zero

[0720] eps = np.finfo(chan.dtype).eps

[0721] normalized_chan = (chan - min_val) / (max_val - min_val + eps)

[0722] normalized_channels.append(normalized_chan)

[0723] # For all pixels, merge the normalized channel pixel values corresponding to different color channels,

[0724] # to obtain the normalized image sample

[0725] normalized_image = cv2.merge(normalized_channels)

[0726] return normalized_image

[0727] # Example: Load an image and normalize it

[0728] image_path = 'path_to_your_image.jpg' # Replace with the path to your image file

[0729] image = cv2.imread(image_path)

[0730] if image is not None:

[0731] normalized_img = normalize_image(image)

[0732] # Display the normalized image

[0733] cv2.imshow('Normalized Image', normalized_img)

[0734] cv2.waitKey(0)

[0735] cv2.destroyAllWindows()

[0736] else:

[0737] print("Error: Unable to load image.")

[0738] In this exemplary code, the normalize_image function resizes the input image to the specified target size using the cv2.resize function. Then, the function separates the image into its color channels (BGR) and calculates the maximum and minimum pixel values for each channel. Next, for each channel, the function calculates the difference between each pixel value and the minimum value of the channel, and divides this difference by the difference between the maximum and minimum values of the channel, resulting in normalized pixel values. The function combines the normalized channels into a new image and returns this image. This normalization process ensures the scale consistency of the image data, which is helpful for subsequent image analysis and processing tasks.

[0739] Optionally, performing image data domain enhancement and amplification processing on the unified scale image samples to obtain amplified image samples includes: modifying the distribution of the unified scale image samples to generate variant unified scale image samples of multiple weather variants such as rain, snow, fog, frost, overcast, and hail; according to different driving scenarios of railway vehicles, generating simulated unified scale image samples matching the different driving scenarios based on the variant unified scale image samples to simulate scenarios such as passing through tunnels, icing of railway vehicle components, and snow falling on railway vehicle components; according to the structural model of railway vehicles, generating positive model image samples similar to the characteristics of the different driving scenarios and negative model sample images that should not be recognized as faults based on the variant unified scale image samples; using a set random convolution kernel to perform convolution operations on the positive model image samples and negative model sample images to obtain amplified image samples.

[0740] In this embodiment, by modifying the distribution of the image samples, variant unified scale image samples of multiple weather variants such as rain, snow, fog, frost, overcast, and hail are generated, which greatly increases the diversity and coverage of the image samples. Immediately afterwards, according to the actual driving scenarios of railway vehicles, based on these variant image samples, simulated unified scale image samples matching different driving scenarios are further generated, such as simulating scenarios where the vehicle passes through a tunnel, vehicle components are iced, and vehicle components are snowed, so as to simulate image data closer to the actual environment. Subsequently, combined with the structural model of railway vehicles, positive model image samples similar to the characteristics of different driving scenarios and negative model sample images that should not be recognized as faults are generated. Using a carefully set random convolution kernel to perform convolution operations on these positive and negative model image samples to obtain a large number of amplified image samples.

[0741] An exemplary code is provided in an embodiment of the present application to simulate the image data domain enhancement and amplification processing flow.

[0742] import cv2

[0743] import numpy as np

[0744] from albumentations import Compose, WeatherAugmentation,HorizontalFlip, RandomBrightnessContrast

[0745] from albumentations.pytorch import ToTensorV2

[0746] # Randomly generated images as unified scale image samples

[0747] height, width = 256, 256

[0748] uniform_scale_image = np.random.randint(0, 255, (height, width, 3), dtype=np.uint8)

[0749] # Weather variant generation

[0750] def weather_augmentation(image):

[0751] weather_transforms = Compose(

[0752] WeatherAugmentation(p=1.0, # Set the probability of applying the transformation

[0753] visualize=False,

[0754] transform_generators=

[0755] 'rain','snow', 'fog', 'cloud', 'brightness', 'contrast','sun_glare'

[0756] ,

[0757] min_visibility=0.2, # Minimum visibility

[0758] max_visibility=1.0),

[0759] ToTensorV2() # Convert the image to a PyTorch tensor if needed )

[0761] augmented_image = weather_transforms(image=image)['image']

[0762] return augmented_image

[0763] # Scene simulation

[0764] def scene_simulation(weather_augmented_image, scene_type):

[0765] # Simulate different scenarios using weather-varied images

[0766] if scene_type == 'tunnel':

[0767] # Simulate tunnel effect

[0768] tunnel_mask = np.zeros_like(weather_augmented_image, dtype=np.uint8)

[0769] tunnel_mask[:, :width / / 2] = weather_augmented_image[:, :width / / 2]# Mask the right half of the image

[0770] simulated_image = tunnel_mask

[0771] elif scene_type == 'icing':

[0772] # Simulate icing on vehicle parts, add white areas or blur effects

[0773] icing_mask = np.zeros_like(weather_augmented_image, dtype=np.uint8)

[0774] icing_mask[:, height / / 4:height / / 2] = [255, 255, 255]# Icing in the upper-middle position of the image

[0775] simulated_image = cv2.addWeighted(weather_augmented_image, 0.5,icing_mask, 0.5, 0)

[0776] elif scene_type =='snowing':

[0777] # Simulate snowfall, add white dots or snowflake effects

[0778] snow_points = np.random.randint(0, height, (100, 2))

[0779] snow_points[:, 1] = np.random.randint(0, width, 100)

[0780] snow_image = np.zeros_like(weather_augmented_image, dtype=np.uint8)

[0781] for point in snow_points:

[0782] snow_image[point[0], point[1]]= [255, 255, 255]

[0783] simulated_image = cv2.addWeighted(weather_augmented_image, 0.8, snow_image, 0.2, 0)

[0784] else:

[0785] simulated_image = weather_augmented_image

[0786] return simulated_image

[0787] # Positive and negative model sample generation

[0788] def generate_model_samples(weather_augmented_image):

[0789] # Generate positive and negative samples directly from the weather-varied image

[0790] positive_sample = weather_augmented_image# Image with assumed fault

[0791] negative_sample = cv2.GaussianBlur(weather_augmented_image, (15, 15),0)# Blurring as the image without fault

[0792] return positive_sample, negative_sample

[0793] # Perform convolution operation using a random convolution kernel to augment samples

[0794] def convolution_augmentation(image, kernel_size=3):

[0795] # Create a random convolution kernel

[0796] kernel = np.random.randn(kernel_size, kernel_size)

[0797] kernel / = np.sum(np.abs(kernel)) # Normalize the convolution kernel

[0798] # Apply the convolution kernel to the image

[0799] augmented_image = cv2.filter2D(image, -1, kernel)

[0800] return augmented_image

[0801] # Process concatenation

[0802] weather_variants = weather_augmentation(uniform_scale_image)

[0803] tunnel_simulation = scene_simulation(weather_variants, 'tunnel')

[0804] icing_simulation = scene_simulation(weather_variants, 'icing')

[0805] snowing_simulation = scene_simulation(weather_variants,'snowing')

[0806] # Generate positive and negative samples for the tunnel scene

[0807] positive_tunnel, negative_tunnel = generate_model_samples(tunnel_simulation)

[0808] # Convolutionally augment the positive and negative samples

[0809] augmented_positive_tunnel = convolution_augmentation(positive_tunnel)

[0810] augmented_negative_tunnel = convolution_augmentation(negative_tunnel)

[0811] In this exemplary code, the weather effect of the image samples of the unified scale is enhanced through the weather_augmentation function for generating weather variations, and images of various weather variations are obtained. Then, using the scene_simulation function, based on these images of weather variations, different driving scenarios such as a railway vehicle passing through a tunnel, parts icing, and snowfall are simulated, and simulated images similar to the actual environment are obtained. Then, through the generate_model_samples function, based on the simulated images and the structural model of the railway vehicle, positive and negative model sample images are generated, representing the situations with and without faults respectively. Through the convolution_augmentation function, the positive and negative model sample images are convolved using random convolution kernels to obtain augmented image samples, which increase the diversity of the training data of the model and help improve the generalization ability of the model.

[0812] Optionally, the obtaining of the TFDS image to be processed and inputting it into the fault component location model to locate the position of the vehicle fault component on the TFDS image to be processed includes: performing grayscale processing on the obtained TFDS image to be processed to obtain a grayscale image to be processed; inputting the grayscale image to be processed into the fault component location model, selecting a number of image candidate regions for different parts, performing affine graphic transformation on the candidate regions, and performing convolution, pooling, and normalization processing to obtain a model input image of a fixed size; the fault component location model extracts features from the model input image to obtain low-dimensional features and high-dimensional features. The low-dimensional features at least include: edge features and contour features, and the high-dimensional features include: semantic features of the fault component; performing fusion processing on the low-dimensional features and the high-dimensional features to obtain fusion features; and locating the position of the vehicle component on the TFDS image to be processed according to the fusion features.

[0813] In this embodiment, the acquired TFDS image to be processed is grayscaled and converted into a grayscale image to be processed. This step helps reduce the computational complexity of image processing and highlight key information. Subsequently, the processed grayscale image is input into the fault component localization model. The model selects candidate regions for different parts of the image and, through a series of operations such as affine graphic transformation, convolution, pooling, and normalization, converts the candidate regions into model input images of a fixed size to meet the processing requirements of the model. The fault component localization model extracts features from these model input images, obtaining both low-dimensional features and high-dimensional features simultaneously. The low-dimensional features cover intuitive information such as edge features and contour features, while the high-dimensional features contain semantic features of the fault components. These features together constitute a complete description of the vehicle's fault components. By fusing the low-dimensional features and high-dimensional features, a comprehensive fused feature is obtained, which contains both local detail information and overall semantic information of the image. Based on this fused feature, the position of the vehicle's fault components on the TFDS image to be processed is located. The entire process is not only efficient and accurate but also capable of adapting to TFDS image processing requirements under different scenarios and conditions, providing strong technical support for vehicle fault detection.

[0814] An exemplary code is provided in an embodiment of this application.

[0815] import cv2

[0816] import numpy as np

[0817] from sklearn.preprocessing import normalize # Use sklearn for normalization

[0818] from some_deep_learning_library import FaultLocalizationModel # Deep learning model library

[0819] # Grayscaling and preprocessing function

[0820] def preprocess_image(image_path):

[0821] image = cv2.imread(image_path, cv2.IMREAD_GRAYSCALE) # Grayscaling processing

[0822] # Add more preprocessing steps, noise reduction, contrast enhancement

[0823] return image

[0824] # Candidate Region Selection, Transformation, and Normalization Functions

[0825] def select_and_transform_regions(gray_image, model_input_size):

[0826] # Select candidate regions

[0827] # Perform affine transformation and convolution on each candidate region. regions = []

[0828] for region in candidate_regions: # candidate_regions are the selected candidate regions

[0829] # Perform affine transformation

[0830] # Crop or resize to model_input_size

[0831] resized_region = cv2.resize(region, (model_input_size, model_input_size))

[0832] regions.append(resized_region)

[0833] # Normalize all regions

[0834] normalized_regions = [normalize(region.reshape(-1, 1), axis=0, norm='l2').ravel() for region in regions]

[0835] return np.array(normalized_regions)

[0836] # Faulty Component Localization Model Class

[0837] class FaultLocalizationModel:

[0838] def __init__(self):

[0839] # Load model weights, etc.

[0840] def extract_features(self, input_images):

[0841] # Model feature extraction function

[0842] # Include operations such as convolution and pooling

[0843] low_dim_features, high_dim_features = self.model.extract_features(input_images)

[0844] return low_dim_features, high_dim_features

[0845] def fuse_features(self, low_dim_features, high_dim_features):

[0846] # Feature fusion function

[0847] # Use concatenation and weighting methods to fuse features

[0848] fused_features = self.model.fuse_features(low_dim_features, high_dim_features)

[0849] return fused_features

[0850] def locate_fault_parts(self, fused_features):

[0851] # Fault part location function

[0852] # Locate fault parts based on fused features

[0853] fault_parts_locations = self.model.predict(fused_features)

[0854] return fault_parts_locations

[0855] # Main process

[0856] def locate_fault_parts_in_image(image_path, model_input_size, fault_localization_model):

[0857] gray_image = preprocess_image(image_path)

[0858] input_images = select_and_transform_regions(gray_image, model_input_size)

[0859] low_dim_features, high_dim_features = fault_localization_model.extract_features(input_images)

[0860] fused_features = fault_localization_model.fuse_features(low_dim_features, high_dim_features)

[0861] fault_parts_locations = fault_localization_model.locate_fault_parts(fused_features)

[0862] return fault_parts_locations

[0863] In this exemplary code, the TFDS image to be processed is converted into a grayscale image through the preprocess_image function. Then, the select_and_transform_regions function is responsible for selecting candidate regions, performing affine transformation on them, adjusting them to the size required for model input, and then normalizing them. The instance fault_localization_model of the FaultLocalizationModel class is used to extract features from the normalized image, including low-dimensional features and high-dimensional features. The extracted features are then fused through the fuse_features method to obtain fused features, and finally the faulty parts are located.

[0864] Optionally, the fault component localization model extracts features from the model input image to obtain low-dimensional features, including: the fault component localization model extracts original low-dimensional features from the model input image through a shallow convolutional network layer; performs multi-scale non-linear transformation on the model input image through an augmentation layer parallelly stacked with the shallow convolutional network layer to obtain multi-scale non-linear features; fuses the original low-dimensional features and the multi-scale non-linear features to obtain the low-dimensional features; the fault component localization model extracts high-dimensional features from the model input image through a deep convolutional network layer parallelly stacked with the shallow convolutional network layer.

[0865] In this embodiment, the shallow convolutional network layer is used to perform preliminary feature extraction on the input model image, capturing the original low-dimensional features of the image, which usually include detailed information such as edges and textures. At the same time, in order to more comprehensively capture the multi-scale information in the image, the model also introduces an augmentation layer parallelly stacked with the shallow convolutional network layer to perform multi-scale non-linear transformation on the input image, thereby generating multi-scale non-linear features. These features can capture the changes and details of the image at different scales and are crucial for the localization of fault components. The model fuses the original low-dimensional features and the multi-scale non-linear features. Through this step, the model can combine the local detailed information and the global multi-scale information of the image to form a more comprehensive and rich low-dimensional feature representation. Such a feature representation not only contains the basic information of the image but also incorporates the change information of the image at different scales, providing strong support for the subsequent localization of fault components. The model extracts deeper features from the input image through a deep convolutional network layer parallelly stacked with the shallow convolutional network layer to obtain high-dimensional features. These high-dimensional features usually contain high-level semantic information of the image, such as the category and shape of the fault component. By combining the low-dimensional features and the high-dimensional features, the model can more accurately identify and locate the fault components in the image, thereby achieving the purpose of fault detection.

[0866] An exemplary code is provided in an embodiment of this application.

[0867] import torch

[0868] import torch.nn as nn

[0869] import torch.nn.functional as F

[0870] # Model class

[0871] class FaultLocalizationModel(nn.Module):

[0872] def __init__(self):

[0873] super(FaultLocalizationModel, self).__init__()

[0874] # Shallow convolutional network layer

[0875] self.shallow_conv_layers = nn.Sequential(

[0876] nn.Conv2d(1, 32, kernel_size=3, stride=1, padding=1),

[0877] nn.ReLU(inplace=True),

[0878] nn.MaxPool2d(kernel_size=2, stride=2) )

[0880] # Augmentation layer multi-scale non-linear transformation

[0881] # Here, convolutional layers of different scales are used to simulate multi-scale feature extraction

[0882] self.augmentation_layers = nn.ModuleList(

[0883] nn.Conv2d(1, 16, kernel_size=5, stride=1, padding=2),

[0884] nn.Conv2d(1, 8, kernel_size=7, stride=1, padding=3),

[0885] # Add more convolutional layers of different scales )

[0887] # Feature fusion layer

[0888] self.feature_fusion = nn.Conv2d(32 + 16 + 8, 64, kernel_size=1) # The dimension of the fused features is 64

[0889] # Deep convolutional network layer

[0890] self.deep_conv_layers = nn.Sequential(

[0891] nn.Conv2d(64, 128, kernel_size=3, stride=1, padding=1),

[0892] nn.ReLU(inplace=True),

[0893] nn.MaxPool2d(kernel_size=2, stride=2),

[0894] # Add more layers to extract high - dimensional features )

[0896] def forward(self, x):

[0897] # Shallow convolutional network layers extract original low - dimensional features

[0898] shallow_features = self.shallow_conv_layers(x)

[0899] # Augmentation layers extract multi - scale non - linear features

[0900] multi_scale_features = [aug_layer(x) for aug_layer in self.augmentation_layers]

[0901] # Take the features of the last scale

[0902] augmented_features = multi_scale_features[-1]

[0903] # Fuse the original low - dimensional features and multi - scale non - linear features

[0904] fused_features = torch.cat([shallow_features, augmented_features], dim = 1)

[0905] fused_features = self.feature_fusion(fused_features)

[0906] # Deep convolutional network layers extract high - dimensional features

[0907] high_dim_features = self.deep_conv_layers(fused_features)

[0908] # Return low-dimensional features (fused features) and high-dimensional features

[0909] return fused_features, high_dim_features

[0910] # Assume the input image (the model input image should be a four-dimensional tensor, including batch_size, channels, height, width)

[0911] # Use a single grayscale image

[0912] input_image = torch.randn(1, 1, 224, 224) # The image size is 224x224

[0913] # Instantiate the model

[0914] model = FaultLocalizationModel()

[0915] # Forward propagation to obtain low-dimensional features and high-dimensional features

[0916] low_dim_features, high_dim_features = model(input_image)

[0917] In this exemplary code, the FaultLocalizationModel extracts features from the input model image through shallow convolutional network layers to obtain the original low-dimensional features containing details such as edges and textures. At the same time, the model performs multi-scale non-linear transformations on the input image through an augmentation layer to extract multi-scale non-linear features, which can capture the changes and details of the image at different scales. Subsequently, the model fuses the original low-dimensional features and multi-scale non-linear features, and combines them into a richer low-dimensional feature representation through a feature fusion layer. The model performs deeper feature extraction on the fused features through deep convolutional network layers to obtain high-dimensional features containing high-level semantic information of the image.

[0918] Optionally, the low-dimensional features and high-dimensional features are fused to obtain fused features, including: mapping the low-dimensional features and the high-dimensional features to the same spatial dimension to obtain aligned low-dimensional mapped features and high-dimensional mapped features; performing hierarchical sampling on the eigenvalues of the low-dimensional mapped features, and performing per-channel eigenvalue fusion on the hierarchically sampled eigenvalues and the high-dimensional mapped features to obtain the fused features.

[0919] In this embodiment, the low-dimensional features and high-dimensional features are mapped to the same spatial dimension to ensure their alignment in dimension, so as to obtain aligned low-dimensional mapped features and high-dimensional mapped features, ensuring the comparability between features of different dimensions and providing a basis for subsequent fusion operations. By performing hierarchical sampling on the eigenvalues of the low-dimensional mapped features, this technique can selectively extract key information from the low-dimensional features and perform per-channel eigenvalue fusion with the corresponding high-dimensional mapped features. Further, this fusion method not only fully utilizes the detailed information in the low-dimensional features but also combines the high-level semantic information in the high-dimensional features, thus obtaining a more comprehensive and rich fused feature. Moreover, this fused feature not only contains the advantages of the original features but also further enhances the feature expression ability through complementarity and enhancement, providing strong support for subsequent task processing.

[0920] An exemplary code is provided in an embodiment of this application to show how to fuse low-dimensional features and high-dimensional features.

[0921] import torch

[0922] import torch.nn as nn

[0923] import torch.nn.functional as F

[0924] # Two features with different dimensions: low_dim_features and high_dim_features

[0925] # Use randomly generated tensors to represent these features

[0926] # Low-dimensional features and high-dimensional features

[0927] low_dim_features = torch.randn(1, 64, 32, 32) # Dimension: [batch_size, channels, height, width]

[0928] high_dim_features = torch.randn(1, 128, 16, 16)# Dimension: [batch_size, channels, height, width]

[0929] # Define a function to map features to the same spatial dimension

[0930] def map_to_same_dimension(feature, target_channels):

[0931] # Use 1x1 convolution to change the number of channels

[0932] return nn.Conv2d(in_channels=feature.shape[1], out_channels=target_channels, kernel_size=1)(feature)

[0933] # Map low-dimensional and high-dimensional features to the same number of channels

[0934] low_dim_mapped = map_to_same_dimension(low_dim_features, 64)

[0935] high_dim_mapped = map_to_same_dimension(high_dim_features, 64)

[0936] # Perform eigenvalue progressive sampling on the low-dimensional mapped features and use average pooling to reduce the spatial size

[0937] low_dim_pooled = F.avg_pool2d(low_dim_mapped, kernel_size=2, stride=2)# Reduce the spatial size by half

[0938] # Ensure that the spatial size of the high-dimensional mapped features matches that of the sampled low-dimensional features

[0939] # If they don't match, interpolation or upsampling techniques may be needed

[0940] # Channel-wise eigenvalue fusion, using addition fusion

[0941] fused_features = low_dim_pooled + high_dim_mapped

[0942] In this exemplary code, two features of different dimensions are defined: low-dimensional features and high-dimensional features. To fuse these two features, they are mapped to the same number of channels through a mapping function, thereby obtaining aligned low-dimensional mapped features and high-dimensional mapped features. Eigenvalue sampling is performed on the low-dimensional mapped features level by level to reduce their spatial size. Average pooling is used here as a simplified example, but more complex sampling strategies may be required in actual applications. Then, ensure that the spatial size of the high-dimensional mapped features matches that of the sampled low-dimensional features. If they do not match, interpolation or upsampling techniques can be used to adjust the size. The two mapped features are combined through per-channel eigenvalue fusion to obtain the fused features.

[0943] Optionally, locating the position of the vehicle component on the TFDS image to be processed according to the fused feature includes: determining multiple optional regions of the vehicle component on the TFDS image to be processed and the confidence levels of different optional regions according to the fused feature; based on the confidence levels, performing regression iteration on the multiple optional regions to generate a positioning box of the vehicle component on the TFDS image to be processed;

[0944] Locating the position of the vehicle component on the TFDS image to be processed according to the positioning box of the vehicle component on the TFDS image to be processed.

[0945] In this embodiment, multiple possible regions of the vehicle component on the TFDS image to be processed are determined according to the fused feature, and each region is given a confidence level. The rich information of the fused feature is utilized to effectively narrow the search range. Immediately afterwards, based on these confidence levels, the technique performs regression iteration on the multiple optional regions. Through continuous optimization and adjustment, an accurate positioning box of the vehicle component is finally generated. This process of regression iteration not only improves the positioning accuracy but also ensures the accuracy and adaptability of the positioning box. According to these positioning boxes, the technique can accurately locate the specific position of the vehicle component on the TFDS image to be processed.

[0946] An embodiment of this application provides an exemplary code showing the process of locating the position of a vehicle component on a TFDS image according to a fused feature.

[0947] import torch

[0948] import torch.nn as nn

[0949] import torchvision.transforms as transforms

[0950] from torchvision.models.detection import FasterRCNN

[0951] from torchvision.models.detection.rpn import AnchorGenerator

[0952] from PIL import Image

[0953] # Pretrained FasterRCNN model

[0954] # Use a randomly generated tensor as the fused feature map

[0955] # Fused feature map

[0956] fusion_feature_map = torch.randn(1, 256, 32, 32) # Dimensions: [batch_size, channels, height, width]

[0957] # Pretrained Faster R-CNN model

[0958] # model = FasterRCNN(...) # Configure specific model parameters

[0959] # Partial functions of Region Proposal Network (RPN) and classification / regression heads

[0960] # 1. Determine multiple candidate regions (i.e., proposed regions) and their confidences

[0961] # Use a function to determine regions and confidences

[0962] def propose_regions(fusion_feature_map, num_proposals=10):

[0963] # Output of RPN, returns proposed regions (bboxes) and corresponding confidences (scores)

[0964] bboxes = torch.randn(num_proposals, 4) * 100 # bbox are coordinates relative to the image size (x1, y1, x2, y2)

[0965] scores = torch.randn(num_proposals) # Confidence scores

[0966] scores = nn.functional.softmax(scores, dim=0) # Apply softmax to ensure the sum of confidences is 1

[0967] return bboxes, scores

[0968] # 2. Iteratively regress multiple candidate regions based on confidence

[0969] # Achieved through an ROI Pooling / Align layer and a classification / regression head

[0970] def regress_and_classify_regions(bboxes, scores, fusion_feature_map):

[0971] # Obtained the final localization bounding box

[0972] # Then optimize the coordinates of the bbox and obtain the final class through the classification / regression head

[0973] # Select the region with the highest confidence as the localization bounding box

[0974] top_idx = scores.argmax()

[0975] final_bbox = bboxes[top_idx]

[0976] return final_bbox

[0977] # 3. Locate the position of the vehicle part based on the localization bounding box

[0978] def locate_vehicle_part(final_bbox, image_size):

[0979] # image_size is the size of the input image (height, width)

[0980] # Convert the bbox coordinates to positions on the image (pixel coordinates)

[0981] x1, y1, x2, y2 = final_bbox * torch.tensor(image_size).float()

[0982] return (x1.item(), y1.item()), (x2.item(), y2.item())

[0983] # Input image size

[0984] image_size = (800, 600)

[0985] # 1. Proposal regions and confidence

[0986] bboxes, scores = propose_regions(fusion_feature_map)

[0987] # 2. Regression iteration and classification

[0988] final_bbox = regress_and_classify_regions(bboxes, scores, fusion_feature_map)

[0989] # 3. Locate the positions of vehicle parts

[0990] part_location = locate_vehicle_part(final_bbox, image_size)

[0991] print(f"The located position of the vehicle part is: top - left coordinate {part_location[0]}, bottom - right coordinate {part_location[1]}")

[0992] In this exemplary code, the function propose_regions constructs the function of the region proposal network, determining multiple regions that may contain vehicle parts and their confidence levels. Then, the function regress_and_classify_regions constructs a more basic version by classifying / regressing the proposals through the classification / regression head.

[0993] Optionally, based on the fusion feature, the positions of the vehicle parts on the to - be - processed TFDS image are located. After that, it further includes: taking a screenshot of the to - be - processed TFDS image based on the positions of the vehicle parts on the to - be - processed TFDS image to generate a sub - image of the vehicle parts.

[0994] In this embodiment, the position of vehicle components on the TFDS image to be processed can be located efficiently and accurately. After successfully locating the vehicle components, the technology can further automatically take an accurate screenshot of the TFDS image based on this position information, thereby generating a sub-image of the vehicle components. Secondly, this function not only greatly improves the efficiency of image processing, but also ensures that the intercepted sub-image can accurately reflect the details of the vehicle components, providing high-quality image data for subsequent analysis and recognition.

[0995] An exemplary code is provided in an embodiment of this application to show how to locate the position of vehicle components on the TFDS image according to the fusion features and take a screenshot of the image based on these positions to generate a sub-image of the vehicle components.

[0996] from PIL import Image

[0997] # TFDS image and the position information of vehicle components

[0998] # Use the Pillow library to load an image file

[0999] image_path = 'path_to_tfds_image.jpg' # Replace with the TFDS image path

[1000] img = Image.open(image_path)

[1001] # The position of vehicle components (coordinates of the upper left corner and the lower right corner)

[1002] part_location = ((100, 200), (300, 400)) # Coordinates of the upper left corner and the lower right corner

[1003] # Based on the position information of vehicle components, take a screenshot of the image to generate a sub-image

[1004] def crop_image_for_vehicle_part(img, part_location):

[1005] left, top, right, bottom = part_location[0][0], part_location[0][1],part_location[1][0], part_location[1][1]

[1006] # Use the crop method of Pillow to crop the image

[1007] cropped_img = img.crop((left, top, right, bottom))

[1008] return cropped_img

[1009] # Perform screenshot operation

[1010] cropped_part = crop_image_for_vehicle_part(img, part_location)

[1011] # Save the sub - image of the vehicle part

[1012] cropped_part_path = 'path_to_save_cropped_image.jpg'# Replace with the path to save

[1013] cropped_part.save(cropped_part_path)

[1014] In this exemplary code, a TFDS image is loaded. Then, the position information of a vehicle part on the image is defined. Next, a function crop_image_for_vehicle_part is defined, which takes the original image and the position information of the vehicle part as inputs and uses the crop method of the Pillow library to crop the image. This function is called to obtain the sub - image of the vehicle part and save it to the specified path.

[1015] The above are only the preferred embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for locating railway vehicle image faults, characterized in that: The positioning method comprises: Obtain labeled railway vehicle image samples and unlabeled TFDS image samples; Preprocessing the labeled railway vehicle image samples using the unlabeled TFDS image samples, and forming the model pre-training image samples with the unlabeled TFDS image samples; Acquire a TFDS image sample, and perform image feature fusion on the TFDS image sample to obtain a TFDS fused image sample; Generate labeled vehicle component image samples according to the vehicle component positioning labels; Preprocessing the labeled vehicle component image samples and merging them with the TFDS image samples to form image samples for model training; According to the model pre-training image samples, semi-automatically annotating some unlabeled vehicle component image samples to construct a model pre-training image set; Constructing a model training image set based on the model training image samples and the vehicle component positioning labels; Based on the model pre-training image set, the target large model is trained to obtain a fault component location pre-training model; Based on the model training image set, the fault component location pre-training model is trained to obtain a fault component location model; Acquire the TFDS image to be processed, and input it into the faulty component positioning model to locate the position of the vehicle faulty component on the TFDS image to be processed; in, The TFDS image samples are image samples of the freight train running fault trackside image detection system; The image feature fusion of the TFDS image samples to obtain the TFDS fused image samples includes: representing each TFDS image sample as a pixel value matrix; extracting feature points of the corresponding TFDS image samples based on the pixel value matrix based on a set feature extraction function; determining corresponding feature points on different TFDS image samples; determining geometric transformation functions of different TFDS image samples, and aligning corresponding feature points on different TFDS image samples based on the geometric transformation function; fusing different TFDS image samples with aligned feature points to obtain the TFDS fused image sample.

2. A method for locating railway vehicle image faults according to claim 1, characterized in that: The positioning method further includes: Digitally capture railway vehicles to obtain digital images; According to the digitized image, a structural model of each component of the vehicle is constructed, and based on the structural model of each component of the vehicle and according to the image samples for model pre-training, some unlabeled vehicle component image samples are semi-automatically annotated to construct a model pre-training image set; The step of digitally collecting the railway vehicle to obtain a digital image includes: Based on the set acquisition strategy, multiple batches of digital acquisition are performed on different railway vehicles of the same model to obtain digital images corresponding to the railway vehicles of the same model. The acquisition strategy is configured based on the acquisition equipment and acquisition method used. The acquisition equipment includes a linear array camera, an area array camera, and a laser scanner. The acquisition method includes shooting with a linear array camera, shooting with an area array camera, and scanning with a laser scanner.

3. A method for locating railway vehicle image faults according to claim 2, characterized in that: The method of performing multiple digital acquisitions on different railway vehicles of the same model based on the set acquisition strategy to obtain digital images corresponding to the railway vehicles of the same model includes: Using the same acquisition method and different acquisition equipment, multiple digital acquisitions are performed on different railway vehicles of the same model to obtain digital images corresponding to the railway vehicles of the same model; The digitized images corresponding to the railway vehicles of the same model are merged to obtain a single digitized image corresponding to a single sampling mode; A plurality of single digitized images corresponding to a single sampling mode are fused to obtain a digitized image corresponding to the same type of railway vehicle under the same sampling mode; According to the digital images corresponding to the same type of railway vehicle under different sampling modes, the digital images corresponding to the same type of railway vehicle are obtained.

4. A method for locating railway vehicle image faults according to claim 3, characterized in that: The method of obtaining the digitized images corresponding to the railway vehicles of the same model according to the digitized images corresponding to the railway vehicles of the same model under different sampling modes comprises: Obtain digital models of vehicle structures and repair parts; The digital images corresponding to the same type of railway vehicle under different sampling modes, the digital model of the vehicle structure and the digital model of the maintenance parts are fused to obtain the digital images corresponding to the same type of railway vehicle.

5. A method for locating railway vehicle image faults according to claim 4, characterized in that: The collection device set consisting of different collection methods and different collection devices is recorded as: Multiple batches of digital acquisition are performed on different railway vehicles of the same model to obtain the corresponding digital images of railway vehicles of the same model, which are recorded as: 。 6. A method for locating railway vehicle image faults according to claim 5, characterized in that: Based on the following formula, for the same type of railway vehicle, the digitized images obtained by using one acquisition method and different acquisition devices are fused and calculated to obtain the single digitized image of the railway vehicle of this model obtained by using different acquisition devices based on this acquisition method, which is recorded as : ,in, is the affine transformation matrix, The collection weight values ​​assigned to different collection devices, weight values for ; Based on the following formula, the digitized images acquired multiple times are fused to obtain a single batch of digitized images of the same type of railway vehicle acquired in the same way, which is recorded as : in, is the affine transformation matrix, β is the weight value assigned to different batches of acquisitions, and the weight value β is ; Based on the following formula, the digitized images of the same type of railway vehicle obtained by different acquisition methods are fused, and the digitized images of the same type of railway vehicle are recorded as : in, is the fusion constant, The weight values ​​assigned to different collection methods, for ; The digital images corresponding to the same type of railway vehicle under different sampling methods, the digital model of the vehicle structure and the digital model of the maintenance parts are fused to obtain the digital images corresponding to the same type of railway vehicle, which are recorded as : ,in, , , The weight values ​​are the digitized images of railway vehicles of the same model. , is the digital model of the vehicle structure. Digital models of repair parts.

7. A method for locating railway vehicle image faults according to claim 6, characterized in that: The positioning method further includes: Establishing a physical structure model of the railway vehicle according to the structural diagram of the railway vehicle; The TFDS historical driving images are fused with the physical structure model of the railway vehicle to generate a vehicle structure digital model.

8. A method for locating railway vehicle image faults according to claim 7, characterized in that: The fusion of the TFDS historical driving image and the physical structure model of the railway vehicle to generate a vehicle structure digital model includes: fusing the TFDS historical driving image with the physical structure model of the railway vehicle to obtain an initial physical structure model of the railway vehicle; Based on the component fine-tuning knowledge description, the initial physical structure model of the railway vehicle is adjusted to generate a vehicle structure digital model.

9. A method for locating railway vehicle image faults according to claim 8, characterized in that: Based on the structural models of the vehicle components, semi-automatically labeling some unlabeled vehicle component image samples according to the model pre-training image samples to construct a model pre-training image set includes: Based on the structural models of the vehicle components, positioning the vehicle components on the image samples used for pre-training of the model; Obtaining the annotation points of vehicle parts on the image samples used for pre-training of the model; According to the positioning of the vehicle parts and the marked points of the vehicle parts, reasoning is performed on some unlabeled vehicle part image samples to mark the vehicle parts on the some unlabeled vehicle part image samples and construct an image set for model pre-training.

10. A method for locating railway vehicle image faults according to claim 9, characterized in that: The method of training the target large model based on the model pre-training image set to obtain a fault component location pre-training model includes: Performing unsupervised training on the target large model based on the unlabeled TFDS image samples in the model pre-training image set to obtain an unsupervised large model; Based on the digitized images of railway vehicles in the model pre-training image set, the unsupervised large model is enhanced and trained to obtain the original large model; Based on the labeled railway vehicle image samples in the model pre-training image set, the original large model is subjected to supervised training to obtain an initial large model; Based on the vehicle component image samples after semi-automatic annotation processing, supervised training is performed on the initial large model to obtain a pre-trained model for locating faulty components.

Citation Information

Patent Citations

  • Van transportation management control method based on Raspberry Pi

    CN111723705A

  • Railway wagon brake shoe fault detection method based on deep learning

    CN114399672A