Train hood front opening and closing damage fault identification method based on Faster R-CNN

By fusing multi-angle polarized images with RGB images and using region-sensitive weight distribution, combined with the Faster R-CNN model, the problem of misjudgment in the fault identification of the front opening and closing mechanism of the train head cover in the existing technology is solved, and high-precision and reliable fault identification is achieved in complex environments.

CN122049523APending Publication Date: 2026-05-15QINGDAO HAITE NEW MATERIAL BOAT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QINGDAO HAITE NEW MATERIAL BOAT CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing methods fail to effectively consider the impact of optical texture fluctuations on the reliability of judgment when identifying faults in the opening and closing mechanism of the train head cover, and lack a region-sensitive mechanism, resulting in a high false recognition rate and an inability to accurately identify structural damage in key parts such as hinges, latches, and slide rails.

Method used

By fusing multi-angle polarized images with synchronous RGB images, surface micro-damage features, component contour features, and structural semantic features are extracted, a region-sensitive weight distribution is constructed, and recognition is performed using a Faster R-CNN model. Combined with texture stability scores, a fault report is output.

Benefits of technology

In complex lighting and strong reflection environments, it improves the stability and accuracy of fault identification, reduces false judgments, enhances the detection capability of key parts, and provides higher identification accuracy and judgment reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122049523A_ABST
    Figure CN122049523A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of graphic data processing, in particular to a Faster R-CNN-based train hood front opening and closing damage fault identification method. The method comprises the following steps: collecting a polygonal polarization image sequence and a synchronous RGB image of a train hood; determining a reflection suppression image through the polygonal polarization image sequence, and generating a fused visual image according to the reflection suppression image and the synchronous RGB image; extracting a surface micro-damage feature, a component contour feature and a structure semantic feature based on the fused visual image, and generating an encrypted feature set; and reconstructing high-resolution detail features based on the encrypted feature set. According to the method, through a multi-source cooperation mechanism of polarization-RGB fusion imaging and confidence-texture stability joint judgment, the tiny damage of the front opening and closing mechanism of the train hood can be stably and accurately recognized, and the fault judgment reliability is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of graphic data processing technology, and in particular to a method for identifying faults in the opening and closing of train head covers based on Faster R-CNN. Background Technology

[0002] The train headliner, a crucial protective component of the high-speed train's front structure, typically relies on hinges, latches, and sliding rails for its opening and closing mechanism. During long-term operation, the headliner's opening and closing mechanism is susceptible to structural damage due to vibration, impact, environmental corrosion, and fatigue loads, including crack propagation, latch notches, latch breakage, and hinge loosening. Faster R-CNN is a deep learning model that uses a Region Proposal Network (RPN) to achieve end-to-end, efficient object detection.

[0003] Existing methods often focus on the classification confidence of deep network outputs, but fail to consider the impact of optical texture fluctuations on judgment reliability. When local reflection anomalies or texture noise occur, the model's high-confidence predictions may still correspond to incorrect results, directly reducing the reliability of the intelligent inspection system. For structurally stressed areas such as hinge areas, latch areas, and slide rail areas, traditional detection models lack regional sensitivity mechanisms, cannot assign feature weights based on the functional differences of components in the opening and closing mechanism, and cannot identify linkage anomalies from a structural logic perspective. Summary of the Invention

[0004] Therefore, it is necessary for the present invention to provide a method for identifying train head cover front opening and closing damage faults based on Faster R-CNN, in order to solve at least one of the above-mentioned technical problems.

[0005] To achieve the above objectives, a method for identifying train head cover opening / closing damage faults based on Faster R-CNN includes the following steps:

[0006] Step S1: Acquire a sequence of multi-angle polarized images and a synchronized RGB image of the train headliner; determine the reflection suppression image through the multi-angle polarized image sequence, and generate a fused visual image based on the reflection suppression image and the synchronized RGB image;

[0007] Step S2: Extract surface micro-damage features, component contour features, and structural semantic features based on the fused visual image, and generate an encrypted feature set; reconstruct high-resolution detail features based on the encrypted feature set;

[0008] Step S3: Construct region-sensitive weight distributions for the hinge area, latch area, and slide rail area based on high-resolution detail features, and output structural reinforcement features;

[0009] Step S4: Input the structural reinforcement features into the pre-trained Faster R-CNN model, output the recognition result information, and calculate the damage confidence.

[0010] Step S5: Construct a texture stability score based on the polarization residual map of the fused visual image, determine a comprehensive judgment score using the damage confidence score and the texture stability score, and output a fault report.

[0011] This invention utilizes polarization-RGB fusion imaging, heterogeneous feature representation, region-sensitive weight modeling, and joint determination of confidence and texture stability to ensure the stability and reliability of damage identification in the opening and closing mechanism of the train headliner under complex lighting conditions, high-reflectivity metal materials, and significant differences in stress structures. Polarization imaging reduces the masking of weak-texture cracks and small notches by specular reflection, allowing the metal surface to retain readable texture even in strong light environments. The detail reconstruction process after fusion with visual enhancement makes micro-damage such as crack tip bifurcation, micro-notches in the latch claws, and deformation of the lock groove boundary more distinguishable in subsequent detection stages. The region-sensitive weight mechanism strengthens the feature representation of key stress-bearing parts such as hinges, latches, and slide rails, enabling the detection network to exhibit differentiated response capabilities to structural logic anomalies and reducing missed detections caused by the failure to represent functional differences of components. The judgment method, which integrates damage confidence and texture stability, can suppress the risk of misjudgment from single-probability outputs under unstable imaging conditions such as strong reflection, periodic jitter texture, and local contamination. This makes the final judgment of faults such as hinge loosening, latch breakage, and track deviation more consistent with the physical damage mechanism. Through the combination of the above-mentioned multi-source enhancement, structural constraints, and reliability verification, this method achieves higher recognition accuracy and judgment stability in the complex service scenarios of high-speed trains, providing reliable support for the intelligent inspection and condition assessment of the headliner opening and closing mechanism. Attached Figure Description

[0012] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0013] Figure 1 This is a flowchart illustrating the steps of a method for identifying train head cover front opening and closing damage faults based on Faster R-CNN according to the present invention.

[0014] Figure 2 A schematic diagram showing the damage to the front opening and closing mechanism of the train's head cover;

[0015] Figure 3 This is a detailed diagram illustrating the Faster R-CNN detection process. Detailed Implementation

[0016] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0017] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0018] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0019] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a method for identifying train headliner front opening / closing damage faults based on Faster R-CNN, the method comprising the following steps:

[0020] Step S1: Acquire a sequence of multi-angle polarized images and a synchronized RGB image of the train headliner; determine the reflection suppression image through the multi-angle polarized image sequence, and generate a fused visual image based on the reflection suppression image and the synchronized RGB image;

[0021] Step S2: Extract surface micro-damage features, component contour features, and structural semantic features based on the fused visual image, and generate an encrypted feature set; reconstruct high-resolution detail features based on the encrypted feature set;

[0022] Step S3: Construct region-sensitive weight distributions for the hinge area, latch area, and slide rail area based on high-resolution detail features, and output structural reinforcement features;

[0023] Step S4: Input the structural reinforcement features into the pre-trained Faster R-CNN model, output the recognition result information, and calculate the damage confidence.

[0024] Step S5: Construct a texture stability score based on the polarization residual map of the fused visual image, determine a comprehensive judgment score using the damage confidence score and the texture stability score, and output a fault report.

[0025] Preferably, in step S1, determining the reflection suppression image through a multi-angle polarization image sequence and generating a fused visual image based on the reflection suppression image and the synchronized RGB image includes:

[0026] Extract the reflected light components and diffuse reflection light components at different polarization angles from the multi-angle polarized image sequence, and calculate the brightness difference matrix;

[0027] Based on the brightness difference matrix, specular reflection interference is suppressed to generate a reflection-suppressed image;

[0028] Extract reflection suppression features from reflection-suppressed images;

[0029] Extract color structure features from synchronized RGB images;

[0030] The reflection suppression features and color structure features are weighted and concatenated, and the weighted concatenation result is optimized to output a fused visual image.

[0031] In this embodiment of the invention, the detection area of ​​the train headliner acquires optical data through a polarization imaging unit mounted on a fixed bracket and a synchronously triggered RGB imaging unit. The polarization imaging unit is configured with four polarization angles: , , , The tolerance for the polarization direction and the transmission direction of the polarizer is limited to... Four polarization images are input into the data acquisition module with fixed numbers to establish a multi-angle polarization image sequence based on pixel positions.

[0032] After the multi-polarized image sequence is input to the optical component separation module, the module uses a pixel-by-pixel comparison method to identify high-brightness and low-brightness pixels at different polarization angles, and determines the relative strength of specular reflection and diffuse reflection components based on the brightness difference. Pairwise differences are calculated for the four brightness values ​​at the same pixel location; differences greater than 15 are considered to indicate specular reflection dominance, and differences less than 15 are considered to indicate diffuse reflection dominance. The results of all pixel difference calculations are recorded in a two-dimensional matrix, which is then input as the brightness difference matrix to the reflection suppression module.

[0033] The reflection suppression module takes the brightness difference matrix as input and reduces the effect of specular reflection by pixel-wise brightness reduction. The reduced brightness data is then used to reconstruct a single-channel image as the reflection suppression image using pixel averaging. After the reflection suppression image is input to the feature extraction module, the feature extraction module extracts reflection suppression features through three layers of convolutional units with fixed parameters. Each convolutional layer uses a fixed ReLU activation function.

[0034] Synchronized RGB images are processed through a separate color structure extraction module. This module decomposes the RGB image into a three-channel matrix, performing edge gradient calculation and color distribution mean difference calculation channel by channel. The resulting channel-wise concatenation forms a color structure feature matrix. This matrix is ​​then weighted and combined with the reflection suppression feature matrix along the channel direction in the feature concatenation module.

[0035] The weighted concatenated features are input into the convolution optimization module. The convolution optimization module first uses a... The convolutional unit completes local feature interactions and compresses the number of channels to 32, then uses one The convolutional unit performs channel adjustment and further compresses the number of channels to 16. The compressed feature matrix is ​​then normalized by a pixel-by-pixel normalization unit.

[0036] Most importantly, the convolution optimization of the weighted splicing result is specifically as follows:

[0037] The weighted splicing result is passed through the first layer. The convolutional layer performs local feature interaction and dimensionality reduction, and performs nonlinear transformation on the output of the first convolutional layer.

[0038] The nonlinear transformation result is passed through the second layer. The convolutional layer performs channel adjustment and feature compression, and outputs the optimized convolution result;

[0039] The convolution optimization results are normalized at the pixel level to generate a fused visual image.

[0040] In this embodiment of the invention, reflection suppression features and color structure features are concatenated in a channel-direction weighted manner according to a fixed weight ratio (reflection suppression features account for 65%, and color structure features account for 35%). This feature tensor is input to the convolution optimization module, and the first processing unit of the convolution optimization module is a single-layer fixed structure. Convolutional layer. This convolutional layer performs position-wise local convolution operations, causing each output channel to be derived from a local convolution of the 48 channels of the input tensor. Information is extracted from the region, and local feature interaction and channel dimensionality reduction transformation are achieved through a linear combination of 32 convolutional kernels. After the convolution operation, the size of the output intermediate tensor is [size missing]. .

[0041] The intermediate tensor, after being output from the convolutional layer, enters the nonlinear transformation unit. The nonlinear transformation unit employs a uniform ReLU processing method. ReLU processing is performed pixel-by-pixel; if the feature value at a point is less than 0, the feature value at that point is set to 0; if the feature value is greater than or equal to 0, it is retained. This processing removes all negative features from the 32-channel intermediate tensor, ensuring that subsequent feature compression is performed based on the non-negative feature space, thus preventing negative features from affecting the compression results between channels. After this nonlinear transformation, the feature tensor size remains [size missing]. .

[0042] The tensor after nonlinear transformation is input to the second processing unit of the convolution optimization module, i.e. Convolutional layer. Because the kernel size is... Therefore, each convolutional kernel performs a linear combination operation once for each spatial location, combining the feature values ​​of the 32 channels according to weights into feature values ​​of the 16 output channels, achieving channel adjustment and feature compression. This process does not change the spatial dimensions, so the output tensor size is... This compression process rearranges and condenses the local features extracted by the previous convolution stage according to the set weights, compressing the feature dimension from 32 channels to 16 channels, so that the subsequent step S2 can have a more distinguishable input feature structure when performing multi-type feature extraction.

[0043] The convolutional feature tensor is input to a pixel-level normalization unit. The pixel-level normalization unit performs normalization on a channel-by-channel basis. During execution, the maximum and minimum values ​​of all pixels within each channel are calculated. After calculation, a normalization transformation is performed on each pixel within that channel. When the maximum and minimum values ​​are equal, to prevent division by zero, all pixel values ​​in that channel are directly normalized to 0.

[0044] Of particular importance is the extraction of surface micro-defect features, component contour features, and structural semantic features based on fused visual images, specifically as follows:

[0045] Construct a multi-scale semantic perception pyramid to separate and fuse normal paint surfaces and potential damage areas in visual images, and determine surface micro-damage features;

[0046] Based on the surface micro-damage features, the width anomaly of the opening and closing gaps and the deformation displacement of the component assembly boundary are quantified to generate the component contour features;

[0047] Based on the component contour features, the latch and hinge components are abstracted into topological nodes. The spatial constraint relationship and linkage logic anomaly between components are judged by the graph attention mechanism, and a high-order structural semantic feature is constructed.

[0048] In this embodiment of the invention, the size of the fused visual image output in step S1 is set to... Where H and W are the image height and width, respectively, and 16 is the normalized number of channels. The fused visual image is first input into a multi-scale semantic perception pyramid building unit. This pyramid structure consists of three fixed-scale downsampling branches, which perform processing at the original scale, 1 / 2 scale, and 1 / 4 scale, respectively. All three inputs use the same convolution kernel size. Feature extraction is performed using convolutional units with 16 channels and a stride of 1. The multi-scale outputs are then stitched together at the top layer of the pyramid along the channel direction to form a uniform size. The feature tensor.

[0049] The feature tensor is input to the semantic region separation unit. The semantic region separation unit divides normal paint surface areas and potential damage areas according to a fixed grayscale threshold range. The average value of each of the 48 channels is calculated. Channels with an average value greater than 0.55 are marked as high-brightness contribution channels, and channels with an average value less than or equal to 0.55 are marked as low-brightness contribution channels. When a pixel's count in the high-brightness contribution channel is greater than or equal to 24, the pixel is classified as a potential damage area; otherwise, it is classified as a normal paint surface area. The region determination results form a binary mask image. The binary mask image and the pyramid feature tensor are superimposed pixel-wise. The 48-channel features at the corresponding positions are retained in the high-brightness mask area, while all feature values ​​in the low-brightness mask area are set to 0, ultimately obtaining the surface micro-damage feature tensor.

[0050] Surface micro-damage features are input into the component profile measurement unit. The component profile measurement unit first extracts the width data of the hood opening and closing slits. The slit widths obtained from all scan lines are averaged using a simple averaging method to calculate the mean width and generate width anomaly features. This width feature is superimposed onto the channel dimension of the surface micro-damage feature tensor, expanding it into a sheet of size using numerical repetition. The width matrix is ​​then spliced ​​with the surface micro-defect features along the channel direction to form... The initial contour feature tensor.

[0051] To determine the deformation displacement of the component assembly boundary, the component contour measurement unit uses a Sobel gradient template (fixed at [-1 0 1; -2 0 2; -1 0 1] and its transpose) to enhance edges in both the horizontal and vertical directions within the fused visual image. The enhanced edge map is then binarized using a threshold of 0.4. The binarized edge map is scanned row by row, and the leftmost and rightmost edge pixel positions of each row are extracted. The difference between these positions and the stored values ​​of the standard assembly boundary reference positions is calculated to form a deformation displacement map. This deformation displacement map is then linearly interpolated and mapped to... After scaling, the component contour feature tensor is concatenated with the aforementioned 49-channel contour features in a channel expansion manner to form a 50-channel component contour feature tensor, which is then fed into the structural semantic construction unit.

[0052] The structural semantic construction unit first extracts the spatial positions of the latch and hinge regions based on the component contour features. This process is achieved using two sets of fixed-region masks: a latch mask covers the lower middle region of the image; and hinge masks cover the upper two sides of the image. The component contour features are averaged region by region under the masking effect, and the average feature vector of each region is used as the topological node feature for that region. The node features are represented by vectors of fixed length 50, where each vector element is derived from the average value of the corresponding channel of the component contour features. The number of latch nodes is set to 1, and the number of hinge nodes is set to 2, for a total of 3 nodes.

[0053] Structural semantic building blocks An adjacency matrix is ​​generated, where each element is calculated using the distance between nodes and the feature difference between nodes. The distance is calculated using Euclidean distance, and the feature difference is calculated as the mean of the absolute values ​​of the channel differences. This adjacency matrix is ​​input to the graph connection calculation unit. This unit performs row-by-row weighting for each node. The weighting method is as follows: the three adjacency values ​​of the row are combined with fixed weights: (lock → hinge left) 0.4, (lock → hinge right) 0.4, (between two nodes in the hinge) 0.2. The weighted connection strength between nodes is used to determine whether the spatial constraint relationship exceeds a preset threshold. The preset threshold is 0.35; exceeding this threshold indicates an anomaly in the linkage logic between two nodes. The weighted relation values ​​of each node are combined with the binary result of whether it is an anomaly to form a node-level semantic vector. The node-level semantic vector has a fixed length of 4 (1 weighted value, 1 anomaly marker, 1 node type, and 1 node position). The node-level semantic vectors of the three nodes are then concatenated sequentially to form a structural semantic feature vector of length 12.

[0054] Preferably, generating the encryption feature set in step S2 includes:

[0055] Construct a cross-scale feature correlation map between surface micro-destruction features, component contour features, and structural semantic features;

[0056] Adaptive fusion of heterogeneous features is performed based on cross-scale feature association graphs to form an initial multi-scale feature set;

[0057] Enhance fault features related to cracks, gaps, and deformation based on the initial multi-scale feature set;

[0058] Based on the differentiated responses of preset fault modes on the feature channels, channel weights are dynamically allocated, and an encrypted feature set strongly correlated with the physical damage mechanism is output.

[0059] In this embodiment of the invention, the surface micro-damage feature tensor size is set to... The component contour feature tensor size is set to The length of the structural semantic feature vector is set to 12. To establish the correlation between the three types of features in the spatial and semantic dimensions, a cross-scale feature correlation map is first constructed. The construction method is as follows: the 12 elements in the structural semantic feature vector are divided into three groups according to type, representing the correlation value of the latch region, the correlation value of the left and right hinge regions, and their anomaly markers, respectively. Within each group, dimensions are mapped to a repeating expansion method. A single-channel matrix. The expansion method involves copying the same value to all pixel positions in the matrix, giving the semantic values ​​of the three regions the same spatial dimension as the surface micro-loss feature tensor. The three expanded matrices are then stitched together along the channel direction to form the size... The semantic extension tensor.

[0060] The surface micro-destruction feature tensor, the component contour feature tensor, and the semantic extension tensor are concatenated along the channel direction. The cross-scale fusion input tensor is then used. Next, a cross-scale feature correlation map is constructed. This correlation map is formed by calculating the joint response values ​​of the three types of features position-by-position. The calculation method is as follows: Channel difference statistics are performed sequentially on the 101 channels after stitching. The value of each channel is subtracted from the minimum value among the 101 channels. When the difference is greater than 0.45, it is recorded as a significant response at that pixel location; if the difference is less than or equal to 0.45, it is considered a normal response. The significant responses of all channels are accumulated pixel-by-pixel to obtain a saliency matrix. The saliency matrix is ​​binarized using a threshold of 45. When the significant response value at a certain location exceeds 45, it is marked as a high-correlation region; other locations are marked as low-correlation regions. The binary map of the high-correlation region serves as the cross-scale feature correlation map, used to indicate the potential common aggregation locations of crack tips, gap anomalies, and structural component logical anomalies.

[0061] The cross-scale feature correlation map is multiplied pixel-by-pixel with the fused input tensor, so that highly correlated regions retain all 101 channels of response values, while the response values ​​of low-correlation regions are compressed to 5% of their original values. This compression ratio is fixed and used to reduce feature interference in non-faulty regions. The multiplication yields an initial multi-scale feature set with a size of [size missing]. .

[0062] To enhance the fault features related to cracks, fissures, and deformation, three types of fault mapping channels were defined in the initial multi-scale feature set. The first type is the crack-sensitive channel, generated by mapping the results of Sobel gradient enhancement, with channel indices set to 20-35. The second type is the fissure width anomaly channel, generated by mapping the scan line width matrix, with channel indices set to 36-50. The third type is the deformation displacement channel, generated by mapping the deformation displacement map, with channel indices set to 51-70. Feature enhancement was performed on each of the three types of channels by multiplying the corresponding channel pixel value by an enhancement coefficient. The enhanced features were then concatenated with the remaining unenhanced channels to maintain their integrity. The tensor serves as the multi-scale feature set after fault enhancement.

[0063] The dynamic channel weight allocation unit adjusts channels according to preset fault modes. These preset fault modes include crack mode, gap anomaly mode, component misalignment mode, and linkage anomaly mode of latches or hinges. The weight allocation rules are as follows: In crack mode, channels 20 to 35 are weighted at 2.0, and the remaining channels at 1.0; in gap anomaly mode, channels 36 to 50 are weighted at 2.0, and the remaining channels at 1.0; in deformation and misalignment mode, channels 51 to 70 are weighted at 1.9; in linkage anomaly mode, the three channels corresponding to the semantic extension tensor (the channels containing the extension matrix) are weighted at 2.3. Weight allocation is performed through channel-by-channel multiplication to keep the weighted channel values ​​within the normalized range. After all channels are weighted, the final result is... The set of encryption features.

[0064] Preferably, the reconstruction of high-resolution detail features based on the encrypted feature set in step S2 specifically involves:

[0065] The encrypted feature set is input into the super-resolution reconstruction network, and the encrypted feature set is upsampled and compressed. The crack edge contour and small gap boundary in the image are restored through feature reconstruction.

[0066] During the feature reconstruction process, local structural features under different receptive fields are extracted sequentially to enhance the crack tip bifurcation mode, micro-tear texture and material fatigue deformation features, and output high-resolution detail features.

[0067] In this embodiment of the invention, the size of the encrypted feature set output in step S2 is set to... Where H and W represent the height and width of the fused visual image, respectively. To ensure that crack edges, notch boundaries, and fine tear textures are fully represented in subsequent detection steps (steps S3 and S4), the encrypted feature set is first input into the upsampling module of the super-resolution reconstruction unit. The upsampling module uses bilinear interpolation to reduce the spatial size of the encrypted feature set from... Upgraded to During upsampling, the feature value of each pixel is obtained by weighted averaging of its four surrounding pixels, with the weight ratio fixed at 1:1:1:1. The size of the tensor after upsampling is... .

[0068] The upsampled tensor is input to the feature compression unit. The feature compression unit contains a convolutional kernel with a size of [missing value]. A convolutional unit with 64 channels and a stride of 1 extracts data at each spatial location. Locally weighted combinations of regions generate a 64-channel feature tensor. This is then passed through a convolutional kernel with a size of [missing information]. A 32-channel compression convolution with a stride of 1 is used to compress 64 channels into 32 channels. Throughout the process, a fixed ReLU method is used for non-linear transformation, meaning all negative values ​​are set to 0, while positive values ​​remain unchanged. This feature compression operation makes the feature dimensions more concentrated, facilitating subsequent region structure enhancement operations.

[0069] The tensor after feature compression is input to the local structure extraction unit. The local structure extraction unit extracts the crack tip, notch boundary, micro-tear texture, and fatigue deformation signal sequentially according to different receptive fields. The first step uses a convolution kernel with a size of... The first step uses a 32-channel convolutional unit to obtain the most basic local edge structure; the second step uses a convolutional kernel with a size of [missing information]. The first step uses a convolutional unit with 32 channels to capture a wider range of crack orientations; the third step uses a convolutional kernel with a size of [missing information]. Low-frequency structural deformation patterns are extracted using 16-channel convolutional units. All three convolutional operations use a stride of 1 and the same ReLU method for nonlinear processing. Due to the different receptive field sizes of the three convolutional units, they respectively enhance the bifurcation point of the crack tip, the sharpening region of the notch boundary, and the deformation texture caused by material fatigue in a layered manner.

[0070] The feature tensors extracted from the three receptive fields are concatenated along the channel direction. The structure aggregation tensor is input to the detail reconstruction fusion unit. The fusion unit integrates features from each receptive field using a weighted summation method, assigning a fixed weight to each receptive field. The summation process is performed channel-by-channel, resulting in an output feature tensor of size [size missing]. .

[0071] Ultimately, in order to achieve the requirements of high-resolution detail representation, The detail tensor is upsampled again using bilinear interpolation to make its spatial size reach .

[0072] Preferably, step S3 includes:

[0073] Obtain the spatial position masks of the hinge area, latch area and slide rail area in the train head cover, and calculate the average feature response intensity of high-resolution detail features in the hinge area, latch area and slide rail area respectively;

[0074] The initial region weights are determined based on the average characteristic response intensity of each region.

[0075] Based on the initial regional weights of each region, the spatial correlation between regions is analyzed, and the spatial correlation is normalized to generate a regional sensitive weight distribution map.

[0076] The high-resolution detail features are multiplied element-wise with the region-sensitive weight distribution map to output the structure-enhancing features.

[0077] In this embodiment of the invention, the high-resolution detail feature size output in step S2 is set to... Where H and W represent the height and width of the fused visual image, respectively, and 32 represents the number of detail channels in the aforementioned reconstruction process. To accurately highlight the key structural stress areas in the opening and closing mechanism of the train's front canopy, it is necessary to construct region-sensitive weight distributions for the hinge area, locking area, and sliding rail area on high-resolution detail features. Therefore, spatial position masks for the three areas are first obtained. The spatial position masks are stored in pixel-level binary matrix form, and the mask matrices are all consistent with the spatial dimensions of the high-resolution detail features, i.e., the size... .

[0078] After obtaining three masks, high-resolution detail features are read channel by channel, and feature values ​​of the corresponding regions are extracted under the influence of the masks. The extraction method is pixel-by-pixel traversal, calculating the feature intensity of all pixels with a value of 1 in the mask for each channel and averaging the results. Taking the hinge region as an example, if the region contains M pixels with a mask value of 1, the average feature response intensity of the hinge region is the sum of all pixels in the region across 32 channels divided by M. The latch region and slide rail region are obtained in the same way to obtain the average feature response intensity. Three average response values ​​are obtained for the three regions, denoted as RH (hinge), RL (latch), and RS (slide rail). To avoid the deviation of response intensity caused by differences in region size, each of the three regions uses the number of 1s in its own mask as the divisor for independent calculation, while maintaining consistent calculation logic.

[0079] After obtaining three average response values, initial region weights are determined based on the response amplitude. The initial region weights are set using a normalized proportional allocation, with each weight ranging from 0 to 1, and the sum of the three weights being 1. This weight combination reflects the relative strength of the detail feature responses of the three regions in the current image.

[0080] Subsequently, spatial correlations between regions are constructed based on the three initial region weights. The spatial correlation uses a fixed construction matrix, which is... The rows and columns correspond to the hinge, latch, and slide rail regions, respectively. The diagonal elements of the matrix are set as the initial weight values ​​for their respective regions, while the off-diagonal elements use the absolute values ​​of the differences in region weights. After constructing the spatial correlation matrix, normalization is required to avoid numerical differences exceeding the standard range.

[0081] To map the spatial correlation of the three regions to the spatial location of the entire image, the three diagonal elements after normalization are copied to generate three images of size [missing information]. The weight planes represent the hinge region weight plane, the latch region weight plane, and the slide rail region weight plane. The values ​​of each weight plane remain unchanged at their corresponding mask positions, and decrease pixel-by-pixel by a decay factor of 0.5 outside the mask to reflect the weak correlation in non-critical regions. The three weight planes are stitched together along the channel dimension to form the weight plane. The distribution map of regional sensitive weights.

[0082] After the region-sensitive weight distribution map is constructed, high-resolution detail features will be... Regional Sensitive Weight Distribution Map Alignment is performed according to the channel expansion method. The expansion method involves repeatedly expanding the 3-channel weight map to 32 channels, so that each detail channel corresponds to a weight channel. After expansion, the size of the aligned region-sensitive weight distribution map is [size missing]. Multiplying the high-resolution detail features element-wise with the 32-channel weight distribution map, the resulting... Tensors are structural reinforcement features.

[0083] Preferably, step S4 includes:

[0084] The structural enhancement features are input into the candidate region network of the pre-trained Faster R-CNN model to generate four types of candidate region sets;

[0085] The four candidate region sets are input into the classification subnet of the Faster R-CNN model, the probability distribution of each candidate region in each damage category is calculated, and the damage category data of each candidate region is output.

[0086] The four candidate region sets are synchronously input into the regression subnet of the Faster R-CNN model to determine the spatial coordinates of each candidate region and output the fault location data of each candidate region.

[0087] Integrate damage category data and fault location data into identification result information;

[0088] The damage confidence level of each detected target is calculated based on the category probability in the identification results.

[0089] In this embodiment of the invention, the size of the structural reinforcement feature is fixed as follows: Where H is the image height, W is the image width, and 32 is the number of detail channels. This structure-enhanced feature is then input into a pre-trained Faster R-CNN candidate region network. The input to the candidate region network uses a fixed-size convolutional window, with the window size set to [value missing]. With a step size set to 1, the structural enhancement features are scanned position by position. The local features obtained from the scan are used to generate a candidate region set through two parallel branches, where the feature scoring branch uses... The convolutional kernel generates a region score for each location as a candidate center, and the region size branch uses... The convolution kernel generates candidate region boundary offsets at different scales. The candidate region sizes are fixed at four categories: small scale (…). Pixels), Medium Scale ( Pixels), large scale ( Pixels), ultra-large scale ( (pixels), the four types of regions are combined into four candidate region sets after boundary adjustment by offset. This process ensures that the output is four independent candidate region sets, each set containing several rectangular regions, and each rectangular region is represented by the coordinates of the upper left and lower right corners.

[0090] After the four candidate region sets are generated, they are input into the classification sub-network of Faster R-CNN. The input of the classification sub-network scales the rectangular regions from the candidate regions into feature blocks of uniform size according to a fixed ratio; the scaling size is set to... Pixels. Feature blocks are processed sequentially through two fixed convolutional layers; the first convolutional kernel is... The number of channels is 64; the second convolutional kernel is... The system has 128 channels. After convolution processing, each candidate region is mapped to a probability distribution of four damage categories: cracks, open / closed gaps, latch deformation, and hinge structure offset. The probability distribution is normalized by summing the values ​​of the four categories and dividing the probability of each category by the sum, resulting in an output where the sum of the four probabilities is 1. This yields the damage category data for each candidate region.

[0091] The four candidate region sets are simultaneously input into the regression subnetwork of Faster R-CNN. The structure of the regression subnetwork is consistent with that of the classification subnetwork, but the output is a region coordinate offset. The offset includes four values: x-direction offset, y-direction offset, width adjustment factor, and height adjustment factor. The offset is applied to the initial coordinates of the candidate region. Taking the initial position of the upper left corner as the reference, the x-offset and y-offset are added to obtain the new upper left corner coordinates; the lower right corner coordinates are recalculated after scaling according to the width adjustment factor and height adjustment factor, thus outputting the fault location data corresponding to each candidate region.

[0092] Damage category data and fault location data are then integrated. The integration method involves establishing a data structure containing four elements: region ID, damage type ID, probability value corresponding to the damage type, and the four coordinate values ​​of the bounding box of the rectangle. In this way, the identification information of all candidate regions is aggregated to form the identification result information.

[0093] The category probabilities in the identification results are used to calculate the damage confidence level. The damage confidence level ranges from 0 to 1, and is calculated by selecting the highest damage category probability from the probability distribution as the damage confidence level for that candidate region.

[0094] Preferably, the structural enhancement features are input into the candidate region network of the pre-trained Faster R-CNN model as follows:

[0095] Based on the prior spatial distribution of the locking area and hinge area in the structural reinforcement features, non-uniform initial anchor point boxes with mixed density arrangement are arranged in the candidate region network, and the dense area of ​​the anchor point box corresponds to the area of ​​concentrated structural stress.

[0096] The first selection subset is generated based on the spatial overlap between the initial anchor point box and the structural reinforcement features;

[0097] A second screening is performed on the first screening subset based on the regional characteristic response intensity to form a primary candidate region set containing potential fault regions.

[0098] Constraints on the outline and boundary continuity of the primary candidate regions are applied to the concentrated superposition of structural features.

[0099] Spatial coordinate and boundary deformation optimization are performed on the constrained primary candidate region set to generate four types of candidate region sets including cracks, notches, misalignments and opening / closing mechanism damage.

[0100] In this embodiment of the invention, the structural strengthening feature has a fixed spatial size after spatial analysis. It includes 32 structural detail channels. To focus the candidate region network on the locking and hinge areas, this embodiment, based on the structural layout of the train's front opening and closing mechanism, sets the center of the locking area as the middle section of the horizontal coordinate of the structural reinforcement feature, and sets the center of the hinge area as the upper edge area of ​​the structural reinforcement feature. Based on this prior spatial distribution, a non-uniform initial anchor point frame with a mixed construction density is deployed in the candidate region network. The size of the anchor point frame is set to three levels: Pixels Pixels and The anchor point spacing is fixed at 4 pixels in densely populated areas and 12 pixels in sparsely populated areas. The anchor points in the latching and hinge areas are densely arranged, while those in other areas are sparsely arranged. This density-mixed arrangement is achieved through a rule of increasing coordinates.

[0101] After constructing the initial anchor boxes, the first round of filtering is performed based on the spatial overlap between the anchor boxes and the structural reinforcement features. The overlap is calculated as follows: the absolute values ​​of all structural reinforcement feature channels within the region containing the anchor box are summed to obtain a value representing the local response intensity; this value is then divided by the region area to obtain the normalized overlap. Anchor boxes with an overlap greater than a set threshold of 0.18 are included in the first filtering subset, while anchor boxes with an overlap less than the threshold are discarded.

[0102] The first selected subset is input into the second selection process. The second selection sorts responses based on the intensity of regional feature responses. The response intensity is calculated by applying a fixed-size [value] to the structural reinforcement features. A local window is defined, and the channel values ​​within the window area are summed using weights. The weights are set sequentially from 1 to 32 according to the channel numbers, so that structural boundary channels (those with earlier numbers) contribute more. A weighted sum is calculated at each candidate region location, and this value is used as the response intensity. Anchor boxes with response intensities higher than 0.25 form a primary candidate region set. This process does not involve any model training; region selection is achieved solely through fixed mathematical operations.

[0103] Structural contour constraints and boundary continuity constraints are superimposed on the primary candidate region set. The structural contour constraints are implemented as follows: using the component contour features from step S2, the distance between the bounding box of each primary candidate region and the component contour line is measured. If the minimum distance between the bounding box and the component contour line exceeds 12 pixels, the region is deemed to have insufficient correlation with the structural component and is discarded. The boundary continuity constraints are implemented as follows: for the edge lines of the primary candidate regions, the gradient direction continuity of the edge lines in the structural reinforcement features is measured. If the continuity is insufficient (continuity index below 0.55), the region is discarded.

[0104] Spatial coordinate and boundary deformation optimization were performed on the primary candidate regions satisfying the dual constraints. Coordinate optimization involved adding offsets derived from the local intensity center to the top-left and bottom-right corners of each candidate region. The offsets were calculated by multiplying the difference between the coordinates of the point of maximum intensity and the geometric center of the region within the candidate region by 0.12, and applying this difference to the x and y coordinates respectively, thus bringing the candidate region closer to the potential damage location. Boundary deformation optimization involved scaling the length of the rectangular sides proportionally to the intensity distribution within the region, with a scaling ratio of 1. (Local maximum strength value) The maximum expansion example is a 6% increase in the region's side length. The optimized candidate regions are divided into four categories according to their failure mode characteristics: cracks, notches, misalignments, and damage to the opening and closing mechanism, forming four candidate region sets.

[0105] Preferably, calculating the damage confidence level of each detected target based on the category probability in the identification result information includes:

[0106] Extract the four damage probability vectors corresponding to each candidate region from the classification subnet output in the recognition results information;

[0107] Calculate the information entropy between each probability direction;

[0108] The maximum probability value in the probability vector is combined with the information entropy. The confidence level is positively correlated with the maximum probability value and negatively correlated with the information entropy. The result of the combined operation is used as the damage confidence level of the corresponding detection target.

[0109] In this embodiment of the invention, after obtaining the identification result information in step S4, the identification result information includes four types of damage probability vectors corresponding to each candidate region. The probability vectors are set to a four-dimensional form, with a fixed order: crack type, open / closed gap type, latch deformation type, and hinge offset type. Taking a certain detection target as an example, the probability vectors can be set to 0.62, 0.21, 0.10, and 0.07. Four values ​​are extracted from the probability vectors.

[0110] After extracting the probability vector, the information entropy of the four-dimensional probability vector is calculated. Information entropy represents the degree of uncertainty of the probability distribution, and its calculation involves three steps: First, perform a natural logarithm operation on each probability value in the probability vector; second, multiply the probability by the corresponding natural logarithm; third, take the negative of the four products and sum them to form the final information entropy value.

[0111] After obtaining the maximum probability value and information entropy, a combination operation is performed. The combination operation uses a linear approach: the maximum probability value is multiplied by a weighting coefficient A, the information entropy is multiplied by a weighting coefficient B, the two are summed, and a constant bias C is added to form the broken confidence level. Coefficient A is set to 0.75, coefficient B to -0.45, and bias C to 0.12. The resulting confidence level is approximately 0.29. If the confidence level is less than 0, it is output as 0; if the confidence level is greater than 1, it is output as 1, ultimately forming a broken confidence level ranging from 0 to 1.

[0112] Preferably, the process of constructing a texture stability score based on the polarization residual map of the fused visual image in step S5 includes:

[0113] Extract the polarization residual map corresponding to the fused visual image;

[0114] The polarization residual map is divided into grids, and the standard deviation of pixel values ​​within each image block is calculated as the local texture fluctuation index of that image block.

[0115] Calculate the average value of the local texture fluctuation index of all image patches as the global texture stability benchmark;

[0116] Identify the target region in the polarization residual map that corresponds to the fault location data;

[0117] Calculate the relative deviation between the local texture fluctuation index and the global texture stability benchmark value of the image patch within each target region;

[0118] A texture stability score is constructed based on the weighted average of the relative deviations of all target regions.

[0119] In this embodiment of the invention, fault location data has been output and damage confidence calculation has been completed in step S4. In step S5, the corresponding polarization residual map is first extracted from the fused visual image. The polarization residual map is composed of the residual components between the reflection suppression image from step S1 and the images at each polarization angle, and its pixel value range is fixed between 0 and 255. The size of the polarization residual map is set to... It maintains consistency with the fused visual image.

[0120] The polarization residual map is divided into grid blocks. The block division method uses a fixed-size grid cell for traversal, with the grid cell size set to [value missing]. The pixels are divided into several image blocks by splitting the image in 16-pixel steps in both the horizontal and vertical directions. Since the size of the polarization residual map is consistent with the fused visual image, the resulting mesh can be obtained... Each image block contains a fixed number of 256 pixels.

[0121] For each image patch, local statistics are performed on its 256 internal pixels. This embodiment uses the standard deviation as a local texture fluctuation index. The standard deviation is calculated as follows: first, the average value of all pixels within the image patch is calculated; then, the difference between each pixel value and this average value is squared, summed, divided by 256, and finally the square root is taken to obtain the local texture fluctuation index. A larger standard deviation value indicates a stronger fluctuation in the polarization residual in that region. This calculation is performed independently for all image patches, generating a complete set of local texture fluctuation indices.

[0122] The average of the local texture fluctuation indices of all the above image patches is used to obtain the global texture stability benchmark value. This benchmark value represents the overall polarization fluctuation level across the entire image range without fault-free feature enhancement.

[0123] The target region corresponding to the polarization residual map is identified using the fault location data output in step S4. Each target region is represented as a rectangle, with its coordinates consisting of the upper left and lower right corners. For each target region, the image patch that completely covers it is selected, and the pre-calculated local texture fluctuation index is extracted from the image patch. If the target region spans multiple image patches, all intersecting image patches are included in the set corresponding to the target region.

[0124] For image patches in all target regions, the relative deviation between their local texture fluctuation index and the global texture stability benchmark is calculated. The relative deviation is calculated by subtracting the local index of the image patch from the global benchmark, and then dividing by the global benchmark to obtain a non-unitized deviation value. A positive deviation means that the texture fluctuation of the region is higher than the global average, and a negative deviation means that the fluctuation of the region is lower than the global average.

[0125] Since the degree of texture fluctuation varies depending on the type of fault, a weighted average of the relative deviations is required to form the final texture stability score. In this embodiment, a fixed weight allocation method is used: cracks are weighted at 0.40, gaps at 0.30, notches at 0.20, and structural misalignments at 0.10. For each detected target, the corresponding weight is first determined based on its fault category. Then, the relative deviations of all image patches within the target area are averaged and multiplied by the fault category weight. The result is the texture stability score for that detected target.

[0126] Preferably, in step S5, a comprehensive judgment score is determined using the damage confidence level and texture stability score, and the output fault report includes:

[0127] A weight allocation strategy is constructed based on damage confidence and texture stability scores;

[0128] In the weighting strategy, when the hinge produces periodic shaking texture and displacement consistency decreases under loose conditions, the texture stability score weight of the hinge loosening fault is greater than the damage confidence score weight; when the latch slot is broken, the latch notch is expanded, or the positioning surface is deformed, the damage confidence score weight of the latch failure fault is greater than the texture stability score weight.

[0129] Based on a weighted allocation strategy, the damage confidence score and texture stability score corresponding to the same detection target are weighted and fused to generate a corresponding comprehensive judgment score;

[0130] Different judgment thresholds are set according to different fault modes. When the comprehensive judgment score of any detected target exceeds the judgment threshold, a fault report of the corresponding fault type is generated.

[0131] In this embodiment of the invention, each detected target outputs a damage confidence score, with a fixed range of 0 to 1. In step S5, the texture stability score constructed from the polarization residual map is also set to 0 to 1. To combine these two indicators into a final comprehensive judgment score, this embodiment first constructs a weight allocation strategy and clarifies the weight ratio of different fault modes during the fusion stage.

[0132] The initial weighting strategy is as follows: the damage confidence weight is set to 0.5, and the texture stability score weight is set to 0.5, ensuring their initial contributions are equal. Subsequently, the weight ratios for different fault types are automatically adjusted based on their mechanistic differences. Structural damage in the latching area manifests as enlarged latch notches, broken slots, and localized deformation of the positioning surface. A significant characteristic of this type of damage is clear crack edges and intact notch shapes, thus the reliability of the damage confidence score is high. To enhance the characteristic contribution of this type of fault, the damage confidence weight is increased to 0.72, and the texture stability score weight is decreased to 0.28. The manifestation of hinge loosening faults differs from latching damage. Loosening causes periodic texture jitter, accompanied by reduced displacement consistency during opening and closing movements. Its visual manifestation is more derived from changes in texture stability, while the damage confidence score only reflects the spatial frame confidence level. Therefore, the texture stability score weight is set to 0.68, and the damage confidence weight is set to 0.32.

[0133] During implementation, the damage confidence score and texture stability score corresponding to the same detection target are weighted and fused according to the above weights. The calculation method is: Comprehensive Judgment Score = (Damage Confidence Score + Texture Stability Score) / (Texture Stability Score + Texture Stability Score) (Breakage confidence weight) + (Texture stability score) (Texture stability score weight). The generated comprehensive score is limited to the range of 0 to 1. If the calculated value exceeds the boundary, it is truncated according to the boundary value.

[0134] After the comprehensive assessment score is generated, differentiated assessment thresholds are set according to different failure modes. The thresholds are determined by the prior structural risk level: the threshold for cracks is set to 0.58; the threshold for notches is set to 0.62; the threshold for latch failures is set to 0.66; and the threshold for loose hinges is set to 0.55. When the comprehensive assessment score of the detected target is higher than the corresponding threshold, the identification result is classified into the corresponding failure type.

[0135] Please see Figure 2 This image shows the appearance and structural design of a train head, including the windshield, wipers, two sets of headlights, and the connecting structure below the head. The middle area is the specific area where the latch is located, which is one of the key component areas of the "train head cover front opening and closing mechanism" and is one of the key areas that need to be detected in the fault identification method based on Faster R-CNN.

[0136] Please see Figure 3This figure illustrates the core process of the Faster R-CNN object detection model, which is mainly divided into two stages: candidate region generation and object classification. The input image first extracts features through a "network" (usually a convolutional neural network), and then undergoes "classifier optimization" to initially screen out candidate regions that may contain objects. The candidate regions are then fed into a "classification network" and combined with a "confidence network" to determine the category of the object within the region. At the same time, "Regression SWork" (boundary box regression) corrects the position of the candidate regions, making the boxes fit the targets more accurately, thus achieving high-precision object detection.

[0137] Final detection: The classification results and the corrected bounding boxes are combined to output the final target detection results (including target category and location).

[0138] Therefore, the embodiments should be considered as exemplary and non-limiting in all respects, and the scope of the invention is not limited by the foregoing description. Thus, all changes falling within the meaning and scope of the equivalents of the application are intended to be included within the scope of the invention.

[0139] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A method for identifying train head cover opening / closing damage faults based on Faster R-CNN, characterized in that, Includes the following steps: Step S1: Acquire a sequence of multi-angle polarized images and a synchronized RGB image of the train headliner; determine the reflection suppression image through the multi-angle polarized image sequence, and generate a fused visual image based on the reflection suppression image and the synchronized RGB image; Step S2: Extract surface micro-damage features, component contour features, and structural semantic features based on the fused visual image, and generate an encrypted feature set; reconstruct high-resolution detail features based on the encrypted feature set; Step S3: Construct region-sensitive weight distributions for the hinge area, latch area, and slide rail area based on high-resolution detail features, and output structural reinforcement features; Step S4: Input the structural reinforcement features into the pre-trained Faster R-CNN model, output the recognition result information, and calculate the damage confidence. Step S5: Construct a texture stability score based on the polarization residual map of the fused visual image, determine a comprehensive judgment score using the damage confidence score and the texture stability score, and output a fault report.

2. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, Step S1 involves determining a reflection suppression image using a multi-angle polarization image sequence, and generating a fused visual image based on the reflection suppression image and the synchronized RGB image, including: Extract the reflected light components and diffuse reflection light components at different polarization angles from the multi-angle polarized image sequence, and calculate the brightness difference matrix; Based on the brightness difference matrix, specular reflection interference is suppressed to generate a reflection-suppressed image; Extract reflection suppression features from reflection-suppressed images; Extract color structure features from synchronized RGB images; The reflection suppression features and color structure features are weighted and concatenated, and the weighted concatenation result is optimized to output a fused visual image.

3. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, Step S2, generating the encryption feature set, includes: Construct a cross-scale feature correlation map between surface micro-destruction features, component contour features, and structural semantic features; Adaptive fusion of heterogeneous features is performed based on cross-scale feature association graphs to form an initial multi-scale feature set; Enhance fault features related to cracks, gaps, and deformation based on the initial multi-scale feature set; Based on the differentiated responses of preset fault modes on the feature channels, channel weights are dynamically allocated, and an encrypted feature set strongly correlated with the physical damage mechanism is output.

4. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, The specific steps in step S2, which involve reconstructing high-resolution detail features based on the encrypted feature set, are as follows: The encrypted feature set is input into the super-resolution reconstruction network, and the encrypted feature set is upsampled and compressed. The crack edge contour and small gap boundary in the image are restored through feature reconstruction. During the feature reconstruction process, local structural features under different receptive fields are extracted sequentially to enhance the crack tip bifurcation mode, micro-tear texture and material fatigue deformation features, and output high-resolution detail features.

5. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, Step S3 includes: Obtain the spatial position masks of the hinge area, latch area and slide rail area in the train head cover, and calculate the average feature response intensity of high-resolution detail features in the hinge area, latch area and slide rail area respectively; The initial region weights are determined based on the average characteristic response intensity of each region. Based on the initial regional weights of each region, the spatial correlation between regions is analyzed, and the spatial correlation is normalized to generate a regional sensitive weight distribution map. The high-resolution detail features are multiplied element-wise with the region-sensitive weight distribution map to output the structure-enhancing features.

6. The method for identifying train headliner front opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, Step S4 includes: The structural enhancement features are input into the candidate region network of the pre-trained Faster R-CNN model to generate four types of candidate region sets; The four candidate region sets are input into the classification subnet of the Faster R-CNN model, the probability distribution of each candidate region in each damage category is calculated, and the damage category data of each candidate region is output. The four candidate region sets are synchronously input into the regression subnet of the Faster R-CNN model to determine the spatial coordinates of each candidate region and output the fault location data of each candidate region. Integrate damage category data and fault location data into identification result information; The damage confidence level of each detected target is calculated based on the category probability in the identification results.

7. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 6, characterized in that, Specifically, the structural enhancement features are input into the candidate region network of the pre-trained Faster R-CNN model as follows: Based on the prior spatial distribution of the locking area and hinge area in the structural reinforcement features, non-uniform initial anchor point boxes with mixed density arrangement are arranged in the candidate region network, and the dense area of ​​the anchor point box corresponds to the area of ​​concentrated structural stress. The first selection subset is generated based on the spatial overlap between the initial anchor point box and the structural reinforcement features; A second screening is performed on the first screening subset based on the regional characteristic response intensity to form a primary candidate region set containing potential fault regions. Constraints on the outline and boundary continuity of the primary candidate regions are applied to the concentrated superposition of structural features. Spatial coordinate and boundary deformation optimization are performed on the constrained primary candidate region set to generate four types of candidate region sets including cracks, notches, misalignments and opening / closing mechanism damage.

8. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 6, characterized in that, The damage confidence level for each detected target is calculated based on the category probability in the identification results, including: Extract the four damage probability vectors corresponding to each candidate region from the classification subnet output in the recognition results information; Calculate the information entropy between each probability direction; The maximum probability value in the probability vector is combined with the information entropy. The confidence level is positively correlated with the maximum probability value and negatively correlated with the information entropy. The result of the combined operation is used as the damage confidence level of the corresponding detection target.

9. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, Step S5, which constructs a texture stability score based on the polarization residual map of the fused visual image, includes: Extract the polarization residual map corresponding to the fused visual image; The polarization residual map is divided into grids, and the standard deviation of pixel values ​​within each image block is calculated as the local texture fluctuation index of that image block. Calculate the average value of the local texture fluctuation index of all image patches as the global texture stability benchmark; Identify the target region in the polarization residual map that corresponds to the fault location data; Calculate the relative deviation between the local texture fluctuation index and the global texture stability benchmark value of the image patch within each target region; A texture stability score is constructed based on the weighted average of the relative deviations of all target regions.

10. The method for identifying train headliner opening / closing damage faults based on Faster R-CNN according to claim 1, characterized in that, In step S5, a comprehensive judgment score is determined using the damage confidence level and texture stability score, and the resulting fault report includes: A weight allocation strategy is constructed based on damage confidence and texture stability scores; In the weighting strategy, when the hinge produces periodic shaking texture and displacement consistency decreases under loose conditions, the texture stability score weight of the hinge loosening fault is greater than the damage confidence score weight; when the latch slot is broken, the latch notch is expanded, or the positioning surface is deformed, the damage confidence score weight of the latch failure fault is greater than the texture stability score weight. Based on a weighted allocation strategy, the damage confidence score and texture stability score corresponding to the same detection target are weighted and fused to generate a corresponding comprehensive judgment score; Different judgment thresholds are set according to different fault modes. When the comprehensive judgment score of any detected target exceeds the judgment threshold, a fault report of the corresponding fault type is generated.