Railway scene image enhancement method based on image feature aggregation model
Through multi-stage processing and feature fusion strategy, the problem of insufficient identification of interest areas in the railway scene image super-segment method in the prior art is solved, and the super-resolution reconstruction quality and model efficiency of railway scene images are improved.
Patent Information
- Application Number
- CN202510990388.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-08-19
- Estimated Expiration
- 2045-07-18
AI Technical Summary
The existing image super-segment method is difficult to effectively distinguish the regions of interest in railway scenes, treat all regions equally, resulting in the expansion of model parameters and failure to fully utilize the characteristics of nearby equipment, which makes the super-segment effect of distant equipment poor.
A multi-stage processing strategy is adopted to extract linear features through the Schal operator, combine morphological corrosion and expansion operators to enhance key linear features, use an asymmetric global attention mechanism to perform feature correlation calculation, and use a rail axial feature fusion mechanism to improve the representation ability of features within the boundary.
The quality of super-resolution reconstruction of railway scene images is improved, the complexity of model training and inference time is reduced, the characteristic sensitivity to railway scenes is enhanced, and the super-segment effect on distant equipment is improved.
Smart Images

Figure CN120510403A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular to a railway scene image enhancement method based on an image feature aggregation model. Background Art
[0002] Images are the primary data source for railway security technology. However, in inclement weather or low-light environments, the image acquisition capabilities of cameras along the railway line are significantly degraded. The resulting low-quality images can cause the security system to lose its ability to detect foreign objects, threatening train safety. Therefore, enhancing low-quality images is currently the primary prerequisite for ensuring the safety level of train operations. Existing image super-resolution methods can be roughly divided into two categories: those based on generative adversarial networks (GANs) and those based on residual and dense learning. Specifically, GANs are prone to generating erroneous textures and suffer from poor real-time performance. While residual and dense learning-based methods avoid the problems of the former, they lack specific improvements for railway scenes. Consequently, they struggle to effectively distinguish regions of interest within railway scenes, treating all regions equally, leading to bloated model parameters. Furthermore, they fail to fully utilize the characteristics of nearby devices, resulting in suboptimal super-resolution for distant devices. The present invention proposes a railway scene image enhancement method based on an image feature aggregation model to address the problem of poor performance caused by the insensitivity of general super-resolution models to railway scene features. Thus, a rail-centered feature enhancement and aggregation super-resolution model is constructed. With the rails - the inherent features of railway scenes - as the center, operations such as boundary feature enhancement, scene content grouping, global feature enhancement, and rail axial feature aggregation are carried out, thereby improving the super-resolution effect of railway scenes. Summary of the Invention
[0003] The purpose of the present invention is to solve the problem in the prior art that it is difficult to effectively distinguish regions of interest in railway scenes and all regions are treated equally, resulting in expansion of model parameters.
[0004] To achieve the above object, the present invention provides a railway scene image enhancement method based on an image feature aggregation model, which is characterized by comprising the following steps: S1. Using a multi-stage processing strategy to enhance the features of the original image, the original image after feature enhancement is equidistantly divided into multiple image blocks of equal size; the divided image blocks are grouped into multiple subgroups based on content clustering rules; S2, taking the subgroup of features within the limit in the grouped subgroup as the core processing object, and performing feature correlation calculation through the asymmetric global attention mechanism; S3. Use the rail axial feature fusion mechanism to improve the ability to characterize the linear structure of the rail within the feature subgroup within the limit.
[0005] As a further improvement of this technical solution, the multi-stage processing strategy working steps are as follows: Step 1: Use the Schall operator to extract line features; Step 2: Use a dual threshold mechanism for noise reduction. Specifically, set a brightness threshold and a region connectivity threshold. Pixels with brightness values less than the brightness threshold and a connected region size greater than the connectivity threshold are considered interference regions that will cause misjudgment when extracting line features. A mask is constructed, and the original line features obtained by the Schall operator are multiplied pixel by pixel with the constructed mask to filter out the interference area features and obtain the denoised image features; Step 3: Enhance the original image features based on the linear feature amplification mechanism of morphological erosion and dilation operators.
[0006] As a further improvement of the present technical solution, the straight line feature is a gradient amplitude calculated by a Schall operator.
[0007] The beneficial effect of the above further scheme is that there are a large number of significant straight line feature elements (such as linear structures such as rails, sleepers, guardrails and columns) in the railway limit area. The effective identification and rational use of these straight line features are key factors in improving the quality of super-resolution reconstruction of railway scene images. Although the existing feature enhancement methods based on deep learning can enhance the feature extraction capabilities of convolutional neural networks in different scenarios to a certain extent, these methods introduce too many additional parameters, significantly increase the training complexity of the model and prolong the inference time. Based on the balance between efficiency and accuracy, this study chooses to use computationally efficient traditional image processing operators to accurately extract straight line features.
[0008] As a further improvement of the present technical solution, the morphological erosion uses a structural element to erode the denoised linear features; the dilation operator uses a structural element to check each pixel position to determine whether to retain the pixel.
[0009] The beneficial effect of the above further scheme is that the morphological erosion and dilation operators are used to further enhance the visual significance of key straight line features, while effectively filtering out the remaining small straight line interference. The morphological erosion and dilation operators, computationally efficient depthwise separable convolution and nonlinear activation functions are alternately connected in a specific sequence. Among them, the morphological erosion and dilation operators can effectively eliminate subtle interference that may be left over by the threshold denoising method; the depthwise separable convolution significantly reduces the number of parameters compared to the standard convolution operation, improving computational efficiency while maintaining the model's expressiveness.
[0010] On the basis of the above technical solution, the present invention can also be improved as follows.
[0011] As a further improvement of the present technical solution, the content clustering rule is as follows: for each divided image block, the number of straight lines is counted, and a number threshold is set, and image blocks with a number of straight lines greater than the number threshold are classified as within-bounded image block subgroups; Calculate the average brightness value of all pixels in each divided image block: add the brightness value of each pixel in the image block, divide it by the total number of pixels in the image block, and get the average brightness value. Perceive the brightness threshold. If the average brightness value is greater than the preset brightness threshold, the image block is classified as a sky part image block subgroup; Image blocks that do not satisfy either the determination condition of the within-bounds image block subgroup or the determination condition of the sky portion image block subgroup are divided into an out-of-bounds image block subgroup.
[0012] As a further improvement to this technical solution, the asymmetric global attention mechanism works as follows: perceiving subgroup features: features of subgroups within the bounds, features of subgroups outside the bounds, flattening the spatial dimensions of each subgroup into a one-dimensional sequence, selecting the one-dimensional sequence features of the subgroups within the bounds as the "query" in the attention calculation; and concatenating the one-dimensional sequence features of all subgroups as the "key" and "value"; The degree of association between the "query" and the "key" is calculated through a dot product operation, and the calculated degree of association is scaled. The result of the degree of association is then normalized and converted into a probability value between 0 and 1, so that the sum of the probabilities of each "query" position associated with the global feature is 1, and the attention weight is obtained. Finally, information useful for the in-bound subgroup is extracted from the "value" according to the attention weight to realize cross-subgroup feature interaction. The one-dimensional sequence features after attention fusion are converted back to the two-dimensional spatial dimension to obtain the updated in-bound subgroup features.
[0013] As a further improvement of this technical solution, the rail axial feature fusion mechanism in S3 includes feature richness statistics and axial attention fusion, which is used to allow clear features in the near distance to assist in enhancing blurred features in the far distance.
[0014] As a further improvement of the present technical solution, the working principle of the feature richness statistics is as follows: for each image block in the bounded image block group, each image block in the bounded image block group is converted from the spatial domain to the frequency domain through a two-dimensional discrete Fourier transform, and the discrete Fourier transform is calculated, and the brightness value of each pixel in the image block is summed up in combination with the exponential function to obtain the frequency domain, presenting the characteristic distribution of the image block in the frequency dimension, and then further calculating the amplitude spectrum and the phase spectrum: by operating the real part and the imaginary part represented in the frequency domain, the specific operation is: squaring them respectively, then summing them, and finally taking the square root, so as to obtain the amplitude spectrum; taking the inverse tangent operation on the ratio of the imaginary part to the real part to obtain the phase spectrum; A comprehensive standard that integrates the two measurement methods is constructed. The square root of the sum of the squares of the amplitude spectra of two image blocks and the differences between the corresponding elements is measured through the Euclidean distance. The similarity measurement is preset and the weighted summation is used to determine the comprehensive measurement standard.
[0015] As a further improvement of the present technical solution, the working principle of the axial attention fusion is: the high-structure information features are linearly transformed respectively to obtain components for attention calculation, and the degree of correlation between the high-structure information features and the low-structure information features is calculated through the attention calculation components, and the high-structure information is most useful for enhancing the low-structure information features. After obtaining the enhanced part of the high-structure information on the low-structure information, it is fused with the original low-structure information features through residual connection. The fusion method is: the enhanced part is directly added to the original low-structure information features, and finally the fused low-structure information features are obtained.
[0016] In addition to the above-described objects, features and advantages, the present invention has other objects, features and advantages. The present invention will be further described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a step diagram of the overall workflow of the present invention. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0019] refer to Figure 1 , a railway scene image enhancement method based on image feature aggregation model, comprising the following steps: S1, using a multi-stage processing strategy to enhance the original image features, and then dividing the original image after feature enhancement into multiple image blocks of equal size at equal intervals, and grouping the image blocks into multiple subgroups based on content clustering rules; The multi-stage processing strategy works as follows: Step 1: Since there are a large number of significant straight line features (such as linear structures such as rails, sleepers, guardrails, and columns) within the bounded area of railway scene images, and the effective identification and rational use of straight line features are key factors in improving the quality of super-resolution reconstruction of railway scene images, existing technologies generally use deep learning feature enhancement methods to a certain extent to enhance the feature extraction capabilities of convolutional neural networks in different scenarios. However, since deep learning feature enhancement methods require the introduction of too many additional parameters, they increase the training complexity of the model and prolong the inference time. Based on the balance between efficiency and accuracy, the present invention uses the Schall operator to extract straight line features: The Schall operator calculates pixel gradients through convolution kernels in both horizontal and vertical directions: the horizontal convolution kernel and vertical convolution kernel Here is an example: , vertical convolution kernel for: , the above horizontal convolution kernel and vertical convolution kernel , which is only an example, and any convolution kernel with a similar structure is essentially applicable to the present invention.
[0020] Gradient calculation: for railway scene images Apply horizontal convolution kernels separately and vertical convolution kernel Perform convolution operation to obtain the gradient amplitude and direction : , ;in, Represents convolution operation, gradient amplitude That is, the initially extracted straight line features ; The Scharnauer operator is a high-pass filter that has been optimized and designed as follows. It can extract straight line features while maintaining rotation invariance, provide higher gradient response sensitivity than traditional edge detection operators such as Sobel, and have higher detection accuracy for target boundary changes.
[0021] Step 2: In order to accurately use the straight line features to determine the exact position of the railway limit, it is necessary to effectively eliminate the interference of the straight line features outside the limit. The most significant interference source comes from the contact wire located in the sky area. Although the contact wire is a key equipment and facility within the railway limit, due to the specific installation angle and shooting angle of the monitoring camera along the railway, the contact wire will still appear in the sky area outside the limit on the original image plane, which will lead to misjudgment in the process of extracting straight line features. Therefore, a dual threshold mechanism is used to reduce noise: a brightness threshold is set to distinguish between high brightness in the sky and low brightness within the limit. , regional connectivity thresholds for distinguishing large catenary areas from small structures within limits , the pixel brightness value is less than the brightness threshold At the same time, the size of the connected area is greater than the connectivity threshold, and the area where the pixel is located is determined to be an interference area that will cause misjudgment when extracting straight line features; and a mask is constructed. , suppress the interference area and retain the features within the limit, so as to ensure that the straight line features of the rails and related track facilities are not interfered with and are accurately preserved. The mask rule is: only retain pixels with "low brightness (within the limit) and large connected areas (non-noise points)", and suppress the interference of the overhead contact network in the sky area. The specific expression is as follows: ; in, is the brightness value of pixel (x,y), is the size of the connected area where the pixel (x, y) is located, is the brightness threshold, is the connectivity threshold.
[0022] Perform precise differential operation on the original image features obtained by Schall operator processing and the image features after adding mask, and use mask to With line features Multiply pixel by pixel to obtain high-quality denoised image features. The expression is: ; Among them, S(x,y) is the image feature after being processed by the Schall operator, is the image feature after denoising (only valid straight lines within the limit are retained).
[0023] After completing the precise extraction and noise suppression of line features, in order to further enhance the visual saliency of key line features while effectively filtering out the remaining small line interference, step three: based on the line feature amplification mechanism of morphological erosion and dilation operators, enhance the extracted boundary boundary features; Morphological corrosion: using structural elements The denoised straight line features Erosion is performed to check each pixel position in the image. When the structural element in the area corresponding to the pixel is completely contained by the denoising feature (that is, the area covered by the structural element belongs to the valid straight line feature area), the pixel is retained; otherwise, it is removed. Any technical means that conforms to the erosion and dilation operations are applicable to this operation step in the present invention. The calculation formulas of the morphological erosion and dilation operators are as follows: Morphological erosion is used to retain only pixels where the structural element is completely contained in the feature area. The formula is: ; The dilation operator is used to retain the pixels that intersect the structural element and the eroded feature area. The formula is: .
[0024] After extracting straight line features, reducing noise, and enhancing the extracted boundary features, the original image is divided into multiple equal-sized image blocks at equal intervals. Based on content clustering rules, the image blocks are intelligently grouped into multiple subgroups with clear semantics. The following example divides the original image into nine equal-sized image blocks at equal intervals and intelligently groups the image blocks into three subgroups with clear semantics. The three subgroups are: an in-bounds image block subgroup (including railway facilities within the boundary), a sky image block subgroup (mainly background areas), and an out-of-bounds image block subgroup (including vegetation in non-critical areas). Content clustering rule: For each divided image block, count the number of straight lines , and set a quantity threshold, and classify the image blocks whose number of straight lines is greater than the quantity threshold as the image block subgroup within the limit; And calculate the average brightness value of all pixels in each divided image block. First add the brightness value of each pixel in the image block, and then divide it by the total number of pixels in the image block to get the average brightness value: ;in is the size of the image block after division.
[0025] Perceptual brightness threshold , the average pixel brightness value Greater than the brightness threshold The image blocks are classified as the sky part image block subgroup.
[0026] S2. Perform fine processing and feature enhancement on image data, taking the bounded feature subgroup as the core processing object and realizing cross-region feature correlation calculation through the asymmetric global attention mechanism; Based on content clustering rules, image blocks are intelligently grouped into multiple subgroups with clear semantics. While this allows for differentiated feature extraction and enhancement operations for different semantic regions, it also leads to potential field of view limitations. That is, when processing a specific subgroup, there may be a lack of global awareness of information in other regions. Therefore, an asymmetric global attention mechanism centered on the railway boundary region is used to address this issue. The asymmetric global attention mechanism is used to break the information isolation between subgroups and build an efficient cross-regional feature interaction channel, so that each subgroup can fully perceive and utilize global context information while maintaining the local feature advantage. The following is an example of the asymmetric global attention mechanism: First, from the railway scene image, perceive the subgroup features: subgroup features within the limit (including key linear structures such as rails and trackside equipment), sky subgroup features (background area, including the sky and distant scenery under the sky), subgroup characteristics outside the limit (such as vegetation around the track, non-critical buildings); To facilitate attention calculation, the spatial dimensions (length × width) of each subgroup are flattened into a one-dimensional sequence, and the obtained 、 、 , converting two-dimensional spatial features into one-dimensional sequence features to facilitate subsequent calculation of associations.
[0027] To highlight the priority of the area within the limit, only the flattened features of the subgroup within the limit are selected. As Query (denoted as Because the boundary is the key area for railway scene super-resolution, it is allowed to dominate the global information query to ensure the priority of feature interaction in the core area; To allow the subgroup within the limit to obtain global information, the flattened features of all subgroups (within the limit, sky, and outside the limit) are concatenated as the Key (denoted as K) and Value (denoted as V); the Key is used to calculate the degree of association with the Query, and the Value is used to carry the global feature information to be fused. In this way, the subgroup within the limit can find useful information from the entire scene.
[0028] Calculate the similarity between Query (Q) and Key (K), and use the dot product operation to measure the degree of correlation between the corresponding elements of the two. The formula is This operation obtains a matrix, in which each element reflects the strength of the correlation between the feature of a certain position of the subgroup within the limit and the feature of each global subgroup, thus building a bridge between the features within the limit and the global features. Because the feature dimension may affect the stability of the calculation, the result is divided by ( Scaling (number of feature channels) can alleviate the problem of gradient disappearance or abnormality and make the attention distribution more reasonable.
[0029] For the above similarity matrix, Function normalization, the formula is ; The similarity values are converted into probabilities between 0 and 1 so that the sum of each row is 1, and the attention weight matrix is obtained. In the attention weight matrix, the features of different positions in the subgroup within the limit will be assigned different weights according to the degree of correlation with the global features, realizing "selective attention" - prioritizing the focus on the features with strong semantic correlation within the limit (such as the correlation between the rails and the guardrail structure) and suppressing irrelevant background noise (such as meaningless textures in the sky). The attention weight matrix is combined with Multiply, that is , according to the weight from the global features Useful information for the subgroups within the limit is extracted to achieve cross-subgroup feature interaction. The subgroups within the limit can be fused with the brightness information of the sky subgroup to assist in distinguishing track boundaries, or fused with the environmental texture information of the subgroups outside the limit to make the super-resolution scene more realistic.
[0030] The one-dimensional sequence features after attention fusion are converted back to the two-dimensional space dimension through the Reshape operation to obtain the updated bounded subgroup features The updated features not only retain the advantages of the original local features within the limit (such as the inherent structure of the rail), but also incorporate global context information (such as the background environment association), providing better basic features for subsequent rail axial feature fusion and super-resolution reconstruction. While ensuring the priority of super-resolution effect in the area within the limit, the overall image super-resolution quality is improved; The asymmetric global attention mechanism enables the in-bounds subgroups to selectively interact with other subgroups, rather than implementing computationally intensive fully connected attention calculations. This has dual advantages: on the one hand, it significantly reduces computational complexity and resource consumption; on the other hand, it ensures that the super-resolution reconstruction effect of the in-bounds area is prioritized.
[0031] In railway scene images, rails usually constitute the main visual elements of the image and show obvious spatial distribution regularity: the rail structure close to the camera side is larger in proportion and the geometric features are clearer and more obvious; as the rails extend to the distance, their features gradually become blurred. The phenomenon of the above-mentioned features gradually attenuating along the axial direction of the track also occurs on other linear railway infrastructure such as guardrails and ballast. Based on this observation, in the present invention, S3, the rail axial feature fusion mechanism is adopted. The rail axial feature fusion mechanism includes feature richness statistics and axial attention fusion, which is used to learn high-quality structural features of nearby rails and apply them to enhance the feature representation of distant rails.
[0032] The feature richness statistic works as follows: Each image block in the bounded image block group is converted from the spatial domain to the frequency domain through a two-dimensional discrete Fourier transform to reveal the distribution characteristics of its frequency components. The discrete Fourier transform is calculated as follows: ,in is the frequency domain representation, is the pixel value of the pixel block, M and N are both equal to 3; Based on discrete Fourier transform, the amplitude spectrum and phase spectrum are further calculated: The amplitude spectrum is obtained by performing operations on the real and imaginary parts of the frequency domain representation: squaring them separately, summing them, and finally taking the square root: ; Take the inverse tangent of the ratio of the imaginary part to the real part to obtain the phase spectrum: ; To comprehensively evaluate the similarity between image blocks, a comprehensive metric based on the Euclidean distance metric and the pre-selected similarity metric is constructed: the square root of the sum of the squares of the difference between the amplitude spectra and the corresponding elements of the two image blocks is measured by the Euclidean distance metric to reflect the numerical difference in the amplitude spectra; a preset similarity metric (based on the ratio of the inner product of the amplitude spectra to the modulus length) is used to measure the similarity between the two image blocks in terms of feature distribution; and a weighted summation is then used to determine the comprehensive metric to judge the similarity between image blocks and further assist in evaluating feature richness. The calculation formula is: ; in, and is the amplitude spectrum of the two image blocks, and is the magnitude spectrum of the two image patches.
[0033] Based on the aforementioned rail axial feature fusion mechanism, the information richness of each image block within the bounds was evaluated. The study also identified the image blocks with the lowest feature richness ranking in the bottom 25% as "low-information blocks." This effectively facilitated the efficient flow of features from high-information blocks to low-information blocks, ensuring the efficient transfer of key features while avoiding unnecessary consumption of computing resources.
[0034] In the processing of axial features of railroad tracks in railway scene images, axial attention fusion is used to ensure that low-structural information features (which can be understood as blurry and less detailed rails and surrounding features in the distance) and high-structural information features (i.e., clear and detailed rails and surrounding features in the near distance) can interact effectively, thereby enhancing blurry features with clear features. Axial attention fusion linearly transforms the high-structural information features to obtain components for attention calculation. The attention calculation components are then used to calculate the correlation between the high-structural information features and the low-structural information features. The high-structural information is then selected to be most useful for enhancing the low-structural information features, thus achieving intelligent feature screening and transfer. After obtaining the enhanced part of the high-structure information on the low-structure information, it is fused with the original low-structure information features through residual connection. The fusion method is: the enhanced part is directly added to the original low-structure information features, and finally the fused low-structure information features are obtained. The corresponding mathematical expression is: ; in, is a feature with low structural information. Represents the i-th high-structure information feature tensor, and MSA is the attention calculation component.
[0035] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and improvements may be made to the present invention without departing from the spirit and scope of the present invention. Such changes and improvements fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A railway scene image enhancement method based on an image feature aggregation model, characterized in that: The following steps are involved: S1. A multi-stage processing strategy is used to enhance the features of the original image. After feature enhancement, the original image is equidistantly divided into multiple image blocks of equal size. Grouping the divided image blocks into multiple subgroups based on content clustering rules; S2, taking the subgroup of features within the limit in the grouped subgroup as the core processing object, and performing feature correlation calculation through the asymmetric global attention mechanism; S3. Use the rail axial feature fusion mechanism to improve the ability to characterize the linear structure of the rail within the feature subgroup within the limit.
2. The railway scene image enhancement method based on the image feature aggregation model according to claim 1, characterized in that: The multi-stage processing strategy works as follows: Step 1: Use the Schall operator to extract line features; Step 2: Use a dual threshold mechanism for noise reduction. Specifically, set a brightness threshold and a region connectivity threshold. Pixels with brightness values less than the brightness threshold and a connected region size greater than the connectivity threshold are considered interference regions that will cause misjudgment when extracting line features. A mask is constructed, and the original line features obtained by the Schall operator are multiplied pixel by pixel with the constructed mask to filter out the interference area features and obtain the denoised image features; Step 3: Enhance the original image features based on the linear feature amplification mechanism of morphological erosion and dilation operators.
3. The railway scene image enhancement method based on the image feature aggregation model according to claim 2, characterized in that: The straight line feature is the gradient amplitude calculated by the Schall operator.
4. The railway scene image enhancement method based on the image feature aggregation model according to claim 2, characterized in that: The morphological corrosion uses structural elements to corrode the denoised linear features; The dilation operator uses a structuring element to check each pixel position to determine whether to keep the pixel.
5. The railway scene image enhancement method based on the image feature aggregation model according to claim 1, characterized in that: The content clustering rule is: setting a quantity threshold, and classifying image blocks with a number of straight lines greater than the quantity threshold as an within-bounds image block subgroup.
6. The railway scene image enhancement method based on the image feature aggregation model according to claim 5, characterized in that: The working principle of the asymmetric global attention mechanism is as follows: perceiving subgroup features: features of subgroups within the bounds, features of subgroups in the sky, and features of subgroups outside the bounds, flattening the spatial dimensions of each subgroup into a one-dimensional sequence, selecting the one-dimensional sequence features of the subgroups within the bounds as the "query" in the attention calculation; and concatenating the one-dimensional sequence features of all subgroups as the "key" and "value"; The degree of association between the "query" and the "key" is calculated through a dot product operation, and the calculated degree of association is scaled. The resulting degree of association is then normalized and converted into a probability value between 0 and 1, so that the sum of the probabilities of each "query" position associated with the global feature is 1, and the attention weight is obtained. Finally, information useful for the in-bound subgroups is extracted from the "value" according to the attention weight, realizing cross-subgroup feature interaction. The one-dimensional sequence features after attention fusion are converted back to the two-dimensional spatial dimension to obtain the updated in-bound subgroup features.
7. The railway scene image enhancement method based on the image feature aggregation model according to claim 6, characterized in that: The rail axial feature fusion mechanism in S3 includes feature richness statistics and axial attention fusion, which is used to allow clear features in the near distance to assist in enhancing blurred features in the far distance.
8. The railway scene image enhancement method based on the image feature aggregation model according to claim 7, characterized in that: The working principle of the feature richness statistics is as follows: for each image block in the bounded image block group, each image block in the bounded image block group is converted from the spatial domain to the frequency domain through a two-dimensional discrete Fourier transform. The discrete Fourier transform is calculated, and the RGB value of each pixel in the image block is summed up with an exponential function to obtain the frequency domain, presenting the feature distribution of the image block in the frequency dimension, and then further calculating the amplitude spectrum and phase spectrum; and constructing a comprehensive standard that integrates the two measurement methods. The square root of the sum of the squares of the difference between the amplitude spectra and the corresponding elements of the two image blocks is measured through the Euclidean distance, and a similarity measure is preset. The weighted summation is used to determine the comprehensive measurement standard.
9. The railway scene image enhancement method based on the image feature aggregation model according to claim 7, characterized in that: The working principle of the axial attention fusion is: linearly transform the high-structure information features respectively to obtain components for attention calculation, and calculate the degree of correlation between the high-structure information features and the low-structure information features through the attention calculation components, and call out the high-structure information that is most useful for enhancing the low-structure information features. After obtaining the enhanced part of the high-structure information on the low-structure information, it is fused with the original low-structure information features through residual connection. The fusion method is: directly add the enhanced part to the original low-structure information features, and finally obtain the fused low-structure information features.
Citation Information
Patent Citations
Multi-information cascade clustering power transmission line detection method
CN112115985A
Grouping non-local attention method based on deep learning
CN116011541A
Railway scene segmentation method based on feature alignment and edge constraint
CN117830633A
Power transmission line infrared target detection method and system based on multistage feature enhancement fusion
CN117911821A
Railway scene understanding method and system based on panoramic segmentation and relation detection
CN118711145A