An image reconstruction method for bridge steel bar corrosion identification
By using multimodal data fusion and physical-guided generative adversarial networks, the problems of insufficient modal fusion and lack of physical constraints in bridge steel reinforcement corrosion detection are solved, and bridge steel reinforcement corrosion image reconstruction with visual realism and physical consistency of corrosion state is achieved.
Patent Information
- Application Number
- CN202511339893.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-09-19
AI Technical Summary
Existing methods for detecting steel corrosion in bridges suffer from low efficiency in single-point detection, reliance on experience in detection results and inability to reflect changes in the internal physical field. Multimodal fusion methods lack spatial alignment, and deep learning-based image reconstruction methods do not incorporate the physical mechanism of corrosion, resulting in reconstruction results that violate actual physical laws.
Multimodal data synchronous acquisition is employed, including temperature field, ultrasonic echo, electromagnetic impedance, etc., combined with physical parameters and environmental data, and pseudo-color images with unified visual realism and physical consistency of the rust state are generated through affine transformation, bilinear interpolation and physical-guided generative adversarial network.
It achieves consistency between the corrosion distribution and the micro-strain field of the structure, improves the reconstruction accuracy, provides intuitive three-dimensional spatial information, solves the problems of insufficient modal fusion and lack of physical constraints, and ensures that the corrosion distribution of the generated image is consistent with the structural strain.
Smart Images

Figure CN120852408B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, in particular to an image reconstruction method for bridge steel bar corrosion identification. BACKGROUND
[0002] In bridge engineering, steel bar corrosion is one of the main factors leading to the decline of structural durability. Traditional detection methods such as half-cell potential method and rebound method have problems such as low single-point detection efficiency and detection results dependent on experience.
[0003] Although image-based detection technology can improve single-point detection efficiency and enhance the objectivity of detection results through visualization, a single modality (such as optical image) is easily affected by environmental light and concrete surface shielding, and cannot directly reflect internal physical field changes (such as micro-strain and eddy current field distribution) caused by steel bar corrosion. Existing multi-modal fusion methods usually use simple data splicing or shallow feature fusion, lacking effective processing of spatial alignment of multi-modal data (such as infrared data and ultrasonic data) and apparent feature data (such as crack orientation and corrosion morphology). Although image reconstruction methods based on deep learning (such as GAN) can generate visually realistic images, they do not incorporate the physical mechanism of steel bar corrosion (such as micro-strain field caused by corrosion expansion), resulting in reconstruction results that may violate actual physical laws and cause abnormal separation of corrosion distribution and structural strain. SUMMARY
[0004] To solve the above technical problems, the present application realizes the following technical scheme:
[0005] An image reconstruction method for bridge steel bar corrosion identification is proposed, comprising the following steps: synchronously collecting multi-modal data, physical parameters and environmental data of the target area; the multi-modal data includes temperature field data, ultrasonic echo data and electromagnetic impedance data; the physical parameters include vibration signal, half-cell potential and micro-strain field data; the environmental data includes environmental temperature, relative humidity and chloride ion concentration on the concrete surface; feature information of each single modality data is extracted to obtain a multi-modal feature map; affine transformation parameters are obtained according to the multi-modal feature map; target grid coordinates are generated according to the affine transformation parameters; the multi-modal feature map is subjected to bilinear interpolation according to the target grid coordinates to obtain an aligned multi-modal feature map; the aligned multi-modal feature map is subjected to channel splicing to obtain a joint feature map; a generator of a physically guided generative adversarial network is used to process the joint feature map to generate a pseudo-color image; a double-path discriminator is used to discriminate the pseudo-color image; when the discrimination result is greater than a threshold value, the generator is optimized by minimizing the loss function, and the final reconstruction image is output.
[0006] Compared with the prior art, the present application has the following advantages and beneficial effects: the present application fuses infrared thermal images, ultrasonic echoes, electromagnetic eddy currents and other multi-modal data, combines a bridge steel bar corrosion image reconstruction method based on physical field constraints, and realizes the unification of visual authenticity and physical consistency of the corrosion state through cascading feature fusion and physical guidance generated adversarial networks, thereby solving the problems of insufficient modal fusion and lack of physical constraints in the prior art. Specifically, by introducing corrosion expansion physical field simulation, the corrosion distribution of the generated image is ensured to be consistent with the structural micro-strain field, thereby avoiding the "mode collapse" problem of traditional GAN. In addition, through joint training of physical model simulation data and real data, the reconstruction accuracy in the noise disturbance scene is significantly improved compared with single data training, and the pseudo-color image simultaneously represents the corrosion degree, depth and structural contour, thereby providing intuitive three-dimensional spatial information for bridge maintenance. BRIEF DESCRIPTION OF DRAWINGS
[0007] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments, and it should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as a limitation on the scope, and for those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.
[0008] Figure 1 A flowchart of an image reconstruction method for bridge steel bar corrosion identification provided by Embodiment 1 of the present application. DETAILED DESCRIPTION
[0009] In order to make the purpose, technical solutions and advantages of the present application more clear and apparent, the following will further describe the present application in combination with embodiments and drawings, and the illustrative embodiments of the present application and their descriptions are only used to explain the present application, and should not be regarded as a limitation on the present application.
[0010] In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present application. However, it is apparent to those skilled in the art that the present application can be implemented without necessarily adopting these specific details. In other embodiments, in order to avoid obscuring the present application, well-known structures, circuits, materials or methods are not specifically described.
[0011] Throughout the specification, reference has been made to "one embodiment," "an embodiment," "one example," or "an example" meaning that a particular feature, structure, or characteristic described in connection with the embodiment or example is included in at least one embodiment of the application. Therefore, appearances of the phrases "one embodiment," "an embodiment," "one example," or "an example" in various places throughout the specification are not necessarily referring to the same embodiment or example. Furthermore, the particular features, structures, or characteristics can be combined in any suitable manner in one or more embodiments or examples. Additionally, the skilled artisan will understand that the drawings provided herein are for purposes of illustration only and that the drawings are not necessarily to scale.
[0012] In the description of the application, the terms "front", "back", "left", "right", "up", "down", "vertical", "horizontal", "high", "low", "inner", "outer", and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, only for the convenience of describing the application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the scope of protection of the application.
[0013] Embodiment 1: Provide an image reconstruction method for bridge reinforcement corrosion identification, comprising the following steps:
[0014] Step 1: Synchronously collect multi-modal data, physical parameters and environmental data of the target area.
[0015] The purpose of this step is to provide multi-dimensional information for subsequent corrosion identification, and to solve the problem that single modal data is greatly affected by environmental interference and cannot reflect internal physical field changes.
[0016] 1. Multi-modal data includes:
[0017] (1) Temperature field data. During the corrosion process of bridge reinforcement, local temperature field changes will be caused due to electrochemical reaction and corrosion expansion. By collecting temperature field data on the surface of the target area concrete, the thermal effect of bridge reinforcement corrosion can be represented, and the corrosion position and corrosion range can be identified.
[0018] (2) Ultrasonic echo data. Bridge steel corrosion will cause cracks, holes or corrosion expansion products inside the concrete, and ultrasonic waves will be attenuated, scattered or reflected due to the inhomogeneity of the medium (such as the change of medium density or the change of medium elastic modulus) during propagation. By collecting the changes of ultrasonic propagation characteristics inside the concrete, internal structural damage caused by bridge steel corrosion can be identified, such as ultrasonic attenuation coefficient, crack orientation distribution and corrosion product filling state. Among them, the ultrasonic attenuation coefficient is the performance of the intensified sound wave energy attenuation caused by the decrease of concrete density in the corrosion area, which can directly reflect the corrosion degree of the bridge steel; the crack orientation distribution is the result obtained by detecting the reflection or diffraction of ultrasonic waves when encountering cracks; the corrosion product filling state is the distribution state of the corrosion product at the steel and concrete interface (such as uniform corrosion or pitting corrosion) indirectly reflected by the time delay and waveform distortion of ultrasonic echo.
[0019] (3) Electromagnetic impedance data. Bridge steel corrosion will cause its cross-sectional area to decrease and the surface to oxidize (forming low-conductivity corrosion products), resulting in increased eddy current path resistance and inductance changes, and thus abnormal electromagnetic impedance (amplitude and phase). By collecting electromagnetic impedance data, the relative impedance change and eddy current phase difference can be directly obtained. Among them, the phase impedance change can be obtained by deducting the background impedance of the non-corrosion area, which can be used to highlight the impedance anomaly of the corrosion area and quantify the corrosion degree; the eddy current phase difference is caused by the difference in dielectric properties of the corrosion product and the steel, which can be used to reflect the corrosion type (such as uniform corrosion or pitting corrosion).
[0020] 2. Physical parameters include:
[0021] (1) Vibration signal. Steel corrosion will cause the cross-sectional area to decrease and the adhesion to the concrete to decrease, and thus the stiffness of the bridge structure to decrease. By collecting vibration signals, the dynamic characteristics of the bridge structure (such as natural frequency, damping ratio, etc.) can be further analyzed, thereby indirectly reflecting the influence of steel corrosion on the stiffness and integrity of the structure.
[0022] (2) Half-cell potential. Steel corrosion will form a galvanic cell reaction, and the half-cell potential reflects the potential difference between the anode and cathode regions of the steel. By cooperating with a reference electrode (such as a copper / copper sulfate electrode), the surface potential of the concrete is measured to qualitatively judge the probability of steel corrosion: the more negative the potential, the higher the probability of corrosion (such as less than -350mV, the probability of corrosion is more than 90%), which can quickly locate the high-risk area of corrosion and assist in the evaluation and monitoring of structural durability and safety.
[0023] (3) Micro-strain field data; the environmental data includes: ambient temperature, relative humidity and chloride ion concentration on the concrete surface. Steel corrosion will produce volume expansion (the volume of corrosion products is about 2-4 times that of iron), causing micro-strain (mainly tensile strain) inside the concrete, and even causing cracks. The micro-strain data field can directly reflect the physical process of bridge steel corrosion, including corrosion expansion effect and structural damage transmission. Among them, the corrosion expansion effect can reflect the strain distribution corresponding to different corrosion forms; the structural damage transmission can indicate the risk of concrete cracking caused by corrosion, which is physically consistent with the crack direction field of ultrasonic echo detection and the infrared thermal spot position.
[0024] 3. Environmental data includes:
[0025] (1) Ambient temperature. Ambient temperature directly affects the electrochemical reaction rate of steel corrosion. Temperature rise will accelerate ion diffusion (such as Cl ion migration) and chemical reaction kinetics, resulting in increased corrosion current density.
[0026] (2) Relative humidity. The essence of steel corrosion is electrochemical corrosion, which requires water as the carrier of the electrolyte solution. Relative humidity directly affects the water content in the concrete pore - on the one hand, when the relative humidity exceeds the critical humidity (usually about 60%-70%), the concrete pore water forms a continuous electrolyte, accelerating the migration of corrosion media such as Cl ions and oxygen, and promoting the corrosion reaction to occur. On the other hand, humidity indirectly changes the corrosion current density by affecting ion conductivity. For example, in a high humidity environment, the migration rate of Cl ions increases, the corrosion current density increases, and the corrosion expansion coefficient increases.
[0027] (3) Chloride ion concentration. Chloride ion is one of the main chemical media that induces steel corrosion, and its role includes: destroying the passivation film: after Cl ions penetrate the concrete protective layer, they will be adsorbed on the surface of the steel, destroy the passivation film, form the anode area of the corrosion cell, and cause local pitting corrosion; accelerate electrochemical reaction: the high migration rate and conductivity of Cl ions enhance the conductivity of the electrolyte solution, accelerate the flow of corrosion current, and increase the corrosion rate.
[0028] Step 2: Preprocess the multi-modal data.
[0029] including the following steps:
[0030] Step 2.1: Non-uniformity correction and bilateral filtering of temperature field data.
[0031] Non-uniformity noise in temperature field data is removed by black body correction and bilateral filtering. The formula for black body correction is: wherein I corr is the infrared image data after black body correction; Iraw Raw temperature field data, i.e. the infrared image data directly obtained from the infrared sensor without any correction; I dark Dark field data, i.e. the output data of the infrared sensor without external light source; K Correction coefficient, used to adjust the gain of the raw data, which can be calculated according to the known blackbody temperature and the characteristics of the sensor; B Correction offset, used to adjust the offset of the raw data, which can be calculated according to the known blackbody temperature and the characteristics of the sensor.
[0032] Further, the formula of bilateral filtering is: Wherein, I out ( p ) is the output value of the bilateral filtering; I ( p ) is the pixel value of the input image at pixel p ; I ( q ) is the pixel value of the input image at pixel q ; Ω is the neighborhood centered at pixel p ; is the spatial weight, usually a Gaussian function, used to measure the spatial distance p between pixel q and pixel ; is the standard deviation of the spatial weight; is the pixel value weight, usually a Gaussian function, used to measure the pixel value difference between pixel p and pixel q ; is the standard deviation of the pixel value weight; W p is the normalization factor, used to ensure that the sum of the weights is 1, .
[0033] Step 2.2: Band-pass filtering and time gain compensation of the ultrasonic echo data.
[0034] Use FIR filter (passband 5-10MHz) to remove high-frequency noise in the ultrasonic echo data; use low-noise, high-precision variable gain amplifier VGA such as AD604 to perform time gain compensation on the ultrasonic echo data.
[0035] Step 2.3: Normalization and background noise subtraction of the electromagnetic impedance data.
[0036] The electromagnetic impedance data can be scaled to a distribution with a mean of 0 and a standard deviation of 1 by mean-standard deviation normalization, and the formula of mean-standard deviation normalization is: wherein, x norm is the normalized electromagnetic impedance data, x is the original electromagnetic impedance data, is the mean of the electromagnetic impedance data, is the standard deviation of the electromagnetic impedance data. Further, the background noise can be deducted from the electromagnetic impedance data by the following formula: x clean =x raw -x noise wherein, x clean is the electromagnetic impedance data after deducting the background noise, x raw is the original electromagnetic impedance data, x noise is the background noise.
[0037] Step 2.4: Fourier transform is performed on the vibration signal to obtain the frequency spectrum information of the vibration signal, and the first three order inherent frequencies and damping ratios of the frequency spectrum information are normalized to obtain the vibration feature vector. It should be noted that Fourier transform is used to convert the vibration signal from time domain to frequency domain, so as to obtain the frequency spectrum information of the signal. The frequency spectrum information can reveal the amplitude and phase of different frequency components in the signal. The first three order inherent frequencies and damping ratios of the frequency spectrum information are extracted, so as to deeply understand the dynamic characteristics of the vibration system, realize the bridge steel bar corrosion diagnosis and state monitoring.
[0038] Step 3: Extracting the feature information of each single modal data to obtain a multi-modal feature map.
[0039] The method is as follows: the following steps are performed on each single modal data respectively:
[0040] Step 3.1: a first convolutional layer of the multi-layer convolutional neural network is used to extract a plurality of local feature information from the single modal data; the convolution kernel size of the first convolutional layer is 5x5, the step length of the first convolutional layer is 2, the output channel number of the first convolutional layer is 64, and the activation function of the first convolutional layer is LeakyReLU function;
[0041] Step 3.2: a second convolutional layer of the multi-layer convolutional neural network is used to extract the associated feature information between the plurality of local feature information; the convolution kernel size of the second convolutional layer is 3x3, the expansion rate or the hollow rate of the second convolutional layer is 2, the output channel number of the second convolutional layer is 128, and the activation function of the second convolutional layer is ReLU function;
[0042] Step 3.3: The associated feature information and the plurality of local feature information are down-sampled by using a max-pooling layer of a multi-layer convolutional neural network, and the feature information obtained by the down-sampling is converted into semantic feature information, to obtain the multi-modal feature map; the pooling kernel size of the max-pooling layer is 2x2, the step of the max-pooling layer is 2, and the output channel number of the max-pooling layer is 256.
[0043] It should be noted that the embodiment extracts feature information of each single modality data by using a multi-layer convolutional neural network. The structure of the multi-layer convolutional neural network is as follows:
[0044] 1. First layer
[0045] It is a convolutional layer, the convolution kernel size of which is 5x5, the step of which is 2, the output channel number of which is 64, and the activation function of which is LeakyReLU function. The function of the first layer is as follows: (1) primary feature extraction - capture local acoustic structures (such as acoustic reflection interface, echo waveform of small-scale crack) in the ultrasonic echo data by using a larger convolution kernel (5x5), reduce the spatial resolution of the feature map (down-sampling), reduce the amount of calculation while retaining the main acoustic features. (2) Nonlinear activation - LeakyReLU avoids the problem of "death" of ReLU neurons in the negative interval, is suitable for processing negative amplitude (such as reflection phase inversion) that may exist in the ultrasonic signal, and enhances the sensitivity of the model to the polarity change of the acoustic wave. (3) Multi-channel feature mapping - output 64-channel feature map, each channel corresponds to different acoustic feature dimensions (such as different frequency components, reflection intensity distribution), and provide rich bottom feature combinations for the subsequent layers.
[0046] 2. Second layer
[0047] For the empty convolution layer, the size of the convolution kernel is 3x3, the empty rate is 2, the number of output channels is 128, and the activation function is the ReLU function. The role of the second layer is: (1) expanding the receptive field - through the empty convolution with the expansion rate of 2, the 3x3 convolution kernel is equivalent to the receptive field range of the 7x7 traditional convolution kernel without increasing the number of parameters, which can capture larger scale acoustic features (such as crack group distribution, overall morphology of corrosion products); (2) preserving spatial resolution - since the step is 1, the feature map size is the same as the output of the first layer (such as input of 128x128, output still of 128x128), avoiding the loss of details caused by excessive down-sampling, suitable for the needs of spatial positioning of internal defects of ultrasonic data; (3) increasing feature dimension - the number of output channels is increased to 128, through more convolution kernels to learn complex acoustic feature combinations (such as sound wave attenuation gradient, phase difference distribution), forming middle-level semantic features (such as "sound wave scattering pattern of rusted area"); (4) ReLU activation - enhancing feature nonlinear expression, suppressing noise while highlighting effective acoustic signals (such as high activation value corresponding to strong reflection echo).
[0048] 3, third layer
[0049] It is a max-pooling layer, the size of the pooling kernel is 2x2, the step is 2, and the number of output channels is 256. The role of the third layer is: (1) spatial down-sampling - through 2x2 max-pooling (step 2) to reduce the feature map size by half (such as from 128x128 to 64x64), further reducing the computational complexity, and through the "maximum value retention" operation to extract the most significant acoustic features (such as the position of the strongest echo, the main crack direction); (2) expanding feature channels - the number of output channels is increased to 256, realized through 1x1 convolution after pooling, the purpose is to further abstract the middle-level features into high-level semantic features (such as "corrosion type features" and "defect severity features"); (3) enhancing translation invariance - the pooling operation makes the features more robust to target position changes, suitable for the scenario where the position of the defect may exist a small offset in ultrasonic detection.
[0050] Step 4: obtaining the affine transformation parameters according to the multi-modal feature map.
[0051] Specifically, receiving the multi-modal feature map, calculating a set of affine transformation parameters through the fully connected layer, including: translation (tx, ty) t x 、 t y ) and rotation angle (a) ). Among them, the translation amount is used to compensate for the horizontal and vertical offset of the sensor in the detection area (such as the installation position deviation of the infrared thermal imager and the ultrasonic probe); the rotation angle is used to correct the viewing angle difference of different modal data (such as the angle between the scanning direction of the ultrasonic probe and the shooting direction of the infrared thermal imager). For example, after the input feature map (such as temperature field data 64x64x64, ultrasonic echo data 64x64x256, electromagnetic impedance data 64x64x64, vibration feature vector 1x1x6, etc.) is compressed by global average pooling or convolution, it is passed through 2 layers of fully connected layers (the number of hidden layer neurons can be set to 256, 64, etc.), and finally 3 parameters are output (translation amount, rotation angle, and scaling factor). t x 、 t y 、 ), forming an affine transformation matrix: .
[0052] Step 5: generating target grid coordinates according to the affine transformation parameters.
[0053] The formula is: . Among them, x and y are the horizontal and vertical coordinates of the point in the original feature map, x’ and y’ are the horizontal and vertical coordinates of the corresponding point in the transformed image, is the rotation angle, t x and t y are affine transformation parameters.
[0054] Step 6: performing bilinear interpolation on the multi-modal feature map according to the target grid coordinates to obtain an aligned multi-modal feature map.
[0055] The original feature map is resampled by bilinear interpolation to achieve pixel-level alignment, avoid pixel distortion caused by traditional interpolation methods, and ensure that the aligned feature map retains the details of the original modal. The size of the aligned feature map is unified to 64x64x256 (channel splicing: 64+256+64+6→380, compressed to 256 channels by 1x1 convolution).
[0056] Step 7: performing channel splicing on the aligned multi-modal feature map to obtain a joint feature map.
[0057] Specifically, the aligned multi-modal feature maps are channel spliced, such as the channel splicing of infrared, ultrasonic, eddy current and vibration feature maps (64+256+64+6=380), to form a high-dimensional feature vector containing multi-modal information. Then, the channel number is compressed to 256 channels through 1x1 convolution, which reduces the amount of calculation while retaining the cross-modal correlation features, providing standardized multi-modal fusion feature input for subsequent network processing.
[0058] Step 8: The generator of the physically guided generative adversarial network processes the joint feature map to generate a false color image.
[0059] The method comprises the following steps:
[0060] Step 8.1: The joint feature map is spliced with physical parameters and environmental data.
[0061] Specifically, the joint feature map (64x64x256) is spliced with the semi-cell potential, micro-strain field data, vibration vector (1x1x6), environmental temperature, relative humidity and chloride ion concentration to form a composite input feature (64x64x256+1x1x(1+1+2+6)). It should be noted that before splicing, the semi-cell potential, micro-strain field data, vibration vector (1x1x6), environmental temperature, relative humidity and chloride ion concentration need to be expanded to spatial dimension feature maps through the broadcast mechanism.
[0062] Step 8.2: The spliced feature data is analyzed using a finite element model to obtain corrosion morphology-related parameters.
[0063] Specifically, the physical simulation sub-network takes the input composite feature as the basis to calculate the physical parameters related to corrosion through a simplified finite element model (FEM), including:
[0064] (1) Concrete elastic modulus
[0065] The concrete elastic modulus is the initial parameter of the model, with a default value of 30 GPa, which can be adjusted according to the actual grade of the concrete in the detection area. This parameter directly affects the calculation results of the subsequent strain field.
[0066] (2) Corrosion expansion coefficient
[0067] The corrosion expansion coefficient can be calculated by the formula: wherein, is the corrosion expansion coefficient; is the constant term; is the cross-section loss rate; C cl is the chloride ion concentration, and the expansion effect is enhanced by a coefficient of 0.1, i.e., the higher the chloride ion concentration, the greater the expansion caused by steel corrosion, and the more obvious the impact on the concrete structure.
[0068] (3) Corrosion current density
[0069] The formula for calculating the corrosion current density is: i corr =i 0 exp ( E half +20nV ), wherein, i corr is the corrosion current density, E half is the half-cell potential, i 0 is the exchange current density, n is the number of electron transfers for the electrode reaction.
[0070] Step 8.3: Classify the corrosion morphology of the bridge reinforcement according to the corrosion morphology correlation parameters.
[0071] The corrosion morphology includes uniform corrosion and pitting corrosion. The corrosion morphology is determined by adding a fully connected layer (Softmax activation) after the channel splicing module, which distinguishes between uniform corrosion and pitting corrosion.
[0072] Step 8.4: Generate a physical field map corresponding to the classification result.
[0073] The physical field map includes a distributed physical field map corresponding to uniform corrosion and a local physical field map corresponding to pitting corrosion. Specifically, according to the classification result of the corrosion morphology, different types of strain fields are generated. For uniform corrosion, a distributed physical field map is output, indicating the overall expansion of the concrete surface due to uniform corrosion; for pitting corrosion, a local physical field map is output, reflecting the local stress concentration caused by the concentration of corrosion products. The generation of the distributed physical field map and the local physical field map is based on the COMSOL simplified version solver, and finally a 16-channel strain distribution map is output, including strain components in different directions.
[0074] Step 8.5: Perform global average pooling on the physical field map to obtain a first feature vector, and perform average pooling on the joint feature map to obtain a second feature vector.
[0075] Specifically, the physical field map size is 64x64x16, which is compressed into a 16-dimensional statistical vector (i.e., the first feature vector described in this embodiment) through a global average pooling operation, which contains statistical information such as the mean and variance of the strain field in each direction, representing the overall characteristics of the physical field. The joint feature map size is 64x64x256, which generates a 256-dimensional feature vector (i.e., the second feature vector described in this embodiment) by performing average pooling in the channel dimension, capturing the distribution of each modal feature in the global range. In addition, the feature vector acquisition unit also directly uses the vibration feature vector as a 6-dimensional feature vector to participate in subsequent gating calculation.
[0076] Step 8.6: Vector and vibration signal vector calculation to obtain gating weight.
[0077] The gating weight calculation formula is: w =σ( W 1· F p + W 2· F f + W 3· F v ), wherein, w is the gating weight; F p is the first feature vector; F f is the second feature vector; F v is the vibration feature vector; W 1 is the weight of the first feature vector, with a dimension of 16x256, which is a matrix that maps the physical field vector to the weight, used to emphasize modal features related to the strain field, such as the channel weight corresponding to ultrasonic crack, eddy current impedance mutation, etc. will be enhanced, W 2 is the weight of the second feature vector, with a dimension of 256x256, which is a self-attention matrix of the joint feature vector, capable of capturing the dependency between different modal features, such as identifying the spatial co-occurrence relationship between infrared hot spots and eddy current impedance anomalies, and enhancing the weight of related feature channels, W 3 is the weight of the vibration feature vector, with a dimension of 6x256, which is a matrix that maps the vibration feature to the weight, ensuring that the structural stiffness change (reflected by the vibration frequency drop rate) is associated with the corrosion feature, so that the generated image matches the corrosion degree with the structural stiffness degradation, and σ is the Sigmoid activation function, which normalizes the weight to the interval (0, 1), thereby suppressing channels unrelated to the current corrosion feature, such as noise modal in non-corrosion areas.
[0078] Step 8.7: Weight each channel of the joint feature map using the gating weight to obtain the fused feature map.
[0079] Specifically, according to the calculated gating weight w , each channel of the joint feature map is weighted to obtain the weighted feature F fused,c =w c ×F f,c . Wherein, F fused,c is the weighted c channel feature map, w c is the gating weight of the c channel, F f,c is the original c channel feature map. By weighting each channel of the joint feature map, the adaptive enhancement of the key feature channel is realized, so that the generator can focus on important information related to steel corrosion in subsequent processing.
[0080] Step 8.8: Use the U-Net architecture to convert the fused feature map into a visualized color image.
[0081] Comprising the following steps:
[0082] Step 8.8.1: Use the encoder to downsample the fused feature map.
[0083] The encoder consists of 5 convolutional layers, each using a 3x3 convolutional kernel with a step size of 2. As the level deepens, the size of the feature map gradually decreases, and the number of channels gradually increases from 64 to 1024. This process continuously extracts high-level semantic features, such as gradually abstracting the overall morphology of the rust area from local rust features.
[0084] Step 8.8.2: Use the decoder to upsample the fused feature map, and concatenate the downsampled result with the upsampled result.
[0085] The decoder includes 4 transposed convolutional layers, each using a 4x4 convolutional kernel with a step size of 2. When performing upsampling operations, each level will be concatenated with the features of the corresponding downsampled layer. For example, the input of the 3rd level upsample is 512 channels, which will be concatenated with the 256 channel features of the 2nd downsampled layer. This jump connection mechanism can preserve the detailed information of the bottom layer (such as the fine structure of the rust edge), while fusing high-level semantic features, effectively avoiding information loss, ensuring the accuracy and integrity of the generated image.
[0086] Step 8.8.3: Map the spliced sample results to a pseudo-color image.
[0087] After the feature processing through downsampling and upsampling, the feature map is mapped to a 3-channel pseudo-color image through a 1x1 convolutional layer, and the image size remains consistent with the input (e.g., 64x64). Before output, the pixel values are normalized to the interval [-1, 1] using the tanh activation function, and then converted to the interval [0, 255] through linear transformation to conform to the standard format for image display. The three channels of the pseudo-color image correspond to different corrosion information: the red channel (R) is used to represent the corrosion area ratio, with pixel values of 0-255 corresponding to corrosion area ratios of 0-30%, and the corrosion area distribution in the overall structure is intuitively displayed through the color depth; the green channel (G) reflects the corrosion depth, with pixel values of 0-255 corresponding to corrosion depths of 0-5mm, and the greener the color, the deeper the corrosion depth, helping users understand the severity of the steel reinforcement corrosion; the blue channel (B) is related to the vibration frequency drop rate Af%, and the higher the pixel brightness, the larger the Af, indicating that the structural stiffness degradation is more obvious, thereby establishing a visual connection between the corrosion degree and the change in structural performance.
[0088] Step 9: Discriminate the pseudo-color image using the dual-path discriminator.
[0089] Before discriminating the pseudo-color image, the pseudo-color image needs to be standardized by adjusting the input pseudo-color image to a fixed input size for the discriminator. For example, a pseudo-color image with a size of H × W × C is adjusted to a size of 256x256xC. During the standardization process, bilinear interpolation or nearest neighbor interpolation is used to maintain the image proportion and avoid geometric distortion. Additionally, the pixel values are normalized to the range [-1, 1] or [0, 1] to facilitate neural network calculations. The standardized pseudo-color image is discriminated, including the following steps:
[0090] Step 9.1: Calculate the gradient magnitude of the pseudo-color image in the x and y directions to obtain a gradient image.
[0091] Specifically, the Sobel operator or Prewitt operator is used to calculate the gradient magnitude of the image in the x and y directions. Taking the Sobel operator as an example, the gradient of the image in the x direction is calculated as and the gradient in the y direction is calculated as where I is the input pseudo-color image. Based on the x-direction gradient G x and the y-direction gradient G y , the formula The gradient image is calculated. Wherein, G(x, y) is the gradient image; represents the gradient of the image in the x direction, reflecting the rate of change of the image in the x direction, that is, the edge intensity of the image in the horizontal direction; represents the gradient of the image in the y direction, reflecting the rate of change of the image in the y direction, that is, the edge intensity of the image in the vertical direction.
[0092] Step 9.2: Weighting the gradient image to obtain a weighted gradient image.
[0093] Specifically, the gradient image is weighted according to the CSF model (such as the band-pass filtering characteristic based on spatial frequency).
[0094] It includes:
[0095] Step 9.2.1: Constructing a CSF model and generating a weighting kernel.
[0096] The expression of the CSF model is . Wherein, A represents the gain of the linear part in the CSF model, affecting the slope of the model in the low frequency region, that is, the sensitivity to low frequency components; B represents the decay rate of the linear part in the CSF model, affecting the falling speed of the model in the high frequency region, that is, the sensitivity decay of the high frequency component; C represents the gain of the Gaussian part in the CSF model, affecting the peak height of the model in the medium frequency region, that is, the sensitivity to medium frequency components; D represents the width of the Gaussian part in the CSF model, affecting the peak width of the model in the medium frequency region, that is, the sensitivity range of the medium frequency component; E represents the center frequency of the Gaussian part in the CSF model, affecting the peak position of the model in the medium frequency region, that is, the peak frequency of the sensitivity to medium frequency components; f is the spatial frequency. The spatial frequency is input into the CSF model to obtain the contrast sensitivity value under the corresponding frequency.
[0097] According to the CSF function, a weighting kernel W(f) in the frequency domain is generated, so that W(f) is proportional to CSF(f) (that is, the frequency sensitive to the human eye, the weighting kernel value is large; the frequency not sensitive to the human eye, the weighting kernel value is small). For example, let W(f)=CSF(f), or normalize CSF(f) to obtain W(f)∈[0,1].
[0098] Step 9.2.2: Weighting the gradient image in the frequency domain.
[0099] The frequency domain representation G(f) of the gradient image is multiplied by the weighting kernel W(f) by frequency component to obtain the weighted frequency domain result. The weighted frequency domain result is inverse Fourier transformed to convert back to the spatial domain to obtain the weighted gradient image.
[0100] Step 9.3: Extract local feature information from the gradient image.
[0101] This step involves extracting the "what" path features in the dual-path feature extraction process. This includes:
[0102] (1) Input layer
[0103] Used as input for the original preprocessed pseudocolor image. I (e.g., 256×256×C).
[0104] (2) Shallow convolutional layer
[0105] The shallow convolution layer is the first layer, used to extract basic visual features such as edges and textures (e.g., color block boundaries and textures in pseudo-color images). Its kernel size is 3×3, stride is 1, output channels are 64, and the activation function is LeakyReLU.
[0106] (3) Middle convolutional layer
[0107] The intermediate convolutional layers, layers 2 and 3, are used to extract mid-level semantic features (such as local structural features like crack direction) through downsampling. Their kernel size is 3×3, stride is 2, and the number of output channels doubles sequentially (i.e., 128 channels in layer 2 and 256 channels in layer 3). The activation function is LeakyReLU.
[0108] (4) Deep convolutional layers
[0109] The deep convolutional layers, layers 4 and 5, are used to extract high-level semantic features (such as abstract concepts like the degree of steel corrosion in bridges). Their kernel size is 3×3, stride is 2, and the number of output channels doubles sequentially (i.e., 512 channels in layer 4 and 1024 channels in layer 5). The activation function is LeakyReLU.
[0110] (5) Output layer
[0111] Used to output a multi-scale feature map set F what ={ f what,1 , f what,2 , f what,3 , f what,4 , f what,5},in, f what,1 Corresponding to the size features after the first convolutional layer, f what,2 Corresponding to the size features after the second convolution layer,f what,3 the size of the feature after the 3rd layer convolution, f what,4 the size of the feature after the 4th layer convolution, f what,5 the size of the feature after the 5th layer convolution. For example, f what,1 the corresponding size is 256x256x64, f what,2 the corresponding size is 256x256x128, f what,3 the corresponding size is 256x256x256, f what,4 the corresponding size is 256x256x512, f what,5 the corresponding size is 256x256x1024.
[0112] Step 9.4: Extract global feature information of the gradient image.
[0113] This step is to perform the where path feature extraction in the dual-path feature extraction. It includes:
[0114] (1) Input layer
[0115] Used to input the contrast sensitivity weighted gradient image.
[0116] (2) Shallow convolution layer
[0117] The shallow convolution layer is the 1st layer, which is used to extract the edge distribution features in the gradient image (such as the outline position of the crack). The convolution kernel size is 3x3, the step is 1, the output channel number is 32, and the activation function is LeakyReLU function.
[0118] (3) Middle convolution layer
[0119] The middle convolution layer is the 2nd-3rd layer, which is used to aggregate local edges by downsampling to form global shapes (such as the overall outline and trend of the crack). The convolution kernel size is 3x3, the step is 2, the output channel number is doubled in turn (i.e. the channel number of the 2nd layer is 64, and the channel number of the 3rd layer is 128), and the activation function is LeakyReLU function.
[0120] (4) Deep convolution layer
[0121] The deep convolution layer is the 4th layer, which is used to extract global layout features (such as the position of the crack in the image and other spatial relationships). The convolution kernel size is 3x3, the step is 2, the output channel number is doubled in turn (i.e. the channel number of the 4th layer is 256), and the activation function is LeakyReLU function.
[0122] (5) Output layer
[0123] Used to output a multi-scale feature map set F where ={ f where,1 , f where2 , f where,3 , f where,4 , f where,5},in, f where,1 Corresponding to the size features after the first convolutional layer, f where,2 Corresponding to the size features after the second convolution layer, f where,3 The size features after the third convolution layer, f where,4 Corresponding to the size features after the 4th convolution layer, f where,5 This corresponds to the size features after the 5th convolutional layer. For example, f where,1 The corresponding dimensions are 256×256×32. f where,2 The corresponding dimensions are 256×256×64. f where,3 The corresponding dimensions are 256×256×128. f where,4 The corresponding dimensions are 256×256×256. f where,5 The corresponding dimensions are 256×256×512.
[0124] Step 9.5: Fuse local feature information and global feature information.
[0125] include:
[0126] Step 9.5.1: Perform cross-path multi-scale feature alignment.
[0127] Will F what and F where Align feature maps of the same size hierarchically to facilitate fusion. For example, ... f what,1 (256×256×64) and f where,1 Align (256×256×32) and set aside f what,2 (256×256×128) andf where,2 (256x256x64) alignment, etc.
[0128] Step 9.5.2: Perform multi-level feature fusion.
[0129] The layer-by-layer fusion mode is adopted. For example, for shallow fusion, first, the fusion features of the first layer and the second layer are concatenated to obtain a feature map of [256x256x(64+32)] = 256x256x96, then the channel number is compressed to 64 through 1x1 convolution, and finally the fusion features are output through the activation function f what,1 and f where,1 are concatenated to obtain a feature map of [256x256x(64+32)] = 256x256x96, then the channel number is compressed to 64 through 1x1 convolution, and finally the fusion features are output through the activation function f fusion,1 For middle-level fusion and deep-level fusion, the above steps are similar, which will not be repeated here.
[0130] Step 9.5.3: Perform global pooling on the fused feature information to obtain a one-dimensional feature vector.
[0131] Specifically, the fusion features of each layer f fusion,1 , f fusion,2 , f fusion,3 , f fusion,4 , f fusion,5 are respectively subjected to global average pooling to obtain vectors v 1, v 2, v 3, v 4, v 5; and the vectors v 1, v 2, v 3, v 4, v 5 are concatenated into v [ v 1, v 2, v 3, v 4, v 5] to form a comprehensive feature vector containing multi-scale information (such as a dimension of 64+128+256+512+1028=1988).
[0132] Step 9.6: Map the one-dimensional feature vector to a discrimination result through a neural network.
[0133] It includes:
[0134] Step 9.6.1: Reduce the dimension of the one-dimensional vector through a fully connected layer and an activation function.
[0135] The purpose of this step is to further extract nonlinear features.
[0136] Step 9.6.2: Determine the authenticity of the input image through the output layer.
[0137] The input image is judged by a single neuron and an activation function, and the output image is the probability that it is a real image. p , p ∈[0,1], when p A value greater than 0.5 indicates a real image; otherwise, it is considered a pseudo-color generated image. It should be noted that for multi-classification scenarios (such as distinguishing the types of pseudo-color images), the number of neurons can be set to equal the number of pseudo-color image categories, and the activation function can be set to Softmax, thereby outputting the probability of each pseudo-color image type (such as the probability distribution of visual images, infrared images, and synthetic pseudo-color).
[0138] Step 10: When the discrimination result is greater than the threshold, optimize the generator by minimizing the loss function and output the final reconstructed image.
[0139] The binary classification cross-entropy loss function is used, and its expression is: L D =-[ ylog ( p )+( 1-y ) log (1- p )],in, L D This represents the cross-entropy loss value for binary classification. y =1 represents the real image label. y =0 indicates a pseudo-color image label. p This represents the probability that the output image is a real image. p ∈[0,1]. In the GAN framework, the discriminator and generator are optimized alternately, with the discriminator aiming to maximize [0,1]. L D The generator is designed to minimize L D This leads to a confrontational game.
[0140] Example 2: Corresponding to the image reconstruction method for identifying bridge steel reinforcement corrosion proposed in Example 1, this example provides an image reconstruction system for identifying bridge steel reinforcement corrosion, comprising:
[0141] The data acquisition module is used to simultaneously acquire multimodal data, physical parameters, and environmental data of the target area; the multimodal data includes: temperature field data, ultrasonic echo data, and electromagnetic impedance data; the physical parameters include: vibration signals, half-cell potential, and micro-strain field data; the environmental data includes: ambient temperature, relative humidity, and chloride ion concentration on the concrete surface;
[0142] a feature extraction module configured to extract feature information of each single modality data to obtain a multi-modality feature map;
[0143] a parameter acquisition module configured to acquire affine transformation parameters according to the multi-modality feature map;
[0144] a coordinate generation module configured to generate target grid coordinates according to the affine transformation parameters;
[0145] a linear interpolation module configured to perform bilinear interpolation on the multi-modality feature map according to the target grid coordinates to obtain an aligned multi-modality feature map;
[0146] a channel concatenation module configured to perform channel concatenation on the aligned multi-modality feature map to obtain a joint feature map;
[0147] an image generation module configured to process the joint feature map by using a generator of a physically-guided generative adversarial network to generate a pseudo-color image;
[0148] an image discrimination module configured to discriminate the pseudo-color image by using a double-path discriminator;
[0149] an image output module configured to optimize the generator by minimizing a loss function when a discrimination result is greater than a threshold value, and output a final reconstructed image.
[0150] Further, the image reconstruction system further comprises a data processing module configured to pre-process the multi-modality data.
[0151] Further, the data processing module comprises:
[0152] a first data processing unit configured to perform non-uniformity correction and bilateral filtering on the temperature field data;
[0153] a second data processing unit configured to perform band-pass filtering and time gain compensation on the ultrasonic echo data;
[0154] a third data processing unit configured to perform normalization processing and background noise subtraction on the electromagnetic impedance data;
[0155] a fourth data processing unit configured to perform Fourier transform on the vibration signal to obtain frequency spectrum information of the vibration signal, and perform normalization processing on the first three order inherent frequencies and damping ratios of the frequency spectrum information to obtain a vibration feature vector.
[0156] Further, the feature extraction module comprises a plurality of independent encoding branches, and one of the encoding branches corresponds to one modality data.
[0157] The encoding branch comprises:
[0158] a first convolutional layer configured to extract a plurality of local feature information from the single-modal data, wherein a convolution kernel size of the first convolutional layer is 5*5, a step size of the first convolutional layer is 2, an output channel number of the first convolutional layer is 64, and an activation function of the first convolutional layer is a LeakyReLU function;
[0159] a second convolutional layer configured to extract associated feature information between the plurality of local feature information, wherein a convolution kernel size of the second convolutional layer is 3*3, an expansion rate or a void rate of the second convolutional layer is 2, an output channel number of the second convolutional layer is 128, and an activation function of the second convolutional layer is a ReLU function;
[0160] a max-pooling layer configured to down-sample the associated feature information and the plurality of local feature information, and convert the down-sampled feature information into semantic feature information to obtain the multi-modal feature map, wherein a pooling kernel size of the max-pooling layer is 2*2, a step size of the max-pooling layer is 2, and an output channel number of the max-pooling layer is 256.
[0161] Further, the image generation module comprises:
[0162] a feature data splicing unit configured to splice the joint feature map with the physical parameters and the environmental data;
[0163] an associated parameter acquisition unit configured to analyze the spliced feature data by using a finite element model to obtain a corrosion morphology associated parameter, wherein the corrosion morphology associated parameter comprises a concrete elastic modulus, a corrosion expansion coefficient, and a corrosion current density;
[0164] a corrosion morphology classification unit configured to classify a corrosion morphology of the bridge reinforcement according to the corrosion morphology associated parameter, wherein the corrosion morphology comprises uniform corrosion and pitting corrosion;
[0165] a strain field generation unit configured to generate a physical field map corresponding to the classification result, wherein the physical field map comprises a distributed physical field map corresponding to the uniform corrosion and a local physical field map corresponding to the pitting corrosion;
[0166] a feature vector acquisition unit configured to perform global average pooling on the physical field map to obtain a first feature vector, and perform average pooling on the joint feature map to obtain a second feature vector;
[0167] a gating weight acquisition unit configured to calculate a gating weight according to the first feature vector, the second feature vector, and the vibration signal vector;
[0168] a channel weighting unit configured to weight each channel of the joint feature map using the gating weight to obtain a fused feature map;
[0169] an image generation unit configured to convert the fused feature map into a visual pseudo-color image using a U-Net architecture.
[0170] Further, the image generation unit comprises:
[0171] an encoder configured to down-sample the fused feature map;
[0172] a decoder configured to up-sample the fused feature map and concatenate the down-sampling result and the up-sampling result;
[0173] an output layer configured to map the concatenated sampling result into the pseudo-color image;
[0174] The pseudo-color image comprises:
[0175] a red channel configured to represent a corrosion area ratio;
[0176] a green channel configured to represent a corrosion depth;
[0177] a blue channel configured to represent a vibration frequency drop rate.
[0178] Further, the image discrimination module comprises:
[0179] a gradient amplitude calculation unit configured to calculate gradients of the pseudo-color image in x and y directions to obtain a gradient image;
[0180] a gradient image weighting unit configured to weight the gradient image to obtain a weighted gradient image;
[0181] a first feature extraction unit configured to extract local feature information of the gradient image;
[0182] a second feature extraction unit configured to extract global feature information of the gradient image;
[0183] a feature fusion unit configured to fuse the local feature information and the global feature information;
[0184] a global pooling unit configured to globally pool the fused feature information to obtain a one-dimensional feature vector;
[0185] a vector mapping unit configured to map the one-dimensional feature vector into a discrimination result.
[0186] Further, the loss function has an expression as follows: L D = -[ ylog( p )+( 1-y ) log (1- p )], wherein, L D is a binary cross-entropy loss value, y = 1 is a real image label, y = 0 is a fake image label, p is a probability that the output image is a real image, p ∈ [0, 1].
[0187] In the embodiment 3, the computer device can further include a power module, a display screen and other necessary components.
[0188] The working process, working details and technical effects of the computer device provided in the embodiment 3 can be referred to the method in the embodiment 1 or any possible method related to the method in the embodiment 1, which will not be described here.
[0189] In the embodiment 4, the computer readable storage medium can be a carrier for storing data, which can include, but is not limited to, floppy disks, optical discs, hard disks, flash memories, USB flash disks and / or memory sticks, and the computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0190] The working process, working details and technical effects of the aforementioned computer readable storage medium provided by the present embodiment can be referred to the method as described in Embodiment 1 or any possible method involving the method as described in Embodiment 1, which will not be repeated here.
[0191] Embodiment 5: The present embodiment provides a computer program product containing instructions, which, when executed on a computer, cause the computer to perform the method as described in Embodiment 1 or any possible method involving the method as described in Embodiment 1. Wherein, the computer can be a general-purpose computer, a special-purpose computer, a computer network or other programmable devices.
[0192] The above specific embodiments further illustrate the purposes, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An image reconstruction method for bridge rebar corrosion identification, characterized in that, The method comprises the following steps: Synchronously collecting multi-modal data, physical parameters and environmental data of a target area; the multi-modal data comprises temperature field data, ultrasonic echo data and electromagnetic impedance data; the physical parameters comprise vibration signals, half-cell potentials and micro-strain field data; the environmental data comprises environmental temperature, relative humidity and chloride ion concentration on the surface of the concrete; Performing non-uniform correction and bilateral filtering on the temperature field data; Performing band-pass filtering and time gain compensation on the ultrasonic echo data; Performing normalization processing and background noise subtraction on the electromagnetic impedance data; Performing Fourier transform on the vibration signals to obtain frequency spectrum information of the vibration signals, performing normalization processing on the first three order inherent frequencies and damping ratios of the frequency spectrum information to obtain a vibration feature vector; Extracting feature information of each single modal data to obtain a multi-modal feature map; Obtaining affine transformation parameters according to the multi-modal feature map; Generating target grid coordinates according to the affine transformation parameters; Performing bilinear interpolation on the multi-modal feature map according to the target grid coordinates to obtain an aligned multi-modal feature map; Performing channel splicing on the aligned multi-modal feature map to obtain a joint feature map; Processing the joint feature map by using a generator of a physically guided generative adversarial network to generate a pseudo-color image; the processing of the joint feature map by using the generator of the physically guided generative adversarial network comprises the following steps: splicing the joint feature map with the physical parameters and the environmental data; analyzing the spliced feature data by using a finite element model to obtain corrosion morphology correlation parameters; the corrosion morphology correlation parameters comprise a concrete elastic modulus, a corrosion expansion coefficient and a corrosion current density; classifying corrosion morphologies of bridge reinforcement according to the corrosion morphology correlation parameters; the corrosion morphologies comprise uniform corrosion and pitting corrosion; generating physical field maps corresponding to the classification results; the physical field maps comprise a distributed physical field map corresponding to the uniform corrosion and a local physical field map corresponding to the pitting corrosion; performing global average pooling on the physical field maps to obtain a first feature vector, performing average pooling on the joint feature map to obtain a second feature vector; calculating a gating weight according to the first feature vector, the second feature vector and the vibration signal vector; performing weighted processing on each channel of the joint feature map by using the gating weight to obtain a fused feature map; converting the fused feature map into a visual pseudo-color image by using a U-Net architecture; Discriminating the pseudo-color image by using a double-path discriminator; When the discrimination result is greater than a threshold value, optimizing the generator by minimizing a loss function and outputting a final reconstructed image.
2. The image reconstruction method for bridge steel corrosion identification according to claim 1, characterized in that, The method for extracting feature information of each single modal data comprises the following steps for each single modal data: Extracting a plurality of local feature information from the single modal data; Extracting correlation feature information between the plurality of local feature information; Performing down-sampling on the correlation feature information and the plurality of local feature information, and converting the feature information obtained by the down-sampling into semantic feature information to obtain the multi-modal feature map.
3. The image reconstruction method for bridge steel bar corrosion identification according to claim 1 or 2, characterized in that, The U-Net architecture is adopted to convert the fused feature map into a visual pseudo-color image, including the following steps: Down-sampling the fused feature map; Up-sampling the fused feature map, and splicing the down-sampling result and the up-sampling result; Mapping the spliced sampling result to the pseudo-color image; The pseudo-color image includes: A red channel for representing the rust area ratio; A green channel for representing the rust depth; A blue channel for representing the vibration frequency drop rate.
4. The image reconstruction method for bridge steel bar corrosion identification according to claim 1 or 2, characterized in that, The pseudo-color image is discriminated by using a double-path discriminator, including the following steps: Calculating the gradient of the pseudo-color image in the x direction and the y direction to obtain a gradient image; Weighting the gradient image to obtain a weighted gradient image; Extracting local feature information of the gradient image; Extracting global feature information of the gradient image; Fusing the local feature information and the global feature information; Global pooling the fused feature information to obtain a one-dimensional feature vector; Mapping the one-dimensional feature vector to a discrimination result.
5. The image reconstruction method for bridge steel corrosion identification according to claim 1 or 2, characterized in that, The expression of the loss function is: L D = [ ylog ( p )+( 1-y ) log (1- p )], wherein, L D is a binary cross-entropy loss value, y = 1 is a real image label, y = 0 is a fake color image label, p is a probability that the output image is a real image, p ∈ [0, 1].
Citation Information
Patent Citations
Multi-modal image reconstruction method based on modal consistency
CN115937055A
Bridge structure hidden damage identification method
CN119043606A