Spraying robot crack recognition and self-adaptive spraying method and equipment based on machine vision and medium

By employing multimodal data fusion and feature transformation techniques, the problem of insufficient generalization ability of convolutional neural networks in bridge crack identification was solved, achieving high-accuracy identification and adaptive spraying of bridge cracks, thus improving spraying quality and efficiency.

CN120900911AActive Publication Date: 2025-11-07SUZHOU AOZHITU INTELLIGENT TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511440458.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing convolutional neural networks have insufficient generalization ability in bridge crack identification due to environmental complexity and multi-scale interference, resulting in unstable identification accuracy and spraying quality.

Method used

By employing a multimodal data fusion mechanism, multimodal data with spatiotemporal alignment, noise suppression, and feature enhancement is constructed through real-time acquisition of color images, depth images, and robot posture and motion information. Combined with a multimodal feature transformation unit, a cross-modal spatial attention fusion module, and reinforcement learning-driven spraying parameter optimization, crack recognition and adaptive spraying are achieved.

Benefits of technology

It significantly improves the accuracy of bridge crack identification and coating quality, enhances adaptability to complex environments, reduces missed and false detections, and ensures the accuracy of coating path planning and coating quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120900911A_ABST
    Figure CN120900911A_ABST
Patent Text Reader

Abstract

The invention discloses a spraying robot crack recognition and self-adaptive spraying method and device based on machine vision and a medium, and relates to the field of spraying, the method comprises the following steps: S1, collecting a multi-modal original data stream in real time, and generating aligned, denoised and enhanced multi-modal data; s2, extracting and transforming a fusion feature pyramid from the multi-modal data; s3, segmenting the fused feature pyramid, and extracting morphological parameters and position coordinates of a crack region; s4, optimized spraying parameters and a local track adjusting instruction are output, and repairable spraying operation is carried out; and S5, real-time visual feedback and quality verification are carried out, and if the spraying quality does not reach the preset standard, strategy adjustment is carried out again. According to the method, the limitation caused by noise propagation during multi-scale feature alignment in a complex bridge environment and feature drift in a dynamic spraying process can be effectively overcome, and the crack recognition accuracy and the spraying repair quality and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of spraying technology, in particular to image data processing, and specifically to a crack identification and adaptive spraying method, device and medium for a spraying robot based on machine vision. BACKGROUND

[0002] As an important part of modern transportation infrastructure, steel structure bridges are prone to surface cracks during long-term service due to the influence of load, environment and material aging. If not detected and repaired in time, crack propagation will threaten the safety of the structure and shorten the service life of the bridge. Traditional manual inspection methods have been difficult to meet multiple indicators such as efficiency, accuracy and safety. In this context, intelligent spraying robot technology based on machine vision has emerged and rapidly become a research hotspot and development direction in the field. This technology integrates advanced image acquisition, intelligent identification and precision control systems to achieve automated and accurate identification of structural surface defects such as cracks, and on this basis, performs adaptive repair spraying operations, thereby significantly improving the intelligent level and overall efficiency of maintenance operations.

[0003] In the prior art, convolutional neural networks (CNN) are often used for image recognition of spraying areas, and the spraying path, nozzle movement speed, spraying pressure and coating flow are determined according to the morphological characteristics of the cracks. For example, the Chinese invention patent CN109146849A discloses a road crack detection method based on convolutional neural networks and image recognition. This method is safe and reliable, fast in detection speed, high in detection efficiency, low in false alarm rate, and does not have night work fatigue phenomenon. It can accurately determine the type of highway defects, and relevant departments can take reasonable measures to repair the road defects in time.

[0004] However, although CNN has strong feature extraction and classification capabilities in image recognition tasks, its direct application in real bridge maintenance scenarios is not ideal due to the complex and variable bridge surface environment, strong light changes, rain and fog weather interference, rust stains, welds, and various noise sources such as coating peeling. A CNN model trained under specific lighting and background conditions often overfits to non-critical features such as background texture and shadows in the training data, i.e., the model learns the "shortcut" in the data rather than the essential features of the cracks. When the environmental conditions change (such as the difference in morning and evening light, new rust morphology), the recognition performance of the model will decrease significantly, leading to missed or false detection of cracks, and thus affecting the accuracy of subsequent spraying path planning and spraying quality. SUMMARY

[0005] The present application overcomes the shortcomings of the prior art and provides a crack identification and adaptive spraying method, device and medium for a spraying robot based on machine vision.

[0006] To achieve the above object, the technical scheme adopted by the present application is as follows: In a first aspect, the present application provides a crack identification and adaptive spraying method for a spraying robot based on machine vision, comprising the following steps: S1, real-time acquisition of a multi-modal original data stream of a bridge structure surface to be detected, and generation of multi-modal data after time-space alignment, noise suppression and feature enhancement; the multi-modal original data stream comprises high-resolution color images, high-precision depth images and posture and motion information of the spraying robot; S2, from the multi-modal data, a multi-modal feature transformation unit is used to extract and transform a fusion feature pyramid with multiple scales for time dynamic changes; the multi-modal feature transformation unit comprises a parallel feature extraction network, a group of intra-modal time sequence feature transformers and a cross-modal spatial attention fusion module; S3, the fusion feature pyramid is segmented by a feature extraction and segmentation unit to identify and extract morphological parameters and position coordinates of a crack region in a three-dimensional space; S4, based on the morphological parameters, the position coordinates and the current posture information of the spraying robot, optimized spraying parameters and local trajectory adjustment instructions are output to perform a repair spraying operation.

[0007] In a preferred embodiment of the present application, in the step of S1, the time-space alignment of the multi-modal original data stream comprises: acquisition of each group of color images and depth images at the same time; registration of the depth images and the color images to the same coordinate system by using pre-calibrated external parameters of each sensor to generate pixel-level aligned RGB-D image pairs; and timestamp alignment of the posture and motion information data of the spraying robot and the RGB-D image pairs to generate uniformly timed data frames.

[0008] In a preferred embodiment of the present application, in the step of S1, the noise suppression of the multi-modal original data stream comprises: application of a non-local mean filter-based image denoising algorithm to the color images; application of a bilateral filter or median filter-based denoising algorithm to the depth images; and application of Kalman filtering or complementary filtering to the posture and motion information data of the spraying robot.

[0009] In a preferred embodiment of the present application, in the step of S1, the feature enhancement of the multi-modal original data stream comprises: application of adaptive histogram equalization to the color images to enhance image contrast and extract gradient information of the images; application of depth gradient calculation to the depth images to highlight small height changes of the surface; and The instantaneous speed and acceleration vector of the spraying robot is calculated based on the posture and motion information of the spraying robot, as the context input of subsequent feature transformation.

[0010] In a preferred embodiment of the present application, in the step of S2, the parallel feature extraction network comprises independent feature extraction paths for color images and depth images, each path adopts a lightweight convolutional neural network as a backbone network for extracting a multi-scale initial feature map set from a single frame of image. The intra-modal temporal feature transformer group comprises at least two parallel temporal transformer modules corresponding to the color image modality and the depth image modality respectively, each temporal transformer module receives a multi-frame, multi-scale feature map sequence from the parallel feature extraction network in the corresponding modality, and generates a multi-scale temporal feature map set stable in the time dimension and robust to feature drift by capturing the dynamic change pattern of crack features in the time dimension through an encoder structure based on the Transformer architecture, comprising a multi-head self-attention mechanism and a feedforward neural network.

[0011] In a preferred embodiment of the present application, in the multi-modal feature transformation unit, the posture and motion information of the spraying robot is further integrated into the input of the intra-modal temporal feature transformer group. Specifically, the posture and motion information of the spraying robot is processed through a one-dimensional convolutional neural network to extract its temporal features, and the temporal features are encoded as conditions and spliced or fused through a gating mechanism with the feature vectors of the color image modality and the depth image modality, which are jointly input into the temporal transformer module to provide additional context information.

[0012] In a preferred embodiment of the present application, the cross-modal spatial attention fusion module is used to receive the multi-scale feature map set of the color image modality and the multi-scale feature map set of the depth image modality from the intra-modal temporal feature transformer group, and for each feature pyramid level, the following operations are performed: The feature map of the color image modality and the feature map of the depth image modality are upsampled or downsampled to align their spatial resolutions. Query, key and value matrices are constructed, and cross-modal spatial attention mechanism is adopted to calculate the similarity between the query and the key to assign attention weights to the value matrix, which is used to realize feature fusion from different modalities and suppress the influence of structured noise through depth information. The cross-modal spatial attention fusion module outputs a multi-scale feature pyramid that fuses color and depth information and is accurately aligned and noise-suppressed in the spatial dimension.

[0013] In a preferred embodiment of the present application, in the step of S3, the feature extraction and segmentation unit adopts a segmentation network based on an encoder-decoder architecture, the encoder part of the segmentation network is composed of or connected to the fusion feature pyramid, and the decoder part receives the multi-scale fusion features output by the encoder, and gradually restores the spatial resolution through upsampling, skip connection and convolution layer, and finally outputs a pixel-level crack probability map or a binary segmentation mask. Through morphological processing and geometric analysis algorithms, morphological parameters of the crack are accurately extracted, including the length, average width, maximum width, direction, connectivity, area, and position coordinates in three-dimensional space of the crack. By matching the segmented crack pixels with the corresponding depth image data, the point cloud representation of the crack in three-dimensional space is obtained, and based on the point cloud data, the actual three-dimensional shape, depth profile information and surface roughness of the crack are calculated.

[0014] The second aspect of the present application provides an electronic device, comprising: at least one processor; and a memory connected in communication with the at least one processor; The memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the machine vision-based crack identification and adaptive spraying method of any one of the above.

[0015] The third aspect of the present application provides a computer readable storage medium, which stores computer instructions for enabling a processor to execute the machine vision-based crack identification and adaptive spraying method of any one of the above.

[0016] The present application solves the defects in the background art, and has the following advantages: (1) The present application provides a machine vision-based crack identification and adaptive spraying method, device and medium for a spraying robot, which introduces a multi-modal data fusion mechanism, synchronously collects and processes color images, depth images and robot pose and motion information, constructs a multi-modal data basis with time and space alignment and denoising enhancement, can comprehensively utilize visual texture and three-dimensional geometric information, significantly enhances the perception ability of crack features, especially in complex backgrounds such as light changes, rust and welds, effectively distinguishes real cracks from structured noise, thereby improving the accuracy of crack detection and further enhancing the accuracy of subsequent spraying path planning and spraying quality.

[0017] (2) In the present application, by designing a modal intra-time feature transformer group in the feature extraction stage, the continuous multi-frame feature sequence is processed by using the encoder structure based on the Transformer, and the dynamic change of the crack feature is modeled in the time dimension through the multi-head self-attention mechanism, which can effectively suppress the feature drift caused by robot movement, light fluctuation and the like, enhance the stability and consistency of the feature in time sequence, and compared with the static FPN structure, the adaptability to the dynamic operation environment can be significantly improved, thereby reducing the missed detection and false detection, and ensuring the reliability of the crack identification result in continuous operation.

[0018] (3) In the present application, by cooperating with the cross-modal spatial attention fusion module, the aligned RGB and depth features are used to construct query, key and value mapping, and intelligent feature fusion and noise suppression are realized through attention weight, which can effectively distinguish the texture similar interference area (such as rust spot, weld) according to the depth information, suppress the response of non-crack structure, and then improve the signal-to-noise ratio of the multi-scale feature pyramid, compared with the traditional multi-scale fusion method, the blind propagation of noise in the traditional FPN can be avoided, and it has stronger anti-interference ability in the complex bridge background, especially suitable for precise identification and geometric parameter extraction of micro cracks, thereby improving the accuracy of crack identification and the quality and efficiency of spraying repair. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings; Figure 1 is the flow chart of the spraying robot crack identification and self-adaptive spraying method based on machine vision of embodiment 1 of the present application; Figure 2 is the structural block diagram of the spraying robot crack identification and self-adaptive spraying method based on machine vision of embodiment 1 of the present application; Figure 3 is the structural block diagram of the multi-modal feature transformation unit of embodiment 1 of the present application; Figure 4 is a structural schematic diagram of an electronic device that can be used to implement embodiment 1 of the present application. DETAILED DESCRIPTION

[0020] With reference to the accompanying drawings: clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0021] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be practiced without the specific details that are set forth in the following description, in other manners different from those described herein, and therefore the scope of the present application is not limited by the specific embodiments disclosed below.

[0022] Summary of the application: To improve the generalization ability of traditional convolutional neural networks in bridge crack identification due to the complexity of the environment and multi-scale interference, some studies attempt to introduce a feature pyramid network (FPN) structure. FPN can extract deep semantic features and shallow detail features of an image in parallel, and realize multi-scale feature fusion through horizontal connection, which can theoretically enhance the perception ability of the model to cracks of different scales, while suppressing background interference using attention mechanism to improve the recognition accuracy in complex backgrounds such as corrosion and welds.

[0023] However, the applicant found that when FPN is directly applied to bridge crack identification and spraying robot operation scenarios, there are still two outstanding problems due to the structural principle: one is the noise propagation mechanism. A large number of structured disturbances on the bridge surface (such as rust spots and weld textures) enter the FPN through shallow features and are propagated and even semantized in the feature fusion process, which is difficult to be effectively suppressed, and instead pollutes the deep features, reducing the purity of crack features. The second is the feature drift problem. In dynamic operation of the spraying robot, factors such as camera pose, paint surface reflection, and environmental lighting can cause image features to drift in time sequence, and the FPN trained statically lacks the ability to adapt to such dynamic changes, resulting in unstable recognition results.

[0024] To solve the above problems, the present application provides a spraying robot crack identification and adaptive spraying method based on machine vision, which introduces multi-modal perception, dynamic time sequence feature transformation, adaptive gated cross-modal attention fusion, and reinforcement learning driven spraying parameter optimization mechanism, which can effectively overcome the limitations caused by noise propagation during multi-scale feature alignment in complex bridge environments and feature drift during dynamic spraying process, and improve the accuracy of crack identification and the quality and efficiency of spraying repair.

[0025] Embodiment 1: As shown in Figure 1 and Figure 2 A spraying robot crack identification and adaptive spraying method based on machine vision, comprising the following steps: S1, real-time acquisition of multi-modal original data stream of the surface of the bridge structure to be detected, and generation of multi-modal data after time-space alignment, noise suppression and feature enhancement; S2, from the multi-modal data, a multi-modal feature transformation unit is used to extract and transform a fusion feature pyramid with multi-scale time dynamic changes; S3, the fusion feature pyramid is segmented by a feature extraction and segmentation unit to identify and extract the morphological parameters of the crack region and the position coordinates in the three-dimensional space; S4, based on the morphological parameters, the position coordinates and the current posture information of the spraying robot, the optimized spraying parameters and the local trajectory adjustment instructions are output to perform the repair spraying operation; S5, during or after the spraying operation, real-time visual feedback and quality verification are performed, and if the spraying quality does not meet the preset standard, the quality evaluation result is fed back as new state information for strategy adjustment.

[0026] In a specific embodiment, in the step of S1, the multi-modal original data stream is collected by a multi-modal perception unit configured on the spraying robot body, which is used to collect the multi-modal original data stream of the surface of the bridge structure to be detected.

[0027] In this embodiment, the data stream covers high-resolution color images, high-precision depth images, and posture and motion information sequences of the spraying robot; the multi-modal perception unit is composed of an RGB color camera, a depth camera, and an inertial measurement unit (IMU).

[0028] Among them, the RGB color camera is preferably an industrial-grade color camera of Basler acA1920-40gc, which is configured to capture image data in the visible light spectrum, with an optical resolution of 1920x1080 pixels and a frame rate of 30 frames / second; to minimize motion blur, the camera integrates a global shutter function to ensure image clarity under high-speed robot motion; the depth camera is preferably an Intel RealSense D435i structured light depth sensor, which is configured to obtain distance information from the surface to be measured to the sensor, with a depth measurement range of 0.1-5 m and a depth measurement accuracy of better than ±1 mm under ideal conditions.

[0029] It should be noted that the depth camera and the RGB color camera achieve millisecond-level time synchronization through a hardware triggering mechanism, ensuring high consistency of each set of RGB-D image pairs in terms of acquisition time.

[0030] Further, the IMU is preferably a six-axis inertial sensor of the model Analog Devices ADIS16470, configured to collect real-time three-axis acceleration data and three-axis angular velocity data around the X, Y, Z axes of the robot, with a data update rate of 200 Hz, for accurately characterizing the instantaneous motion state of the robot, including linear velocity, angular velocity, and attitude including pitch, yaw, and roll.

[0031] Specifically, all sensors of the multi-modal perception unit are fixedly installed on the end effector of the spraying robot and calibrated for internal and external parameters, wherein the camera internal parameter calibration adopts Zhang's calibration method, and the external parameter calibration adopts multi-camera joint calibration based on a checkerboard or a self-defined calibration board, to ensure high consistency of the sensor data in space and time, laying a foundation for subsequent data fusion.

[0032] In the present embodiment, the specific steps of the spatio-temporal alignment processing of the multi-modal raw data stream are as follows: 1. Start.

[0033] 2. Based on the timestamp information of the RGB color camera and the depth camera, perform time synchronization at the hardware level or the software level to ensure that each set of color images and depth images are collected within less than 5 milliseconds; 3. Use the pre-calibrated external parameters of each sensor, including the rotation matrix and the translation vector, to register the depth image to the coordinate system of the color image, to generate a pixel-level aligned RGB-D image pair; 4. Align the timestamp of the IMU-acquired attitude and motion information data with the timestamp of the RGB-D image pair, and use linear interpolation or nearest neighbor interpolation method to resample the IMU data to generate a unified time sequence data frame, to ensure the temporal synchronization of all modal information.

[0034] 5. End.

[0035] The specific steps of the noise suppression processing of the multi-modal raw data stream are as follows: 1. Start.

[0036] 2. Apply a non-local mean filtering-based image denoising algorithm to the color image; wherein the search window size is set to 21x21 pixels, the similarity calculation window size is set to 7x7 pixels, and the standard deviation of the Gaussian weighting function is set to 10, to effectively remove high-frequency random noise in the image while maximizing the clarity of edge structures such as cracks.

[0037] 3. Apply a bilateral filtering-based denoising algorithm to the depth image; The kernel size is set to 9×9 pixels, the color space standard deviation sigma_color is set to 75, and the spatial domain standard deviation sigma_space is set to 75. These settings are used to smooth random fluctuations in depth data and effectively preserve depth edge information, avoiding blurring of crack boundaries.

[0038] 4. Apply extended Kalman filtering to the IMU data; The state vector includes the robot's position, velocity, attitude (quaternion), and gyroscope and accelerometer biases. By fusing accelerometer and gyroscope measurements, more accurate and stable robot attitude and motion information can be estimated, which can suppress the inherent Gaussian white noise and drift error of the sensors.

[0039] 5. End.

[0040] The specific steps for feature enhancement processing of multimodal raw data streams are as follows: 1. Begin.

[0041] 2. Apply the CLAHE algorithm for adaptive histogram equalization to color images; The contrast limit was set to 2.0 and the grid size was set to 8x8 to enhance the local contrast of the image and make the cracks more obvious under different lighting conditions. Simultaneously, gradient information of the image is extracted, and edge features are extracted using the Canny operator, with a low threshold set to 50 and a high threshold set to 150 to highlight the boundaries of the cracks.

[0042] 3. Apply depth gradient calculation to the depth image, using the Sobel or Scharr operator to calculate the gradient of the depth image in the x and y directions to highlight small surface variations, such as depressions or protrusions that often indicate the presence of cracks. 4. Based on the IMU data processed by Kalman filtering, calculate the robot's instantaneous linear velocity, angular velocity, and linear acceleration vector, and use these motion state parameters as context input for subsequent feature transformation to provide the model with motion information of the painting robot.

[0043] 5. End.

[0044] like Figure 3 As shown, in a specific embodiment, in step S2, the multimodal feature transformation unit is used to extract and transform multi-scale feature representations that are dynamically changing over time from continuous multimodal data frames, and to solve the problems of noise propagation and feature drift during multi-scale feature alignment. It includes: a parallel feature extraction network, an intramodal temporal feature transformer group, and a cross-modal spatial attention fusion module.

[0045] In this embodiment, the parallel feature extraction network contains independent feature extraction paths for different modal data: one path processes RGB color image sequences, and the other path processes depth image sequences; each path uses a lightweight and efficient convolutional neural network as the backbone network.

[0046] Exemplarily, based on a version of MobileNetV3-Small architecture optimized for specific tasks and pruned, rich features are extracted while ensuring computational efficiency; for the RGB image path, the input is a three-channel color image frame; for the depth image path, the input is a single-channel depth image frame; wherein the backbone network outputs a set of multi-scale initial feature maps, containing 4 feature levels with decreasing resolution and increasing semantic information (for example, feature maps with resolutions of 1 / 4, 1 / 8, 1 / 16, and 1 / 32 relative to the input resolution); for an input image of 1920x1080, the output feature map sizes are 480x270, 240x135, 120x67, and 60x33, respectively.

[0047] Specifically, the output feature maps capture low-level texture features containing edges and corner points, middle-level semantic features containing local structure of cracks, and high-level abstract features containing overall morphology and context of cracks, respectively; the parallel feature extraction network can generate a series of multi-scale feature maps with rich semantic information for each input frame, providing a basis for subsequent temporal transformation and modal fusion.

[0048] In this embodiment, the intra-modal temporal feature transformer group includes two parallel temporal transformer modules, corresponding to the RGB color image modality and the depth image modality, respectively; each modality's temporal transformer module receives a sequence of continuous multi-scale feature maps from the parallel feature extraction network for its corresponding modality.

[0049] Further, for each RGB or depth modality, the input is a multi-scale feature map sequence of length T (e.g., T=8 frames), each feature map is first flattened into a series of feature vectors.

[0050] Exemplarily, for a feature map with size HxWxC, it will be flattened into HxW C-dimensional feature vectors; in order to preserve temporal information, each feature vector is also added with a learnable position encoding, which can indicate the temporal position of the feature vector in the sequence; the core of the temporal transformer module is an encoder structure based on the Transformer architecture, which contains multi-head self-attention mechanism and feedforward neural network inside.

[0051] Specifically, the multi-head self-attention mechanism allows the model to simultaneously focus on the features of different positions in the sequence in the time dimension, thereby capturing the dynamic change pattern of the crack features between consecutive frames; for example, when the spraying robot moves or the lighting conditions change, the visual appearance of the crack may appear a slight drift or local occlusion, and the self-attention mechanism effectively suppresses the influence of noise frames or unstable frames (such as image blur frames caused by transient occlusion or violent motion) by calculating the query (Query), key (Key) and value (Value) matrices and assigning dynamic weights to the value matrix according to the similarity between the query and the key, thereby enhancing the crack features that appear continuously and stably in the time dimension; after processing by the temporal transformer module, each modality generates a set of multi-scale temporal feature maps that are more stable in the time dimension and are robust to feature drift, which can effectively alleviate the feature drift problem faced by the static FPN in the dynamic spraying process and ensure the continuity and consistency of the crack features.

[0052] As a preferred embodiment, in the extraction and transformation of multi-modal data, the spraying robot pose and motion information sequence collected by the IMU can be further integrated into the input of the intra-modal temporal feature transformer group.

[0053] Specifically, the IMU data sequence including three-axis acceleration and three-axis angular velocity is first processed by a small one-dimensional convolutional neural network (1D CNN), which contains three convolutional layers, each followed by a ReLU activation function and batch normalization, for extracting high-level temporal features thereof; The extracted IMU temporal features are encoded as conditions, which can be concatenated with the visual feature vectors or fused through a gating mechanism for each visual feature vector in the intra-modal temporal transformer.

[0054] For example, in the gating mechanism, the IMU features first generate a gating vector through a small multi-layer perceptron, which is element-wise multiplied with the visual feature vector to dynamically adjust the influence of IMU information on the transformation of visual features. This conditional encoding can provide additional context information for the self-attention mechanism, such as the instantaneous motion speed, acceleration and rotation direction of the robot, which helps the model better understand and compensate for visual feature drift caused by dynamic robot motion, especially in high-speed or non-uniform motion scenarios, thereby improving the accuracy and stability of feature representation.

[0055] In this embodiment, the cross-modal spatial attention fusion module is used to fuse the multi-scale feature maps of the color image modality and the multi-scale feature maps of the depth image modality from the intra-modal temporal feature transformer group, which have been optimized in the time dimension, through an intelligent, spatially aware attention mechanism to solve the noise propagation problem when aligning multi-scale features in a complex bridge environment.

[0056] Further, the cross-modal spatial attention fusion module performs the following operations for each feature pyramid level: Bilinear or nearest neighbor interpolation operation is performed on the feature maps of the color image modality and the feature maps of the depth image modality to align their spatial resolutions precisely so as to perform pixel-wise level fusion.

[0057] Query, key and value matrices are constructed; in one specific design, the processed RGB feature maps are taken as the query Q , the depth feature maps are taken as the key K and the value V , and the specific operations are as follows: the RGB modality query matrix Q rgb = FFN Q ( F rgb,aligned ), the depth modality key matrix K depth = FFN K ( F depth,aligned ), and the depth modality value matrix V depth = FFN V ( F depth,aligned ), where FFN represents a feedforward network; F rgb,aligned represents the aligned RGB feature maps; F depth,aligned represents the aligned depth feature maps.

[0058] The cross-modal spatial attention mechanism is adopted to calculate the similarity between the query Q and the key K through a dot product operation, and the result is scaled and Softmax normalized to obtain an attention weight map; the weight map then acts on the value V to generate a fused feature representation.

[0059] Exemplarily, the depth information can provide height information for structured noises such as welds, rust patches, and paint peeling on the surface of a bridge, while a real crack usually appears as a tiny depression or gap. When a certain region in the RGB feature map contains texture features similar to a crack, but the corresponding depth information shows that the region is not significantly depressed, the attention weight of the depth modality will be assigned a lower value, thereby suppressing the misidentification of the region as a crack in the RGB modality. Conversely, when both RGB and depth information strongly indicate the presence of a crack (e.g., the RGB texture shows a thin line structure, and the depth map presents a significant linear depression in the region), the attention weight will be amplified, achieving effective fusion and enhancement of the features of the two modalities.

[0060] It can be understood that the cross-modal spatial attention fusion module outputs a set of multi-scale feature pyramids that fuse RGB and depth information and are precisely aligned and noise-suppressed in the spatial dimension. The features of the pyramid have higher discriminability for real cracks.

[0061] As a preferred embodiment, in the cross-modal spatial attention fusion module, the attention mechanism can further employ a learnable gating network to dynamically adjust the contribution weights of different modal features.

[0062] Specifically, the gating network receives feature maps (such as feature summary representations obtained by global average pooling) from each modality, and predicts a set of gating coefficients between 0 and 1 through a small multi-layer perceptron (MLP).

[0063] Exemplarily, for the RGB modality feature map F rgb,aligned and the depth modality feature map F depth,aligned , the concatenated input vector of the gating network is then processed by the MLP to process the concatenated features z to output two scalar contribution weights of the RGB modality in the fused features g rgb and the contribution weight of the depth modality in the fused features g depth , which are mapped to the range [0, 1] by the Sigmoid activation function, and then subjected to element-level or channel-level multiplication operations with the feature maps of each modality to achieve weighting of the feature information.

[0064] For example, when the ambient lighting conditions are good, the RGB texture information is rich and reliable, but the depth information may be limited or noisy due to uneven surface material reflectivity, the gating network can learn and increase the weight of g rgb , while reducing g depththe weight of the depth information; while in the environment with complex illumination and serious shadow interference, but more reliable depth information, the weight of the depth information is increased g depth The gating network automatically learns how to allocate the optimal modal weight according to the input feature content and environmental conditions in the training process through end-to-end optimization, further enhancing the adaptability of feature fusion.

[0065] In a specific embodiment, in the step S3, the feature extraction and segmentation unit adopts a segmentation network based on an encoder-decoder architecture, preferably a DeepLabV3+ model; wherein the encoder part of the segmentation network is composed of or connected with the fusion feature pyramid, for example, receiving features from each level of the pyramid through a skip connection.

[0066] In this embodiment, the encoder uses a pre-trained backbone network and integrates a dilated spatial pyramid pooling module to capture multi-scale context information; the decoder part receives the multi-scale fusion features output by the encoder, and gradually restores the spatial resolution of the feature map through upsampling, skip connection and convolution layer, and finally outputs a pixel-level crack probability map or a binary segmentation mask.

[0067] Wherein, the backbone network is selected from one of ResNet-101 or EfficientNet-B4; the pixel-level crack probability map is, for example, each pixel value is between 0 and 1, indicating the probability that the pixel belongs to the crack; the binary segmentation mask is converted into a binary graph by setting a threshold, such as 0.5, 1 indicating a crack and 0 indicating a background.

[0068] It can be understood that the feature extraction and segmentation unit can accurately identify and segment the crack area on the surface of the bridge, and on this basis, accurately extract the morphological parameters of the crack through a series of morphological processing and geometric analysis algorithms such as opening operation, closing operation and thinning algorithm; wherein the morphological parameters include but are not limited to the length, average width, maximum width, direction, connectivity, area and position coordinates in three-dimensional space of the crack.

[0069] Specifically, the length of the crack is extracted by applying a skeletonization algorithm to the binary segmentation mask to extract the crack centerline, then calculating its pixel length, and converting it into actual physical length combined with the camera calibration parameters; the width of the crack is determined by calculating the average or maximum Euclidean distance from the crack skeleton line to the crack edge, and is also converted into physical units; the direction of the crack is determined by performing principal component analysis on the crack pixel set, or by calculating the tangent direction of each point on the crack skeleton line to determine its main direction; the connectivity of the crack is determined by analyzing its topological structure to distinguish single cracks, branch cracks or crack networks.

[0070] The skeleton algorithm is preferably a Zhang-Suen thinning algorithm; and the camera calibration parameters include focal length and pixel size.

[0071] Further, the feature extraction and segmentation unit can also integrate depth information to reconstruct the crack in three-dimensional geometry. Specifically, by matching the segmented crack pixels obtained from the binary segmentation mask with the corresponding depth map data, the point cloud representation of the crack in three-dimensional space can be obtained. Based on this point cloud data, the actual three-dimensional shape of the crack, the depth profile information including the average depth and maximum depth of the crack, and the surface roughness can be further calculated.

[0072] For example, by locally fitting the crack point cloud, the roughness of the surface around the crack can be quantified, or by calculating the average depth difference between the points inside the crack region and the points in the surrounding non-crack region, the depth information of the crack can be obtained. The accurate three-dimensional geometric feature data provides more accurate three-dimensional spatial positioning and geometric matching data for subsequent adaptive spraying parameter control, enabling the spraying operation to truly achieve precise repair of the crack defect.

[0073] In a specific embodiment, in the step of S4, the adaptive spraying parameter control unit receives the crack morphology parameters output by the multi-modal feature transformation unit and the current pose information of the spraying robot.

[0074] The crack morphology parameters include but are not limited to length, width, direction, connectivity, three-dimensional position and depth information; and the current pose information of the spraying robot is obtained by fusing IMU data and robot body encoder data.

[0075] Specifically, the adaptive spraying parameter control unit adopts a reinforcement learning paradigm to dynamically optimize the spraying parameters and motion trajectory of the spraying robot. The adaptive spraying parameter control unit includes a reinforcement learning agent, an environment model, and a reward function.

[0076] In this embodiment, the reinforcement learning agent is a policy network and a value network based on a deep neural network, preferably using the Soft Actor-Critic (SAC) algorithm architecture; the policy network is a multi-layer perceptron (MLP) that receives the current state as input and outputs continuous spraying actions, i.e., spraying parameters and local trajectory adjustment instructions including spraying pressure, coating flow, spraying speed, spraying gun attitude, and spraying gun to surface distance, and the output layer uses a tanh activation function to map the action value to a predefined action space range; the value network is two Q networks for double Q learning in the SAC algorithm, which, like the policy network, is an MLP that receives the current state and action as input and evaluates the expected return (Q value) of executing the specific action in this state to guide the learning direction of the policy network.

[0077] In this embodiment, the input of the environment model includes the geometric features of the currently detected crack, the current pose of the spraying robot, the distance and angle between the spraying gun and the target surface, the current ambient light intensity, the target surface material properties, and the quality evaluation results of the previous spraying attempt; the environment model is a simulator based on physical simulation and empirical data, which is used to simulate the physical characteristics in the spraying process, including paint fluid mechanics, deposition efficiency, drying time, coating adhesion, and evaluation of spraying effect by the visual feedback system; the environment model is used to update the environment state according to the action of the agent and calculate the reward.

[0078] wherein the geometric features include, for example, length 200 mm, average width 2 mm, maximum depth 1.5 mm, three-dimensional position coordinates, main direction, curvature, etc.; the current pose of the spraying robot includes position [X, Y, Z] and orientation [Roll, Pitch, Yaw]; the material properties include concrete, steel, water absorption, surface roughness; the paint fluid mechanics includes paint atomization effect, spraying mode, droplet size distribution; the deposition efficiency includes coating thickness and uniformity.

[0079] In this embodiment, the reward function is used to quantify the quality and efficiency of the spraying operation and guide the learning direction of the reinforcement learning agent, and the reward function is designed as follows: when the spraying operation can achieve complete coverage of the crack area, uniform coating thickness, good adhesion, and minimal material waste, the reward value reaches the maximum.

[0080] Specifically, the reward value is calculated by a plurality of indicators, including the difference between the actual spraying coverage and the target coverage, the material consumption, the path execution deviation, the overspray area beyond the crack area, and the coating adhesion score; each indicator is provided with a corresponding weight coefficient, and optionally, the coverage weight is 5.0, the material waste weight is 2.0, the path deviation weight is 1.0, the overspray weight is 3.0, and the adhesion weight is 4.0, which are set according to expert experience and preliminary experiments to reasonably balance the importance between different optimization objectives.

[0081] wherein the target coverage is generally set to 100%, the actual coverage is obtained by comparing the images before and after spraying or analyzing the depth information through the visual feedback system. The material consumption is in milliliters, the path deviation refers to the average offset distance between the actual path and the ideal path, the overspray area refers to the area of the non-target area that is mistakenly sprayed, and the adhesion score is a comprehensive score between 0 and 1, which is calculated from indirect indicators such as coating thickness, uniformity, and surface roughness.

[0082] Further, through continuous optimization of the reward function, the reinforcement learning agent can gradually master an efficient spraying strategy, which can adaptively adjust the spraying parameters, including spraying pressure 0.1-0.6 MPa, flow rate 10-50 ml / min, moving speed 50-200 mm / s, spray gun posture, and spray gun distance from the surface 100-300 mm, according to the real-time detected crack morphology, robot state and environmental conditions, to achieve accurate, efficient and high-quality crack repair. The adaptive spraying parameter control unit can dynamically adjust the spraying strategy according to the real-time detected crack characteristics, robot motion state and environmental changes, as well as the visual feedback after spraying, so as to overcome the limitations of static preset of spraying parameters in traditional methods and realize adaptive spraying.

[0083] wherein the spray gun posture includes pitch, yaw and roll angles within ±45°.

[0084] As a preferred embodiment, the reinforcement learning agent of the adaptive spraying parameter control unit can be initialized in combination with expert experience or traditional PID control strategy at the initial stage of training, so as to accelerate convergence and provide a stable initial spraying strategy.

[0085] Illustratively, a set of lookup table rules based on crack width and depth is preset, and a PID controller is used to preliminarily determine the spraying flow rate, pressure and speed; the reinforcement learning agent can be trained offline in a highly simulated bridge surface and spraying environment; the simulation environment is built using Unity3D or Gazebo physical engine, which accurately simulates the geometric characteristics of the bridge surface, the crack morphology, the rheological behavior of the paint, the working characteristics of the spraying equipment, and the imaging process of the visual sensor, to ensure the effectiveness of the offline training strategy; after obtaining the preliminary strategy model through offline training, a small amount of actual spraying test is performed for online fine-tuning and deployment, and the strategy is further optimized through interaction data with the real environment to adapt to unpredictable factors in actual spraying operations.

[0086] wherein the crack morphology includes length, width, depth and curvature; the rheological behavior of the paint includes viscosity and surface tension; the working characteristics include spray cone angle and spray force; and the imaging process includes illumination, shadow and reflection.

[0087] Further, the optimized spraying parameters and local trajectory adjustment instructions output by the adaptive spraying parameter control unit are sent to the spraying robot actuator via the spraying robot motion control system to drive the spraying gun to perform repair spraying operations.

[0088] The spraying robot executive mechanism comprises a multi-degree-of-freedom mechanical arm and a spraying gun integrated with a spraying nozzle. The mechanical arm is a six-axis industrial robot, such as a KUKA KR CYBERTECH nano series mechanical arm, which has the characteristics of high precision and high repeatability positioning accuracy, and can accurately execute a spraying trajectory. The spraying gun is an air atomizing spraying gun, the spraying pressure of which can be accurately adjusted within the range of 0.1-0.6 MPa through a proportional valve, the coating flow rate can be accurately adjusted within the range of 10-50 ml / min through a peristaltic pump or a metering pump, and the nozzle outlet diameter is 0.5-1.5 mm, which can be replaced or dynamically adjusted according to the width and depth of the crack to achieve the best coating coverage effect and spraying precision; for example, for a micro-crack with a width less than 1 mm, a 0.5 mm diameter nozzle can be selected for fine spraying at low pressure and low flow rate; for a structural crack with a width greater than 3 mm, a 1.2 mm or 1.5 mm diameter nozzle is selected, which is matched with a higher pressure and flow rate to ensure sufficient filling and coverage.

[0089] In a specific embodiment, in the step of S5, during the spraying operation, the spraying robot uses the multi-modal perception unit for real-time visual feedback and quality verification. The visual feedback is configured to collect and analyze images of the sprayed area again after the spraying is completed or during the spraying process. By comparing the image data before and after spraying, or using depth information to evaluate the coating thickness and uniformity, the spraying quality can be evaluated in real time, such as coverage completeness, coating uniformity, presence of sagging, orange peel or insufficient thickness, etc.

[0090] In the step of S5, the comparison of the image data before and after spraying uses image difference or feature matching technology; the evaluation of the coating thickness and uniformity is performed by comparing the depth value changes of the sprayed area and the unsprayed area.

[0091] Specifically, if the coverage rate is found to be less than 95 % or the coating thickness uniformity standard deviation is greater than 0.1 mm, i.e., the preset standard is not met, the quality evaluation result will be fed back to the reinforcement learning agent of the adaptive spraying parameter control unit as new state information. The reinforcement learning agent will adjust the strategy based on this feedback, generate a re-spraying instruction or correct the current spraying parameters until the spraying quality meets the preset requirements. Through the closed-loop feedback mechanism, the reliability of the spraying operation is ensured, and the need for manual intervention and rework is minimized.

[0092] Experimental group: This experimental group is used to verify the performance of the spraying robot crack identification and adaptive spraying method based on machine vision in Example 1 in actual bridge crack repair.

[0093] The application scenario is an experimental platform simulating a concrete bridge, the sprayed material is a polymer-modified mortar with a maximum aggregate particle size of ≤4.75 mm and an initial flow value of ≥290 mm, which is purchased from Sichuan Jin Hongyuan New Materials; the experimental platform is pre-fabricated with various forms of cracks; the characteristics of a typical crack are shown in Table 1.

[0094] Table 1: Characteristics of a typical crack

[0095] Experimental steps: 1. The multi-modal perception unit carried by the spraying robot scans the crack area along the preset path.

[0096] 2. The multi-modal data stream is processed in real time and enters the multi-modal feature transformation unit to generate a fusion feature pyramid.

[0097] 3. The feature extraction and segmentation unit performs pixel-level segmentation on the fusion feature pyramid and extracts the morphological parameters and three-dimensional geometric information of the crack.

[0098] 4. The crack parameters and the pose of the spraying robot are input into the adaptive spraying parameter control unit, and the reinforcement learning agent outputs the optimized spraying parameters (spraying pressure 0.35 MPa, coating flow 35 ml / min, spraying speed 120 mm / s, spraying gun pitch angle 5°, distance from surface 180 mm) and trajectory adjustment instructions based on the current state.

[0099] 5. The spraying robot execution mechanism performs the spraying operation, and after the spraying is completed, the multi-modal perception unit collects the sprayed area image again for quality evaluation, and the evaluation result is fed back to the reinforcement learning agent. If it does not meet the standard, it will be re-sprayed or the parameters will be fine-tuned.

[0100] Control group: This control group is used to show the performance of the traditional feature pyramid network (FPN) based spraying robot crack recognition and fixed parameter spraying method in the same experimental scenario.

[0101] Method description: This control group uses a crack segmentation network based on FPN (PANet), which only uses RGB color images as input and does not contain time series feature transformation and cross-modal attention fusion mechanism. After crack segmentation, the spraying parameters are preset according to the empirical rule (based on crack average width and depth lookup table determination), which does not have dynamic adaptive ability and no reinforcement learning feedback closed loop.

[0102] Hardware configuration: The same as the experimental group, but the depth camera and IMU data are not used for feature fusion in the crack identification stage, only for auxiliary robot positioning. The spray gun parameters are set to fixed values, with a spray pressure of 0.4 MPa, a coating flow rate of 40 ml / min, a spray speed of 150 mm / s, a spray gun pitch angle of 0°, and a distance from the surface of 200 mm.

[0103] The experimental results of the above experimental group and control group are shown in Table 2.

[0104] Table 2: Experimental results of the experimental group and the control group

[0105] From Table 2: From the above comparative data, it can be seen that the proposed method has achieved significant improvement in crack detection F1-score, reaching 0.965, which is 9.41% higher than the control group of 0.882; this is mainly due to the rich information provided by the multi-modal perception unit, the robustness of the dynamic time-series multi-modal feature transformation unit to feature drift, and the effective suppression of noise by the cross-modal spatial attention fusion module, especially under complex lighting and background interference, which can more accurately identify real cracks.

[0106] In terms of spray performance, the spray coverage of the proposed method is as high as 98.7%, far superior to the control group of 85.3%, with an increase of 15.71%; which directly reflects the excellent ability of the reinforcement learning agent in the adaptive spray parameter control unit to dynamically adjust the spray parameters according to the crack morphology and robot state; in terms of coating uniformity, the standard deviation of the coating thickness of the proposed method is only 0.08 mm, which is 61.90% lower than the control group of 0.21 mm, indicating that the proposed method can achieve more precise and uniform coating deposition, at the same time, the material utilization rate is increased from 75.8% to 92.1%, reducing 21.50% of material waste, showing the economic benefits of the invention.

[0107] Although the spray time of the proposed method is slightly longer than that of the control group, increasing the time cost by about 25%, mainly because the reinforcement learning agent needs to perform more iterations of optimization in the decision-making process, and may introduce minor trajectory adjustments and more precise parameter changes during the spraying process; however, considering the comprehensive requirements of crack repair tasks for quality and efficiency, the proposed method significantly improves key quality indicators such as crack identification accuracy, spray coverage, coating uniformity, and material utilization, while slightly increasing the time cost, which is completely acceptable, and from the long-term maintenance and repair effect, the comprehensive benefits it brings far exceed the time investment.

[0108] Example 2: Figure 4An electronic device structure diagram that can be used to implement the embodiment 1 of the present application is shown. The electronic device 10 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the present application described and / or claimed in this document.

[0109] As shown in Figure 4 The electronic device 10 includes at least one processor 11, and a memory, such as a Read-Only Memory (ROM) 12, a Random Access Memory (RAM) 13, etc., connected to the at least one processor 11 in communication, where the memory stores a computer program executable by the at least one processor 11, and the computer program is executed by the at least one processor 11 to enable the at least one processor 11 to perform the method provided by the present application.

[0110] Further, the processor 11 can perform various appropriate actions and processes according to the computer program stored in the Read-Only Memory (ROM) 12 or loaded into the Random Access Memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0111] Various components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunications networks.

[0112] Further, the processor 11 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), various specialized Artificial Intelligence (AI) computing chips, various processors running machine learning model algorithms, a Digital Signal Process (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, e.g., the method of resource management for a database.

[0113] In some specific embodiments, the method of resource management for a database can be implemented as a computer program tangibly embodied in a computer readable storage medium, e.g., the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded onto the RAM 13 and executed by the processor 11, one or more steps of the method of resource management for a database described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the method of resource management for a database by other any appropriate means, e.g., by means of firmware.

[0114] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a Field Programmable Gate Array (FPGA), an Application Specific Integrated Circuit (ASIC), an Application Specific Standard Product (ASSP), a System on Chip (SOC), a Complex Programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0115] Computer programs used to implement the present methods can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed by the processor of the machine, implements the functions / acts specified in the flow diagrams and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0116] In the context of the present application, the computer readable storage medium stores computer instructions for causing a processor to implement the method of resource management of a database provided by the present application when executed. The computer readable storage medium can be a tangible medium which can contain or store the computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, the computer readable storage medium can be a machine readable signal medium. More specific examples of the machine readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0117] To provide for interaction with a user, the systems and techniques described here can be implemented on an electronic device having a display device (e.g., a Cathode Ray Tube (CRT) or a Liquid Crystal Display (LCD) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0118] The systems and techniques described herein can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described herein, or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0119] Optionally, the computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. A server can be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS) services.

[0120] In light of the above, it should be appreciated that many modifications and variations to exemplary embodiments of the present application are possible in light of the above teachings. It is, therefore, to be understood that within the scope of the present application there is a full equivalency of all features between the claims and the specification. Moreover, any combination of the above-described elements in all possible variations thereof is encompassed by the application unless otherwise indicated herein or otherwise clearly contradicted by context.

[0121] In addition, it should be understood that, although the description herein is based upon embodiments, not every embodiment contains only one independent technical solution, and the description herein is only for the sake of clarity, and those skilled in the art should understand the description as a whole, and the technical solutions in each embodiment can be properly combined to form other embodiments that those skilled in the art can understand.

Claims

1. A machine vision-based crack identification and adaptive spraying method for a spraying robot, characterized in that, The method comprises the following steps: S1, collecting a multi-modal original data stream of a bridge structure surface to be detected in real time, and generating multi-modal data after time-space alignment, noise suppression, and feature enhancement; the multi-modal original data stream comprises high-resolution color images, high-precision depth images, and posture and motion information of a spraying robot; S2, from the multi-modal data, a multi-modal feature transformation unit is used to extract and transform a fusion feature pyramid with multi-scale time dynamic changes; the multi-modal feature transformation unit comprises a parallel feature extraction network, a group of intra-modal time sequence feature transformers, and a cross-modal spatial attention fusion module; S3, the fusion feature pyramid is segmented by a feature extraction and segmentation unit to identify and extract morphological parameters and position coordinates of a crack region in a three-dimensional space; S4, based on the morphological parameters, the position coordinates, and the current posture information of the spraying robot, optimized spraying parameters and local trajectory adjustment instructions are output to perform a repair spraying operation.

2. The machine vision-based crack identification and adaptive spraying method for a spraying robot according to claim 1, characterized in that: In the step S1, the time-space alignment of the multi-modal original data stream comprises: collecting each set of color images and depth images at the same time; aligning the depth images and the color images to the same coordinate system by using pre-calibrated sensor external parameters to generate pixel-level aligned RGB-D image pairs; and aligning the posture and motion information data of the spraying robot and the RGB-D image pairs by time stamp to generate uniformly sequenced data frames. 3.The machine vision based crack identification and adaptive spraying method of a spraying robot according to claim 1, characterized in that: In the step S1, the noise suppression of the multi-modal original data stream comprises: applying a non-local mean filter-based image denoising algorithm to the color images; applying a bilateral filter or median filter-based denoising algorithm to the depth images; and applying Kalman filtering or complementary filtering to the posture and motion information data of the spraying robot.

4. The machine vision based crack identification and adaptive painting method for a painting robot according to claim 1, characterized in that: In the step S1, the feature enhancement of the multi-modal original data stream comprises: applying adaptive histogram equalization to the color images to enhance image contrast and extract gradient information of the images; applying depth gradient calculation to the depth images to highlight height changes of the surface; and calculating instantaneous speed and acceleration vectors of the spraying robot based on the posture and motion information data of the spraying robot as context input for subsequent feature transformation.

5. The machine vision based crack identification and adaptive painting method for a painting robot according to claim 1, characterized in that: In the step S2, the parallel feature extraction network comprises independent feature extraction paths for color images and depth images, each path using a lightweight convolutional neural network as a backbone network to extract multi-scale initial feature maps from a single image frame; the group of intra-modal time sequence feature transformers comprises at least two parallel time sequence transformer modules corresponding to the color image modality and the depth image modality, each time sequence transformer module receiving a multi-frame, multi-scale feature map sequence of its corresponding modality from the parallel feature extraction network, and generating a multi-scale time sequence feature map set stable in the time dimension and robust to feature drift by capturing dynamic change patterns of crack features in the time dimension through an encoder structure based on a Transformer architecture, comprising a multi-head self-attention mechanism and a feedforward neural network.

6. The machine vision based crack identification and adaptive painting method of claim 5, wherein: In the multi-modal feature transformation unit, the pose and motion information of the spraying robot is further integrated into the input of the intra-modal temporal feature transformer group; Specifically, the pose and motion information of the spraying robot is processed by a one-dimensional convolutional neural network to extract its temporal features, And the temporal features are encoded as conditions, spliced with the feature vectors of the color image modal and the depth image modal, or fused through a gating mechanism, and input into the temporal transformer module to provide additional context information.

7. The machine vision based crack identification and adaptive painting method of claim 5, wherein: The cross-modal spatial attention fusion module is configured to receive the color image modal multi-scale feature map set and the depth image modal multi-scale feature map set from the intra-modal temporal feature transformer group, and perform the following operations for each feature pyramid level: The feature map of the color image modal and the feature map of the depth image modal are up-sampled or down-sampled to align their spatial resolutions; A query, key and value matrix is constructed, and a cross-modal spatial attention mechanism is adopted to calculate the similarity between the query and the key to assign attention weights to the value matrix, so as to realize feature fusion from different modalities and suppress the influence of structured noise through depth information; The cross-modal spatial attention fusion module outputs a set of multi-scale feature pyramids that fuse color and depth information and are accurately aligned in the spatial dimension and have noise suppression. 8.The machine vision based crack identification and adaptive painting method of a painting robot according to claim 1, wherein: In the step of S3, the feature extraction and segmentation unit adopts a segmentation network based on an encoder-decoder architecture, the encoder part of the segmentation network is composed of or connected to the fusion feature pyramid, and the decoder part receives the multi-scale fusion features output by the encoder and gradually restores the spatial resolution through up-sampling, skip connection and convolutional layer, and finally outputs a pixel-level crack probability map or a binary segmentation mask; Through morphological processing and geometric analysis algorithms, morphological parameters of the crack are accurately extracted, including the length, average width, maximum width, direction, connectivity, area, and position coordinates in three-dimensional space of the crack; By matching the segmented crack pixels with the corresponding depth image data, the point cloud representation of the crack in three-dimensional space is obtained, and based on the point cloud data, the actual three-dimensional shape, depth profile information, and surface roughness of the crack are calculated.

9. An electronic device, comprising: comprise: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the machine vision-based spraying robot crack identification and adaptive spraying method of any one of claims 1-8.

10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to implement the machine vision-based spraying robot crack identification and adaptive spraying method of any one of claims 1-8 when executed. The computer readable storage medium stores computer instructions for causing the processor to implement the machine vision-based spraying robot crack identification and adaptive spraying method of any one of claims 1-8 when executed.

Citation Information

Patent Citations

  • A pavement crack detection method based on convolution neural network and image recognition

    CN109146849A

  • Tunnel crack automatic repairing method, device and system and readable storage medium

    CN115182747A

  • Device and method for identifying and repairing underwater cracks of concrete structure

    CN118756632A

  • Bridge structure crack position identification method based on deep learning and computer vision

    CN119295660A

  • Bridge pier column crack automatic detection system based on wall-climbing robot

    CN119492741A