Bridge hidden part patrol method based on AI visual technology

The bridge hidden parts inspection system based on AI vision technology solves the problem that traditional manual inspection methods are difficult to fully cover the hidden parts of bridges. It realizes real-time monitoring of hidden parts of bridges and early warning of potential dangers, improves inspection efficiency and data credibility, and reduces labor costs.

CN120708143APending Publication Date: 2025-09-26HOHAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510805444.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Traditional manual inspection methods are difficult to fully cover the hidden parts of bridges, resulting in blind spots in detection, low efficiency and high risks, poor timeliness of data collection, and difficulty meeting the precise, efficient and intelligent management and control standards of modern bridges.

Method used

A bridge hidden part inspection system based on AI vision technology is used. Image information is collected through edge inspection terminals, and network control terminals are configured to perform target detection and event type classification. Event alarm discrimination is performed using AI analysis algorithms, and alarm information is transmitted through wireless communication modules. Deep learning algorithms are used to extract object features and classify events.

Benefits of technology

It realizes real-time monitoring of hidden parts of bridges and early warning of potential dangers, improves inspection efficiency and data credibility, reduces labor costs, and improves the safety of bridge operation and management and data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708143A_ABST
    Figure CN120708143A_ABST
Patent Text Reader

Abstract

The invention provides a bridge hidden part inspection method based on an AI vision technology, and aims to solve the problems of high labor cost, low efficiency, multiple detection blind areas and non-intelligent data processing of traditional inspection. According to the method, image information is collected through an edge patrol terminal, target detection and event type division are performed by using a data processing terminal and a convolutional neural network, and event alarm and recording are triggered. The method is characterized in that technologies such as a convolutional neural network, edge calculation and data fusion are integrated, and means such as lightweight model design, heterogeneous calculation scheduling, dynamic power consumption management and network transmission optimization are combined, so that the deployment efficiency, economy and environmental adaptability are remarkably improved while high-precision detection is ensured, and the method is suitable for large-scale popularization and application. The method is suitable for bridge hidden part patrol in a complex field environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of inspection management technology, and specifically to a management method for inspecting hidden parts of bridges, specifically a bridge hidden part inspection system based on AI vision technology. Background Art

[0002] With the continuous advancement of transportation infrastructure construction in my country, the number of bridges entering service is increasing. Among the many factors that ensure the safe operation of bridges, hidden safety hazards in bridge areas have become a major factor affecting bridge safety. To ensure the proper maintenance of bridges and the safe operation of their surroundings, bridge maintenance and management departments deploy relevant equipment or periodically dispatch personnel to inspect the condition of bridges and their surroundings. Bridge inspections typically involve on-site inspections by staff, who record the bridge's operating conditions and the impact on the surrounding environment. Bridge inspections are a crucial and fundamental task in ensuring safe bridge operation and are of great significance to bridge maintenance and management. Inspections of hidden areas present a technical challenge in bridge maintenance and management. Traditional visual inspection methods, limited by manpower, the environment, and equipment, struggle to fully cover blind spots such as piers and beam bottoms. Furthermore, data collection is inefficient and highly subjective. The collected data is highly time-sensitive, and failure to promptly report to management departments can impact bridge operations and management.

[0003] Manual bridge inspections are difficult to fully cover, including hidden areas. This creates blind spots in the bridge and surrounding areas, preventing the timely detection of hidden dangers. Furthermore, traditional manual inspection methods are inefficient, pose high risks when working at height, and result in blind spots, preventing the timely detection of hidden dangers. Visual inspections are influenced by personnel experience, and records of bridge operating conditions and surrounding conditions are prone to subjective errors. Paper records are difficult to digitize and archive, limiting data transmission efficiency and hindering subsequent analysis.

[0004] The existing inspection model for hidden areas of bridges is no longer adaptable to the sophisticated, efficient, and intelligent management standards of modern bridges. Furthermore, industry regulators are demanding greater transparency and traceability in inspection processes. To accurately capture key data such as inspector location information, task execution dynamics, and the bridge's surrounding environment, there is an urgent need to restructure the management and maintenance system based on digital technology. By integrating next-generation information technologies such as the Internet of Things (IoT) perception networks, mobile internet communication protocols, and GIS (Geographic Information Systems), a smart inspection platform featuring full-process traceability, intelligent data analysis, and multi-terminal collaborative decision-making will be constructed, systematically optimizing bridge operation and maintenance efficiency throughout its lifecycle.

[0005] In existing technologies, intelligent bridge hidden area inspection systems are typically equipped with inspection terminal devices. Their technical solutions primarily include inspection information collection modules, data storage units, and command transmission components. However, due to the existing technical architecture's heavy reliance on manual inspection operations, significant technical deficiencies remain in reducing inspection labor costs, improving operational efficiency, and implementing real-time abnormal condition detection. Therefore, there is an urgent need to design an electronic inspection system with integrated intelligent assistance functions. By combining standardized process control with intelligent analysis algorithms, this system can achieve the technical goals of standardizing inspection operations, optimizing efficiency, and enhancing data credibility. Summary of the Invention

[0006] The purpose of the present invention is to provide a cost-effective and easy-to-use method for inspecting hidden parts of bridges based on AI vision technology in order to solve the problems mentioned in the above background technology.

[0007] Convolutional neural networks are a type of feedforward neural network that includes convolutional calculations and has a deep structure. Convolutional neural networks are proposed based on the biological mechanism of receptive fields. Convolutional neural networks are specifically used to process data with a grid-like structure. For example, time series data (which can be considered a one-dimensional grid formed by regular sampling on the time axis) and image data (which can be considered a two-dimensional pixel grid). The convolutional neural network used in this embodiment processes image data;

[0008] The above-mentioned purpose of this application is achieved through the following technical solutions:

[0009] S1: Configure on-site equipment through the edge inspection terminal, collect image information, and configure the network control terminal;

[0010] S2: Processing the scene image information through the data processing terminal, performing target detection on the scene image, and classifying the object category information in the image;

[0011] S3: Through the data discrimination module, the event type of the target in the detected image is classified;

[0012] S4: Through AI analysis algorithms, the classified events are classified according to the preset event types for event alarm discrimination;

[0013] S5: trigger an event alarm, transmit the alarm information to the network control terminal, and record the alarm log information;

[0014] Optionally, step S1 includes:

[0015] S11: Access edge inspection terminals for real-time status monitoring, configure network control terminal parameters, check system and equipment operating status, and use the equipment to obtain real-time data from hidden areas under the bridge;

[0016] S12: Establishing communication between the data acquisition module and the edge inspection module through the device communication module via network signals;

[0017] S13: Monitor the device's operating status, obtain video stream data, and control the data acquisition device to a preset angle through the control algorithm pre-embedded in the device's communication module, convert the data acquisition viewing angle with one click, or adjust the data acquisition device's rotation angle as needed;

[0018] S14: The data acquisition module collects image information through the pre-connected camera and transmits the data to the data processing terminal in real time based on the communication module;

[0019] Optional step S11 includes:

[0020] S11a: Adjust the threshold parameters for triggering event alarms in the AI ​​analysis terminal. The specific formula is as follows:

[0021]

[0022] T(t) is the dynamic threshold, T0 is the basic threshold, α is the confidence feedback gain, the value range is 0.2~0.5, which is used to control the sensitivity of the confidence, and β is the environment-confidence coupling coefficient, the value range is 1.0~3.0, which is used to adjust the nonlinear relationship between confidence and environment.

[0023] is the confidence change gradient, y is used for the current frame detection confidence variance,

[0024]

[0025] The gradient of environmental factor changes represents the difference in environmental status.

[0026]

[0027] γ environment compensation coefficient, the value range is -0.1 to 0.1, negative value means lowering the threshold when the environment is bad,

[0028] ΔE(t) is the real-time environmental degradation degree,

[0029]

[0030] Where E(t) = w1·V t +w2·L t +w3·W t

[0031] V1 is the visibility index, which comes from the meteorological API, L t is the light intensity, from the light sensor, W t is the precipitation intensity, derived from radar data, w i is the weight parameter, with initial values ​​of w1=0.5, w2=0.3, w3=0.2;

[0032] S11b: Through the user rights management module, all user accounts using the system are managed in a unified manner, including user rights and accessible areas.

[0033] S11c: The log query module is used to display the alarm log information recorded by the AI ​​analysis terminal. The alarm log content information is queried according to time and event category. The event display and processing module is used to display the alarm information when the event is triggered, and provide event triggering time information and event triggering location information;

[0034] Optional step S13 includes:

[0035] S13a: The device status control function monitors the operating status of the data acquisition device and the quality of the data collected in real time. If any device operation failure or data collection anomaly is found, the fault information is uploaded through the device communication module and recorded in the log;

[0036] S13b: The video stream acquisition function intercepts the image and video stream information in the data according to the requirements corresponding to different functions of the data processing terminal to provide data pre-processing function for the data processing module;

[0037] Optionally, step S2 includes:

[0038] S21: The data processing terminal analyzes and processes the collected real-time data to obtain results, and the data storage terminal records the data processing results and adds timestamp information to the log;

[0039] S22: The data processing terminal internally deploys a data processing algorithm, wherein the data detection algorithm processes the real-time data and obtains a processing result;

[0040] S23: The data storage terminal stores different types of processing result information in different storage areas, and different log information is recorded in different storage units;

[0041] Optionally, step S3 includes:

[0042] S31: The data detection module performs data preprocessing based on the data detection model. Specifically, the feature extraction model network structure includes a bottom-up path and a top-down path, and uses horizontal connections to achieve multi-scale feature fusion processing of different targets in the data.

[0043] S32: The data in the bottom-up path passes through the backbone network layer by layer to extract data feature information. Specifically, the backbone network is composed of four groups of convolutional neural networks, each of which includes at least one convolution layer, an activation function, a pooling layer, and a normalization layer. The activation function is a relu function.

[0044] S33: After each set of convolution, a convolutional layer feature map is generated. The upper convolution output is connected to the lower convolution input data to generate feature information. The top-down path starts from the highest layer feature and gradually samples and restores the feature map of spatial resolution through upsampling.

[0045] P i =Upsample(P i+i );

[0046] Upsample(x)=Resize(x,scale=2);

[0047] P i is the feature map of the i-th layer,

[0048] Upsampling uses bilinear interpolation,

[0049]

[0050] w ij is the weighted distance, Q ij are neighboring pixels,

[0051] S34: lateral connection fuses the bottom-up path features with the upsampled feature channels of the top-down path;

[0052] S35: Classify the objects in the image, draw detection boxes, and divide the event types according to the recognition results;

[0053] Optionally, step S34 includes:

[0054] S34a: A skip connection is established between the network's first input and tail output. The skip connection includes a dimension adaptation module, which includes a 1x1 convolution.

[0055] S34b: Adjust the input data dimension information and output data through 1x1 convolution to make their dimensions consistent;

[0056] S34c: After the data is processed by the convolutional neural network, it is fused with the data after the dimension change.

[0057] P i fused =(Pi+C′ i )

[0058] Pi is the top-down feature map, C′i is a bottom-up feature map;

[0059] S34d: output after relu activation function;

[0060] Optional step S35 includes:

[0061] S35a: The feature input layer generates feature maps of different resolutions based on the feature map, and uses the generated feature map information as input information to the first convolutional layer to obtain a feature vector;

[0062] S35b: Input the feature vectors into the second convolutional layer to further integrate the context information, classify the objects in the original data based on the judgment of the classification layer, and determine the number of anchor boxes, bounding box coordinates and classification categories of each grid;

[0063] S35c: The processing layer independently activates the category probabilities contained in the prediction results of each anchor box through the sigmoid function. Specifically, the probability of each category is calculated independently, allowing the target to be labeled as multiple categories;

[0064] S35d: After activation, the overlapping prediction boxes are merged, only the prediction boxes with confidence higher than the threshold are retained and the category labels and probability information are displayed;

[0065] Optionally, step S4 includes:

[0066] S41: The data identification module classifies the processed data based on the screen category information returned by the data processing terminal, and classifies the processed data according to pre-set events;

[0067] S42: Event types are divided into four categories: personnel and vehicle intrusion, illegal stacking, illegal operation and parking, and fire and smoke. The data discrimination module obtains the event classification results, which are passed as input information to the data analysis module. The AI ​​analysis module determines whether the alarm condition is triggered according to the event corresponding discrimination mechanism;

[0068] Optionally, step S42 includes:

[0069] S42a: If the data processing terminal identifies a person or vehicle, the data identification module classifies the data into a person or vehicle intrusion event. If the person or vehicle remains in the image for more than a preset time, it is determined to be a person or vehicle intrusion event, triggering an event alarm. The on-site sound and light alarm device is activated, and the alarm information is transmitted to the network management terminal and recorded in the alarm log;

[0070] S42b: If the data processing terminal identifies a stack or illegal construction, it is determined to be an illegal stacking event. If the stack is not cleared after a preset time, an event alarm is triggered, an alarm log is recorded, and the alarm information is transmitted to the network management terminal;

[0071] S42c: If a special operation vehicle or operation machinery is identified, it is judged as an illegal operation or parking event. If the characteristic vehicle or machinery is continuously located in the image for more than a preset time, the on-site sound and light alarm device is triggered, the alarm log is recorded, and the alarm information is transmitted to the network management terminal in real time;

[0072] S42d: If fire and smoke are detected, it is determined to be a fire and smoke event, and the sound and light alarm device is triggered to drive away people and vehicles on the scene, transmit the alarm information to the network management terminal, and record the alarm log;

[0073] Optionally, step S5 includes:

[0074] S51: After the event alarm is triggered, the on-site detection screen information is displayed;

[0075] S52: After the AI ​​analysis terminal triggers an event alarm, an audible and visual alarm is immediately triggered on site. The device alarm function is implemented by the on-site audible and visual alarm device;

[0076] S53: Save the alarm information and record the alarm log.

[0077] An electronic device, characterized by an edge patrol unit, a data processing unit, a data storage unit and an AI analysis unit. The edge patrol unit is used to collect bridge site video and image data, monitor equipment operating status, build network communication and configure usage settings and parameters. The storage unit is used to store collected image data information, data processing results, alarm logs and equipment status abnormality logs. The wireless communication unit is used to communicate with other devices. The AI ​​analysis unit executes pre-set instruction operations according to a pre-set event plan, so that the electronic device executes the method described in any one of claims 1 to 8.

[0078] The technical solution provided by this application has the following beneficial effects:

[0079] Edge inspection equipment is used to detect and locate the acquired on-site images of the bridge area; the bridge area and its surrounding environment are monitored and analyzed to determine the operating status of the bridge and changes in its surrounding environment; image information of the surrounding environment of the bridge area and the conditions under the bridge columns and in hidden areas are recorded; by configuring device control parameters and setting the device observation angle, video stream data is acquired, and the acquired video stream data is pushed to the data processing unit with the help of the communication module. With the help of edge computing devices, the inherent defects of the traditional centralized cloud processing model in terms of real-time performance, reliability and environmental adaptability are solved through a distributed computing architecture. In outdoor areas with limited communication and network conditions, by deploying edge inspection equipment, it is possible to effectively identify and monitor the use of the areas under the bridge and in hidden parts. Based on the current operating status of the bridge, the data processing unit and the AI ​​analysis unit are combined to analyze potential dangerous conditions on site, respond in advance and take warning measures to avoid dangerous events.

[0080] The deep learning algorithm is used to take the video stream data obtained by the edge patrol unit as the data source. The learning mechanism of signal forward propagation and error reverse adjustment is used. Through multiple iterative learning, a data feature pyramid is constructed. The forward feature extraction is combined with the reverse data feature upsampling method to extract object features. After multiple iterative adaptations, a deep learning model is successfully built to provide further support for AI analysis.

[0081] The object data information obtained through data processing is combined with the object classification results to perform real-time event category conversion and determine the analysis results; combined with the location and area occupied by the object in the on-site image, the event category to which the object belongs is determined; based on the event classification information, the preset event processing plan is pointed to, and the AI ​​analysis unit independently completes the event alarm, processing, recording and feedback work. BRIEF DESCRIPTION OF THE DRAWINGS

[0082] Figure 1 Schematic diagram of the structure of the bridge hidden part inspection system provided in the embodiment of the present application

[0083] Figure 2 Flowchart of a method for inspecting hidden parts of a bridge shown in some embodiments of the present application

[0084] Figure 3 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Specific implementation methods

[0085] In order to have a clearer understanding of the technical features, purposes and effects of this application, the specific implementation methods of this application are now described in detail with reference to the accompanying drawings.

[0086] The implementation of this application provides a method for inspecting hidden parts of bridges based on AI vision technology.

[0087] Please refer to Figure 1 , Figure 1 This is a step diagram of a driver attention analysis method using eye tracking technology in an embodiment of the present application, including:

[0088] S1: Configure field equipment through edge inspection units, collect image information, and configure network control units;

[0089] S2: Processing the scene image information through the data processing unit, performing target detection on the scene image, and classifying the object category information in the image;

[0090] S3: Through the data discrimination module, the event type of the target in the detected image is classified;

[0091] S4: Through AI analysis algorithms, the classified events are classified according to the preset event types for event alarm discrimination;

[0092] S5: trigger an event alarm, transmit the alarm information to the network control terminal, and record the alarm log information;

[0093] Step S1 includes:

[0094] S11: Access edge inspection units for real-time status monitoring, configure network control unit parameters, check system and equipment operating status, and use the equipment to obtain real-time data from hidden areas under the bridge;

[0095] S12: Establishing communication between the data acquisition module and the edge inspection module through the device communication module via network signals;

[0096] S13: Monitor the device's operating status, obtain video stream data, and control the data acquisition device to a preset angle through the control algorithm pre-embedded in the device's communication module, convert the data acquisition viewing angle with one click, or adjust the data acquisition device's rotation angle as needed;

[0097] S14: The data acquisition module collects image information through the pre-connected camera and transmits the data to the data processing unit in real time based on the communication module;

[0098] Specifically, during the edge patrol unit deployment phase, the edge computing device is connected to the on-site data acquisition device, and the MQTT (Message Queue Telemetry Transport Protocol) protocol is used to realize communication between the edge and the cloud. The data packet is encapsulated in JSON format, including timestamp, device ID, and data type; the main link is the 5G network, and the backup link is the wireless network module, which is used to transmit key alarm information.

[0099] Step S11 includes:

[0100] S11a: Adjust the threshold parameters for triggering event alarms in the AI ​​analysis terminal. The specific formula is as follows:

[0101]

[0102] T(t) is the dynamic threshold, T0 is the basic threshold, α is the confidence feedback gain, the value range is 0.2~0.5, which is used to control the sensitivity of the confidence, and β is the environment-confidence coupling coefficient, the value range is 1.0~3.0, which is used to adjust the nonlinear relationship between confidence and environment.

[0103] is the confidence change gradient, y is used for the current frame detection confidence variance,

[0104]

[0105] The gradient of environmental factor changes represents the difference in environmental status.

[0106]

[0107] γ environment compensation coefficient, the value range is -0.1 to 0.1, negative value means lowering the threshold when the environment is bad,

[0108] ΔE(t) is the real-time environmental degradation degree,

[0109]

[0110] Where E(t) = w1·V t +w2·L t +w3·W t

[0111] V1 is the visibility index, which comes from the meteorological API, L t is the light intensity, from the light sensor, W t is the precipitation intensity, derived from radar data, w i is the weight parameter, with initial values ​​of w1=0.5, w2=0.3, w3=0.2;

[0112] S11b: Through the user rights management module, all user accounts using the system are managed in a unified manner, including user rights and accessible areas.

[0113] S11c: The log query module is used to display the alarm log information recorded by the AI ​​analysis unit, and the alarm log content information is queried according to time and event category. The event display and processing module is used to display the alarm information when the event is triggered, and provide event triggering time information and event triggering location information;

[0114] Step S13 includes:

[0115] S13a: The device status control function monitors the operating status of the data acquisition device and the quality of the data collected in real time. If any device operation failure or data collection anomaly is found, the fault information is uploaded through the device communication module and recorded in the log;

[0116] S13b: The video stream acquisition function intercepts the image and video stream information in the data according to the requirements corresponding to different functions of the data processing unit to provide data preprocessing function for the data processing module;

[0117] Specifically, a PID (proportional, integral, and differential) control algorithm is deployed in the edge computing device, 6 commonly used acquisition angles are preset, and one-click switching is supported; the gimbal is remotely controlled through a network control terminal, and the camera pitch and yaw angle data are fed back in real time.

[0118] Specifically, the video stream bit rate is dynamically adjusted according to the network bandwidth; one key image frame is extracted every 10 seconds for static scenes, and full frame rate recording is started for dynamic scenes and marked as "high priority data".

[0119] Specifically, the data acquisition equipment will collect real-time images and video information, mainly from surveillance cameras under the bridge and processed by the data processing unit. This part of the data collected mainly collects on-site images and analyzes the types of objects in the images.

[0120] Step S2 includes:

[0121] S21: The data processing unit analyzes and processes the collected real-time data to obtain a result, and the data storage unit is used to record the data processing result and add timestamp information to the log;

[0122] S22: The data processing unit internally deploys a data processing algorithm, wherein the data detection algorithm processes the real-time data and obtains a processing result;

[0123] S23: The data storage unit stores different types of processing result information in different storage areas, and different log information is recorded in different storage units;

[0124] Specifically, the data storage unit is a computer-readable medium, which is divided into three storage subspaces. The data storage subspaces are divided into first, second and third data storage subspaces. The data storage schemes include a first storage scheme, a second storage scheme and a third storage scheme. Here, the data processing unit only uses the first and second data storage schemes. The first data storage subspace stores static scene data, and the second storage subspace stores dynamic scene data. Because dynamic storage data can better reflect the occurrence of events and the urgency of accidents than static data, dynamic storage data is regarded as "high-priority data" and the second storage scheme is adopted for it. The second storage scheme is to collect data in real time, add events and occurrence location tags, save complete video data and record log information. The third storage subspace is used to store event alarm log records of the processing results of the AI ​​analysis unit. The log records include the location and time of the event, the event alarm trigger time, and the event response time.

[0125] Optionally, step S3 includes:

[0126] S31: The data detection module performs data preprocessing based on the data detection model. Specifically, the feature extraction model network structure includes a bottom-up path and a top-down path, and uses horizontal connections to achieve multi-scale feature fusion processing of different targets in the data.

[0127] S32: The data in the bottom-up path passes through the backbone network layer by layer to extract data feature information. Specifically, the backbone network is composed of four groups of convolutional neural networks, each of which includes at least one convolution layer, an activation function, a pooling layer, and a normalization layer. The activation function is a relu function.

[0128] S33: After each set of convolutions, a convolutional layer feature map is generated. The upper convolution output is connected to the lower convolution input data to generate feature information. The top-down path starts from the highest layer feature and gradually samples and restores the feature map of spatial resolution through upsampling.

[0129] P i =Upsample(P i+i );

[0130] Upsample(x)=Resize(x,scale=2);

[0131] P i is the feature map of the i-th layer,

[0132] Bilinear interpolation is used above.

[0133]

[0134] w ij is the weighted distance, Qij are neighboring pixels,

[0135] S34: Horizontal connection fuses the bottom-up path features with the upsampled features of the top-down path;

[0136] S35: Classify the objects in the image, draw detection boxes, and divide the event types according to the recognition results;

[0137] Optionally, step S34 includes:

[0138] S34a: A skip connection is established between the network's first input and tail output. The skip connection includes a dimension adaptation module, which includes a 1x1 convolution.

[0139] S34b: Adjust the input data dimension information to match the output data dimension through 1x1 convolution;

[0140] S34c: After the data is processed by the convolutional neural network, it is fused with the data after the dimension change.

[0141] P i fused =(Pi+C′ i )

[0142] Pi is the top-down feature map, C′ i is a bottom-up feature map;

[0143] S34d: output after relu activation function;

[0144] Specifically, the lateral connection fuses the bottom-up path's low-level features (high resolution) with the top-down path's high-level features (strong semantics). First, a 1×1 convolution operation is performed on the low-level features, adjusting their number of channels to align with the high-level features, thus aligning the top-down and bottom-up feature information channels. The upsampled high-level features are then element-wise added to the aligned low-level features to complete feature fusion. The fused feature map is then smoothed using a 3×3 convolution to eliminate the phenomenon where high-frequency components of the signal are mistakenly interpreted as low-frequency components due to insufficient sampling frequency or improper processing during the upsampling process. This process effectively integrates feature information from different levels.

[0145] Step S35 includes:

[0146] S35a: The feature input layer generates feature maps of different resolutions based on the feature map, and uses the generated feature map information as input information to the first convolutional layer to obtain a feature vector;

[0147] S35b: Input the feature vectors into the second convolutional layer to further integrate the context information, classify the objects in the original data based on the judgment of the classification layer, and determine the number of anchor boxes, bounding box coordinates and classification categories of each grid;

[0148] S35c: The processing layer independently activates the category probabilities contained in the prediction results of each anchor box through the sigmoid function, and calculates the probability of each category independently, allowing the target to be marked as multiple categories;

[0149] S35d: After activation, the overlapping prediction boxes are merged, only the prediction boxes with confidence higher than the threshold are retained and the category labels and probability information are displayed;

[0150] Specifically, in order to generate the anchor box prediction for each grid, the feature map is first input, and then the 3×3 convolution of the first convolutional layer is performed to adjust the number of channels and fuse local features to obtain the basic prediction features. In the second convolutional layer, a further 3×3 convolution is performed, and the hollow convolution is used to expand the receptive field to fuse context information. The final prediction parameters are generated through 1×1 convolution. Each anchor box corresponds to a preset number of categories, including bounding box offset, confidence, and multiple category probabilities.

[0151] Specifically, through independent category activation, the sigmoid function is applied to each of the identified categories, and the probability of each category is calculated independently, allowing multi-label classification; the confidence is activated by sigmoid, and the bounding box coordinates are decoded based on the transformation of the anchor box. Finally, the predicted value of each anchor box is decoded into the format of [x_center, y_center, width, height, confidence, class_1, ..., class_C]. Specifically, this is the format of a target detection prediction box, which contains the x and y coordinates of the detection box center point, the width of the detection box, the confidence that the target object is contained in the detection box, and the probability value of each category;

[0152] Specifically, through non-maximum suppression and filtering, the prediction results of all anchor boxes are input, and low-confidence predictions are filtered out according to the confidence threshold. Then, the predictions are grouped by category and non-maximum suppression is applied to each category to merge overlapping detection boxes. The final output is a list of detection results, where each detection box contains the information [x_min, y_min, x_max, y_max, class_id, confidence], that is, a list containing the coordinates of the detection box corners, category, and confidence information.

[0153] Step S4 includes:

[0154] S41: the data identification module classifies the image category information returned by the data processing unit and classifies the processed data according to pre-set events;

[0155] S42: Event types are divided into four categories: personnel and vehicle intrusion, illegal stacking, illegal operation and parking, and fire and smoke. The data discrimination module obtains the event classification results, which are passed as input information to the data analysis module. The AI ​​analysis module determines whether the alarm condition is triggered according to the event corresponding discrimination mechanism;

[0156] Step S42 includes:

[0157] S42a: If the data processing unit identifies a person or vehicle, the data identification module classifies the data into a person or vehicle intrusion event. If the person or vehicle remains in the image for more than a preset time, it is determined to be a person or vehicle intrusion event, triggering an event alarm. The on-site sound and light alarm device is activated, and the alarm information is transmitted to the network management unit and recorded in the alarm log;

[0158] S42b: If the data processing unit identifies a stack or illegal construction, it is determined to be an illegal stacking event. If the stack is not cleared after a preset time, an event alarm is triggered, an alarm log is recorded, and the alarm information is transmitted to the network management unit;

[0159] S42c: If a special operation vehicle or operation machinery is identified, it is determined to be an illegal operation or parking event. If the characteristic vehicle or machinery is continuously located in the image for more than a preset time, the on-site sound and light alarm device is triggered, the alarm log is recorded, and the alarm information is transmitted to the network management unit in real time;

[0160] S42d: If fire and smoke are detected, it is determined to be a fire and smoke event, and the sound and light alarm device is triggered to drive away personnel and vehicles on the scene, transmit the alarm information to the network management unit, and record the alarm log;

[0161] Step S5 includes:

[0162] S51: After the event alarm is triggered, the on-site detection screen information is displayed;

[0163] S52: After the AI ​​analysis unit triggers the event alarm, an audible and visual alarm is immediately triggered on site. The device alarm function is implemented by the on-site audible and visual alarm device;

[0164] S53: Save the alarm information and record the alarm log.

[0165] End: The system has finished running

[0166] This system has achieved significant enhancements in the accuracy and response speed of incident and accident warnings during bridge operation and maintenance, especially in the discovery, response, and handling of safety hazards in the surrounding areas of hidden parts of bridges, showing higher stability and reliability. Through the coordinated optimization of algorithm architecture and computing processes, this system has significantly improved deployment efficiency and economy while maintaining high-precision detection performance. It adopts a lightweight model design, performs edge computing adaptation optimization, adopts a heterogeneous computing scheduling strategy of CPU and NPU collaborative reasoning, customizes operator acceleration solutions based on the hardware characteristics of edge devices, implements dynamic power consumption management, and adjusts the equipment operating status in real time according to the detection task load. Overall, the present invention has created a novel, efficient and economical solution aimed at reducing bridge operating costs and improving the work efficiency of inspectors. Its unique technical advantages and broad application potential indicate that it will play a vital role in the field of bridge safety in the future.

[0167] It should be made clear that the device shown in the aforementioned implementation case only divides and explains the various functional modules as examples when realizing its functions. In actual application scenarios, it can be flexibly adjusted according to actual needs, and the functions can be assigned to different functional modules for completion, that is, the internal structure of the device can be divided into different functional modules according to actual needs to realize all or part of the functions described above. In addition, the device in the aforementioned implementation case and the method case for its implementation are based on the same design concept. For its specific implementation details, please refer to the method implementation case, which will not be repeated here.

Claims

1. The method for inspecting hidden parts of bridges based on AI vision technology is characterized by: A method for inspecting hidden bridge parts based on AI vision technology is used. The system includes an edge inspection unit, a data processing unit, a data storage unit, an AI analysis unit, and a network management unit. The method includes the following steps: S1: Configure field equipment through edge inspection units, collect image information, and configure network control units; S2: Processing the scene image information through the data processing unit, performing target detection on the scene image, and classifying the object category information in the image; S3: Through the data discrimination module, the event type of the target in the detected image is classified; S4: Through AI analysis algorithms, the classified events are classified according to the preset event types for event alarm discrimination; S5: trigger an event alarm, transmit the alarm information to the network control terminal, and record the alarm log information; 2. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 1, characterized in that: The edge inspection unit includes a data acquisition module, a device communication module, and an edge inspection module; step S1 includes: S11: Access edge inspection units for real-time status monitoring, configure network control unit parameters, check system and equipment operating status, and use the equipment to obtain real-time data from hidden areas under the bridge; S12: Establishing communication between the data acquisition module and the edge inspection module through the device communication module via network signals; S13: Monitor the device's operating status, obtain video stream data, and control the data acquisition device to a preset angle through the control algorithm pre-embedded in the device's communication module, convert the data acquisition viewing angle with one click, or adjust the data acquisition device's rotation angle as needed; S14: The data acquisition module collects image information through the pre-connected camera and transmits the data to the data processing unit in real time based on the communication module; 3. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 1, which involves an equipment management and status monitoring module, is characterized in that: Step S11 includes: S11a: Adjust the threshold parameters for triggering event alarms in the AI ​​analysis unit. The specific formula is as follows: T(t) is the dynamic threshold, T0 is the basic threshold, α is the confidence feedback gain, the value range is 0.2~0.5, which is used to control the sensitivity of the confidence, and β is the environment-confidence coupling coefficient, the value range is 1.0~3.0, which is used to adjust the nonlinear relationship between confidence and environment. is the confidence change gradient, y is used for the current frame detection confidence variance, The gradient of environmental factor changes represents the difference in environmental status. γ environment compensation coefficient, the value range is -0.1 to 0.1, negative value means lowering the threshold when the environment is bad, ΔE(t) is the real-time environmental degradation degree, Where E(t) = w1·V t +w2·L t +w3·W t V1 is the visibility index, which comes from the meteorological API, L t is the light intensity, from the light sensor, W t is the precipitation intensity, derived from radar data, w i is the weight parameter, with initial values ​​of w1=0.5, w2=0.3, w3=0.2; S11b: Through the user rights management module, all user accounts using the system are managed in a unified manner, including user rights and accessible areas. S11c: The log query module is used to display the alarm log information recorded by the AI ​​analysis unit, and the alarm log content information is queried according to time and event category. The event display and processing module is used to display the alarm information when the event is triggered, and provide event triggering time information and event triggering location information; 4. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 1, characterized in that: The edge inspection unit has the functions of device status control, video stream acquisition, and device alarm; step S13 includes: S13a: The device status control function monitors the operating status of the data acquisition device and the quality of the data collected in real time. If any device operation failure or data collection anomaly is found, the fault information is uploaded through the device communication module and recorded in the log; S13b: The video stream acquisition function intercepts the image and video stream information in the data according to the requirements corresponding to different functions of the data processing unit to provide data preprocessing function for the data processing module; 5. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 2, characterized in that: Step S2 includes: S21: The data processing unit analyzes and processes the collected real-time data to obtain a result, and the data storage unit is used to record the data processing result and add timestamp information to the log; S22: The data processing unit internally deploys a data processing algorithm, wherein the data detection algorithm processes the real-time data and obtains a processing result; S23: The data storage unit stores different types of processing result information in different storage areas, and different log information is recorded in different storage units; 6. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 3, wherein the data processing unit comprises a feature extraction module, a data classification module, and a data processing module, wherein: Step S3 includes: S31: The data detection module performs data preprocessing based on the data detection model. Specifically, the feature extraction model network structure includes a bottom-up path and a top-down path, and uses horizontal connections to achieve multi-scale feature fusion processing of different targets in the data. S32: The data in the bottom-up path passes through the backbone network layer by layer to extract data feature information. Specifically, the backbone network is composed of four groups of convolutional neural networks, each of which includes at least one convolution layer, an activation function, a pooling layer, and a normalization layer. The activation function is a relu function. S33: After each set of convolutions, a convolutional layer feature map is generated. The upper convolution output is connected to the lower convolution input data to generate feature information. The top-down path starts from the highest layer feature and gradually samples and restores the feature map of spatial resolution through upsampling. P i =Upsample(P i+i ); Upsample(x)=Resize(x,scale=2); P i is the feature map of the i-th layer, Upsampling uses bilinear interpolation, w ij is the weighted distance, Q ij are neighboring pixels, S34: Horizontal connection fuses the bottom-up path features with the upsampled features of the top-down path; S35: Classify the objects in the image, draw detection boxes, and divide the event types according to the recognition results; 7. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 3, characterized in that: Step S34 includes: S34a: A skip connection is established between the network's first input and tail output. The skip connection includes a dimension adaptation module, which includes a 1x1 convolution. S34b: Adjust the input data dimension information to match the output data dimension through 1x1 convolution; S34c: After the data is processed by the convolutional neural network, it is fused with the data after the dimension change. Pi is the top-down feature map, C′ i is a bottom-up feature map; S34d: The data fusion process adds the main path output and the skip connection processing result, and then passes through the relu activation function; 8. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 3, wherein the data classification module comprises a feature input layer, a first convolution layer, a second convolution layer, a classification layer, and a post-processing layer, characterized in that: Step S35 includes: S35a: The feature input layer generates feature maps of different resolutions based on the feature map, and uses the generated feature map information as input information to the first convolutional layer to obtain a feature vector; S35b: Input the feature vectors into the second convolutional layer to further integrate the context information, classify the objects in the original data based on the judgment of the classification layer, and determine the number of anchor boxes, bounding box coordinates and classification categories of each grid; S35c: The processing layer independently activates the category probabilities contained in the prediction results of each anchor box through the sigmoid function. Specifically, the probability of each category is calculated independently, allowing the target to be labeled as multiple categories; S35d: After activation, the overlapping prediction boxes are merged, only the prediction boxes with confidence higher than the threshold are retained and the category labels and probability information are displayed; 9. The method for inspecting hidden parts of bridges based on AI vision technology according to claim 3, wherein the AI ​​analysis unit comprises a data discrimination module and an AI analysis module, wherein step S4 comprises: S41: the data identification module classifies the image category information returned by the data processing unit and classifies the processed data according to pre-set events; S42: Event types are divided into four categories: personnel and vehicle intrusion, illegal stacking, illegal operation and parking, and fire and smoke. The data discrimination module obtains the event classification results, which are passed as input information to the data analysis module. The AI ​​analysis module determines whether the alarm condition is triggered according to the event corresponding discrimination mechanism; 10. The method for inspecting hidden parts of a bridge based on AI vision technology according to claim 4, wherein step S42 comprises: S42a: If the data processing unit identifies a person or vehicle, the data identification module classifies the data into a person or vehicle intrusion event. If the person or vehicle remains in the image for more than a preset time, it is determined to be a person or vehicle intrusion event, triggering an event alarm. The on-site sound and light alarm device is activated, and the alarm information is transmitted to the network management unit and recorded in the alarm log; S42b: If the data processing unit identifies a stack or illegal construction, it is determined to be an illegal stacking event. If the stack is not cleared after a preset time, an event alarm is triggered, an alarm log is recorded, and the alarm information is transmitted to the network management unit; S42c: If a special operation vehicle or operation machinery is identified, it is determined to be an illegal operation or parking event. If the characteristic vehicle or machinery is continuously located in the image for more than a preset time, the on-site sound and light alarm device is triggered, the alarm log is recorded, and the alarm information is transmitted to the network management unit in real time; S42d: If fire and smoke are detected, it is determined to be a fire and smoke event, and the sound and light alarm device is triggered to drive away personnel and vehicles on the scene, transmit the alarm information to the network management unit, and record the alarm log; 11. The method for inspecting hidden parts of bridges based on AI visual technology according to claim 3, wherein the network management unit includes a device management and status monitoring module, an AI model parameter configuration module, a user authority management module, a log query module, and an event alarm display module, characterized in that: Step S5 includes: S51: After the event alarm is triggered, the on-site detection screen information is displayed; S52: After the AI ​​analysis unit triggers the event alarm, an audible and visual alarm is immediately triggered on site. The device alarm function is implemented by the on-site audible and visual alarm device; S53: Save the alarm information and record the alarm log.