Highway intelligent inspection multi-terminal cooperative management method and system based on edge computing

By employing a multi-terminal collaborative management approach that combines edge computing and 5G communication, the problems of high video stream latency and insufficient terminal collaboration in existing technologies have been solved. This enables real-time anomaly identification and precise collaborative tracking in the intelligent highway inspection system, thereby improving the accuracy of disease assessment.

CN120499349BActive Publication Date: 2025-11-21中交资产管理有限公司 +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510978047.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-21
Estimated Expiration
2045-07-16

AI Technical Summary

Technical Problem

In existing intelligent highway inspection systems, the high latency of video streams collected by terminals and processed in the cloud cannot meet the requirements for real-time anomaly identification. Fixed inspection cameras and mobile inspection vehicles do not achieve dynamic coordination, resulting in incomplete spatial feature collection of abnormal objects and affecting the accuracy of disease level assessment.

Method used

A multi-terminal collaborative management method based on edge computing is adopted. The inspection video stream is uploaded to the edge computing node through 5G communication for dynamic task segmentation and distributed video frame processing. Anomalies are identified by a pre-trained highway inspection object recognition model, and multi-angle collaborative tracking is triggered by the associated inspection terminal group through the 5G-MQTT communication network.

Benefits of technology

It reduces processing latency, improves the real-time performance and accuracy of anomaly identification, enables dynamic collaborative tracking across multiple terminals, provides efficient data support, and offers precise data support for highway maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499349B_ABST
    Figure CN120499349B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on edge computing's highway intelligent inspection multi-terminal collaborative management method and system, it is related to artificial intelligence field, including: through 5G communication, the inspection video stream of multiple inspection terminals along highway is uploaded to edge computing node;Edge computing node cluster carries out dynamic task fragmentation to video stream, generates distributed video frame processing queue;Extract current inspection video frame, obtain first abnormal object feature and abnormal recognition result using pre-trained highway inspection object recognition model;Determine associated inspection terminal group according to spatial position information in abnormal result, issue collaborative tracking instruction through 5G-MQTT communication network, trigger associated terminal group to execute multi-angle collaborative tracking and feedback tracking video stream.The application utilizes edge computing to reduce processing delay, combined with 5G-MQTT to realize multi-terminal dynamic collaboration, improve the real-time performance, accuracy and collaborative tracking efficiency of abnormal identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and more specifically, to a multi-terminal collaborative management method and system for intelligent highway inspection based on edge computing. Background Technology

[0002] Existing intelligent highway inspection systems mostly adopt a terminal-to-cloud processing mode, which results in high latency in video stream transmission to the cloud, failing to meet the real-time identification requirements for anomalies such as road surface cracks and obstacles. Furthermore, fixed inspection cameras and mobile inspection vehicles do not achieve dynamic coordination. After a fixed terminal detects an anomaly, the mobile terminal cannot obtain accurate location information in a timely manner for tracking, leading to incomplete spatial feature collection of the abnormal object and affecting the accuracy of the damage level assessment. Summary of the Invention

[0003] The purpose of this invention is to provide a multi-terminal collaborative management method and system for intelligent highway inspection based on edge computing.

[0004] In a first aspect, embodiments of the present invention provide a multi-terminal collaborative management method for intelligent highway inspection based on edge computing, comprising:

[0005] Based on multiple inspection terminals deployed along the highway, inspection video streams are uploaded in real time to the corresponding edge computing nodes via 5G communication channels.

[0006] The received inspection video stream is dynamically segmented based on an edge computing node cluster to generate a distributed video frame processing queue.

[0007] Extract the current inspection video frame from the distributed video frame processing queue. The current inspection video frame is the video frame data that needs to be identified as an abnormal inspection object.

[0008] The current inspection video frame is loaded into a pre-trained highway inspection object recognition model. The current inspection video frame is then identified through the original input unit of the highway inspection object recognition model to obtain the first abnormal object feature of the current inspection video frame.

[0009] Based on the first abnormal object feature of the current inspection video frame, the abnormal inspection object identification result of the current inspection video frame is obtained;

[0010] Based on the spatial location information in the abnormal inspection object identification results, the associated inspection terminal group is determined, and a collaborative tracking instruction is issued to the associated inspection terminal group through the 5G-MQTT communication network established between edge computing nodes;

[0011] Based on the collaborative tracking instruction, the associated inspection terminal group is triggered to perform multi-angle collaborative tracking, and the generated tracking video stream is fed back to the edge computing node cluster.

[0012] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being used to execute the method described in the first aspect.

[0013] Compared to existing technologies, the beneficial effects of this invention include: employing a multi-terminal collaborative management method and system for intelligent highway inspection based on edge computing, as disclosed in this invention, the inspection video streams of multiple inspection terminals along the highway are uploaded to edge computing nodes via 5G communication; the edge computing node cluster dynamically segments the video streams to generate a distributed video frame processing queue; the current inspection video frame is extracted, and a pre-trained highway inspection object recognition model (original input unit) is used to identify the first abnormal object features and the abnormality recognition result; based on the spatial location information in the abnormality result, the associated inspection terminal group is determined, and a collaborative tracking command is issued via the 5G-MQTT communication network, triggering the associated terminal group to execute multi-angle collaborative tracking and feedback the tracking video stream. This invention utilizes edge computing to reduce processing latency and combines 5G-MQTT to achieve dynamic multi-terminal collaboration, improving the real-time performance, accuracy, and collaborative tracking efficiency of abnormality recognition, providing efficient data support for highway maintenance decisions. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating the steps of the multi-terminal collaborative management method for intelligent highway inspection based on edge computing provided in an embodiment of the present invention;

[0016] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0018] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0019] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the multi-terminal collaborative management method for intelligent highway inspection based on edge computing provided in this embodiment. The following is a detailed description of this multi-terminal collaborative management method for intelligent highway inspection based on edge computing.

[0020] Step S201: Based on multiple inspection terminals deployed along the highway, the inspection video stream is uploaded to the corresponding edge computing node in real time through the 5G communication channel;

[0021] Step S202: Dynamically segment the received inspection video stream based on the edge computing node cluster to generate a distributed video frame processing queue.

[0022] Step S203: Extract the current inspection video frame from the distributed video frame processing queue. The current inspection video frame is the video frame data that needs to be identified as an abnormal inspection object.

[0023] Step S204: Load the current inspection video frame into the pre-trained highway inspection object recognition model, and use the original input unit of the highway inspection object recognition model to identify the current inspection video frame to obtain the first abnormal object feature of the current inspection video frame.

[0024] Step S205: Based on the first abnormal object feature of the current inspection video frame, obtain the abnormal inspection object identification result of the current inspection video frame;

[0025] Step S206: Based on the spatial location information in the abnormal inspection object identification result, determine the associated inspection terminal group, and send a collaborative tracking instruction to the associated inspection terminal group through the 5G-MQTT communication network established between edge computing nodes;

[0026] Step S207: Based on the collaborative tracking instruction, the associated inspection terminal group is triggered to perform multi-angle collaborative tracking, and the generated tracking video stream is fed back to the edge computing node cluster.

[0027] In this embodiment of the invention, for example, the server serves as the core execution entity of the multi-terminal collaborative management method for intelligent highway inspection based on edge computing. It relies on multiple types of inspection terminals deployed along a two-way four-lane highway section (one fixed high-definition inspection camera deployed every kilometer along the roadside, totaling four units, numbered terminals 1 to 4; two mobile inspection vehicles equipped with GPS and 4K cameras are deployed in the middle of the section, numbered terminals 5 and 6), an edge computing node cluster (nodes A, B, and C), and a 5G communication network to achieve closed-loop management of the entire process from inspection video acquisition and edge intelligent processing to multi-terminal collaborative tracking. All terminals establish low-latency connections (latency ≤ 50ms) with their corresponding edge computing nodes (e.g., terminals 1-4 are associated with node A, and terminals 5-6 are associated with node B) via 5G communication modules to ensure real-time transmission of video streams. One day, Terminal 1 (a fixed camera deployed near the start of the road segment) begins routine inspections, uploading a real-time 1080P resolution video stream at 25 frames per second, containing a precise timestamp, a unique terminal identifier, and current GPS coordinates. Terminal 5 (a mobile inspection vehicle) travels at a constant speed along the main lane, uploading a 4K resolution video stream of the road ahead at 30 frames per second, clearly capturing road details. Terminal 6 (the mobile inspection vehicle) travels along the opposite lane, uploading a 1080P resolution video stream of the road surface to the side at 25 frames per second, covering a wider field of view. The server receives the video streams from all terminals through edge computing nodes A and B, and categorizes and stores them according to terminal identifiers and timestamps, laying the foundation for subsequent processing.

[0028] Subsequently, the server monitors the processing load status of the edge computing node cluster in real time: Node A currently has a CPU utilization of 70% and 35% memory remaining (the preset load threshold is 80%, which is within the processing capacity); Node B has a CPU utilization of 85% (exceeding the threshold, so no new tasks will be assigned); and Node C has a CPU utilization of 60% and 40% memory remaining (sufficient processing capacity). For the video stream uploaded by Terminal 1, the server, combining node load and video stream frame rate (25 frames / second), determines an adaptive slicing strategy of "every 10 frames as a task slice," ensuring node processing efficiency while avoiding overhead caused by excessively small slices. The server divides the video stream into multiple task slices, such as Slice 1 (frames 1-10), Slice 2 (frames 11-20), and Slice 3 (frames 21-30). Slice 1 is then assigned to Node A (lower load, prioritized for processing), Slice 2 is assigned to Node C (stronger processing capacity, higher priority than Slice 1), and Slice 3 is temporarily stored in a distributed video frame processing queue (the queue dynamically adjusts task priorities according to the real-time processing capacity of the nodes to ensure that high-load nodes are not over-assigned). After receiving the video fragments, nodes A and C immediately initiate the video frame decoding and preprocessing process (including grayscale conversion, noise reduction, and size normalization) to prepare for subsequent anomaly identification.

[0029] When node C processes segment 2 (frames 11-20), the server detects a significant grayscale anomaly in the right lane of the road surface in frame 15 (timestamp: 2024-XX-XXXX:XX:XX.XXX) using the frame difference preprocessing module. Compared to the previous frame (frame 14), the grayscale value change in this area exceeds a preset threshold of 30, indicating that the frame contains a suspected abnormal object (preliminarily speculated to be a road surface crack). The server extracts frame 15 from the distributed video frame processing queue, marks it as "current inspection video frame - frame 15", and pushes it to the input queue of the pre-trained highway inspection object recognition model to initiate the anomaly recognition process.

[0030] The highway inspection object recognition model is a lightweight deep learning model built on MobileNetV3 on the server, aiming to balance recognition accuracy and edge computing efficiency. It includes a raw input unit and an enhanced input unit (only the raw input unit is used for preliminary recognition here). The raw input unit consists of a cascaded lightweight feature extraction component and a location mapping component: the lightweight feature extraction component extracts features from frame 15 through the Conv2D layer and DepthwiseConv layer of MobileNetV3, and outputs the first edge feature map (size: 32×32×5, where "5" corresponds to the 5 common highway anomaly types identified by the model training: road surface cracks, obstacles, damaged markings, potholes, and roadbed settlement). Each dimension (corresponding to one anomaly type) contains 16×16 gridded detection units, and the unit value represents the probability that the region belongs to the corresponding anomaly type. The location mapping component performs spatial localization processing on the first edge feature map: First, it applies the Sigmoid activation function to each cell of each anomaly type dimension (e.g., "road crack") to generate an activation distribution map for that dimension (cells with activation values ​​≥ 0.8 are identified as suspected anomaly regions); then, it merges the activation distribution maps of the five dimensions to construct the first multi-channel spatial response map (size: 32×32×5); next, it performs extreme value analysis on the activation distribution map of the "road crack" dimension to find the 10 cells with the highest activation values, and takes their center coordinates as the reference anchor point (intra-frame coordinates: (X1, Y1)); subsequently, it performs Gaussian neighborhood diffusion on the reference anchor point. Modeling (variance σ=5) yields the spatial probability distribution parameters for this dimension (mean μ=(X1,Y1), variance σ²=25). Based on these parameters, an anomaly probability distribution map for "road surface cracks" is generated (areas with a probability value ≥0.9 are considered anomalous areas, with intra-frame coordinate ranges from (X2,Y2) to (X3,Y3)). Finally, the spatial positioning domain (intra-frame coordinate range) of the "road surface cracks" is transformed with the GPS coordinates of terminal 1 (latitude φ1, longitude λ1) and camera parameters (focal length f=4mm, installation height h=8m) to obtain the true GPS coordinates of the anomalous area (latitude φ2, longitude λ2). Through this process, the server obtains the first anomalous object feature of the current inspection video frame: the anomaly type is clearly "road surface cracks," the location features include the true GPS coordinates (latitude φ2, longitude λ2) and the intra-frame coordinate range, the shape feature is linear (length 1.2 meters, width 0.1 meters), and the area feature is 0.12 square meters.

[0031] Based on the characteristics of the first abnormal object, the server, in conjunction with the "Technical Specification for Highway Maintenance" (JTGH10-2009), determines the anomaly type as "longitudinal crack in the road surface" based on shape characteristics (linear) and location characteristics (right lane of the road surface); based on area characteristics (0.12 square meters) and length characteristics (1.2 meters), the damage level is determined to be Level II (moderate severity, requiring repair within 7 days); the spatial location is clearly defined as the actual GPS coordinates (φ2 North latitude, λ2 East longitude) and the corresponding road segment station number (e.g., KX+XX). Finally, the anomaly inspection object identification results include: anomaly type (longitudinal crack in the road surface), spatial location (GPS coordinates + road segment station number), damage level (Level II), shape parameters (length / width), area parameters, and timestamp (consistent with frame 15).

[0032] Next, the server extracts the GPS coordinates (φ2 N, λ2 E) of the target object from the abnormal inspection object identification results, and calculates the real-time relative distances between all inspection terminals deployed along the highway and these coordinates: Terminal 2 (fixed camera, deployed 1 km east of Terminal 1, GPS: φ3 N, λ3 E) is 0.8 km away; Terminal 3 (fixed camera, deployed 1 km east of Terminal 2, GPS: φ4 N, λ4 E) is 0.9 km away; Terminal 5 (mobile inspection vehicle, currently 0.7 km north of the target object, GPS: φ5 N, λ5 E) is 0.7 km away; and Terminal 6 (mobile inspection vehicle, currently 0.9 km south of the target object, GPS: φ6 N, λ6 E) is 0.9 km away. The preset distance threshold is 1 km, therefore, Terminal 2, Terminal 3, Terminal 5, and Terminal 6 are selected to form an associated inspection terminal group (all within the threshold range, capable of collaborative tracking).

[0033] The server issues collaborative tracking commands to the associated inspection terminal group through a 5G-MQTT communication network established between edge computing nodes (based on the low-latency MQTT protocol of 5G, QoS level 3, ensuring reliable transmission of commands). The command content includes: target object information (GPS coordinates: latitude φ2, longitude λ2; anomaly type: longitudinal crack in road surface; damage level: level 2), tracking requirements (capture at least two orthogonal perspectives, such as front and side, oblique front and side, to ensure multi-dimensional feature collection), synchronization clock stamp (2024-XX-XXXX:XX:XX.XXX, to ensure all terminals execute the command simultaneously and avoid timing deviations), and priority (high, marked as an urgent task, the terminal must suspend other non-urgent tasks and execute tracking first).

[0034] Upon receiving the instruction, the associated inspection terminal group strictly followed the synchronized clock stamp to initiate collaborative tracking: Terminal 2 (fixed camera) adjusted its horizontal angle to +10° (changing from frontal shooting to oblique frontal shooting), focusing on the endpoint of the crack in the target object to capture the direction of crack extension; Terminal 3 (fixed camera) adjusted its pitch angle to -15° (changing from frontal shooting to side shooting), focusing on the side of the crack to capture crack depth information; Terminal 5 (mobile inspection vehicle) adjusted its driving direction to the right lane where the target object was located, reduced its speed to 20km / h, and adjusted the camera view to the side front (45° angle with the road surface) to capture the angle between the crack and the driving direction and the surrounding environment; Terminal 6 (mobile inspection vehicle) turned around from the opposite lane, slowly drove along the lane where the target object was located, and adjusted the camera view to oblique frontal (60° angle with the road surface) to supplement the capture of details on the other side of the crack. During the tracking process, the terminals exchange perspective data in real time via 5G communication: Terminal 5 (the mobile inspection vehicle) sends video frames from its side-front view (including the relative positions of the crack and adjacent markings) to Terminals 2 and 3, providing a size reference for the fixed camera; Terminal 3 (side view) sends crack depth information (obtained through shadow analysis, approximately 5 cm) to Terminals 5 and 6, helping the mobile inspection vehicle adjust its driving route to maintain the optimal shooting distance; Terminal 2 (oblique front view) sends the coordinates of the crack endpoint to Terminal 6, helping it adjust the camera angle to accurately capture the crack's extension direction. The server dynamically updates the collaborative tracking strategy based on the perspective data exchanged between the terminals: for example, Terminal 5 adjusts its driving distance from 1 meter to 0.8 meters based on the depth information from Terminal 3, improving the resolution of the side-front view; Terminal 6 adjusts the camera's horizontal angle from +15° to +10° based on the endpoint coordinates of Terminal 2, ensuring the crack endpoint is fully within the field of view.

[0035] All tracking video streams generated by the associated inspection terminals are transmitted back to the edge computing node cluster via the 5G communication channel: Terminal 2 uploads a 1080P resolution, 25 frames per second oblique frontal view video stream; Terminal 3 uploads a 1080P resolution, 25 frames per second side view video stream; Terminal 5 uploads a 4K resolution, 30 frames per second moving side-frontal view video stream; and Terminal 6 uploads a 1080P resolution, 25 frames per second moving oblique frontal view video stream. The server marks these video streams as high-priority data streams (QoS level 3), triggering the edge computing node cluster (nodes A and C) to perform real-time aggregation processing: First, the side view of terminal 3 and the moving front view of terminal 5 are fused in 3D to generate a 3D model of the crack, clearly showing the 3D features of the crack such as length, width, and depth; Second, the oblique front view of terminal 2 and the moving oblique front view of terminal 6 are stitched together to generate a panoramic view of the crack, fully presenting the extension direction of the crack and the surrounding environment (such as the traffic conditions of adjacent lanes, the position of roadside guardrails, etc.); Finally, the tracking video streams are associated and stored with the anomaly identification results (such as associating the 3D model with the damage level and repair suggestions), providing comprehensive and accurate data support for subsequent maintenance decisions.

[0036] In this embodiment of the invention, the dynamic task segmentation of the received inspection video stream based on the edge computing node cluster to generate a distributed video frame processing queue can be implemented through the following example.

[0037] By monitoring the real-time processing load status and remaining computing resources of each node in the edge computing node cluster, the frame slicing strategy of the current video stream is determined.

[0038] The received inspection video stream is adaptively sliced ​​according to the frame slicing strategy to generate multiple video frame task segments.

[0039] The multiple video frame task fragments are assigned to nodes whose load status is below a preset threshold to form a distributed video frame processing queue, wherein the task fragments in the distributed video frame processing queue are dynamically prioritized according to the node's processing capacity.

[0040] In this embodiment of the invention, for example, the server, as the execution entity, first monitors the real-time processing load status and remaining computing resources of the edge computing node cluster (nodes A, B, and C): Node A currently has a CPU utilization of 70% and 35% memory remaining (the preset load threshold is 80%, which is within a processable state); Node B has a CPU utilization of 85% (exceeding the threshold and not suitable for assigning new tasks); Node C has a CPU utilization of 60% and 40% memory remaining (sufficient processing capacity). Based on the above monitoring results, the server, combined with the inspection video stream parameters (1080P resolution, 25 frames / second) uploaded by terminal 1, determines an adaptive frame slicing strategy of "every 10 frames as a task slice". This strategy ensures both node processing efficiency (avoiding task scheduling overhead caused by excessively small slices) and matches the current computing capacity of the nodes (nodes A and C can handle the processing pressure of 10 frames / slice). Subsequently, the server slices the video stream from terminal 1 according to this strategy, generating multiple video frame task slices such as slice 1 (frames 1-10), slice 2 (frames 11-20), and slice 3 (frames 21-30). Next, the server assigns the task slices to nodes with load values ​​below a preset threshold: node B is not assigned due to excessive load; node A, with lower load, is assigned slice 1; node C, with stronger processing capabilities (lower CPU utilization and more remaining memory), is assigned slice 2, and its priority is set higher than slice 1 (ensuring that nodes with stronger processing capabilities process more tasks first); slice 3 is temporarily stored in a distributed video frame processing queue. The task slices in the queue dynamically adjust their priorities based on the real-time processing capabilities of the nodes. If node A's CPU utilization subsequently rises to 85% (reaching the threshold), the server will move slice 3 from node A's pending list to node C and increase its priority, ensuring that high-load nodes are not over-assigned and maximizing resource utilization. Through the above process, the server generates a distributed video frame processing queue that is dynamically adjusted based on node load, achieving efficient segmentation and allocation of the inspection video stream.

[0041] In this embodiment of the invention, the step of determining the associated inspection terminal group based on the spatial location information in the abnormal inspection object identification result, and issuing a collaborative tracking instruction to the associated inspection terminal group through the 5G-MQTT communication network established between edge computing nodes, can be implemented through the following example.

[0042] Extract the GPS coordinate information of the target object from the abnormal inspection object identification results;

[0043] Based on the GPS coordinate information, calculate the real-time relative distance between each inspection terminal deployed along the highway and the target object;

[0044] Select multiple inspection terminals whose real-time relative distance is less than a preset distance threshold to form the associated inspection terminal group;

[0045] The downlink instruction carrying the location of the target object and the tracking priority is sent to the associated inspection terminal group through the 5G-MQTT communication network. The downlink instruction includes a synchronization clock stamp to ensure the timing consistency of instruction execution.

[0046] In an embodiment of the present invention, for example, the server, as the execution subject, first extracts the GPS coordinate information of the target object from the abnormal inspection object identification result. The coordinates are 30.1234° north latitude and 120.5678° east longitude (corresponding to the center position of the longitudinal crack in the road surface on the right lane at chainage K12+340). Next, based on the GPS coordinates, the server calculates the real-time relative distances between all inspection terminals deployed along the highway and the target object: Terminal 2 (fixed camera, deployed 1 km east of Terminal 1, current GPS: 30.1240°N, 120.5680°E) is 0.8 km from the target; Terminal 3 (fixed camera, deployed 1 km east of Terminal 2, current GPS: 30.1250°N, 120.5690°E) is 0.9 km away; Terminal 5 (mobile inspection vehicle, currently 0.7 km north of the target, GPS: 30.1230°N, 120.5670°E) is 0.7 km away; Terminal 6 (mobile inspection vehicle, currently 0.9 km south of the target, GPS: 30.1220°N, 120.5660°E) is 0.9 km away. Subsequently, the server selects terminals 2, 3, 5, and 6, whose real-time relative distance is less than a preset threshold (1 km), to form a group of associated inspection terminals (all within effective tracking range and capable of collaboration). Finally, the server sends downlink collaborative tracking instructions to the associated inspection terminal group through a 5G-MQTT communication network established between edge computing nodes (based on the low-latency MQTT protocol of 5G, QoS level 3, ensuring reliable instruction transmission). The instructions include: the GPS coordinates of the target object (30.1234°N, 120.5678°E), the anomaly type (longitudinal crack in the road surface), the damage level (level 2), the tracking priority (high, marked as an emergency task), and a synchronization clock stamp (2024-06-15 14:30:00.000). This clock stamp ensures that all terminals start tracking at the same time, avoiding timing deviations and guaranteeing consistency in multi-angle collaboration.

[0047] In this embodiment of the invention, the step of triggering the associated inspection terminal group to perform multi-angle collaborative tracking based on the collaborative tracking instruction and feeding back the generated tracking video stream to the edge computing node cluster can be implemented through the following example.

[0048] According to the collaborative tracking instruction, multiple inspection terminals in the associated inspection terminal group are controlled to synchronously adjust their camera viewing angles to capture at least two orthogonal viewing angles of the target object;

[0049] Real-time exchange of perspective data among multiple inspection terminals, and updating of collaborative tracking strategy based on the dynamic position of the target object;

[0050] The tracking video streams generated by each inspection terminal are transmitted back to the edge computing node cluster via a 5G communication channel. The tracking video streams are marked as high-priority data streams and trigger real-time aggregation processing by the edge computing node cluster.

[0051] In an embodiment of the present invention, for example, after the server, as the executing entity, issues a collaborative tracking instruction to the associated inspection terminal group (terminal 2, terminal 3, terminal 5, terminal 6), it first triggers each terminal to synchronously adjust the camera angle to capture an orthogonal angle. Terminal 2 (fixed camera, deployed 0.8 km west of the target) adjusts its horizontal angle to +10° as instructed, switching from a frontal to a diagonal frontal view, focusing on the eastern end of the crack to capture its extension direction; Terminal 3 (fixed camera, deployed 0.9 km south of the target) adjusts its pitch angle to -15°, switching from a frontal to a side view, aiming at the southern edge of the crack to obtain depth information; Terminal 5 (mobile inspection vehicle, located 0.7 km north of the target) immediately changes lanes to the right lane, reduces speed to 20 km / h, and adjusts the 4K camera at the front of the vehicle to the side front (45° angle with the road surface), covering the northern side of the crack and the surrounding environment; Terminal 6 (mobile inspection vehicle, located 0.9 km south of the target) turns around from the opposite lane, slowly drives along the target lane, and adjusts its camera to a diagonal frontal view (60° angle with the road surface) to supplement the capture of details at the western end of the crack. All terminals strictly followed the synchronization clock stamp (2024-06-15 14:30:00.000) in the instruction to ensure simultaneous tracking. During tracking, each terminal exchanged perspective data in real time via the 5G network: Terminal 5 sent a 4K video frame from the side front (containing the relative position of the crack and the adjacent 1.5-meter mark as a size reference) to Terminals 2 and 3; Terminal 3 obtained a crack depth of approximately 5 centimeters through shadow analysis and sent this data to Terminals 5 and 6; Terminal 2 identified the intra-frame coordinates of the eastern end of the crack (X=600, Y=800) and sent it to Terminal 6. The server updated the coordination strategy based on this data: Terminal 5 adjusted the driving distance from 1 meter to 0.8 meters based on the depth information from Terminal 3 to improve the clarity of the side front perspective; Terminal 6 adjusted the horizontal angle of the camera from +15° to +10° based on the end point coordinates of Terminal 2 to ensure the end point was fully in the frame; Terminal 2 adjusted the focal length based on the mark reference from Terminal 5 to reduce the width measurement error from ±0.05 meters to ±0.02 meters. The tracking video streams generated by each terminal are transmitted back to the edge node cluster via 5G: Terminal 2 uploads a 1080P / 25fps oblique frontal stream, Terminal 3 uploads a 1080P / 25fps side stream, Terminal 5 uploads a 4K / 30fps moving side-front stream, and Terminal 6 uploads a 1080P / 25fps moving oblique frontal stream. The server marks these streams as high priority (QoS level 3) to ensure reliable low-latency transmission. The edge node cluster immediately initiates real-time aggregation: Node A performs 3D fusion of the side stream from Terminal 3 and the side-front stream from Terminal 5 to generate a 3D model of the crack (1.2 meters long, 0.1 meters wide, and 5 centimeters deep); Node C stitches together the oblique frontal stream from Terminal 2 and the oblique frontal stream from Terminal 6 to generate a panoramic view (the crack has an angle of 30° with the road axis, and the surrounding area includes the right-side guardrail and traffic conditions in adjacent lanes).Finally, the server will track and store the video stream and the anomaly identification results (level 2 disease, repair recommended within 7 days) to provide accurate data support for maintenance decisions.

[0052] In this embodiment of the invention, the highway inspection object identification model is obtained in the following ways, and can be implemented through the following examples.

[0053] An initial model is obtained, which includes a raw input unit and an enhanced input unit. Each of the raw input unit and the enhanced input unit is configured with a cascaded lightweight feature extraction component and a location mapping component.

[0054] The lightweight feature extraction component in the original input unit extracts features from the inspection video frame instance to obtain a first edge feature map. The position mapping component in the original input unit then marks the spatial location domain of the inspection anomaly type based on the first edge feature map and aligns it to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance.

[0055] The lightweight feature extraction component in the enhancement input unit extracts features from the enhanced video frame instance to obtain a second edge feature map. The position mapping component in the enhancement input unit then uses the second edge feature map to mark the spatial location domain of the inspected anomaly type and aligns it to the enhanced video frame instance, obtaining the second anomaly object feature corresponding to the enhanced video frame instance. The enhanced video frame instance is video frame data enhanced based on the inspected video frame instance.

[0056] Mask modeling is performed based on the first edge feature map, the first abnormal object feature, the second edge feature map, and the second abnormal object feature to obtain the mask modeling error;

[0057] The model parameters of the initial model are updated based on the mask modeling error to obtain the highway inspection object recognition model that has been trained.

[0058] In this embodiment of the invention, for example, an object recognition model for intelligent highway inspection is obtained by constructing an initial model and performing iterative training. The following example, using a video frame from a two-way four-lane highway section (frame 15, 1080P resolution, including longitudinal cracks in the road surface, GPS coordinates of terminal 1 are 30.1230°N, 120.5670°E, timestamp 2024-06-15 14:25:00.600), illustrates the model acquisition process in detail: The server first loads the pre-designed initial model, which adopts a dual-input parallel structure, including two branches: the original input unit and the enhanced input unit. The two units have identical structures, both consisting of cascaded lightweight feature extraction and location mapping components: The lightweight feature extraction component, built on MobileNetV3, includes three Conv2D layers (3×3 kernels, stride 1, padding=1) and two DepthwiseConv layers (3×3 kernels, stride 1), used to extract multi-dimensional edge features from video frames, balancing feature extraction capability and computational efficiency; the location mapping component includes a Sigmoid activation function, a Gaussian neighborhood diffusion module, and a coordinate transformation module, used to map the feature map to the spatial location of the video frame, achieving accurate localization of abnormal regions. The initial model's output objective is to identify five types of highway anomalies (road surface cracks, obstacles, damaged road markings, potholes, and roadbed settlement), and output the spatial location, shape, and area features of the abnormal regions. The server inputs frame 15 into the original input unit and performs the following processing: Lightweight feature extraction: The component first normalizes frame 15 to a size of 224×224 (adapting to the model input), extracts low-level edge features (such as the outline of cracks) through a Conv2D layer, and then extracts high-level semantic features (such as the texture of cracks) through a DepthwiseConv layer, finally outputting the first edge feature map (size 32×32×5). Each channel corresponds to an anomaly type (channel 1: road surface crack, channel 2: obstacle, channel 3: damaged road markings, channel 4: pothole, channel 5: roadbed settlement), and the value of each 32×32 grid cell represents the probability (range 0-1) that the region belongs to the corresponding anomaly type. For example, the cell value at coordinates (110, 170) (after normalization) in channel 1 is 0.92, indicating that the probability of this region belonging to a road surface crack is extremely high.Location Mapping: The component performs spatial localization for channel 1 (road surface cracks): A Sigmoid activation function is applied to each cell to generate an activation distribution map (cells with activation values ​​≥ 0.8 are considered suspected abnormal areas); the activation distribution maps of the five channels are fused to construct the first multi-channel spatial response map (size 32×32×5); extreme value analysis is performed on the activation distribution map of channel 1 to find the 10 cells with the highest activation values, and their center coordinates are taken as the reference anchor point (normalized coordinates (110, 170), corresponding to the original 1080P frame coordinates (550, 850)); Gaussian neighborhood diffusion modeling (variance σ=5) is performed on the reference anchor point to obtain the spatial probability distribution parameters of this channel (mean μ=( 110, 170), variance σ²=25); Based on the spatial probability distribution parameters, an abnormal probability distribution map is generated (areas with a probability value ≥0.9 are judged as abnormal areas, corresponding to the normalized coordinate range (100, 160)-(120, 180), and the coordinate range restored to the original 1080P frame is (520, 820)-(580, 880)); The abnormal area is aligned to the original frame 15, and combined with the GPS coordinates of terminal 1 (30.1230°N, 120.5670°E) and camera parameters (focal length 4mm, installation height 8m), the true GPS coordinates (30.1234°N, 120.5678°E) are calculated through perspective transformation. Through the above processing, the server obtains the first abnormal object features: the abnormal type is "longitudinal crack in the road surface", the location features include the real GPS coordinates (30.1234°N, 120.5678°E) and the original frame coordinate range (520, 820)-(580, 880), the shape features are linear (length 1.2 meters, width 0.1 meters), and the area features are 0.12 square meters. To improve the model's ability to identify abnormal details, the server enhances the abnormal region of frame 15 (original frame coordinates 520, 820)-(580, 880): first, the abnormal region is locally cropped (to obtain a 120×120 image); then, the ESRGAN super-resolution model is used to enlarge the cropped image to 448×448, generating an enhanced frame video frame instance (preserving the texture details of the crack edge, such as the loose particles on both sides of the crack, improving the clarity of details).The enhanced frame instance is input into the enhanced input unit, and the processing flow is the same as that of the original input unit: Lightweight feature extraction: The component extracts the second edge feature map (size 64×64×5, the feature map size is larger because the enhanced frame has a higher resolution), and the probability distribution of the anomaly type for each channel (e.g., the cell value of coordinate (220,340) in channel 1 is 0.95, which is higher than the corresponding position in the original frame); Position mapping: The component performs activation, extreme value analysis, and Gaussian diffusion on channel 1 to generate an anomaly probability distribution map (coordinates 200,320)-(240,360) in the enhanced frame, which is restored to the original frame coordinates and is consistent with the first anomaly object feature), and aligns it to the enhanced frame to obtain the second anomaly object feature: the anomaly type, position, shape, and area are completely consistent with the first anomaly object feature, but the texture features of the crack edge are clearer (e.g., the measurement error of the crack width is reduced from ±0.05 meters to ±0.02 meters). To ensure consistency in feature extraction of anomalous regions in the original and enhanced frames, the server performs mask modeling: Mask generation: A first mask (32×32×5) is generated using the spatial probability distribution parameters of the first anomalous object feature (mean μ=(110,170), variance σ²=25), where the mask value for each channel is the spatial probability (range 0-1) of the corresponding position in that channel; a second mask (64×64×5) is generated using the spatial probability distribution parameters of the second anomalous object feature (mean μ=(220,340), variance σ²=25). Mask application: The first mask is multiplied by the first edge feature map to obtain the masked first feature map (32×32×5); the second mask is multiplied by the second edge feature map to obtain the masked second feature map (64×64×5). Error calculation: The masked second feature map is downsampled to a size of 32×32 (compared to the first feature map). Figure 1The L2 loss (mean squared error) of the two frames is calculated to obtain the mask modeling error (e.g., 0.05). This error represents the consistency of features in the abnormal regions in the original frame and the enhanced frame. The smaller the error, the more stable the model is in extracting features from the abnormal regions. The server uses the mask modeling error (0.05) to update all parameters of the initial model through the backpropagation algorithm (Adam optimizer, learning rate 0.001): adjusting the weights of the Conv2D layer and DepthwiseConv layer of the lightweight feature extraction component in the original input unit and the enhanced input unit (enhancing the ability to extract features from crack edges); optimizing the parameters of the Gaussian neighborhood diffusion module of the location mapping component (improving the accuracy of spatial positioning). The above process is repeated, using 100,000 inspection video frame instances (including various anomaly types such as road surface cracks, obstacles, damaged markings, potholes, and roadbed settlement) for training, calculating the mask modeling error every 1,000 training frames. When the error converges to below 0.01 (indicating that the model has stabilized in extracting features from the abnormal regions), training is stopped, and the trained highway inspection object recognition model is obtained. This implementation uses a dual-input unit (raw input + enhanced input) to process video frames and combines mask modeling to ensure feature consistency. The resulting model can efficiently process regular inspection video streams and accurately identify abnormal details, providing reliable algorithmic support for intelligent highway inspection.

[0059] In this embodiment of the invention, the following implementation methods are also provided.

[0060] The location features of multiple abnormal types in the first abnormal object feature are fused by channel fusion to obtain the first feature descriptor after channel fusion.

[0061] Based on the first feature descriptor, anomaly type identification is performed to obtain the first anomaly identification result;

[0062] The first anomaly type identification error is obtained based on the deviation between the first anomaly identification result and the anomaly target value of the inspection video frame instance.

[0063] The location features of multiple anomaly types in the second anomaly object feature are channel fused to obtain the channel fused second feature descriptor.

[0064] Based on the second feature descriptor, anomaly type identification is performed to obtain the second anomaly identification result;

[0065] The second anomaly type identification error is obtained based on the deviation between the second anomaly identification result and the anomaly target value of the inspection video frame instance.

[0066] The step of updating the model parameters of the initial model based on the mask modeling error includes:

[0067] The model parameters of the initial model are updated based on the mask modeling error obtained by summing the first anomaly type identification error and the second anomaly type identification error.

[0068] In this embodiment of the invention, for example, the server, as the execution entity, after acquiring the first abnormal object feature and the second abnormal object feature, optimizes and updates the model parameters through channel fusion, abnormal type identification, and multi-error accumulation. The following describes the process in detail using the scenario of frame 15 (longitudinal road surface crack, labeled "longitudinal road surface crack", one-hot vector [1,0,0,0,0]): For the first abnormal object feature (containing location features of 5 abnormal types: road surface crack, obstacle, damaged road marking, pothole, and roadbed settlement), the server first performs channel fusion: extracting the location feature vector for each abnormal type (road surface crack: spatial probability distribution mean μ=(110,170), variance σ²=25, intra-frame coordinate range (520,820)-(580,880)). 0); Obstacles: mean μ=(0,0), variance σ²=0, empty intra-frame coordinate range; the same applies to other anomaly types); An attention mechanism is used, and weights are assigned according to the activation values ​​of each anomaly type in the first edge feature map (the highest activation value for the road surface crack channel is 0.92, the highest for the obstacle channel is 0.01, and the rest are all <0.01) (road surface crack weight 0.98, the rest are all <0.01); The position feature vectors of the five anomaly types are weighted and concatenated to obtain the first feature descriptor (dimension 128, of which the position feature of the road surface crack accounts for 98%). Subsequently, anomaly type identification is performed based on the first feature descriptor: the first feature descriptor is input into the fully connected classification layer (activation function is Softmax) at the tail of the model, and the probability distribution of each anomaly type is output: road surface crack 0.95, obstacle 0.02, road marking damage 0.01, pothole 0.01, roadbed settlement 0.01; the "longitudinal crack of the road surface" corresponding to the highest probability is taken as the first anomaly identification result. Calculate the identification error of the first anomaly type: The anomaly target value of the inspection video frame instance (frame 15) is the labeled "longitudinal crack in the road surface" (corresponding to the one-hot vector [1,0,0,0,0]); the deviation is calculated using the cross-entropy loss function: (error 1 = -log(0.95) ≈ 0.051) (the smaller the error, the higher the identification accuracy).For the second abnormal object feature (the abnormal feature of the enhanced frame instance, which is consistent with the abnormal type and location of the first abnormal object feature, but with clearer details), the server performs the same processing flow as the first abnormal object feature: Channel fusion: Extract the location feature vectors of 5 abnormal types (the mean of the spatial probability distribution of road surface cracks is μ=(220,340), the variance is σ²=25, and the intra-frame coordinate range is consistent with the first feature), and weight concatenate them through the attention mechanism (road surface crack weight 0.99) to obtain the second feature descriptor (dimension 128); Abnormal type identification: Input the second feature descriptor into the fully connected classification layer, and output the probability distribution: road surface cracks 0.98, and the other abnormal types are all <0.01. The second abnormal identification result is still "road surface longitudinal crack"; Calculate the second abnormal type identification error: Use the cross-entropy loss function, (error 2=-log(0.98)≈0.020) (because the details of the enhanced frame are clearer, the error is less than the first abnormal type identification error). The server sums the first anomaly type identification error (0.051), the second anomaly type identification error (0.020), and the previously calculated mask modeling error (0.05, the feature consistency error between the original frame and the enhanced frame) to obtain the total error: (0.051 + 0.020 + 0.05 = 0.121). Then, the Adam optimizer (learning rate 0.001) is used to backpropagate the total error to all parameters of the initial model: adjusting the weights of the Conv2D and DepthwiseConv layers in the lightweight feature extraction component (enhancing the extraction capability of road crack edge textures, such as the features of loose particles on both sides of the crack); optimizing the parameters of the Gaussian neighborhood diffusion module in the location mapping component (reducing the variance σ to 4, improving the accuracy of spatial positioning, and making the intra-frame coordinate range of the anomaly region closer to the real annotation); and updating the weights of the fully connected classification layer (enhancing the ability to distinguish the "longitudinal road crack" type and reducing the probability of misclassification for other anomaly types). This implementation integrates location features of multiple anomaly types through channel fusion, outputs anomaly type probabilities in the classification layer, calculates the recognition error using cross-entropy loss, and finally sums the recognition error and mask modeling error as the total error to update the model. This process improves the model's accuracy in identifying anomaly types while ensuring the consistency of features between the original and enhanced frames, making the model more adaptable to the actual needs of highway inspection.

[0069] In this embodiment of the invention, the first edge feature map, the first abnormal object feature, the second edge feature map and the second abnormal object feature each include feature codes of multiple dimensions. The number of multiple dimensions is the number of abnormal types identified by the initial model training. The multiple dimensions correspond one-to-one with multiple abnormal types.

[0070] The step of performing mask modeling based on the first edge feature map, the first abnormal object feature, the second edge feature map, and the second abnormal object feature to obtain the mask modeling error can be implemented through the following example.

[0071] A first feature alignment error is obtained based on the first abnormal object feature and the second edge feature map; and a second feature alignment error is obtained based on the second abnormal object feature and the first edge feature map; wherein, the first feature alignment error represents the feature alignment error between abnormal types with the first abnormal type as the positioning subject; the second feature alignment error represents the feature alignment error between abnormal types with the second abnormal type as the positioning subject; the first abnormal type and the second abnormal type are two abnormal types that need to be spatially separated and identified among the plurality of abnormal types;

[0072] The mask modeling error is obtained by summing the first feature alignment error and the second feature alignment error.

[0073] In this embodiment of the invention, executively speaking, the server, as the execution entity, after acquiring the first edge feature map (original inspection video frame 15, 1080P resolution, including longitudinal cracks in the road surface), the first abnormal object feature (spatial location and shape features of cracks in the original frame), the second edge feature map (enhanced frame 15, 448×448 resolution, including cracks and newly added road potholes), and the second abnormal object feature (spatial location features of cracks and potholes in the enhanced frame), performs mask modeling through feature alignment error calculation to ensure the model's ability to distinguish features of anomaly types (longitudinal cracks in the road surface, potholes in the road surface) that require spatial separation. The following details this in conjunction with a specific scenario: The initial model training identifies 5 types of highway anomalies, corresponding to 5 feature dimensions (the number of dimensions is consistent with the number of anomaly types, and they correspond one-to-one): Dimension 1: Longitudinal cracks in the road surface (first anomaly type, requiring spatial separation from potholes); Dimension 2: Obstacles; Dimension 3: Damaged road markings; Dimension 4: Potholes in the road surface (second anomaly type, requiring spatial separation from cracks); Dimension 5: Subgrade settlement. First edge feature map (original frame): Probability distribution of each anomaly type in the original frame for each dimension (dimensional 1 has the highest activation value, indicating that the original frame mainly contains longitudinal cracks in the road surface; dimension 4 has extremely low activation value, with no effective pothole features); First anomaly object feature: Spatial location feature of each anomaly type for each dimension (dimensional 1 contains spatial location information of cracks: coordinate range of the right lane in the original frame and the actual GPS coordinates; dimension 4 has no effective location feature); Second edge feature map (enhanced frame): Probability distribution of each anomaly type in the enhanced frame for each dimension (dimensional 1 still has the highest activation value, and the activation value of dimension 4 has significantly increased, indicating that potholes have been added to the enhanced frame); Second anomaly object feature: Spatial location feature of each anomaly type for each dimension (dimensional 1 contains spatial location information of cracks (consistent with the original frame); dimension 4 contains spatial location information of potholes: coordinate range of the side of the crack in the enhanced frame and the actual GPS coordinates). The first feature alignment error is used to measure the degree of separation between the spatial features of each anomaly type and the crack in the enhanced frame (second edge feature map) when the longitudinal crack in the road surface (dimensional 1) is the localization subject (i.e., whether the location of the crack can be correctly distinguished from other anomaly types).The server performs the following steps: Extracting the spatial features of the located subject: Obtaining the spatial location information of dimension 1 (longitudinal crack in the road surface) from the first abnormal object features, the coordinate range of the right lane in the original frame (520, 820) to (580, 880) and the real GPS coordinates (30.1234°N, 120.5678°E); Aligning the feature map size of the enhanced frame: Adjusting the second edge feature map (64×64×5) of the enhanced frame to the same size as the first edge feature map (32×32×5) of the original frame to ensure the comparability of spatial locations; Calculating the spatial deviation between each dimension and the subject: For each dimension (1-5) of the second edge feature map of the enhanced frame, calculating the degree of deviation between the corresponding abnormal type spatial location and the spatial location of the located subject (crack) (the more the spatial locations overlap, the smaller the deviation; conversely, the smaller the overlap, the larger the deviation). For example: Dimension 1 (crack): Spatial location is completely consistent with the positioning subject, with a deviation of 0; Dimension 4 (pothole): Spatial location is next to the crack in the enhanced frame (coordinate range 600, 700 to 650, 750), with a significant deviation from the spatial location of the crack; Dimensions 2, 3, and 5: No effective anomaly type features, with minimal deviation; The first feature alignment error is obtained by summing the spatial deviations of each dimension: The first feature alignment error is obtained by summing the spatial deviations of each dimension (mainly contributed by the pothole in dimension 4; the larger the deviation, the weaker the model's ability to spatially separate potholes from cracks). The second feature alignment error is used to measure the degree of separation between the spatial features of each anomaly type and the pothole in the original frame (first edge feature map) when the pothole (dimensional 4) is the positioning subject (i.e., whether the location of the pothole can be correctly distinguished from other anomaly types). The server performs the following steps: Extracting the spatial features of the positioning subject: Obtaining the spatial location information of dimension 4 (pothole) from the second abnormal object features, enhancing the coordinate range (600, 700) to (650, 750) next to the crack in the frame and the real GPS coordinates (30.1236°N, 120.5680°E); Aligning the feature map size of the original frame: Adjusting the first edge feature map (32×32×5) of the original frame to the same size as the second edge feature map (64×64×5) of the enhanced frame to ensure the comparability of spatial locations; Calculating the spatial deviation between each dimension and the subject: For each dimension (1-5) of the first edge feature map of the original frame, calculating the degree of deviation between the corresponding abnormal type spatial location and the spatial location of the positioning subject (pothole).For example: Dimension 4 (potholes): Spatial location is completely consistent with the main object being located, with a deviation of 0; Dimension 1 (cracks): Spatial location is in the right lane within the original frame (coordinate range 520,820 to 580,880), with a significant deviation from the spatial location of potholes; Dimensions 2, 3, and 5: No effective anomaly type features, with minimal deviation; The second feature alignment error is obtained by summing the spatial deviations of each dimension: The second feature alignment error is obtained by summing the spatial deviations of each dimension (mainly contributed by cracks in dimension 1; the larger the deviation, the weaker the model's ability to spatially separate cracks and potholes). The server sums the first feature alignment error (spatial separation error with cracks as the main object) and the second feature alignment error (spatial separation error with potholes as the main object) to obtain the mask modeling error. This error reflects the model's ability to distinguish the features of the two anomaly types (longitudinal cracks and potholes) that need to be spatially separated. The larger the error, the more the spatial features of the two overlap, making it difficult for the model to separate them correctly; the smaller the error, the more independent the spatial features of the two, and the more accurately the model can distinguish them. This implementation method uses the anomaly types requiring spatial separation as the localization subject, calculates the spatial feature deviations of each anomaly type in the enhanced frame and the original frame, and sums the two to obtain the mask modeling error. This process ensures that the model can correctly distinguish between the two easily confused anomaly types: longitudinal cracks and potholes in the road surface, thus improving the model's spatial recognition capability.

[0074] In this embodiment of the invention, the step of obtaining the first feature alignment error based on the first abnormal object feature and the second edge feature map can be implemented through the following example.

[0075] Based on the feature code corresponding to the first abnormal type in the first abnormal object feature and the feature code corresponding to the second abnormal type in the second edge feature map, the spatial response difference value between the first abnormal type and the second abnormal type is obtained with the first abnormal type as the localization subject.

[0076] The first feature alignment error is obtained by weighting the spatial response difference values ​​of the first anomaly type and multiple anomaly types using a standardization function, with the first anomaly type as the localization subject.

[0077] The step of obtaining the second feature alignment error based on the second abnormal object features and the first edge feature map includes:

[0078] Based on the feature code corresponding to the second abnormal type in the second abnormal object features and the feature code corresponding to the first abnormal type in the first edge feature map, the spatial response difference value between the second abnormal type and the first abnormal type is obtained with the second abnormal type as the localization subject.

[0079] By using a standardization function, the spatial response differences between the second anomaly type and multiple anomaly types are weighted and calculated to obtain the second feature alignment error, with the second anomaly type as the localization subject.

[0080] In this embodiment of the invention, for example, the server, as the execution entity, after acquiring the first abnormal object features (spatial features of longitudinal cracks in the road surface of the original inspection video frame 15), the second edge feature map (pothole feature distribution in the enhanced frame 15), the second abnormal object features (spatial features of potholes in the enhanced frame 15), and the first edge feature map (crack feature distribution in the original frame 15), calculates the first feature alignment error (with cracks as the main component) and the second feature alignment error (with potholes as the main component) through spatial response difference evaluation and standardized weighting, ensuring the model's ability to distinguish the features of abnormal types (longitudinal cracks in the road surface and potholes in the road surface) that require spatial separation. The following describes this in detail with a specific scenario: The first feature alignment error is used to measure the degree of spatial feature separation between each abnormal type and the crack in the enhanced frame when the longitudinal crack in the road surface (the first abnormal type, dimension 1) is the core (i.e., whether the model can correctly distinguish the position of the crack from other abnormalities). The server performs the following steps: Extracting spatial features of the locating subject and target type: From the first anomalous object feature map (original frame 15), the feature encoding of the longitudinal crack (dimension 1) is extracted, including its spatial location information within the original frame (right lane coordinate range: 520,820 to 580,880) and spatial distribution concentration (small variance indicates concentrated crack location); From the second edge feature map (enhanced frame 15), the feature encoding of the pothole (second anomalous type, dimension 4) is extracted, including its spatial location information within the enhanced frame (crack-side coordinate range: 600,700 to 650,750) and spatial distribution concentration (slightly larger variance indicates slightly more dispersed pothole range). Calculating the spatial response difference between cracks and potholes: The server evaluates the spatial distribution overlap and calculates the spatial response difference between cracks and potholes (the larger the value, the less spatial overlap between the two, and the easier it is for the model to separate them). For example, a crack is located in the right lane of the original frame, while a pothole is located beside the crack in the enhanced frame; the positional deviation is significant, and the difference value is large (indicating that the model can initially distinguish between the two). Calculate the spatial response difference between cracks and all anomaly types: For the five dimensions (corresponding to five anomaly types) of the enhanced frame's second edge feature map, calculate the spatial response difference with cracks one by one: Dimension 1 (cracks): Spatially identical, difference value 0 (no separation requirement); Dimension 4 (potholes): Significant positional deviation, large difference value (requires focused separation); Dimension 2 (obstacles), Dimension 3 (damaged road markings), Dimension 5 (subgrade settlement): No effective activation (no anomalies), extremely large difference value (no need to pay attention). Standardize and weight to obtain the first feature alignment error: The server uses a standardization function (such as Softmax) to convert the five difference values ​​into weights. The larger the difference value, the higher the weight (indicating that the separation of this type from cracks is more important). For example: Dimension 4 (potholes) has a large difference value and a high weight (approximately 40%); Dimensions 2, 3, and 5 have extremely large difference values, but because there are no effective anomalies, their weights are low (approximately 60% in total); Dimension 1 has a weight of 0 (no separation requirement).Subsequently, the difference values ​​are multiplied by the weights and summed to obtain the first feature alignment error (mainly contributed by the difference values ​​of potholes, ensuring that the model focuses on potholes that need spatial separation). The second feature alignment error is used to measure the degree of spatial feature separation between each anomaly type and potholes in the original frame when potholes (second anomaly type, dimension 4) are the core (i.e., whether the model can correctly distinguish the location of potholes from other anomalies). The server performs the following steps: extracting spatial features of the localization subject and target type: extracting the feature encoding of potholes (dimensional 4) from the second anomaly object features (enhanced frame 15), including its spatial location information in the enhanced frame (coordinate range of crack side: 600, 700 to 650, 750) and spatial distribution concentration (slightly large variance); extracting the feature encoding of longitudinal cracks (first anomaly type, dimension 1) from the first edge feature map (original frame 15), including its spatial location information in the original frame (coordinate range of right lane: 520, 820 to 580, 880) and spatial distribution concentration (small variance). Calculate the spatial response difference between pits and cracks: Similarly, assess the spatial overlap by calculating the spatial response difference between pits and cracks (the larger the value, the less spatial overlap there is, and the easier it is for the model to separate them). For example, a pit may be located next to a crack in the enhanced frame, while the crack's position is fixed in the original frame; the positional deviation is significant, resulting in a large difference (indicating the model can initially distinguish between the two). Calculate the spatial response difference between pits and all anomaly types: For each of the five dimensions of the first edge feature map of the original frame, calculate the spatial response difference with the pits: Dimension 4 (pits): Spatially identical, difference value 0 (no separation requirement); Dimension 1 (cracks): Significant positional deviation, large difference value (requires focused separation); Dimensions 2, 3, and 5: No effective activation, extremely large difference values ​​(no need to focus on). Standardize and weight the second feature alignment error: Use a standardization function to convert the five difference values ​​into weights; the larger the difference value, the higher the weight (indicating the greater the importance of separating this type from the pit). For example, dimension 1 (cracks) has a large difference value and a high weight (approximately 70%); dimensions 2, 3, and 5 have extremely large differences, but because there are no effective anomalies, their weights are low (approximately 30% in total); dimension 4 has a weight of 0 (no separation requirement). Then, the difference values ​​are multiplied by their weights and summed to obtain the second feature alignment error (mainly contributed by the difference values ​​of cracks, ensuring the model focuses on cracks requiring spatial separation). The first feature alignment error (primarily cracks) reflects the spatial separation capability of pits and cracks in the enhanced frame, while the second feature alignment error (primarily pits) reflects the spatial separation capability of cracks and pits in the original frame. The sum of these two errors yields the mask modeling error (e.g., 17.27). The larger this error, the stronger the model's ability to distinguish features of anomaly types requiring spatial separation (cracks and pits).The server uses this error to update the model parameters (such as adjusting the weights of the feature extraction component and optimizing the parameters of the spatial mapping module), further improving the spatial identification accuracy of the model and ensuring that easily confused anomaly types (such as cracks and potholes) can be accurately located and clearly classified during highway inspections, thus meeting actual maintenance needs.

[0081] In this embodiment of the invention, the step of obtaining the mask modeling error based on the second feature alignment error accumulated from the first feature alignment error can be implemented through the following example.

[0082] The first feature alignment error accumulation result obtained with the target anomaly type as the second anomaly type and the second feature alignment error are accumulated to obtain the feature alignment error accumulation result corresponding to the target anomaly type; the target anomaly type is any one of the plurality of anomaly types.

[0083] The mask modeling error is obtained by averaging the sum of the feature alignment errors corresponding to multiple anomaly types.

[0084] In this embodiment of the invention, for example, the server, as the execution entity, after obtaining the first feature alignment error (spatial separation error with each anomaly type as the localization subject) and the second feature alignment error (spatial separation error with each anomaly type as the target type), obtains the mask modeling error by accumulating the target anomaly types and averaging them, ensuring that the model's spatial recognition capability for all anomaly types is balanced. The following details the scenario with five anomaly types (longitudinal cracks in the road surface, obstacles, damaged road markings, potholes, and roadbed settlement): The server treats all five anomaly types as target anomaly types (each anomaly type needs to be spatially separated from other types), and calculates the accumulated feature alignment error for each target one by one: Target 1: Longitudinal cracks in the road surface (dimension 1) (needs to be separated from potholes); Target 2: Potholes in the road surface (dimension 4) (needs to be separated from cracks); Target 3: Obstacles (dimension 2) (no effective activation); Target 4: Damaged road markings (dimension 3) (no effective activation); Target 5: Roadbed settlement (dimension 5) (no effective activation). For each target anomaly type, the server accumulates the first feature alignment error (separation error when locating the subject) with the target as the first anomaly type and the second feature alignment error (separation error when the target type is the second anomaly type) to obtain the accumulated feature alignment error result for that target: Target 1 (longitudinal crack in the road surface, dimension 1): First feature alignment error (separation error between each type and the crack in the enhanced frame with the crack as the subject of location): 9.63 (from previous scenes, mainly contributed by potholes); Second feature alignment error (separation error between the crack and the pothole when the crack is the second anomaly type, i.e., when the subject of location is the pothole): 7.64 (separation error between the pothole and the crack in previous scenes); Accumulated result: 9.63 + 7.64 = 17.27. Target 2 (potholes, dimension 4): First feature alignment error (separation error between potholes and other types in the original frame, with potholes as the primary location element): 8.21 (hypothetical value, mainly contributed by cracks); Second feature alignment error (separation error between potholes and cracks when potholes are the second anomaly type, i.e., when cracks are the primary location element): 6.92 (hypothetical value, separation error between cracks and potholes); Cumulative result: 8.21 + 6.92 = 15.13. Targets 3-5 (obstacles, damaged road markings, roadbed settlement): No effective activation (no anomalies), first and second feature alignment errors are both 0; Cumulative result: 0 + 0 = 0 (for each target). The server calculates the arithmetic mean of the cumulative feature alignment errors for the five target anomaly types, resulting in mask modeling error = (17.27 + 15.13 + 0 + 0 + 0) / 5 = 6.48. The mask modeling error (6.48) reflects the model's ability to spatially identify all anomaly types in a balanced way. The larger the error, the stronger the model's ability to distinguish features of anomaly types that require spatial separation (such as cracks and pits), and the less burden it has on types that do not have effective activation (such as obstacles).The server uses this error to update the model parameters (such as adjusting the weights of the feature extraction component and optimizing the parameters of the spatial mapping module) to ensure that the model has balanced spatial localization and classification accuracy for all anomaly types, thus meeting the actual needs of "comprehensive coverage and accurate identification" in highway inspection.

[0085] In this embodiment of the invention, the first edge feature map includes feature encodings of multiple dimensions, the number of multiple dimensions being the number of anomaly types identified by the initial model training, and the multiple dimensions corresponding one-to-one with multiple anomaly types; each dimension in the first edge feature map corresponds to a detection unit with gridded partitions.

[0086] The step of marking the spatial location domain of the inspection anomaly type based on the first edge feature map by the position mapping component in the original input unit and aligning it to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance can be implemented through the following example.

[0087] For each dimension, the activation function is applied cell by cell in the corresponding gridded partition to obtain the activation distribution map for that dimension.

[0088] A first multi-channel spatial response map is constructed based on the activation distribution maps corresponding to the multiple dimensions; the first multi-channel spatial response map includes activation distribution maps corresponding to multiple dimensions.

[0089] For each dimension of the first multi-channel spatial response map, the spatial coordinates of the extreme value region are determined as the reference anchor point;

[0090] For each dimension's baseline anchor point, neighborhood diffusion modeling is performed to obtain the spatial probability distribution parameters corresponding to each dimension;

[0091] Generate an anomaly probability distribution map for the corresponding anomaly type in each dimension based on the spatial probability distribution parameters of each dimension.

[0092] Locate the spatial location domain of the anomaly type based on the anomaly probability distribution map of each anomaly type;

[0093] The spatial positioning domains of multiple anomaly types are respectively aligned to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance; the first anomaly object feature includes the position features of multiple anomaly types.

[0094] In this embodiment of the invention, for example, the server, acting as the execution entity, processes the inspection video frame instance (frame 15, 1080P resolution, containing longitudinal cracks in the road surface, with the GPS coordinates of terminal 1 at 30.1230°N, 120.5670°E) in the original input unit. Relying on the first edge feature map (32×32×5 pixels, corresponding to 5 anomaly types, each dimension being a 32×32 gridded detection unit), the server uses a location mapping component to align the spatial positioning domain markers of the anomaly types with the video frame, ultimately obtaining the first anomaly object features. The following details this in conjunction with a specific scenario: the 5 dimensions of the first edge feature map (dimensional 1: longitudinal cracks in the road surface, dimension 2: obstacles, dimension 3: damaged road markings, dimension 4: potholes in the road surface, dimension 5: roadbed settlement) all correspond to 32×32 gridded partitions (each unit represents an area of ​​approximately 34×34 pixels in the video frame, 1080P / 32≈34). The server applies the Sigmoid activation function to each cell of the gridded partitions in each dimension, converting cell values ​​(probability values ​​between 0 and 1) into activation values ​​(0-1, with higher values ​​indicating a greater probability that the area belongs to the corresponding anomaly type), thus obtaining the activation distribution map for each dimension. For example, in the gridded partition of dimension 1 (longitudinal cracks in the road surface), the cell activation value corresponding to the crack in the right lane of frame 15 is relatively high (e.g., the activation value of grid (11,17) is 0.92); the cell activation values ​​of dimensions 2-5 (obstacles, damaged road markings, etc.) are extremely low (all <0.1, no valid anomalies). The server combines the activation distribution maps of the five dimensions in channel order to construct the first multi-channel spatial response map (size 32×32×5). Each channel of this map corresponds to an activation distribution of one anomaly type. For example, the activation distribution of channel 1 (dimension 1) is concentrated around grid (11,17) (corresponding to the crack location); the activation distribution of channels 2-5 is almost all 0 (no valid anomalies). For each dimension of the first multi-channel spatial response map, the server uses the Non-Maximum Suppression (NMS) algorithm to detect extreme regions (i.e., continuous cell regions with the highest activation values) in the activation distribution map, and takes the center coordinates of the region as the reference anchor point (representing the core location of the anomaly type). For example, in the activation distribution map of dimension 1, the extreme region is the grid (10-12, 16-18) (a continuous 3×3 cell, with activation values ​​all > 0.8), and the center coordinates are (11, 17) (the row and column numbers of the 32×32 grid); dimensions 2-5 have no valid extreme regions, and the reference anchor point is (0, 0) (invalid coordinates). The server performs Gaussian neighborhood diffusion modeling on the reference anchor point of each dimension (simulating the spatial distribution range of the anomaly type), and calculates the spatial probability distribution parameters (including mean and variance) based on the coordinates and activation values ​​of the reference anchor point. For example, the mean of the reference anchor point (11, 17) in dimension 1 is (11, 17) (i.e., the center coordinates of the extreme region); the variance is 25 (set according to the diffusion range of the activation values, the larger the value, the more dispersed the anomaly region).The server generates an anomaly probability distribution map (32×32) for each dimension's corresponding anomaly type based on the spatial probability distribution parameters. For example, in the anomaly probability distribution map of dimension 1, the probability values ​​of cells around the baseline anchor point (11,17) are high (e.g., the probability value of grid (10-12,16-18) is >0.9); the probability values ​​of cells far from the baseline anchor point are low (e.g., <0.1), indicating that the probability of this area being a crack is extremely low. Based on the anomaly probability distribution map of each anomaly type, the server selects cell regions with probability values ​​higher than a preset threshold (e.g., 0.9) as the spatial location domain for that anomaly type (i.e., the coverage area of ​​the anomaly type within the gridded partition). For example, the spatial location domain for dimension 1 is grid (10-12,16-18) (the row and column range of a 32×32 grid); dimensions 2-5 have no regions with probability values ​​higher than 0.9, so their spatial location domains are empty. The server aligns the spatial positioning domains of multiple anomaly types to the inspection video frame instances respectively: Grid coordinates to pixel coordinates: Through inverse normalization transformation (converting 32×32 grid coordinates to pixel coordinates of 1080P video frames), for example, the pixel coordinate range of frame 15 corresponding to grid (10-12, 16-18) is (520, 820) to (580, 880) (1080P / 32×10=337.5→approximately 520, 1080P / 32×16=540→approximately 820, and the maximum value is calculated similarly); Pixel coordinates to GPS coordinates: Combining the camera parameters of inspection terminal 1 (focal length 4mm, installation height 8m) and GPS coordinates (30.1230°N, 120.5670°E), the true GPS coordinates of the spatial positioning domain (such as 30.1234°N, 120.5678°E) are calculated through perspective transformation. Ultimately, the server obtains the first anomaly object features, including: the location features of dimension 1 (longitudinal cracks in the road surface): intra-frame pixel coordinate range (520, 820) - (580, 880), and the actual GPS coordinates (30.1234°N, 120.5678°E); and the location features of dimensions 2-5: empty (no valid anomalies). This implementation achieves precise spatial positioning of anomaly types through gridded activation, multi-channel response, extreme value anchoring, neighborhood diffusion, probability distribution, spatial positioning, and video frame alignment. Through this process, the server extracts the location features of longitudinal cracks in the road surface from the original inspection video frames, providing crucial data support for subsequent anomaly identification and collaborative tracking.

[0095] In this embodiment of the invention, the enhanced video frame instance is generated based on the inspection video frame instance, which can be implemented through the following example.

[0096] Based on the anomaly probability distribution map of each anomaly type obtained by processing the original input unit, the enhanced sample is generated for the inspection video frame instance using the target enhancement generation strategy corresponding to the enhanced input unit, thereby obtaining the enhanced video frame instance.

[0097] In this embodiment of the invention, for example, the server, as the execution entity, after processing the inspection video frame instance (frame 15, 1080P resolution, containing longitudinal cracks in the road surface) in the original input unit, relies on the anomaly probability distribution map of each anomaly type (from the position mapping component of the original input unit) and adopts the target enhancement generation strategy corresponding to the enhanced input unit to generate enhanced video frame instances (used to improve the model's ability to identify anomaly details). The following is a detailed description in conjunction with a specific scenario: The server first extracts the anomaly probability distribution map of 5 dimensions (corresponding to 5 anomaly types) obtained after the original input unit processes frame 15. Among them, the anomaly probability distribution map of dimension 1 (longitudinal cracks in the road surface) shows that: within the pixel coordinate range (520, 820) to (580, 880) of the right lane in frame 15, the probability value of all units is higher than the preset threshold (0.9), forming a high-probability anomaly area (representing the core location of the crack); the anomaly probability distribution maps of dimensions 2-5 (obstacles, damaged markings, etc.) have no high-probability areas (no valid anomalies). The server employs a target enhancement generation strategy corresponding to the enhanced input unit—"local region cropping and region super-resolution reconstruction" (enhancing details in high-probability anomalous regions). The specific steps are as follows: Local Region Cropping: Based on the high-probability anomalous regions in dimension 1 (520,820 to 580,880), frame 15 is precisely cropped to obtain a local image of 120x120 pixels (containing the core area of ​​the crack, such as the crack's edge, width, and surrounding loose particles); Region Super-Resolution Reconstruction: The ESRGAN super-resolution model is used to enhance the details of the cropped local image, enlarging it to 448x448 pixels (preserving the texture features of the crack edges, such as the granular damage on both sides of the crack, improving detail clarity); Background Blending: The super-resolution local image is seamlessly blended with the original background of frame 15 (excluding the high-probability regions) (maintaining consistent lighting and color with the original frame, enhancing only the details of the crack region). Through these steps, the server generates an enhanced video frame instance (still 1080P in size). The enhanced frame displays clearer details in high-probability anomaly regions (cracks) (e.g., the measurement error of crack width is reduced from ±0.05 meters to ±0.02 meters), while the background remains consistent with the original frame (avoiding the introduction of irrelevant noise). This implementation locates high-probability anomaly regions using the anomaly probability distribution map of the original input unit and generates enhanced frames using a targeted enhancement strategy (local cropping + super-resolution) for the enhanced input unit. This preserves the background information of the original frame while improving the clarity of details in the anomaly regions. Through this process, the server provides more accurate training samples for the enhanced input unit, which helps improve the model's feature extraction accuracy and spatial localization capabilities for anomaly types.

[0098] In this embodiment of the invention, the initial model includes at least two enhancement input units, namely a first enhancement input unit and a second enhancement input unit; wherein, the first enhancement input unit corresponds to a first target enhancement generation strategy, and the second enhancement input unit corresponds to a second target enhancement generation strategy;

[0099] The enhanced video frame instance can be generated by enhancing the inspection video frame instance, as shown in the following example.

[0100] The abnormal probability distribution map of each abnormal type of the inspection video frame instance is obtained by processing the original input unit, and the inspection video frame instance is processed by the first target enhancement generation strategy to obtain the first enhanced video frame instance; the first target enhancement generation strategy includes: local region cropping and region super-resolution reconstruction;

[0101] Based on the anomaly probability distribution map of each anomaly type of the first enhanced frame video frame instance obtained by processing the first enhanced frame video frame instance by the first enhanced input unit, the first enhanced frame video frame instance or the inspection video frame instance is processed by the second target enhancement generation strategy to obtain the second enhanced frame video frame instance; the second target enhancement generation strategy includes: local region cropping, region super-resolution reconstruction and neighborhood diffusion modeling.

[0102] In this embodiment of the invention, executively speaking, the server, as the execution entity, relies on the two enhancement input units (first enhancement input unit and second enhancement input unit) of the initial model and the corresponding target enhancement generation strategy. Starting from the inspection video frame instance (frame 15, 1080P resolution, containing longitudinal cracks in the road surface, the GPS coordinates of terminal 1 are 30.1230°N, 120.5670°E), it gradually generates the first enhancement frame and the second enhancement frame to improve the model's ability to recognize abnormal details and context. The following describes the specific scenario in detail: The first target enhancement generation strategy of the first enhancement input unit is "local region cropping and regional super-resolution reconstruction", which aims to enhance the detail clarity of abnormal areas. The server performs the following steps: Obtain the abnormal probability distribution map of the original input unit: After processing frame 15, the original input unit outputs an abnormal probability distribution map in 5 dimensions (corresponding to 5 abnormal types). The anomaly probability distribution map for dimension 1 (longitudinal cracks in the road surface) shows that: within the pixel coordinate range (520, 820) to (580, 880) of the right lane in frame 15, the probability values ​​of all units are higher than the preset threshold (0.9), forming a high-probability anomaly core region (representing the main location of the crack, such as the central axis of the crack); dimensions 2-5 have no high-probability regions (no valid anomalies). Local region cropping: Based on the high-probability anomaly core region of dimension 1, the server performs precise cropping on frame 15, extracting a local image of size 120×120 pixels (containing the core region of the crack, such as the edge contour, width, and surrounding 10 pixels of loose particles). The cropped local image retains only the key details of the crack, excluding irrelevant background (such as the normal road surface of the left lane). Region super-resolution reconstruction: The server uses the ESRGAN super-resolution model to enhance the details of the cropped local image, enlarging it to 448×448 pixels (a 4-fold increase in resolution). After super-resolution processing, the texture features of the crack edges (such as granular damage on both sides of the crack and depressions inside the crack) are clearer. For example, the measurement error of the crack width is reduced from ±0.05 meters in the original frame to ±0.02 meters. Background fusion: The server seamlessly fuses the super-resolution local image with the original background of frame 15 (excluding high-probability areas, such as the left lane and roadside guardrail) (by color matching and lighting adjustment, keeping the background consistent with the original frame). The fused image is the first enhanced frame video frame instance (still 1080P in size), whose detail clarity in the core crack area is significantly improved, while the background remains unchanged. The second enhancement input unit's second target enhancement generation strategy is "local region cropping, regional super-resolution reconstruction, and neighborhood diffusion modeling," which aims to enhance the contextual information of the abnormal region (i.e., the correlation features between the crack and the surrounding environment). The server performs the following steps: Obtain the anomaly probability distribution map of the first enhancement input unit: After processing the first enhanced frame, the first enhancement input unit outputs the anomaly probability distribution map of each anomaly type in the first enhanced frame.Because the core crack area in the first enhanced frame has clearer details, the anomaly probability distribution map of dimension 1 (longitudinal cracks in the road surface) shows that the high-probability area extends to the pixel coordinate range (530, 830) to (570, 870) (slightly smaller than the core area of ​​the original frame, but more precise), and the probability value of the neighboring area (10 pixels outside the core area) is improved (e.g., 0.7-0.8, representing minor damage around the crack). Local area cropping (including the neighboring area): Based on the anomaly probability distribution map of the first enhanced frame, the server crops the first enhanced frame (instead of the original frame, because the core area of ​​the first enhanced frame is more precise), extracting a local image of size 140×140 pixels (containing the core crack area (530, 830) - (570, 870) and a surrounding 10-pixel neighboring area). The cropped local image not only contains the core details of the crack, but also covers the transition area between the crack and the surrounding environment (such as the boundary between the crack edge and the normal road surface). Region Super-Resolution Reconstruction (Neighborhood Enhancement): The server performs super-resolution reconstruction on the cropped local image (enlarged to 448×448 pixels), focusing on enhancing details in the neighborhood region (such as loose particles at the crack edge and minor road surface depressions). After super-resolution processing, the features of the neighborhood region (such as particle size and depression depth) are clearer, helping the model learn the contextual information of "how the crack spreads from the core region to the periphery". Neighborhood Diffusion Modeling: The server performs neighborhood diffusion modeling on the super-resolution local image (simulating the characteristics of crack diffusion to the periphery through Gaussian blur and edge-preserving filtering), diffusing the features of the crack core region (such as edge contours) to the neighborhood region (10 pixels outside the core region). In the diffused image, the feature transition between the crack core region and the neighborhood region is more natural, and the model can better learn the "spatial expansion pattern of the crack". Background Fusion: The server seamlessly fuses the processed local image with the original background of the first enhanced frame to generate a second enhanced frame video frame instance (1080P size). The enhanced frame provides clear details in the core crack region and rich contextual information in the neighborhood region (such as the relationship between the crack edge and the surrounding road surface), offering the model more comprehensive training samples. This implementation uses a progressive enhancement strategy with two enhancement input units to gradually improve the detail clarity and contextual information of the abnormal region: the first enhanced frame, through "local cropping + super-resolution," focuses on the details of the core crack region, helping the model learn the crack's shape, size, and other features; the second enhanced frame, through "cropping with neighborhood + super-resolution + neighborhood diffusion," focuses on the relationship between the crack and the surrounding environment, helping the model learn the crack's expansion pattern and contextual features (such as how the crack affects the surrounding road surface). Through this process, the server provides the model with multi-dimensional enhanced samples, which helps improve the model's feature extraction accuracy and spatial generalization ability for abnormal types.

[0103] In this embodiment of the invention, the step of obtaining the first anomaly type identification error based on the deviation between the first anomaly identification result and the anomaly target value of the inspection video frame instance can be implemented through the following example.

[0104] Obtain the error probability evaluation value of the initial model trained in the previous training cycle for the target anomaly type; the target anomaly type is one of the anomaly types identified by the initial model during training;

[0105] Based on the temporal dependence of the error probability evaluation value of the same anomaly type trained in consecutive training cycles, the error probability evaluation value of the initial model trained in the current training cycle for the target anomaly type is obtained.

[0106] Determine whether the first anomaly identification result of each inspection video frame instance in the current training cycle meets the preset identification standard; the first anomaly identification result of the inspection video frame instance in the current training cycle includes the error probability evaluation value of multiple anomaly types of the inspection video frame instance;

[0107] Delete inspection video frame instances that do not meet the preset recognition criteria. Based on the deviation between the first anomaly recognition result of the inspection video frame instance that meets the preset recognition criteria and the anomaly target value of the inspection video frame instance, obtain the first anomaly type recognition error of the current training cycle.

[0108] In an embodiment of the invention, executively speaking, the server, acting as the execution entity, trains the initial model in the 100th training cycle (the current cycle). For the target anomaly type (longitudinal road surface crack, dimension 1), it obtains the first anomaly type identification error (reflecting the current cycle's model's accuracy in identifying this type) through error propagation from previous cycles, temporal dependency calculation, instance selection, and bias calculation. The following details this in conjunction with a specific scenario: The server first extracts the error probability assessment value of the initial model for the target anomaly type (longitudinal road surface crack) from the previous training cycle (the 99th cycle)—this value is the average cross-entropy loss of all training instances (1000 frames) in the 99th cycle, calculated to be 0.08 (indicating that the model's identification error for longitudinal road surface cracks in the 99th cycle is 8%, with an accuracy of 92%). Based on the temporal dependency characteristics of continuous training cycles (using a weighted moving average method, with a weight of 0.7 for the current cycle and 0.3 for the previous cycle), the server calculates the initial error probability assessment value for the current cycle (cycle 100): [Current initial error probability = 0.7 × current cycle prediction error + 0.3 × previous cycle error probability]; where "current cycle prediction error" is the average cross-entropy loss of the first 50 frames of the 100th cycle (0.06), therefore: [Current initial error probability = 0.7 × 0.06 + 0.3 × 0.08 = 0.042 + 0.024 = 0.066]; this value indicates that the initial identification error of the model for longitudinal cracks in the road surface in the current cycle is approximately 6.6% (accuracy approximately 93.4%). For each of the 100 inspection video frame instances in the current period (all containing longitudinal cracks in the road surface, labeled as "longitudinal cracks in the road surface", one-hot vector [1,0,0,0,0]), the server checks whether the first anomaly identification result (the error probability evaluation value of each anomaly type output by the original input unit) meets the preset identification standard (error probability ≤ 0.1, i.e., accuracy ≥ 90%): 1. Instance 1 (Frame 15): In the first anomaly identification result, the error probability evaluation value of the longitudinal crack in the road surface is 0.05 (cross-entropy loss), which meets the standard (0.05 ≤ 0.1); 2. Instance 2 (Frame 20): The error probability evaluation value is 0.07, which meets the standard; 3. Instance 3 (Frame 30): The error probability evaluation value is 0.12, which does not meet the standard (0.12 > 0.1); 4... (the remaining 97 instances): Among them, the error probability of 8 instances is > 0.1, which does not meet the standard. Finally, the server deletes the 9 instances that do not meet the standard (Instance 3 and the other 8), and retains 91 instances that meet the standard (Instance 1, Instance 2, etc.).For 91 instances that meet the criteria, the server calculates the deviation (i.e., cross-entropy loss) between the first anomaly identification result (error probability assessment value of longitudinal road cracks) and the anomaly target value (labeled "longitudinal road cracks", corresponding to error probability 0) for each instance, and takes the average as the first anomaly type identification error for the current period: for example, the deviation for instance 1 is 0.05, the deviation for instance 2 is 0.07, the deviation for instance 4 is 0.06, ..., the deviation for instance 100 is 0.08, and the sum of the deviations of all instances that meet the criteria is 5.46 (the average deviation of 91 instances is 5.46 ÷ 91 ≈ 0.06). This implementation obtains the initial error assessment for the current period through error propagation (weighted moving average) in the previous period, retains high-quality instances through preset standard screening (error probability ≤ 0.1), and finally calculates the average deviation of instances that meet the criteria to obtain the first anomaly type identification error (0.06). This error reflects the current period model's identification accuracy (94%) for the target anomaly type (longitudinal road cracks), which is an improvement over the previous period (92%). Through this process, the server ensures the stability (screening low-quality instances) and convergence (propagating time-dependent errors) of model training.

[0109] In this embodiment of the invention, if the current training period is the first training period, the error probability evaluation value of the initial model trained in the current training period for the target anomaly type is the initial confidence benchmark value; if the current training period is not the first training period, the step of obtaining the error probability evaluation value of the initial model trained in the current training period for the target anomaly type based on the temporal dependency characteristics of the error probability evaluation values ​​trained in consecutive training periods for the same anomaly type can be implemented through the following example.

[0110] Obtain the average value of the error probability assessment value of the target anomaly type for each inspection video frame instance in the previous training cycle.

[0111] Based on the historical confidence inertia factor and the previous training cycle, the initial model is trained to evaluate the error probability of the target anomaly type, and the historical confidence component is obtained.

[0112] The observation mean component is obtained based on the average value of the error probability assessment value of the target anomaly type of each inspection video frame instance in the preceding training cycle, according to the observation correction factor; wherein, the historical confidence inertia factor and the observation correction factor satisfy the unit normalization constraint.

[0113] The sum of the historical confidence component and the observed mean component is used as the error probability assessment value for training the initial model for the target anomaly type in the current training cycle.

[0114] In this embodiment of the invention, the exemplary objective is to identify the "crack" anomaly type (belonging to one of K anomaly types) in video frames. The training process is divided into cycles (each cycle consists of processing 1000 video frames). The server needs to calculate the error probability assessment value of the model for the "crack" type in each cycle (used to measure the error confidence of the model when predicting "cracks," with a value range of 0~1, and a larger value indicates a higher error probability). First training cycle (t=1): Initial confidence baseline value; when the server starts the first training cycle, since there is no historical training data, the initial confidence baseline value is directly adopted (preset to 0.5, representing "neutral confidence," that is, the probability of "crack" prediction error in the initial state of the model is 50%). Server actions: The server reads the training configuration file, determines that the current period is the first training cycle, automatically sets the error probability assessment value of the "crack" anomaly type to 0.5, and writes it to the training log of this cycle. For training periods other than the first one (t=2): calculations are performed based on temporal dependence characteristics. When entering the second training period, the server needs to calculate the error probability assessment value for the current period based on the training data of the previous period (t=1), combined with the historical confidence inertia factor (λ=0.7, representing the weight of retaining historical evaluation values) and the observation correction factor (μ=0.3, representing the correction weight of the current observation data, satisfying the unit normalization constraint of λ+μ=1). The server retrieves the training log of the first period and extracts the error probability assessment value of the "crack" anomaly type for all 1000 inspection video frame instances in that period (assuming the assessment values ​​for each frame are 0.45, 0.52, 0.58, ..., 0.61, for a total of 1000 values). Server Actions: The server averages the error probability assessment values ​​of 1000 "cracks" to obtain the mean of the preceding period: (0.45+0.52+0.58+…+0.61) / 1000=0.55 (meaning that in the first period, the average error probability of the model predicting "cracks" is 55%). The server reads the error probability assessment value of "cracks" in the first period (i.e., the assessment value at t=1, which is 0.5), multiplies it by the historical confidence inertia factor (λ=0.7), and obtains the historical confidence component (representing the continuation of historical assessment results). Calculation process: Historical confidence component = previous period assessment value × historical confidence inertia factor = 0.5 × 0.7 = 0.35 (meaning 70% of the historical confidence is retained, corresponding to an error probability of 35%). The server multiplies the mean of the preceding period (0.55) by the observation correction factor (μ=0.3) to obtain the observation mean component (representing the correction of the current observation data, i.e., adjusting the confidence based on the actual performance of the previous period). Calculation process: Observation mean component = previous period mean × observation correction factor = 0.55 × 0.3 = 0.165 (i.e., introducing 30% current observation correction, corresponding to an error probability of 16.5%). The server adds the historical confidence component (0.35) to the observation mean component (0.165) to obtain the error probability assessment value of the "crack" anomaly type in the second period.Calculation process: Current period evaluation value = historical confidence component + observed mean component = 0.35 + 0.165 = 0.515 (that is, in the second period, the model's prediction error probability of "crack" is 51.5%, slightly higher than 0.5 in the first period, indicating that the model's prediction error for "crack" has slightly increased). First period: The server directly uses the initial value (0.5) as the error probability evaluation value of "crack", without needing historical data. Non-first periods (such as the second period): The server calculates the current period evaluation value (0.515) by using "historical confidence component (preserving history) + observed mean component (correcting the current)". This method preserves the historical performance inertia of the model and adjusts the confidence level through current observation data, realizing time-dependent error probability evaluation, which is more in line with the dynamic change law of model training. This scenario takes the "crack" anomaly type as an example to fully demonstrate the detailed process of the server calculating the error probability evaluation value of the target anomaly type under different training periods, which fits the actual scenario of highway inspection model training, and the logic is clear and quantifiable.

[0115] In this embodiment of the invention, the target inspection video frame instance is one of the multiple inspection video frame instances loaded into the initial model in the same training batch during the current training cycle; the determination of whether the first anomaly identification result of the target inspection video frame instance in the current training cycle meets the preset identification standard can be implemented through the following example.

[0116] Determine whether the highest error probability assessment value among the error probability assessment values ​​of multiple anomaly types of the target inspection video frame instance in the current training cycle is greater than the error probability threshold of the current training cycle. If it is greater than the error probability threshold of the current training cycle, then determine that the first anomaly identification result of the target inspection video frame instance in the current training cycle meets the preset identification standard.

[0117] In this embodiment of the invention, for example, the server is performing the second training cycle of the highway inspection object recognition model, and is currently loading the third batch of 100 inspection video frame instances (100 frames per batch). The target inspection video frame instance is the 50th frame in this batch, which contains a highway surface and includes two anomaly types: "cracks" (obvious longitudinal cracks) and "potholes" (minor depressions). The server needs to determine whether the first anomaly recognition result of this frame (the model's prediction result for "cracks" and "potholes") meets the preset recognition criteria (i.e., whether the error probability exceeds the current cycle threshold). The server reads the 100 frames of video from the second cycle, third batch, from the storage cluster, parses the pixel data (1920×1080 resolution) of the 50th frame, and calls the model forward inference interface to obtain the anomaly type recognition result of this frame—the model predicts a probability of 0.85 for "cracks" (85% confidence) and a probability of 0.6 for "potholes" (60% confidence). According to the method of claim 14, the server retrieves the error probability assessment values ​​for the second period: the error probability assessment value for the "crack" anomaly type is 0.7 (calculated from the historical confidence component of the first period 0.35 + the observed mean component 0.35, representing a 70% error probability when the model predicts "crack"); the error probability assessment value for the "pothole" anomaly type is 0.5 (calculated from the historical confidence component of the first period 0.25 + the observed mean component 0.25, representing a 50% error probability when the model predicts "pothole"). The server retrieves the error probability threshold for the second period (preset to 0.6, calculated from the average of the highest error probabilities of all frames in the first period, representing the critical error probability value that needs attention in the current period). The server extracts the highest value of 0.7 from the error probability assessment values ​​of "crack" (0.7) and "pothole" (0.5), and compares it with the current period threshold of 0.6: 0.7 > 0.6, indicating that the model's prediction error probability for "crack" in this frame exceeds the critical value and needs to be focused on. The server marks the first anomaly identification result of the frame as meeting the preset identification criteria (i.e., the prediction error of "crack" needs further optimization) and writes this result to the training log for subsequent adjustment of model parameters (such as increasing the training weight of "crack" samples). The server determines whether the anomaly identification result of the target frame meets the criteria through the process of "loading the target frame → calculating the anomaly type error probability → extracting the highest value → comparing with the current period threshold". This process focuses on the high error type predicted by the model (such as "crack"), helping the server accurately locate the direction that needs optimization and improve the identification accuracy of the highway inspection model.

[0118] In this embodiment of the invention, the error probability threshold of the current training cycle is obtained through the following process, which can be implemented through the following example.

[0119] Get the highest error probability evaluation value among the multiple abnormality type error probability evaluation values ​​of each inspection video frame instance loaded into the initial model in the same training batch during the current training cycle.

[0120] The arithmetic mean of the highest error probability evaluation values ​​corresponding to the multiple inspection video frame instances is used to obtain the error probability threshold for the current training cycle.

[0121] In this embodiment of the invention, for example, the server is performing the third training cycle of the highway inspection object recognition model, and has currently loaded the second batch of 100 inspection video frame instances (each batch is fixed at 100 frames). Each video frame needs to identify three anomaly types: "cracks," "potholes," and "loose guardrails." The server has calculated the error probability evaluation value (value range 0~1, the larger the value, the higher the error probability of the model predicting that type) for each frame using the method of claim 14. Now, it is necessary to calculate the error probability threshold for the current cycle (the third cycle) to determine whether the recognition results of subsequent frames meet the standard. The server retrieves the 100 video frame data (resolution 1280×720) of the second batch of the third cycle from the distributed storage system, and simultaneously reads the error probability evaluation values ​​for the three anomaly types corresponding to each frame in this batch (stored in the JSON file of the training log). For example: Frame 1 (longitudinal cracks in the road surface): "Crack" error probability 0.7, "Pothole" 0.5, "Loose guardrail" 0.6; Frame 2 (shallow potholes in the road surface): "Crack" 0.6, "Pothole" 0.8, "Loose guardrail" 0.5; Frame 3 (slightly loose guardrail): "Crack" 0.4, "Pothole" 0.3, "Loose guardrail" 0.9; ... Frame 100 (no obvious anomalies): "Crack" 0.2, "Pothole" 0.1, "Loose guardrail" 0.3. The server extracts the maximum value of the error probability evaluation values ​​of the three anomaly types for each frame (i.e., takes the value of the anomaly type with the highest prediction error probability in that frame): Frame 1 maximum value: 0.7 ("Crack"); Frame 2 maximum value: 0.8 ("Pothole"); Frame 3 maximum value: 0.9 ("Loose guardrail"); ... Frame 100 maximum value: 0.3 ("Loose guardrail"). The server collects the highest error probability evaluation values ​​from 100 frames (100 values ​​in total: 0.7, 0.8, 0.9, ..., 0.3), and calculates the arithmetic mean: Error probability threshold = (-0.7 + 0.8 + 0.9 + ... + 0.3). The server uses the calculated value of 0.65 as the error probability threshold for the 3rd training cycle, writes it to the configuration file for that cycle (path: / train / cycle_3 / threshold.json), and synchronizes it to the global variables of the model training framework (such as PyTorch). This is used to subsequently determine whether the recognition results of other batches of frames within that cycle meet the preset standard (i.e., whether the highest error probability of a certain frame exceeds 0.65). The server obtains the error probability threshold (0.65) for the current training cycle through the process of "loading batch data → extracting the highest error probability of each frame → calculating the arithmetic mean". This threshold reflects the overall error level of the current batch of models for various anomalies. The server can then compare the highest error probability of a frame with this threshold to accurately identify the high-error anomaly types that need optimization (such as the "loose guardrail" error of 0.9 > 0.65 in frame 3, which requires a key adjustment to the training weights of this type), thereby improving the model's inspection accuracy.

[0122] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned multi-terminal collaborative management method and system for intelligent highway inspection based on edge computing. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0123] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. A multi-terminal collaborative management method for intelligent highway inspection based on edge computing, characterized in that, include: Based on multiple inspection terminals deployed along the highway, inspection video streams are uploaded in real time to the corresponding edge computing nodes via 5G communication channels. The received inspection video stream is dynamically segmented based on an edge computing node cluster to generate a distributed video frame processing queue. Extract the current inspection video frame from the distributed video frame processing queue. The current inspection video frame is the video frame data that needs to be identified as an abnormal inspection object. The current inspection video frame is loaded into a pre-trained highway inspection object recognition model. The current inspection video frame is then identified through the original input unit of the highway inspection object recognition model to obtain the first abnormal object feature of the current inspection video frame. Based on the first abnormal object feature of the current inspection video frame, the abnormal inspection object identification result of the current inspection video frame is obtained; Based on the spatial location information in the abnormal inspection object identification results, the associated inspection terminal group is determined, and a collaborative tracking instruction is issued to the associated inspection terminal group through the 5G-MQTT communication network established between edge computing nodes; Based on the collaborative tracking instruction, the associated inspection terminal group is triggered to perform multi-angle collaborative tracking, and the generated tracking video stream is fed back to the edge computing node cluster. The highway inspection object identification model is obtained through the following methods: An initial model is obtained, which includes a raw input unit and an enhanced input unit. Each of the raw input unit and the enhanced input unit is configured with a cascaded lightweight feature extraction component and a location mapping component. The lightweight feature extraction component in the original input unit extracts features from the inspection video frame instance to obtain a first edge feature map. The position mapping component in the original input unit then marks the spatial location domain of the inspection anomaly type based on the first edge feature map and aligns it to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance. The lightweight feature extraction component in the enhancement input unit extracts features from the enhanced video frame instance to obtain a second edge feature map. The position mapping component in the enhancement input unit then uses the second edge feature map to mark the spatial location domain of the inspected anomaly type and aligns it to the enhanced video frame instance, obtaining the second anomaly object feature corresponding to the enhanced video frame instance. The enhanced video frame instance is video frame data enhanced based on the inspected video frame instance. Mask modeling is performed based on the first edge feature map, the first abnormal object feature, the second edge feature map, and the second abnormal object feature to obtain the mask modeling error; The model parameters of the initial model are updated based on the mask modeling error to obtain the highway inspection object recognition model that has been trained.

2. The method according to claim 1, characterized in that, The edge computing node cluster dynamically segments the received inspection video stream to generate a distributed video frame processing queue, including: By monitoring the real-time processing load status and remaining computing resources of each node in the edge computing node cluster, the frame slicing strategy of the current video stream is determined. The received inspection video stream is adaptively sliced ​​according to the frame slicing strategy to generate multiple video frame task segments. The multiple video frame task fragments are assigned to nodes whose load status is below a preset threshold to form a distributed video frame processing queue, wherein the task fragments in the distributed video frame processing queue are dynamically prioritized according to the node's processing capacity.

3. The method according to claim 1, characterized in that, The step of determining the associated inspection terminal group based on the spatial location information in the abnormal inspection object identification result, and issuing collaborative tracking instructions to the associated inspection terminal group through the 5G-MQTT communication network established between edge computing nodes, includes: Extract the GPS coordinate information of the target object from the abnormal inspection object identification results; Based on the GPS coordinate information, calculate the real-time relative distance between each inspection terminal deployed along the highway and the target object; Select multiple inspection terminals whose real-time relative distance is less than a preset distance threshold to form the associated inspection terminal group; The downlink instruction carrying the location of the target object and the tracking priority is sent to the associated inspection terminal group through the 5G-MQTT communication network. The downlink instruction includes a synchronization clock stamp to ensure the timing consistency of instruction execution.

4. The method according to claim 3, characterized in that, The step of triggering the associated inspection terminal group to perform multi-angle collaborative tracking based on the collaborative tracking command, and feeding back the generated tracking video stream to the edge computing node cluster, includes: According to the collaborative tracking instruction, multiple inspection terminals in the associated inspection terminal group are controlled to synchronously adjust their camera viewing angles to capture at least two orthogonal viewing angles of the target object; Real-time exchange of perspective data among multiple inspection terminals, and updating of collaborative tracking strategy based on the dynamic position of the target object; The tracking video streams generated by each inspection terminal are transmitted back to the edge computing node cluster via a 5G communication channel. The tracking video streams are marked as high-priority data streams and trigger real-time aggregation processing by the edge computing node cluster.

5. The method according to claim 1, characterized in that, The method further includes: The location features of multiple abnormal types in the first abnormal object feature are fused by channel to obtain the first feature descriptor after channel fusion. Based on the first feature descriptor, anomaly type identification is performed to obtain the first anomaly identification result; Obtain the error probability evaluation value of the initial model trained in the previous training cycle for the target anomaly type; the target anomaly type is one of the anomaly types identified by the initial model during training; Based on the temporal dependence of the error probability evaluation value of the same anomaly type trained in consecutive training cycles, the error probability evaluation value of the initial model trained in the current training cycle for the target anomaly type is obtained. Determine whether the first anomaly identification result of each inspection video frame instance in the current training cycle meets the preset identification standard; the first anomaly identification result of the inspection video frame instance in the current training cycle includes the error probability evaluation value of multiple anomaly types of the inspection video frame instance; Delete inspection video frame instances that do not meet the preset recognition standards. Based on the deviation between the first anomaly recognition result of the inspection video frame instance that meets the preset recognition standards and the anomaly target value of the inspection video frame instance, obtain the first anomaly type recognition error of the current training cycle. The location features of multiple abnormal types in the second abnormal object feature are fused by channel to obtain the channel-fused second feature descriptor; Based on the second feature descriptor, anomaly type identification is performed to obtain the second anomaly identification result; The second anomaly type identification error is obtained based on the deviation between the second anomaly identification result and the anomaly target value of the inspection video frame instance. The step of updating the model parameters of the initial model based on the mask modeling error includes: The model parameters of the initial model are updated based on the mask modeling error obtained by summing the first anomaly type identification error and the second anomaly type identification error.

6. The method according to claim 1, characterized in that, The first edge feature map, the first abnormal object feature, the second edge feature map, and the second abnormal object feature all include feature codes of multiple dimensions. The number of multiple dimensions is the number of abnormal types identified by the initial model training, and the multiple dimensions correspond one-to-one with multiple abnormal types. The step of performing mask modeling based on the first edge feature map, the first abnormal object feature, the second edge feature map, and the second abnormal object feature to obtain the mask modeling error includes: Based on the feature code corresponding to the first abnormal type in the first abnormal object feature and the feature code corresponding to the second abnormal type in the second edge feature map, the spatial response difference value between the first abnormal type and the second abnormal type is obtained with the first abnormal type as the localization subject. By using a standardization function, the spatial response difference values ​​of the first anomaly type and multiple anomaly types are weighted and calculated respectively to obtain the first feature alignment error, with the first anomaly type as the localization subject. Based on the feature code corresponding to the second abnormal type in the second abnormal object features and the feature code corresponding to the first abnormal type in the first edge feature map, the spatial response difference value between the second abnormal type and the first abnormal type is obtained with the second abnormal type as the localization subject. By using a standardization function to calculate the spatial response differences between the second anomaly type and multiple anomaly types, a second feature alignment error is obtained. The first feature alignment error represents the feature alignment error between anomaly types when the first anomaly type is the localization subject; the second feature alignment error represents the feature alignment error between anomaly types when the second anomaly type is the localization subject; the first anomaly type and the second anomaly type are two anomaly types among the multiple anomaly types that require spatial separation and identification. The first feature alignment error accumulation result obtained with the target anomaly type as the second anomaly type and the second feature alignment error are accumulated to obtain the feature alignment error accumulation result corresponding to the target anomaly type; the target anomaly type is any one of the plurality of anomaly types. The mask modeling error is obtained by averaging the sum of the feature alignment errors corresponding to multiple anomaly types.

7. The method according to claim 1, characterized in that, The first edge feature map includes feature encodings in multiple dimensions, the number of which is the number of anomaly types identified by the initial model training, and each of the multiple dimensions corresponds one-to-one with a multiple anomaly type; each dimension in the first edge feature map corresponds to a detection unit with gridded partitions. The step of the position mapping component in the original input unit performing spatial localization domain marking of the inspection anomaly type based on the first edge feature map and aligning it to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance includes: For each dimension, the activation function is applied cell by cell in the corresponding gridded partition to obtain the activation distribution map for that dimension. A first multi-channel spatial response map is constructed based on the activation distribution maps corresponding to the multiple dimensions; the first multi-channel spatial response map includes activation distribution maps corresponding to multiple dimensions. For each dimension of the first multi-channel spatial response map, the spatial coordinates of the extreme value region are determined as the reference anchor point; For each dimension's baseline anchor point, neighborhood diffusion modeling is performed to obtain the spatial probability distribution parameters corresponding to each dimension; Generate an anomaly probability distribution map for the corresponding anomaly type in each dimension based on the spatial probability distribution parameters of each dimension. Locate the spatial location domain of the anomaly type based on the anomaly probability distribution map of each anomaly type; The spatial positioning domains of multiple anomaly types are respectively aligned to the inspection video frame instance to obtain the first anomaly object feature corresponding to the inspection video frame instance; the first anomaly object feature includes the position features of multiple anomaly types.

8. The method according to claim 7, characterized in that, The initial model includes at least two augmentation input units, namely a first augmentation input unit and a second augmentation input unit; wherein the first augmentation input unit corresponds to a first target augmentation generation strategy, and the second augmentation input unit corresponds to a second target augmentation generation strategy; The enhanced video frame instance is generated based on the inspection video frame instance, including: The abnormal probability distribution map of each abnormal type of the inspection video frame instance is obtained by processing the original input unit, and the inspection video frame instance is processed by the first target enhancement generation strategy to obtain the first enhanced video frame instance; the first target enhancement generation strategy includes: local region cropping and region super-resolution reconstruction; Based on the anomaly probability distribution map of each anomaly type of the first enhanced frame video frame instance obtained by processing the first enhanced frame video frame instance by the first enhanced input unit, the first enhanced frame video frame instance or the inspection video frame instance is processed by the second target enhancement generation strategy to obtain the second enhanced frame video frame instance; the second target enhancement generation strategy includes: local region cropping, region super-resolution reconstruction and neighborhood diffusion modeling.

9. A server system, characterized in that, Includes a server, the server being used to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Pavement disease detection method and system

    CN115049984A

  • Deep reinforcement learning cooperative scheduling method and device for heterogeneous computing resources

    CN117909044A

  • Intelligent identification method and system for electricity-related public safety hidden trouble based on edge AI

    CN118247697A