A road recognition method and device based on a visual large model and a storage medium
By using a visual large-scale model recognition method that combines edge and cloud computing, road target samples in complex scenarios are dynamically filtered and uploaded for precise calibration. This solves the problem of low recognition accuracy of mobile patrol terminals in complex scenarios and achieves efficient and low-cost recognition of road defects and events.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UESTC (SHENZHEN) ADVANCED RES INST
- Filing Date
- 2026-05-12
- Publication Date
- 2026-06-09
AI Technical Summary
Mobile patrol terminals have low accuracy in recognizing road targets in complex scenarios, making it difficult to meet the high-precision requirements of intelligent patrols.
A road recognition method based on a large visual model is adopted. Through the collaborative work of the edge and the cloud, the edge performs preliminary screening and judgment, dynamically selects calibration samples and uploads them to the cloud for accurate analysis, and combines them with the large visual model for calibration processing to output a structured road report.
It improves the recognition accuracy and precision of edge devices in complex and dynamic scenarios, and realizes efficient and low-cost road defects and event recognition, while taking into account real-time performance and high precision, and reducing network bandwidth and cloud inference costs.
Smart Images

Figure CN122176671A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of road defect and event recognition technology, and in particular to a road recognition method, device and storage medium based on a large visual model. Background Technology
[0002] With the rapid development of smart cities and intelligent transportation systems, the maintenance of road infrastructure and the real-time perception of traffic incidents have become crucial to ensuring the efficiency of urban road operations. Traditional inspections of road defects (such as cracks, potholes, and debris) and traffic incidents (such as congestion, accidents, and abnormal parking) primarily rely on manual patrols using patrol vehicles or fixed-point monitoring using cameras. Manual patrols suffer from low efficiency, high missed detection rates, and significant susceptibility to subjective factors. While fixed-point monitoring offers broad coverage, it has blind spots and cannot flexibly address the need for detailed inspections over long distances. To overcome the limitations of relying solely on manual patrols and fixed-point monitoring, mobile patrol terminals equipped with high-definition cameras (such as intelligent patrol vehicles, drones, and vehicle-mounted edge computing devices) have become increasingly widespread in recent years, driving the transformation of road patrols from "passive reporting and response" to "proactive, comprehensive perception."
[0003] However, in real-world road patrol scenarios, complex conditions such as variable outdoor lighting, frequent severe weather (such as rain, snow, and fog), and blurred or distorted road images place extremely high demands on the accuracy of road target recognition. The lightweight algorithms mounted on mobile patrol terminals have limited recognition capabilities in such complex scenarios, making it difficult to accurately identify various complex targets. This results in a significant drop in the recognition accuracy of these terminals, failing to effectively meet the high-precision requirements of actual intelligent road patrol operations. Summary of the Invention
[0004] This application provides a road recognition method, device, and storage medium based on a large visual model, which solves the technical problem of low accuracy in road target recognition by edge devices such as mobile patrol terminals in the prior art. It enables accurate recognition of road defects and events, especially road targets in complex dynamic scenes. Edge devices can perform efficient and low-cost recognition of road targets, improving the recognition accuracy and precision of edge devices.
[0005] In a first aspect, embodiments of the present invention provide a road recognition method based on a large visual model, applied at the edge, the method comprising: Each road surface image in the road surface image stream is acquired, and the road surface images are processed by a target detection model to obtain a target set and target information for each target in the target set. The target information includes: target number, coordinates of the target bounding box, prediction category, original confidence score, original Logits vector, target image patch, context image, timestamp, location information, and road scene auxiliary information. In each target, based on the target information, it is determined whether the target is a calibration sample; If the target is a calibration sample, then the target information of the target is used as the sample information of the calibration sample, and the calibration sample and the sample information of the calibration sample are sent to the cloud, so that the cloud can perform calibration processing on the calibration sample through a large visual model to obtain the calibration information of the calibration sample, and then output the structured road report of the calibration sample. The calibration information includes: target category, target confidence and supplementary attributes. If the target is an uncalibrated sample, a structured road report for the target is obtained and output based on the target information of the target.
[0006] Based on the same inventive concept, in a second aspect, the present invention also provides a road recognition method based on a large visual model, applied in the cloud, the method comprising: Obtain the calibration sample sent from the edge and the sample information of the calibration sample; Based on the sample information, the VQA cue words of the calibration sample are obtained. Then, through the visual big model, the calibration information of the calibration sample is obtained according to the sample information and the VQA cue words. The calibration information includes: target category, target confidence and supplementary attributes. The calibration information and sample information of the calibration sample are fused together to obtain and output the structured road report of the calibration sample.
[0007] Based on the same inventive concept, in a third aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the road recognition method based on a large visual model as described in the second aspect.
[0008] Based on the same inventive concept, in a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the road recognition method based on a large visual model as described in the second aspect.
[0009] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages: In this embodiment, after acquiring each road surface image from the road surface image stream and processing the road surface images using a target detection model to obtain a target set and target information for each target in the target set, each target needs to be judged. For each target, based on the target information, it is determined whether the target is a calibration sample. Here, a dynamic screening strategy is used to set a dedicated independent judgment for the detection and recognition results of the lightweight target detection model at the edge, avoiding false detections by the target detection model. Through multi-dimensional judgment conditions, it comprehensively judges whether the target is a calibration sample and whether it needs to be calibrated in the cloud, improving the accuracy, precision, and reliability of the target detection model. It also breaks the traditional binary opposition of "full cloud upload" or "full local processing," and is different from the conventional collaborative approach of coarse detection at the edge and fine detection in the cloud. Instead, it accurately identifies high-value calibration samples that cannot be reliably judged at the edge and uploads them to the cloud as needed, providing high-quality input for subsequent semantic calibration and closed-loop training, achieving a balance between bandwidth, computing power, and accuracy.
[0010] If the target is a calibration sample, its target information is used as the sample information of the calibration sample. The calibration sample and its sample information are then sent to the cloud, whereby the cloud uses a large visual model to calibrate the calibration sample and output its calibration information. This calibration information includes the target category, target confidence level, and supplementary attributes. The edge device uploads dynamically filtered targets that cannot be reliably determined to the cloud. The cloud then performs precise analysis on these targets (i.e., calibration samples) and outputs their calibration information. This approach ensures real-time performance at the edge device while maintaining the high-precision processing capabilities of the cloud.
[0011] If the target is an uncalibrated sample, a structured road report for the target is obtained and output based on the target information. Targets detected and identified by the target detection model are highly stable and reliable; a structured road report is directly output based on the target information of these targets, allowing staff to directly utilize the relevant information and improving the overall efficiency of the method.
[0012] Thus, the road recognition method in this embodiment can accurately identify road defects and events, especially road targets in complex dynamic scenarios. The edge device can efficiently and cost-effectively identify road targets, improving the recognition accuracy and precision of the edge device. Attached Figure Description
[0013] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference figures denote the same parts throughout the drawings. In the drawings: Figure 1 This invention illustrates a step-by-step flowchart of a road recognition method based on a large visual model applied to the edge of the device, as described in an embodiment of the present invention. Figure 2 The diagram illustrates the steps of a road recognition method based on a large visual model applied in the cloud, as described in an embodiment of the present invention. Detailed Implementation
[0014] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0015] Example 1 The first embodiment of the present invention provides a road recognition method based on a large visual model, which is applied to the edge, such as... Figure 1 As shown, the identification method includes: S101: Acquire each road image from the road image stream, and perform detection processing on the road images through the target detection model to obtain the target set and the target information of each target in the target set. The target information includes: target number, coordinates of the target bounding box, prediction category, original confidence score, original Logits vector, target image patch, context image, timestamp, location information, and road scene auxiliary information. S102, for each target, determine whether the target is a calibration sample based on the target information of the target; S103, if the target is a calibration sample, the target information of the target is used as the sample information of the calibration sample, and the calibration sample and the sample information of the calibration sample are sent to the cloud so that the cloud can perform calibration processing on the calibration sample through the visual big model to obtain the calibration information of the calibration sample, and then output the structured road report of the calibration sample. The calibration information includes: target category, target confidence and supplementary attributes. S104, if the target is an uncalibrated sample, then obtain and output the target's structured road report based on the target's target information.
[0016] It should be noted that the road recognition method based on a large visual model in this embodiment is applied to the edge. The edge refers to a mobile patrol terminal, which includes, but is not limited to, intelligent patrol vehicles, drones, and vehicle-mounted edge computing devices. The cloud includes, but is not limited to, personal computers, tablets, or customized devices set up according to actual needs.
[0017] In this embodiment, after acquiring each road image from the road image stream and processing it using a target detection model to obtain a target set and target information for each target in the target set, each target needs to be judged. For each target, based on its target information, it is determined whether the target is a calibration sample. Here, a dynamic filtering strategy is used to perform hierarchical and segmented processing on the recognition results at the edge, dividing the targets into regular targets that can be stably output directly by the edge and difficult targets that require further enhancement processing in the cloud. Through this judgment process, on the one hand, it avoids uploading all targets to the cloud indiscriminately, which would waste network bandwidth, overburden cloud inference, and increase system response latency; on the other hand, it also avoids relying solely on the local model at the edge to process all targets, which would lead to false detections, missed detections, and unstable outputs in complex scenarios. Thus, a collaborative processing mechanism of "rapid screening at the edge + focused calibration in the cloud" can be formed at the overall system level, allowing edge resources to be prioritized for real-time inspection and cloud resources to be prioritized for processing high-value difficult targets, thereby improving the overall resource utilization efficiency, recognition stability, and engineering feasibility of the system, and providing a high-quality input foundation for subsequent semantic calibration, confidence calibration, and closed-loop training. By employing multi-dimensional judgment criteria, it comprehensively determines whether a target is a calibration sample and whether it needs to be calibrated in the cloud, thereby improving the accuracy, precision, and reliability of the target detection model. It also breaks down the traditional binary opposition of "full cloud deployment" or "full local processing," and differs from the conventional collaborative approach of coarse detection at the edge and fine detection in the cloud. Instead, it accurately identifies high-value calibration samples that cannot be reliably determined at the edge and uploads them to the cloud as needed, providing high-quality input for subsequent semantic calibration and closed-loop training, achieving a balance between bandwidth, computing power, and accuracy.
[0018] If the target is a calibration sample, its target information is used as the calibration sample's sample information. Both the calibration sample and its sample information are sent to the cloud, allowing the cloud to calibrate the sample using a large-scale visual model and output calibration information. This calibration information includes the target category, target confidence level, and supplementary attributes. At the edge, dynamically filtered targets that cannot be reliably determined are uploaded to the cloud. The cloud then performs precise analysis on these targets (i.e., calibration samples) and outputs their calibration information. This ensures both real-time performance at the edge and high-precision processing capabilities at the cloud. Furthermore, leveraging the strong cross-scene semantic understanding, fine-grained discrimination, and complex background analysis capabilities of the large-scale visual model, high-precision enhancement processing is performed on targets that cannot be reliably determined at the edge, resulting in higher-confidence target categories, target confidence levels, and supplementary attributes. Compared to the initial identification results directly output from the edge, the calibration information processed in the cloud not only improves the accuracy of identifying difficult targets under complex lighting, inclement weather, background interference, and long-tail scenarios, but also generates standardized results with correction criteria and supplementary attributes. This makes the structured road report more complete, more interpretable, and more suitable for directly serving risk classification, manual review, and work order dispatch. Furthermore, the cloud calibration results can be further refined into high-quality samples for subsequent model distillation training and version iteration, thereby enabling the system to have continuous optimization and long-term evolution capabilities.
[0019] If the target is an uncalibrated sample, a structured road report for the target is obtained and output based on the target information. Targets detected and identified by the target detection model are highly stable and reliable. A structured road report is directly output based on the target information of these targets, allowing staff to directly use the relevant information of these targets and improving the overall efficiency of the method.
[0020] Thus, the road recognition method in this embodiment can accurately identify road defects and events, especially road targets in complex dynamic scenarios. The edge device can efficiently and cost-effectively identify road targets, improving the recognition accuracy and precision of the edge device.
[0021] Below, in conjunction with Figure 1 This document details the specific implementation steps of the road recognition method based on a large visual model provided in Embodiment 1: First, step S101 is executed to acquire each road surface image in the road surface image stream, and the road surface images are processed by the target detection model to obtain the target set and the target information of each target in the target set. The target information includes: target number, coordinates of the target bounding box, prediction category, original confidence score, original Logits vector, target image patch, context image, timestamp, location information and road scene auxiliary information.
[0022] Specifically, the edge device acquires real-time road image streams via image acquisition modules such as cameras, and obtains image information for each road image in the stream. Image information includes: timestamps (i.e., the time the road image was captured), location information, and road scene auxiliary information. Road scene auxiliary information includes: the edge device's shooting posture, vehicle speed, light intensity and weather conditions for capturing the road image, lane information, and road condition information. To improve the real-time performance of the edge device, the acquired road images can undergo preprocessing such as resolution adjustment, deblurring enhancement, distortion correction, or region of interest cropping before being input into a pre-trained lightweight object detection model (such as an improved YOLOv11, EdgeDet, or other detection networks suitable for edge devices) for fast local inference.
[0023] By using a target detection model, each road surface image is processed to obtain a target set of N targets. The target information includes: target number, target bounding box coordinates, prediction category, raw confidence score, raw Logits vector, target image patch, context image, timestamp, location information, and road scene auxiliary information.
[0024] Among them, in the standard set In the diagram, i represents the sequence of targets in the target set, and b... i Let c be the coordinates of the target bounding box. i For predicting the category, s i This represents the original confidence level. i =(z i,1 , z i,2 , ..., z i,K ) represents the original Logits vector output by the classification branch before Softmax. i,1 z is the original Logits value corresponding to the first category in the original Logits vector. i,2 This represents the original Logits value corresponding to the second category in the original Logits vector, and so on. i,K This represents the original Logits value corresponding to the Kth category in the original Logits vector. id i Here, K represents the target number, K represents the total number of categories, and k represents the specific categories within that total number of categories. The total number of categories includes: cracks, potholes, water accumulation, repair marks, spilled debris, illegally parked vehicles, and other road defects or abnormal events. The predicted category is the category of the target detected and identified by the target detection model.
[0025] The edge can also extract the corresponding target image patch based on the target bounding box. i The context image is obtained by expanding outwards from the target bounding box at a preset ratio. iThis is used for subsequent cloud-based semantic calibration. The preset ratio can be set according to actual needs. Then, the target image patch is applied. i Context Image i The coordinates b of the target bounding box i Predicted category c i Original confidence level s i The original Logits vector z i Target ID i The system establishes associated records with corresponding timestamps, location information, and road scene auxiliary information, providing an input basis for subsequent uncertainty judgment, confidence calibration, result fusion, and closed-loop distillation.
[0026] Next, step S102 is executed, in which, based on the target information of each target, it is determined whether the target is a calibration sample.
[0027] Specifically, step S102 involves performing uncertainty assessment and a distribution mechanism for each target. To balance real-time performance at the edge and high accuracy in the cloud under limited bandwidth and computing power, an "uncertainty distribution mechanism" is designed to filter the detection results (i.e., targets) output in step S101. For each target i, the distribution decision value U is calculated. i and based on U i Decide whether the target should be uploaded to the cloud as a calibration sample.
[0028] For each target i, the confidence uncertainty score is obtained based on the target's original confidence level. The category confusion weakness score is obtained based on the target's predicted category. The morphological anomaly score is obtained based on the target's bounding box. Based on the target's confidence uncertainty score, category confusion weakness score, and morphological anomaly score, the target's distribution decision value is obtained, as shown in the following formula (1): (1); Where Uconf(i) is the confidence uncertainty score of target i, Uweak(i) is the category confusion weakness score of target i, Umorph(i) is the morphological anomaly score of target i, α is the weight coefficient corresponding to the confidence uncertainty score, β is the weight coefficient corresponding to the category confusion weakness score, and γ is the weight coefficient corresponding to the morphological anomaly score.
[0029] If the target's distribution decision value is not less than the distribution threshold, the target is identified as a calibration sample. Calibration samples are those that require calibration processing in the cloud; the distribution threshold can be set according to actual needs. If the target's distribution decision value is less than the distribution threshold, the target is identified as a non-calibration target.
[0030] 1) Confidence Uncertainty Term Judgment: The specific process of obtaining the confidence uncertainty term score of the target based on the original confidence level is as follows: One example is a mechanism of 1 and 0. If the original confidence level of the target is s... i If the target is located within the hesitation interval (e.g., [0.4, 0.7]), it indicates that the target is within the decision boundary of the target detection model. Therefore, the target's confidence uncertainty score is set to 1, denoted as Uconf(i) = 1. The hesitation interval can be set according to actual needs. If the target's original confidence score s i If the target is not located within the hesitation interval (e.g., [0.4, 0.7]), it indicates that the target is a stable detection result of the target detection model. Therefore, the target's confidence uncertainty score is set to 0, denoted as Uconf(i) = 0. Another example is the continuous scoring mechanism. When the target's original confidence score is within the hesitation interval, the confidence uncertainty score is determined based on the distance between the original confidence score and the center value of the hesitation interval. The closer the original confidence score is to the center value of the hesitation interval, the higher the confidence uncertainty score. Uconf(i) = 1 - |s i -s0| / d, where s0 is the center value of the hesitation interval and d is the normalization coefficient.
[0031] 2) Weakness Category Judgment: The specific process of obtaining the target's category confusion weakness score based on the target's predicted category is as follows: One example is a mechanism of 1 and 0, checking the target's predicted category c. i Does it fall within the weak category range of the predefined confusion category pair set Pweak? Pweak = {(patch marks, craters), (shadows, cracks), (reflections, puddles), ...}. If c i If a target belongs to any category in a specific confusion category pair within the confusion category pair set, then the target's category confusion weakness score is set to 1, denoted as Uweak(i) = 1. If c i If a target does not belong to the confusion category pair set, its category confusion weakness score is set to 0, denoted as Uweak(i) = 0. The confusion category pair set Pweak can be dynamically updated based on historical misclassification statistics, confusion matrices under different scenarios, or manual empirical rules. The confusion matrix under different scenarios is used to statistically analyze the misclassification distribution relationship between each true category and the predicted category under specific scenario conditions. For example, the confusion probability of "pothole" versus "water accumulation" in a rainy day scenario, the confusion probability of "crack" versus "shadow" in a bright light scenario, and the confusion probability of "repair marks" versus "pothole" in a construction scenario. When the confusion probability of a category pair in the corresponding scenario is higher than a preset threshold η, the category pair is added to the confusion category pair set Pweak to improve the accuracy of screening difficult samples.
[0032] Another example is a specific scoring mechanism, which obtains the confusion matrix of the scene corresponding to the target based on the road scene auxiliary information of the target information. Then, based on the confusion matrix of the corresponding scene, the category confusion weakness score of the target is obtained, that is, the probability value of the confusion matrix of the corresponding scene.
[0033] 3) Shape rule judgment: The specific process of obtaining the shape anomaly score of the target based on the target bounding box is as follows: Based on the coordinates of the target bounding box, obtain the aspect ratio (Raspect) and area (A) of the target bounding box. i Fill rate F i This allows us to demonstrate the degree of deviation between the target and the historical prior template through the aspect ratio Raspect, area Ai, and fill rate Fi. If Raspect, Ai... i or F i Beyond the predicted category c i For the corresponding standard geometric statistical range, the target's morphological anomaly score is set to 1, denoted as Umorph(i)=1; otherwise, the target's morphological anomaly score is set to 0, denoted as Umorph(i)=0. Furthermore, a comprehensive judgment can be made by combining the target's edge continuity (such as cracks), texture roughness (such as cracks and pits), regional reflective features (such as water accumulation), or the target's stability in consecutive frames to obtain the specific numerical value of the target's morphological anomaly score.
[0034] The edge terminal adopts a distribution strategy that prioritizes "distribution decision value judgment and supplements it with strong triggering conditions." Specifically, when U i When the value is greater than or equal to the distribution threshold τu, the target is marked as a calibration sample and uploaded to the cloud. The sample information of this calibration sample (i.e., the target information) is also simultaneously uploaded to the cloud. Furthermore, for predefined high-risk situations, this can be disregarded by U. i Restrictions trigger uploads directly. For targets that do not trigger uploads, i.e., U... i For targets with a distribution threshold τu, if the target is a non-calibrated sample, the target information at the edge end is directly accepted and the process proceeds to step S104. A high-risk scenario is as follows: if the confidence uncertainty score is within the confidence threshold range and the predicted category belongs to the confused category pair set, such as when the original confidence is in the hesitation range and the predicted category belongs to the confused category pair, then the target is determined as a calibration sample. The confidence threshold can be set according to actual needs. Alternatively, if the morphological anomaly score is within the morphological anomaly threshold range, then the target is determined as a calibration sample. This avoids excessive uploading due to relying solely on any single judgment condition and also avoids missing key and difficult samples due to relying solely on the weighted total score.
[0035] In this embodiment, the edge-end uses an end-to-cloud collaborative mechanism based on uncertainty distribution. By combining multi-dimensional information such as confidence uncertainty, weakness categories, and morphological rules, it dynamically filters high-value samples that the edge end cannot reliably determine and uploads them to the cloud. This ensures real-time performance at the edge end while accommodating the high-precision processing capabilities of the cloud, and significantly reduces bandwidth consumption and cloud inference costs while maintaining edge-end recognition accuracy. This also avoids the problems of high bandwidth consumption, high inference latency, and expensive cloud computing power associated with directly uploading all images or videos to the cloud for large-scale model inference. Thus, it effectively reduces the amount of invalid data transmission and unnecessary large-scale model calls, achieving a dynamic balance between real-time edge processing and high-precision cloud-based discrimination under limited network bandwidth and cloud computing power, improving the engineering feasibility and operational economy of large-scale road patrol systems. Compared to the full cloud upload approach, the road recognition method in this embodiment is more suitable for road patrol scenarios with high real-time requirements, large deployment scale, and limited resources.
[0036] It should also be noted that in weak network environments, the edge device can compress, crop, or sort the uploaded calibration samples according to the current network bandwidth, device load, and task priority, and prioritize uploading the sample information of the top M calibration samples with the largest UI, so as to reduce communication overhead and ensure the real-time inspection capability of the edge device.
[0037] If the target is a calibration sample, step S103 is executed, indicating that the target needs further semantic calibration processing in the cloud. The target information of the target is used as the sample information of the calibration sample, and the calibration sample and its sample information are sent to the cloud. The cloud then uses a large visual model to calibrate the calibration sample, obtaining its calibration information, and outputting a structured road report. The calibration information includes: target category, target confidence, and supplementary attributes. It should be noted that the specific process of the cloud calibrating the calibration sample to obtain its calibration information and then outputting the structured road report is the cloud-based road recognition method in Example 2, which will not be elaborated here; please refer directly to the content of Example 2.
[0038] If the target is a non-calibrated sample, step S104 is executed, indicating that the target has high stability and high reliability under the edge-end judgment conditions. Based on the target information, a structured road report for the target is obtained and output. The structured road report includes: target number, edge-end number, acquisition time (i.e., the target's timestamp), location information (e.g., road location), lane information, final category, final confidence level, calibration confidence level, target bounding box / segmentation region, on-site image URL, semantic judgment basis, risk level, damage length level, severity level, whether it affects traffic, suspected cause label, textual description, whether it has been corrected, category before correction, category after correction, and work order suggested actions. The structured road report can be uploaded to the intelligent traffic management cloud platform via HTTPS, MQTT, or other secure communication methods, and / or directly output to the edge end. The structured road report is used for visualization, alarm linkage, and maintenance work order dispatch. Since the edge-end generates structured road reports directly based on target information without sending the target information to the cloud for semantic calibration using a large visual model, some items in the generated structured road reports can be left blank or marked "None". For example, fields such as calibration confidence level, disease length level, severity level, whether it affects traffic, suspected cause label, textual description, whether it has been corrected, and the category before and after correction can all be left blank or marked "None".
[0039] In addition, the object detection model at the edge can also be updated and upgraded according to the upgrade package sent by the cloud within a set time period. The process of updating and upgrading the object detection model is carried out using OTA (Over-The-Air) technology. Specifically, under the set network environment, the edge device receives the upgrade package sent by the cloud. Specifically, when the set conditions such as network quality, remaining battery power, and device idle time are met (i.e., the set network environment) are met, the edge device initiates a version check request to the cloud. After the cloud approves the request, it sends a version update download response to the edge device. If the latest update version exists, the edge device downloads the full or differential upgrade package and completes signature verification, hash verification, and compatibility verification locally. The set network environment can also be specifically configured according to actual needs.
[0040] After the upgrade package passes verification, an updated version of the target detection model is obtained and written to a backup partition or backup running directory. Next, the updated version needs to undergo preset tests, including: smoke test, sampling inference, memory usage check test, and inference latency check test. These preset tests can also be customized according to actual needs.
[0041] After the updated version passes the preset test, the current version of the object detection model is updated to the updated version to realize the update and upgrade of the object detection model. It can also send feedback to the cloud to show that the upgrade was successful.
[0042] In this embodiment, the edge device continuously accumulates high-quality difficult example samples / training samples calibrated in the cloud to perform knowledge distillation and iterative optimization on the edge target detection model. Combined with a secure OTA upgrade mechanism that includes signature verification, smoke testing, and anomaly rollback capabilities, this enables the edge target detection model to continuously evolve and adaptively enhance for new and dynamic scenarios. Thus, by accurately identifying road defects and events through the target detection model, especially road targets in complex and dynamic scenarios, the edge device can perform efficient and low-cost road target identification, improving the accuracy, precision, and generalization ability of edge devices.
[0043] Furthermore, if the edge device experiences inference failure, latency exceeding limits, abnormal resource consumption, or a drop in the accuracy of the target detection model exceeding a set threshold after an upgrade, it indicates an anomaly in the updated version. The edge device will automatically roll back to the previous version and upload failure logs, device status, and abnormal samples to the cloud, allowing the cloud to retrain and re-release the trained upgrade package. The threshold for the drop in accuracy can be set according to actual needs. Through this secure and rollback-enabled OTA mechanism, the continuous evolution of the target detection model on the edge device can be guaranteed without affecting the continuous operation of road patrol services, dynamically improving the accuracy, precision, and generalization ability of the target detection model.
[0044] Example 2 Based on the same inventive concept, the second embodiment of the present invention also provides a road recognition method based on a large visual model, applied in the cloud, such as... Figure 2 As shown, the method includes: S201, Obtain the calibration sample and sample information of the calibration sample sent by the edge terminal; S202. Based on the sample information, obtain the VQA cue words of the calibration sample. Then, through the visual big model, obtain the calibration information of the calibration sample according to the sample information and VQA cue words. The calibration information includes: target category, target confidence and supplementary attributes. S203 fuses the calibration information and sample information of the calibration sample to obtain and output a structured road report of the calibration sample.
[0045] In this embodiment, after acquiring the calibration samples and their sample information sent by the edge device, the cloud obtains the VQA cue words for the calibration samples based on the sample information. Then, through a large visual model, the target category and supplementary attributes of the calibration samples are obtained based on the sample information and VQA cue words. Furthermore, the target confidence score of the calibration samples is obtained based on the sample information. Based on the target category, supplementary attributes, and target confidence score, the calibration information of the calibration samples is obtained. This further transforms the initial identification results of difficult targets by the edge device into highly reliable, interpretable, and uniformly representative calibration results. Category recalibration reduces the risk of misjudgment by the edge device under complex lighting, severe weather, background interference, and long-tail scenarios. Confidence recalibration makes the output score closer to the reliability of the true prediction, thus avoiding the high-score misjudgment problem that may occur when making business decisions solely based on the original scores from the edge device. Therefore, the results calibrated by the cloud not only outperform the initial results from the edge device in terms of recognition accuracy, but also are more suitable as the basis for structured road reports, risk classification, manual review, and subsequent model iterations in terms of stability, consistency, and decision-making capability.
[0046] Next, the calibration information and sample information of the calibration samples are fused to obtain and output a structured road report for the calibration samples. The basic target information formed during the edge-end acquisition phase is unified with the high-confidence category results, confidence results, and supplementary attributes formed during the cloud calibration phase. This ensures that the same target produces consistent, complete, and traceable output results across multiple dimensions, including location, category, confidence level, attributes, and correction basis. This fusion process avoids the separation between the original edge-end detection results and the cloud calibration results, improving the consistency, completeness, and interpretability of the structured road report. Simultaneously, the fused results can more directly serve risk classification, manual review, alarm linkage, and work order dispatch, improving the efficiency of converting identification results into a closed-loop business process for actual maintenance management and traffic incident handling.
[0047] Below, in conjunction with Figure 1 This document details the specific implementation steps of the road recognition method based on a large visual model provided in Embodiment 2: First, step S201 is executed to obtain the calibration samples and their sample information sent from the edge device. Next, step S202 is executed to obtain the VQA cue words for the calibration samples based on the sample information. Then, using a large visual model, the calibration information for the calibration samples is obtained based on the sample information and the VQA cue words. The calibration information includes: target category, target confidence, and supplementary attributes.
[0048] Specifically, after receiving the calibration samples, the cloud constructs a semantic calibration strategy for each sample, consisting of "image patch + context image + initial edge judgment result + scene prior". It then calls the Visual Large Model (VLM) to perform secondary discrimination based on visual question answering, which is used to complete category correction, attribute completion, and false detection removal. The semantic calibration strategy is based on the target image patch. i Context Image i The predicted category c of the edge output i Target bounding box b i Original confidence level s i timestamp t i It is formed by location information (GI) and optional road scene auxiliary information such as weather, lighting, vehicle speed, and lane information. By simultaneously utilizing local texture and global road environment information, misjudgments caused by relying solely on local appearance can be reduced.
[0049] The specific process of obtaining VQA cue words for calibration samples based on sample information is as follows: the cloud uses the predicted category c from the sample information of the calibration samples sent from the edge device. i The system retrieves the fields for "candidate category set," "easily confused negative sample description," "key distinguishing attributes," and "structured output template" corresponding to the predicted category from the prompt word library. These fields are then assembled into a Visual Question Answering Prompt (VQA) prompt. The prompt word library is maintained according to the disease / event category; for example, for "cracks," easily confused negative sample descriptions such as shadows, seams, and repair lines are configured, and for "potholes," easily confused negative sample descriptions such as water accumulation, manhole covers, and asphalt patches are configured.
[0050] The Visual Question Answering (VQA) prompt includes the following fields: a) Task instruction field, used to instruct the large-scale visual model to determine the true category of the calibration sample; b) Candidate category field, used to limit the output to a preset disease / event set; c) Negative sample exclusion field, used to describe the non-target object most easily confused with the target of the calibration sample and the basis for excluding the object; d) Discrimination feature field, used to list semantic features such as whether the edges are neat, whether the texture is continuous, whether there is reflection, whether it extends linearly, and whether it crosses lane markings; e) Output template field, used to instruct the large-scale visual model to output the target category, candidate categories, discrimination criteria, attribute labels, and target confidence (i.e., semantic confidence) in a set structured format. The structured format can be set according to actual needs.
[0051] It should be noted that, to accommodate newly added defects, abnormal events, or unknown targets in open road scenarios, the candidate category field is preferably constructed using a "restricted candidate set + open abnormal exit" approach. Generally, the large visual model is prioritized to perform fine-grained discrimination within the candidate category set. When the calibration sample does not match any candidate category in the candidate category set, or the semantic consistency score of the calibration sample is lower than the semantic credibility threshold τ... semantic When this happens, the system allows outputting labels such as "Other Categories," "Unknown Categories," or "Abnormal Targets," triggering manual review, sample accumulation, or subsequent model expansion processes. This avoids the situation where novel diseases or abnormal events in open scenarios are forcibly classified into incorrect categories due to overly strict restrictions on the candidate category set. Semantic credibility threshold τ semantic It can be set according to actual needs. For example, when the edge initially identifies the calibration sample as a "pit," the large visual model constructs the following structured VQA cue words: a) Task instruction field: Please combine the image patch of the calibration sample with the context image to determine the true category of the region of the image patch of the calibration sample, and explain the basis for the judgment.
[0052] b) Candidate Category Field: The candidate category set for the calibration sample is "pothole, patch, water accumulation, manhole cover". If the calibration sample does not match any of the candidate categories, "unknown category" will be output.
[0053] c) Negative Sample Exclusion Field: Please focus on excluding non-target objects that are easily confused with potholes. Non-target objects include: puddles, reflective areas on the road surface, manhole covers, and repair patches. Specifically, if the surface of the image patch in the calibration sample has obvious specular reflection, unstable edges that change with lighting, continuous texture, and no collapse depth characteristics, it is more likely to be puddles or reflective areas. If the area has regular boundaries, an approximately rectangular or circular outline, and uniform material texture, it is more likely to be a manhole cover or repair patch, rather than a pothole.
[0054] d) Identification feature fields: Please pay close attention to the following semantic features: whether the edges are broken and irregular; whether the internal texture is collapsed, fragmented or pitted; whether there is continuous specular reflection on the surface of the image block; whether the outline is regular; whether there is obvious break with the surrounding asphalt texture; whether it is accompanied by local effusion, shadows or patch boundaries.
[0055] e) Output Template Fields: Please output the results in the following structured format: ① Target Category; ② Candidate Category; ③ Whether to suggest correcting the prediction results at the edge; ④ Main Judgment Criteria; ⑤ Attribute Labels; ⑥ Target Confidence; ⑦ If it does not belong to the candidate category set, should it be marked as "Unknown Category"? Attribute labels are used to describe auxiliary discrimination information of the calibration sample region other than the target category. Auxiliary discrimination information includes labels such as edge morphology, texture features, optical features, severity, traffic impact, and suspected causes. For example, edge fragmentation, regular contour, surface reflection, texture breakage, high severity, traffic impact, suspected water accumulation, suspected repair marks, etc.
[0056] By constructing VQA cue words from calibration samples, the cue word examples fully cover five components: task instructions, candidate categories, negative sample exclusion, discriminative features, and structured output templates. Furthermore, by setting an "unknown category" exit point, the adaptability to newly added diseases and abnormal events in open scenarios is enhanced, significantly improving the comparability, stability, and interpretability of the large-scale visual model among fine-grained confused categories.
[0057] The specific process of obtaining calibration information for calibration samples using a large visual model, based on sample information and VQA prompts, is as follows: target image patch... i Context Image i Sample information and VQA prompts are input into the Vision-Language Model (VLM) to obtain structured response results. The response results include the category determined by the Vision-Language Model (VLM). vlm The probability p that the i-th calibration sample belongs to each candidate class k vlm,i (k) Attribute tag set A i And the explanatory text R i Based on the response results, specifically the candidate category probability, attribute matching degree, and conflict item penalty, the semantic consistency score (Ssemantic) of the calibration sample is calculated as shown in the following formula (2).
[0058] (2); Here, the candidate category set is denoted as K, which is the total set of categories, and k represents a candidate category in the candidate category set K, where k∈K, T k A standard description template for category k; F k N represents the set of key discriminative attributes for class k, used to characterize the typical attribute features that a calibration sample should possess when it belongs to class k; k F represents the set of negative sample features for class k, used to characterize conflicting properties that are more likely to occur when the calibration sample does not belong to class k. k and Nk It can be pre-built from domain prior knowledge, statistical results of historical labeled samples, or a category feature knowledge base maintained in conjunction with the prompt word library. To interpret text R i Similarity between the standard description template and category k For attribute tag set A i The degree of attribute matching between the key discriminative attribute set of category k and the category k. For attribute tag set A i The penalty for conflicts between the negative sample feature set of class k and the target class k, where α, β, γ, and δ are the weight coefficients of the corresponding terms, such as α being the candidate class probability p. vlm,i (k) corresponds to the weight coefficient.
[0059] The semantic consistency score is not based solely on the single-class probability output by the large visual model, but rather considers four types of information: first, the predicted probability p of the large visual model for the current candidate class k. vlm,i (k). Second, the semantic similarity between the large model's explanatory text and the standard description template of the candidate category. Third, the degree of matching between the large model's output attribute label and the key discriminative attributes of the candidate category. Fourth, the degree of conflict between the large model's output attribute label and the negative sample features of the candidate category. By weighting and summing the above items and introducing a conflict penalty term, the semantic consistency between the i-th calibration sample and the candidate category k can be determined more robustly, thereby avoiding misjudgments caused by class correction based solely on probability values.
[0060] If the difference between the semantic consistency score of the determined category and the semantic consistency score of the predicted category in the large visual model is greater than the correction threshold, i.e. S semantic,i (c vlm )-S semantic,i (c i If the value is greater than the correction threshold ΔS, and the class probability p vlm,i (c vlm () greater than the semantic credibility threshold τ semantic Then the predicted class c of the calibration sample will be... i Revised to c new =c vlm c new This is called the corrected category or calibrated category, which is the category determined by the visual large model and used as the target category for the calibration sample. If the difference between the semantic consistency score of the visual large model's determined category and the semantic consistency score of the predicted category is no greater than the correction threshold, i.e., S... semantic,i (c vlm )-S semantic,i (c i ) is not greater than the correction threshold ΔS, or p vlm,i (cvlm () Not greater than the semantic credibility threshold τ semantic Then the predicted category c at the edge will be... i The target category for the calibration sample is the predicted category c that retains the edge. i Correction threshold ΔS and semantic credibility threshold τ semantic These settings can be configured according to actual needs. If the large visual model outputs an "unknown category" label, or if the semantic consistency scores of all candidate categories are below the semantic credibility threshold τ... semantic If the result is not found, the calibration sample will be marked as an open-class abnormal sample and will proceed to manual review, result retention, or subsequent sample settling process.
[0061] While completing the category correction, the cloud can also output supplementary attributes of the calibration samples. For example, supplementary attributes include: disease length level, severity level, whether it affects traffic, suspected cause label and text description, and write "whether to correct", "the category before correction, i.e. the predicted category", "the category after correction, i.e. the target category" and "the basis for correction" into the calibration information, providing traceable supervision signals for subsequent result fusion and model distillation.
[0062] In this embodiment, to address the issue of false positives and false negatives that easily occur in edge-based lightweight target detection models under complex lighting, background interference, long-tailed diseases, and abnormal event scenarios, a semantic calibration process based on a large visual model driven by structured visual question-answering prompts is employed. After the problematic samples (i.e., calibration samples) selected at the edge are uploaded to the cloud, the large visual model is guided by structured VQA prompts to perform secondary semantic discrimination. This achieves prediction category correction, attribute completion, and false positive elimination of the recognition results at the edge, thereby significantly improving the accuracy of disease and event recognition in complex dynamic scenarios.
[0063] Furthermore, a visual question-answering prompt word is constructed in the cloud for each calibration sample, containing candidate categories, negative sample exclusion descriptions, key identification attributes, and a structured output template. This allows the large visual model to perform fine-grained discrimination within a constrained candidate semantic space, rather than generalizing to a free-form answer. This approach is particularly suitable for easily confused scenarios such as "cracks / shadows," "potholes / patches," "water accumulation / reflection," and "repair marks / damaged areas," significantly reducing the probability of misjudgment under complex lighting, rainy conditions, tunnels, construction traces, and background interference. Simultaneously, cloud-based semantic calibration can supplement the output with attribute information such as the severity of the damage, the degree of traffic impact, suspected cause labels, and correction criteria, thereby enhancing the accuracy, interpretability, and business usability of the recognition results. Thus, through cloud-based semantic calibration and a structured VQA prompt word mechanism, the recognition accuracy and robustness in complex dynamic scenarios are significantly improved, effectively reducing the false positive and false negative rates.
[0064] The process of obtaining the target confidence score of the calibration samples employs a calibration strategy performed on the confidence score of the calibration samples using a large visual model. Specifically, for the calibration samples that have undergone semantic calibration, further confidence calibration is performed to address the issue of "overconfidence" in the target detection model at the edge, ensuring that the target confidence score output in the cloud more accurately reflects the reliability of the recognition results and provides a unified and reliable quantitative basis for subsequent risk classification, priority ranking of manual review, and work order dispatch.
[0065] At the edge, during inference, the original Logits vector z output by the object detection model's head classification branch before Softmax is preserved. i =(z i,1 , z i,2 , ..., z i,K ), and together with the predicted category c i Target bounding box b i The sample information is uploaded to the cloud. The cloud uses an independent calibration dataset Dcal to pre-learn the optimal temperature parameter T. T This is obtained by minimizing the negative log-likelihood loss: (3); Among them, z j Let y represent the original Logits vector of the j-th calibration sample in the calibration dataset. j Let represent the true class label of the j-th calibration sample in the calibration dataset, K be the total set of classes, and T be the temperature scaling parameter, where T > 0. This parameter scales the original Logits vector to adjust the smoothness of the Softmax output probability distribution. The predicted probability corresponding to the true class of a calibration sample in the calibration dataset is the true class label y assigned to the j-th calibration sample by the model after temperature scaling and Softmax normalization of its original Logits vector. j The predicted probability can be further expressed as formula (4): (4); Where m represents the category traversal index in the category set K.
[0066] For any calibration sample i, the original Logits vector of calibration sample i is adjusted according to the optimal temperature parameter T. Scaling is performed to obtain the post-calibration confidence of the calibration sample belonging to each candidate category k. When the candidate category set K is determined, the post-calibration confidence corresponding to k∈K can be further taken for subsequent fusion calculations. As shown in the following formula (5): (5); Or it can be expressed as formula (6): (6); Where, p cali,i (k) represents the post-calibration confidence that the i-th calibration sample belongs to the k-th category, T These are the optimal temperature parameters learned based on the calibration dataset Dcal.
[0067] If the target category is the predicted category c of the new sample information i Then, based on the predicted category, the calibrated confidence level p is obtained. cali,i (c i ), and the post-calibration confidence level p cali,i (c i The target confidence level is determined as follows: If the target category is the calibrated category c... new Then, based on the calibrated category, the calibrated confidence level p is obtained. cali,i (c new ), and the post-calibration confidence level p cali,i (c new The target confidence level is determined as follows. The calibrated category is the corrected category c after semantic calibration using the large visual model. new Therefore, without multi-source fusion, p cali,i (c ) is determined as the target confidence level. Where, c =ci or c =c new .
[0068] Furthermore, when the predicted category belongs to a set of confusing category pairs, the calibrated confidence p of the target category is used. cali,i (c The class probability p of the target class obtained through the large visual model. vlm,i (c The overall confidence level is obtained as shown in formula (7) below. Furthermore, the overall confidence level is determined as the target confidence level.
[0069] (7); Where μ is the fusion weight, μ∈[0,1], c sfinal represents the target category, and sfinal represents the overall confidence level.
[0070] When using the above fusion method, the overall confidence score sfinal is determined as the target confidence score. After the above processing, the calibrated confidence score no longer only represents the internal scoring of the model, but can be directly used for risk warning classification, priority ranking of manual review, and subsequent work order dispatch threshold judgment.
[0071] SFinal can be divided into three levels based on actual business needs: high credibility, medium credibility, and pending manual review. When sfinal ≥ the first target confidence threshold τ1, the calibration sample's result after being evaluated by the large visual model is a high-confidence result. When the second target confidence threshold τ2 ≤ sfinal < the first target confidence threshold τ1, the calibration sample's result after being evaluated by the large visual model is a medium-confidence result. When sfinal < the second target confidence threshold τ2, the calibration sample's result after being evaluated by the large visual model is a low-confidence result. The second target confidence threshold is less than the first target confidence threshold. The first and second target confidence thresholds can be set according to actual needs. By aligning the calibrated confidence with the true accuracy, the problem of mis-assignment caused by "high-score misjudgment" in traditional detection models can be avoided.
[0072] To address the common issues of overconfidence and mismatch between output scores and true accuracy in target detection models, this embodiment employs a cloud-based confidence calibration method. This method applies temperature scaling to the raw Logits vector output by the edge-based target detection model, calibrating it to make the target confidence score, after semantic calibration using the large visual model, closer to the true prediction. This improves the recognition accuracy and reliability of the calibrated samples, providing a reliable basis for subsequent risk grading, manual review prioritization, and work order dispatch. Furthermore, this confidence calibration mechanism transforms the target confidence score from a "model score" into a "credibility indicator usable for risk grading and business decision-making." Temperature scaling calibration of the raw Logits vector output by the edge-based target detection model also aligns it with the semantic probability (i.e., the category probability p of the large visual model) in the cloud. vlm,i Further combining these methods yields a calibrated confidence level or overall credibility that more closely approximates the true prediction accuracy, i.e., the target confidence level. This effectively alleviates the "overconfidence" problem commonly found in existing target detection models, making high-scoring results more reliable and low-scoring results more discriminative. Consequently, it provides a unified and reliable quantitative basis for risk warning classification, priority ranking of manual review, alarm threshold judgment, and dispatch strategy formulation, reducing false alarms and mishandling caused by "high-scoring misjudgments."
[0073] Semantic calibration of the calibration samples is performed using a large visual model to obtain the target category, supplementary attributes, and target confidence. In the sample information of the calibration samples, the category is updated to the target category, the confidence is updated to the target confidence, and supplementary attributes are added to obtain the calibration information of the calibration samples.
[0074] Then, step S203 is executed to fuse the calibration information and sample information of the calibration sample to obtain and output the structured road report of the calibration sample.
[0075] Specifically, based on the target ID, timestamp difference, intersection-over-union (IoU) ratio of the target bounding boxes, distance to the center point of the target bounding boxes, and location information, sample information and calibration information are correlated and fused to obtain fused information of the calibration samples. The category of the fused information is the target category, and the confidence level of the fused information is the target confidence level. Specifically, the cloud or edge collaborative service, for targets in the same inspection trajectory, correlates sample information and calibration information based on the target ID, timestamp difference, IoU ratio of the target bounding boxes, distance to the center point of the target bounding boxes, and location information. After correlation, the calibration information and sample information are formed into a unified fusion unit, i.e., fused information. Within this fusion unit, the final category, final confidence level, final bounding box location, and final attribute set are further determined.
[0076] For the final category, if the target category of the calibration sample obtained after calibration in the cloud is the calibrated category c... new This indicates that the cloud platform corrects the category of the calibration sample, and the calibrated category c will be... new As the final category. If the target category of the calibration sample obtained after calibration in the cloud is the predicted category c. i This indicates that the cloud did not correct the category of the calibration sample, and therefore the predicted category c at the edge is retained. i The predicted category c i As the final category.
[0077] For the final confidence level, the target confidence level (i.e., the calibrated confidence level or the comprehensive confidence level) output by the cloud-based large visual model is preferentially used as the basis for the final confidence level. Furthermore, the original confidence level s at the edge is retained. i As a traceability field.
[0078] For bounding box location, if the large visual model does not output a calibrated bounding box after semantic calibration of the calibration samples, the target bounding box at the edge is directly used as the final bounding box. If the large visual model outputs a calibrated bounding box after semantic calibration of the calibration samples, it means that a more accurate calibrated target region or segmentation result is returned synchronously from the cloud. In this case, the target bounding box at the edge and the calibrated bounding box from the cloud are fused according to preset weights to obtain the final bounding box location b. final .
[0079] For the final attribute set, if the calibration samples output supplementary attributes such as disease length level, severity level, whether they affect traffic, suspected cause label, textual description, and correction basis after semantic calibration by the large visual model, these supplementary attributes are directly added to the final attribute set. Preferably, for attribute information belonging to the same target, the attribute information output after cloud semantic calibration is used first. When the cloud does not output corresponding attribute information, the attribute information at the edge is retained or marked as null for later supplementation. The fused attribute set (i.e., the final attribute set) can be denoted as: .
[0080] Among them, A edge Represents the attribute information at the edge, A cloud This represents the attribute information output after cloud-based semantic calibration, i.e., supplementary attributes. Fuse(·) represents the attribute fusion function. If A cloud If a corresponding attribute exists, then A will be used first. cloud The attribute value in Acloud; if the corresponding attribute does not exist in Acloud, the fallback will use Acloud. edge The attribute values in.
[0081] After completing the above-mentioned association with the target and multi-source fusion, the final category c is obtained. final Final position b final Final attribute set A final And the final confidence level.
[0082] Subsequently, business segments are categorized based on the final confidence level: 1) When the target confidence level (i.e., the final confidence level) is greater than or equal to the first target confidence level threshold τ1, a structured road report is generated directly based on the fused information, and the structured road report is output. That is, the target is treated as a high-confidence event and directly enters the automatic reporting process, giving priority to using the cloud-corrected category, confidence level, and supplementary attributes to generate a formal event report.
[0083] 2) When the second target confidence threshold τ2 ≤ target confidence (i.e. final confidence) < the first target confidence threshold τ1, the supplementary attributes are retained, and the fusion score of the candidate category is obtained according to the candidate category in the fusion information. The candidate category corresponding to the maximum fusion score is determined as the final category of the calibration sample, and the maximum fusion score is determined as the final confidence of the final category. Based on the supplementary attributes, final category, final confidence, and added manual review markers, a structured road report is generated and output.
[0084] Specifically, when the second target confidence threshold τ2 ≤ target confidence (i.e., final confidence) < the first target confidence threshold τ1, the target is classified as a moderately credible suspected event. At this point, to balance recognition accuracy and business continuity, supplementary fusion can be performed on the category probabilities at the edge and the category probabilities of the large visual model in the cloud, while retaining the supplementary attributes from the cloud, to obtain the final category and final confidence. A fusion score is calculated for each candidate category k in the candidate category set Ω: (8); in, p is the fusion score for candidate category k. cali (k), p vlm `(k)` and `AttrMatch(k)` represent the post-calibration confidence, visual large model class probability, and attribute matching degree of the candidate class `k` corresponding to the current calibration sample, respectively; for simplicity, the sample index `i` is omitted here. `ρ1`, `ρ2`, and `ρ3` are the weight coefficients of the corresponding items, and `AttrMatch(k)` represents the matching degree between the supplementary attribute and the candidate class `k`. The candidate class with the largest fusion score is taken as the final class `c`. final As shown in formula (9): (9).
[0085] The corresponding final confidence level s fuse Represented as: (10).
[0086] After the above-mentioned supplementary fusion is adopted, if c final Correction category c in the cloud new If consistent, then maintain the result output from the cloud. If c final Predicted category c at the edge i If the results are consistent and the fusion score is higher, the target category and its corresponding confidence level at the edge are saved as the output conclusion for the current target. If there is still a discrepancy, the target category and its corresponding confidence level at the edge are retained as the final identification conclusion for the current target, and the verification result from the cloud is written into the structured road report as auxiliary judgment information, while retaining the "cloud verification not fully confirmed" label. For the moderately reliable results in this case, it is preferable to retain the severity level, whether it affects traffic, suspected cause label, and textual description of the edge output, while adding the "awaiting manual verification" or "awaiting secondary confirmation" label for subsequent processing by the business platform.
[0087] 3) When the target confidence level (i.e., the final confidence level) is less than the second target confidence level threshold τ2, the fused information is output to the manual inspection queue for staff to review and verify. Specifically, the target is judged as a low-confidence result and enters the manual sampling, delayed confirmation, or pending review queue. In this case, the fused information may only retain basic identification information and evidence data, without directly triggering automatic dispatch. The basic identification information includes at least the target number, edge number, timestamp, location information, lane information, target bounding box or segmented region, predicted category of the edge (i.e., target category), and original confidence level of the edge. The evidence data includes at least one or more of the following: on-site image or video clip, target image patch Patchi, context image Contexti, attribute labels, semantic determination criteria, candidate category probability, fused score, and correction suggestions. By retaining the above basic identification information and evidence data, a traceable basis can be provided for manual review, delayed confirmation, and subsequent model optimization.
[0088] After the above processing, the final confidence score no longer only represents the internal score of the model, but together with the category results, attribute results and business processing logic, it constitutes a comprehensive decision-making basis that can be directly used for risk warning classification, priority ranking of manual review and subsequent work order dispatch.
[0089] To address the challenge of unifying edge detection results with cloud-based calibration results for direct service to business systems, this embodiment provides a result fusion and structured upload mechanism. By spatiotemporally correlating, multi-source fusion, and integrating credibility of edge detection results and cloud-based calibration results, a structured event report is generated, including target location, calibration category, credibility, correction basis, and work order suggested actions. This improves the efficiency of road maintenance and traffic incident handling. Furthermore, compared to traditional solutions that only output defect categories or detection boxes, this embodiment's fusion method reduces manual secondary processing, increases the automation of road maintenance work order dispatch, abnormal event alarm linkage, and inspection result archiving, and enables the identification results to be directly converted into executable business objects, thereby significantly improving the efficiency of converting road inspection results into actual business actions.
[0090] The cloud also features a distillation closed-loop update mechanism, establishing a closed-loop evolutionary mechanism of "cloud calibration—sample accumulation—model distillation—OTA deployment—edge feedback." The cloud-based distillation closed-loop update process is as follows: Obtain a set of difficult examples. Using knowledge distillation methods, train the student model based on the teacher model to obtain the updated version parameters of the student model. Encapsulate the updated version parameters to obtain an upgrade package, and send the upgrade package to the edge. It should also be noted that the student model is the object detection model at the edge. The teacher model can be a large-scale vision model in the cloud, or a model with a larger number of parameters depending on actual needs.
[0091] There are five specific approaches to obtaining updated version parameters of the student model by distilling the teacher model using the knowledge distillation method: Option A The cloud platform periodically selects high-value samples from historical calibration samples to construct a set of Hard Examples to improve the capabilities of the edge-end model. The Hard Examples set includes: a) samples that were incorrectly predicted by the edge-end model but were confirmed correct after semantic calibration by the cloud platform; b) samples with high initial confidence levels on the edge-end model but ultimately classified as false positives; c) long-tail disease / event samples that appear multiple times and are representative of the scene; and d) newly added scene samples under seasonal, weather, and diurnal variations. The cloud platform uses the category probabilities, attribute labels, and text explanations output by the large visual model as teacher signals (Soft Labels), and combines this with the original annotations or verification results to perform knowledge distillation training on the lightweight object detection model at the edge-end.
[0092] For the nth difficult example in the difficult example set, let the classification Logits vector output by the teacher model be denoted as . The student model outputs the classification Logits vector as follows: The distillation temperature parameter τ is used. d The outputs of the teacher model and the student model are softened to obtain the prediction probabilities of the teacher model and the student model, respectively, as shown in formulas (11) and (12): (11); (12); in, This represents the probability that the nth hard case sample belongs to the kth category of the teacher's soft label. This represents the student's soft label probability that the nth difficult example belongs to the kth class, i.e., the predicted probability that the student model judges the nth difficult example to belong to the kth class. Distillation temperature parameter τ d This is used to adjust the smoothness of the prediction probability distributions of the teacher model and the student model, so that the student model can learn the relative preference relationships of the teacher model for each category.
[0093] Distillation loss L KD The Kullback-Leibler divergence (KL divergence) between the teacher and student distributions is expressed as: (13).
[0094] Meanwhile, to ensure that the student model is still constrained by the true labels, a hard-label classification supervision loss is introduced into the student model. If the true label of the nth sample is represented in one-hot encoding form as... The regular classification probability of the student model is expressed as: The classification loss is then expressed as: (14).
[0095] in, It can be obtained by Softmax calculation from the classification output of the student model without distillation temperature scaling.
[0096] Furthermore, for the target localization branch in the target detection task, let the predicted bounding box of the student model be... The corresponding supervised bounding box is Then the bounding box regression loss can be expressed as: (15); in, This represents the smoothed L1 loss function, used to measure the positional deviation between the predicted bounding box and the supervised bounding box.
[0097] Alternatively, the generalized intersection-union loss can be expressed as: (16); Wherein, GIoU represents the generalized intersection-union ratio between the predicted bounding box and the supervised bounding box. The larger the GIoU, the higher the overlap and geometric consistency between the predicted and supervised boxes, and therefore the smaller the corresponding loss.
[0098] The total loss function for distillation training (i.e., the first total loss function) It can be represented as: + + (17); in, This represents the distillation loss value. For classification loss value, L represents the bounding box regression loss or the bounding box intersection-union ratio loss. attr L represents the attribute supervision loss value. sem The semantic consistency loss value corresponding to the text interpretation. The weighting coefficient for distillation loss values. These are the weighting coefficients for the classification loss values. These are the weighting coefficients for the bounding box regression loss value or the weighting coefficients for the bounding box intersection-union ratio loss value. These are the weighting coefficients for the attribute supervision loss values. represents the weighting coefficients for the semantic consistency loss value. Through the above joint optimization, it can be ensured that the model at the edge not only learns the category preferences and soft discriminative boundaries of teacher models such as large visual models, but also retains the ability to fit the real annotations and the ability to accurately regress the target location.
[0099] During training, the total loss function of distillation training. When the training cycle reaches a preset cycle threshold, or the decrease in the total loss function is less than a preset threshold ε, the current round of distillation training can be considered complete, and the updated version parameters of the student model can be obtained. Both the preset cycle threshold and the preset threshold ε can be set according to actual needs.
[0100] Solution A uses the class probabilities, attribute labels, and textual explanations output by the large visual model in the cloud as a distillation method for the teacher signal, and performs periodic distillation training on the lightweight model at the edge. This main solution can simultaneously utilize class information, attribute information, and semantic explanation information, enabling the edge model to not only learn the discrimination results of the teacher model, but also its high-order semantic understanding of complex scenes, long-tail defects, and easily confused targets. Furthermore, the teacher signal has a clear structure, facilitating integration with subsequent distillation training and OTA upgrade processes, and exhibits good feasibility and stability.
[0101] Option B In Scheme B, the teacher signal, in addition to category probabilities, further includes multi-dimensional information such as attribute labels, text interpretations, salient regions, segmentation masks, or intermediate semantic features, and trains the student model using joint distillation loss. Compared to distillation using only category probabilities, Scheme B enables the student model to simultaneously learn the teacher model's category discrimination logic, attribute representation ability, spatial attention regions, and intermediate semantic representation ability.
[0102] Step ①: Select high-value samples from the historical calibration samples in the cloud to construct the distillation training set D. hard The distillation training set can use the same data source and basic screening criteria as Scheme A, both derived from high-value difficult examples in historical calibration samples. The difference lies in that Scheme B selects samples that simultaneously possess one or more additional teacher information, including attribute labels, semantic embeddings corresponding to textual interpretations, salient regions or segmentation masks, and intermediate semantic features, to support multi-dimensional teacher signal joint distillation. Difficult example samples may include: category-corrected samples, high-confidence false positive samples at the edge, long-tail disease samples, complex background samples, and newly added scene samples.
[0103] Step ②: For the nth training sample in the distillation training set, the visual large model outputs a multi-dimensional teacher signal, which includes at least: class probability p n teacher Attribute label vector a n teacher The semantic embedding vector e corresponding to the text interpretation n teacher , salient region or segmentation mask M n teacher and optional intermediate semantic features fn teacher .
[0104] Step ③: The student model at the edge outputs the corresponding result for the same sample. The output result for this training sample includes: classification probability p. n student Attribute prediction vector a n student intermediate semantic features f n student And the salient region or segmentation result M n student .
[0105] Step 4: Construct the classification distillation loss, attribute distillation loss, semantic feature alignment loss, salient region alignment loss, hard label classification loss, and bounding box regression loss respectively.
[0106] Step 5: Perform a weighted summation of the various losses to obtain the total loss function. The parameters of the student model are then updated using gradient descent or other optimization methods.
[0107] Step 6, the total loss function of distillation training When the training cycle reaches the preset cycle threshold, or the decrease in the total loss function is less than the preset threshold ε, the current round of distillation training can be determined to be completed, and the updated version parameters of the student model can be obtained.
[0108] Specifically, for the nth sample, the classification probability of the teacher model is... Classification probabilities of the student model They are represented as follows: (18); (19); Classification Distillation Loss Represented as: (20).
[0109] Property Distillation Loss Represented as: (twenty one).
[0110] Where R is the total number of attribute tags, and r is the index of the attribute tag. and Let represent the output vectors of the teacher model and the student model on the r-th attribute label, respectively.
[0111] Semantic feature alignment loss Represented as: (twenty two).
[0112] Significant region or segmentation mask alignment loss Represented as: (twenty three).
[0113] In formula (23), the subscript 2 represents the L2 norm.
[0114] Or it can be expressed as: (twenty four).
[0115] In addition, to ensure that the student model is still constrained by the true labels, hard-label classification loss is used. and bounding box regression loss They are represented as follows: (25); (26) or (27).
[0116] Second total loss function for: (28); in, , , , , and These are the weighting coefficients for the corresponding losses.
[0117] In Solution B, the student model not only learns the category preferences of the teacher model but also learns the teacher model's discrimination logic regarding target morphology, texture, spatial regions of interest, and semantic interpretation. Therefore, it exhibits stronger adaptability to complex backgrounds, long-tailed defects, and abnormal targets in open scenes, making it particularly suitable for scenarios requiring further enhancement of the semantic representation capabilities of edge models and robustness to complex scenes. Solution B strikes a trade-off between the richness of teacher signals and training complexity, making it more suitable for deployment scenarios with ample cloud training resources and high requirements for complex scene recognition capabilities.
[0118] Option C: High-confidence sample screening distillation Scheme C does not directly use all the output results of the large cloud-based visual model for distillation training. Instead, it first evaluates the credibility of the teacher model's output results and retains only high-confidence, high-consistency samples as distillation targets. The core of this scheme lies in reducing the negative impact of low-quality teacher signals and noisy labels on student model updates through a sample selection mechanism, thereby improving the stability and reliability of distillation training.
[0119] Step ①: Extract the candidate sample set Dcand from the semantic calibration results in the cloud. The candidate samples include: category-corrected samples, open-class anomaly samples, high-confidence false detection samples from the edge, and complex scene samples that have been manually verified.
[0120] Step ②: For the nth sample in the candidate sample set, calculate the teacher credibility score Q of the sample by combining the class probability of the teacher model, semantic consistency score, attribute matching degree, negative sample conflict degree, and optional human-machine review results. i .
[0121] Step ③: Teachers with credibility scores higher than the preset threshold τ are considered trustworthy. t Alternatively, the teacher credibility of the candidate samples can be ranked from highest to lowest, and the top M candidate samples can be included in the high-confidence distillation training set D. distill Among them, the preset threshold τ t The specific values of M can be set according to actual needs.
[0122] Step ④, only for the high-confidence distillation training set D distill The training samples are used for distillation training, and samples with low teacher confidence scores are not directly involved in distillation updates.
[0123] Step 5: For the selected samples (i.e., the high-confidence distillation training set D)... distill The training samples are used to construct classification distillation loss, hard label classification loss, and bounding box regression loss, and the version parameters of the student model are updated accordingly.
[0124] Step 6, the total loss function of distillation training When the training cycle reaches the preset cycle threshold, or the decrease in the total loss function is less than the preset threshold ε, the current round of distillation training can be determined to be completed, and the updated version parameters of the student model can be obtained.
[0125] Specifically, the teacher credibility score Q of the nth training sample n Represented as: (29); in, This indicates that the teacher model applies to the target category c. The predicted probability, Let Match(A) represent the semantic consistency score of the nth sample. n ,F c ) indicates the attribute label and target category c The degree of attribute matching between key identification attributes, Conflict(A n N c ) indicates the attribute label and target category c The degree of conflict between negative sample features , , and These are the corresponding weighting coefficients. For example, For the teacher model, the target category c The weighting coefficients corresponding to the predicted probabilities.
[0126] When Q n ≥τ t At that time, the training sample x n Included in the high-confidence distillation training set D distill ,Right now Alternatively, use a sorting and filtering method: That is, the top M samples with the highest teacher credibility scores are selected as the distillation training objects.
[0127] For the nth sample after screening, the distillation temperature parameter τ can still be used. d Construct the classification probabilities of the teacher model and the student model: (30); (31).
[0128] Only for the high-confidence distillation training set D distill Calculation of categorical distillation loss in samples for: (32).
[0129] Meanwhile, to ensure that the student model retains its ability to learn from real annotations, the distillation training set D is used. distill The selected samples can then be further analyzed to calculate the classification loss based on their corresponding hard labels. With bounding box regression loss : (33); (34) or (35).
[0130] Third total loss function Represented as: (36); in, , , These are the weighting coefficients for the corresponding loss terms.
[0131] Scheme C effectively filters out low-quality teacher signals and noisy samples, reducing the risk of negative transfer from erroneous distillation, thereby improving the stability of distillation training and the reliability of model updates. It is particularly suitable for open scenarios, complex backgrounds, and situations where teacher outputs may be unstable. In comparison, Scheme C achieves a more balanced effect between utilizing sufficient teacher signals and training stability, reducing reliance on extremely high-confidence screening conditions while maintaining high update efficiency. Therefore, Scheme C is more suitable for continuous online update scenarios.
[0132] After the AC (Acoustic Coding) scheme completes training, the cloud generates version parameters for an updated version (i.e., a new version) of the object detection model at the edge. These version parameters include: various thresholds associated with the object detection model, a preset threshold η, a confidence threshold range, a morphological anomaly threshold range, a distribution threshold, a set of confused category pairs, a prompt word template version number, and post-processing parameter configurations. The cloud then distributes the updated version parameters to the edge via OTA (Over-The-Air) technology.
[0133] The AC (Acceptance and Response) solution preferably adopts a closed-loop update approach based on knowledge distillation. This involves using teacher signals—such as category probability distributions, attribute labels, and textual interpretations—generated by a large cloud-based visual model to periodically distill and update the lightweight edge model. This is combined with a secure and reversible OTA (Over-The-Air) mechanism to ensure continuous model evolution. The advantages of AC are: unified control of the model update process in the cloud, enabling the edge model to continuously learn the discrimination logic of the large visual model for complex scenes, while maintaining the centralization, consistency, and good version controllability of the update process. Furthermore, the training objectives, implementation paths, and deployment methods in the cloud are relatively clear, facilitating integration with edge object detection models, version management systems, and OTA upgrade mechanisms, resulting in good engineering feasibility and system stability.
[0134] To enhance the adaptability of centralized update paths in scenarios involving distributed deployment across multiple devices or strong privacy constraints, a solution DE is provided for model update and upgrade schemes at the edge.
[0135] Option D Solution D abandons knowledge distillation as the primary update method and instead employs an active learning framework. The cloud selects the most informative, representative, or uncertain samples from historical calibration results and sends them back to the edge. The edge then performs incremental learning or online learning locally. The core of this solution is that instead of having the edge model learn all teacher outputs, it prioritizes learning from a small number of "most worthwhile" difficult examples to reduce update costs and improve sample utilization efficiency.
[0136] The process of Plan D includes the following steps: Step ①: Construct a candidate sample set D in the cloud from historical semantic calibration results, manual review results, and edge misjudgment records. cand D cand The candidate samples include: high-confidence false detection samples at the edge, conflicting judgment samples between the edge and the cloud, long-tail category samples, complex background samples, and open-class anomaly samples.
[0137] Step 2: For the candidate sample set D cand For each candidate sample, calculate the information score I. n The information content score can comprehensively consider factors such as the uncertainty of the edge model prediction, the degree of difference between edge and cloud discrimination, sample representativeness, and the historical frequency of the sample.
[0138] Step 3: Sort the candidate samples in descending order based on their information content scores, and select those with information content higher than the preset threshold τ. I The samples, or the candidate samples ranked in the top M positions, constitute the active learning training set D. active M can be set to a specific value according to actual needs.
[0139] Step 4: The edge device receives the active learning training set D from the cloud. active And will actively learn from the training set D active With a small number of historical stable samples or original replay samples D replay Together they form the incremental training set D inc This is to mitigate the catastrophic forgetting problem that may occur during the incremental learning process of the model.
[0140] Step 5: Utilize the incremental training set D at the edge. inc Fine-tuning or small-batch online updates of the local model yields the updated model parameters, i.e., the version parameters of the updated version.
[0141] Step 6: When the incremental training rounds reach a preset round threshold, or when at least one of the following conditions is met: the mAP (mean Average Precision) of the edge-end model on the validation sample set is higher than the preset mAP threshold, the F1 score (harmonic mean of precision and recall) is higher than the preset F1 threshold, the false positive rate decreases more than the preset false positive threshold, the false negative rate decreases more than the preset false negative threshold, and the recall rate for long-tail disease identification is higher than the preset recall threshold, the incremental training update for this round is completed, and an updated version of the edge-end object detection model is generated. The preset mAP threshold, preset F1 threshold, preset false positive threshold, preset false negative threshold, and preset recall threshold can all be set according to actual needs.
[0142] Candidate sample set D candInformation score of the nth candidate sample Represented as: (37); Among them, U n Δ represents the uncertainty score of the nth candidate sample. n R represents the degree of difference between the recognition results at the edge and the calibration results in the cloud. n F represents the representative score of the nth candidate sample. n The information content score represents the frequency of occurrence or scene coverage value of the nth candidate sample during the historical inspection process, with ω1, ω2, ω3, and ω4 being the corresponding weight coefficients. The higher the information content score, the more worthy the sample is of priority for active learning and updating.
[0143] Among them, the uncertainty score U of candidate sample n n The entropy predicted by the model at the edge is expressed as: (38); Where, p n (k) represents the predicted probability of the edge model that the nth candidate sample belongs to the kth class, U n The larger the value, the higher the uncertainty of the candidate sample.
[0144] The difference Δ between the recognition results at the edge and the calibration results in the cloud n Defined as: When c n ≠c n cloud At that time, Δ n = 1. When c n = c n cloud At that time, Δ n = 0. c n c is the predicted category of the recognition result at the edge. n cloud This is a correction category for the recognition results in the cloud.
[0145] Alternatively, it can be defined as a distance function based on the difference in probability distributions, for example: (39); Where, p n edge p represents the model prediction probability distribution at the edge. n cloud This represents the category probability distribution corresponding to the semantic calibration results in the cloud. KL() represents the KL divergence, i.e., the relative entropy.
[0146] Representative score R nThe similarity between the candidate sample features and other samples in the current sample library is determined based on their coverage relationship, for example, by defining it as the similarity between the sample feature vector and the cluster center: R n = Sim( f n , μ l (40); Where Sim() is the similarity function, f n μ represents the feature vector of candidate sample n. l This represents the cluster center of the cluster to which candidate sample n belongs. R n The larger the value, the more representative the candidate sample is of the typical characteristics of its sample cluster.
[0147] It should also be noted that the most informative sample is... Not less than the information content threshold (e.g., τ) I The most informative samples are those most valuable for subsequent model updates at the edge. An informative score is calculated by comprehensively considering the sample's uncertainty, the difference between edge-end recognition results and cloud calibration results, sample representativeness, and scene coverage value. Samples are then selected based on this informative score. A higher informative score indicates that the sample is more worthy of priority for active learning updates.
[0148] The most representative sample is R. n Samples with a representativeness threshold or higher. The most representative samples are those that can represent the common characteristics of a certain type of scenario, a certain type of disease pattern, or a certain type of misjudgment pattern. Clustering can be performed based on the characteristics of candidate samples, and then the representativeness score can be determined based on the similarity between the candidate sample and its cluster center. The higher the representativeness score, the more representative the sample is of the common characteristics of its sample cluster.
[0149] The sample with the greatest uncertainty is U n Samples with an uncertainty threshold not less than the threshold value. The most uncertain samples represent those at the edge of the model output with low stability of prediction results, which can be measured by the entropy value of the predicted probability distribution, the difference between the probability of the largest class and the probability of the second largest class, or the width of the confidence interval.
[0150] The information content threshold, representativeness threshold, and uncertainty threshold can be set according to actual needs.
[0151] After selecting the active learning training set D active After that, the sample The following selection rules may be adopted: (41) or (42).
[0152] To reduce the forgetting effect in incremental learning, it is preferable to use the replay sample set D.replay With active learning training set D active Together they form the incremental training set D inc : (43).
[0153] During the local incremental training phase, the loss function of the model at the edge... Represented as: (44); in, Indicates the loss under classified supervision. This represents the bounding box regression loss. This represents a parameter constraint or historical knowledge retention term, used to suppress excessive forgetting of previously learned knowledge during incremental updates.
[0154] Furthermore, the parameter constraint term is expressed as: (45). Where θ represents the currently updated model parameters at the edge, θ old This represents the parameters of the old model before the update.
[0155] Scheme D can update using only a small number of the most informative and representative samples, thus significantly reducing the amount of training data transmitted and the cost of edge updates, and improving sample utilization efficiency. It is particularly suitable for scenarios where edge devices have limited computing power, require high-frequency updates, but are not suitable for full retraining. Furthermore, by introducing replay samples and parameter constraints, it can alleviate the catastrophic forgetting problem to some extent. In contrast, Scheme D utilizes the teacher signal from a large cloud model for centralized distillation, which can more fully inherit the discriminative ability of the teacher model while ensuring the consistency of the overall update direction. Therefore, it has greater advantages in long-term stable iteration and global semantic transfer.
[0156] Option E Solution E does not directly upload the raw samples collected by each edge device to the cloud for unified training. Instead, it adopts a federated learning framework. This framework allows multiple edge devices to train or fine-tune their models locally based on their own collected data, uploading only local model parameter updates, gradient information, or model differences to the cloud. The cloud then aggregates these data to form a global model, which is then distributed back to the edge devices. The core of this solution is to enhance data privacy protection and distributed deployment capabilities while ensuring collaborative model evolution through a "data-without-the-edge, parameter-to-the-cloud" approach.
[0157] The process of Plan E includes the following steps: Step 1: Initialize global model parameters θ in the cloud. 0 and θ 0 It was distributed to multiple edge devices participating in federal training.
[0158] Step 2: During the r-th round of federated training, the g-th edge receives the current global model parameters. And based on the local dataset D g Perform F rounds of local training to obtain locally updated model parameters. .
[0159] Step 3: At each edge, do not upload the original images, videos, or labeled data; only upload the parameter differences of the local model. Up to the cloud.
[0160] Step 4: The cloud performs weighted aggregation of local model updates based on the sample size, device quality, or update reliability of each edge device to obtain new global model parameters. .
[0161] Step 5: The cloud will redeploy the updated global model to each edge device to begin the next round of federated training.
[0162] Step 6: When the improvement in global model accuracy in the cloud falls below the preset accuracy threshold, complete this round of federated updates and use the resulting global model as the new version model for subsequent deployment or canary release. The preset accuracy threshold is set according to actual needs.
[0163] The local parameter update of the g-th edge after the r-th round of federated training (i.e., the parameter difference of the local model). ) is represented as: (46); in, This represents the global model parameters before the r-th round of federated training. This represents the model parameters of the g-th edge after local training is completed, i.e., the locally updated model parameters.
[0164] The cloud can use a federated average (FedAvg) method to aggregate local update results from multiple edge endpoints. The aggregation formula is expressed as follows: (47); Where N represents the total number of edge nodes participating in federated training, l is the index of the edge node participating in federated training, and n g This represents the number of local training samples at the g-th edge. This represents the global model parameters for the new round after aggregation.
[0165] If aggregation is performed using model differencing, it is represented as: (48).
[0166] For each edge, the local training objective function of the edge is... Represented as: (49); Among them, D g Let l(x) represent the local training set of the g-th edge. n ;θ) represents sample x n The single-sample loss function corresponding to the model parameters θ.
[0167] If the data distribution differs significantly between the edge models, a regularization term can be added to the local objective function to constrain the deviation between the local and global models. For example: (50); Where μ is the regularization coefficient, used to reduce the impact of local model updates on the stability of the global model under non-independent and identically distributed data conditions. Let θ represent the local training objective function for the g-th edge after introducing the proximal regularization term; r This represents the current global model parameters distributed from the cloud during the r-th round of federated training; ||·||_2^2 represents the squared L2 norm.
[0168] Preferably, to further improve the reliability of the aggregation results, the cloud can also assign an updated credibility weight w to the g-th edge. g The weights can be dynamically determined based on the quality of local samples at the edge, model stability, historical error levels, or network state. Therefore, the aggregation formula can be extended to: (51). Among them, the following conditions are met: (52).
[0169] Solution E enables collaborative evolution among multiple edge devices without uploading raw data, reducing privacy risks and communication burdens associated with centralized data uploads. It is particularly suitable for distributed deployment environments involving multiple regions, multiple devices, and sensitive scenario data. Simultaneously, federated learning allows different edge devices to learn features in their respective local scenarios, which are then aggregated in the cloud to form a global model, thereby enhancing the system's coverage of multiple scenario distributions.
[0170] In summary, Scheme D is more suitable for lightweight model update scenarios where edge computing power is limited but frequent, small-step iterative updates are required. Alternative Scheme E is more suitable for networked collaborative update scenarios that emphasize privacy protection, multi-device collaboration, and distributed deployment capabilities. In comparison, the primary scheme currently preferred by this invention achieves a more balanced effect among update consistency, full utilization of teacher semantics, version control capabilities, deployment costs, and engineering feasibility. Therefore, it is more suitable as the preferred model update method for long-term online operation and large-scale deployment of road defect and event recognition systems.
[0171] In addition, the OTA (Over-The-Air) delivery process in the cloud is as follows: 1) The cloud-based model management service encapsulates the version parameters of the updated version, forming an upgrade package that includes the weight file, configuration file, version description file, and hash value of the target detection model. 2) The upgrade package digest is digitally signed using a private key to prevent tampering during transmission. 3) Under the condition that the network environment meets the requirements of network quality, remaining battery power, and device idle time, the edge device initiates a version check request to the cloud. 4) If the edge device detects an updated version in the cloud, it downloads the full or differential upgrade package and performs signature verification, hash verification, and compatibility verification locally. 5) After the verification passes, the edge device writes the version parameters of the updated version to a backup partition or backup running directory. 6) The edge device performs a smoke test locally, including sampling inference, memory usage check, and inference latency check. 7) After the test passes, the edge device switches the inference service to the updated version model and reports a successful upgrade receipt to the cloud.
[0172] OTA (Over-The-Air) releases can employ a canary release strategy, initially deploying the updated version to a small number of edge devices and monitoring their recognition accuracy, latency, resource consumption, and stability. Once preset operating conditions are met, the rollout is gradually expanded. These preset operating conditions can be set according to actual needs. For example, preset operating conditions may include one or more of the following: the updated model's recognition accuracy on the canary devices reaches a preset accuracy requirement, or the accuracy decrease compared to the original model does not exceed a preset threshold; average inference latency, peak inference latency, memory consumption, and processor consumption do not exceed corresponding preset thresholds; there are no model loading failures, inference anomalies, service crashes, or abnormal rollbacks within a preset continuous running time; and the upgrade success rate is higher than a preset success rate threshold, and the anomaly alarm rate is lower than a preset alarm threshold. If the preset operating conditions are met, the updated version can be gradually rolled out to more edge devices.
[0173] To address the shortcomings of existing solutions, such as a lack of continuous evolution capabilities and difficulty in adapting to environmental changes and emerging diseases during long-term operation, this embodiment provides a distillation-based closed-loop update method based on large-scale model teacher signals. By continuously accumulating high-quality, challenging example samples calibrated and confirmed in the cloud, knowledge distillation and iterative optimization are performed on the edge-end model. Combined with a secure OTA upgrade mechanism possessing signature verification, smoke testing, and anomaly rollback capabilities, the edge recognition model achieves continuous evolution and adaptive enhancement to new scenarios. Compared to solutions relying on fixed initial calibration files, one-time training, or offline updates, this embodiment can adapt more quickly to seasonal changes, lighting variations, new diseases, and new scenario events, extending the system's effective lifespan, reducing manual maintenance and repeated deployment costs, and improving long-term operational stability and the system's sustainable evolution capabilities.
[0174] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages: 1. Cloud-based semantic calibration mechanism based on structured VQA prompts This embodiment proposes a process for automatically retrieving candidate category sets, negative sample descriptions, key discriminative attributes, and output template prompt words based on the initial category judgment at the edge. This enables the large visual model to complete secondary discrimination within a controlled semantic space, and combines semantic consistency scoring to achieve category correction and attribute completion. This mechanism differs from existing schemes that rely solely on fixed standard answer comparisons, and also from interactive defect segmentation schemes that rely on point or box prompts. It can perform dynamic semantic error correction for newly added defects, complex backgrounds, and easily confused targets in open road scenes, and is the core technical foundation for achieving high-precision recognition in this invention.
[0175] 2. Online Confidence Alignment and Risk Classification Methods for Open Scenarios This embodiment proposes a temperature-based confidence calibration method based on the original Logits of an edge model. This method can be further integrated with cloud-based semantic probabilities to obtain a comprehensive confidence level that can be directly mapped to risk classification and business threshold control. This method addresses the problem of traditional detection models having "high scores but not necessarily high reliability," transforming confidence level from a simple internal model score into a credible indicator that supports business decisions. It provides a unified quantitative basis for result fusion, alarm classification, and manual review and ranking.
[0176] 3. Spatiotemporal fusion and structured upload mechanism for business closed loop This embodiment associates edge detection results with cloud calibration results using target number, timestamp, spatial location (i.e., location information), and bounding box overlap. It prioritizes cloud-based correction categories and calibration confidence levels, and uploads the correction basis, risk level, image evidence, attribute labels, and work order suggested actions to the platform in the form of a structured event report. This mechanism directly transforms the recognition algorithm results into executable business objects, achieving an effective connection from "recognition" to "handling," which is a key innovative aspect of this invention, distinguishing it from traditional solutions that only output detection results.
[0177] 4. Continuous evolution architecture based on large model teacher signals, distillation training, and secure OTA. This embodiment designs a complete update chain of "difficult example accumulation—teacher distillation—version packaging—signature verification—canary release—anomaly rollback". Its core lies in continuously improving the edge model using soft tags, attribute tags, and explanatory information generated by the cloud-based visual large model, and securely distributing model parameters and configurations through a verifiable and rollback-enabled OTA mechanism. While ensuring the continuous operation of online inspection services, this allows the system to adapt to new scenarios, new diseases, and environmental changes over the long term, achieving a balance between continuous evolution and stable operation.
[0178] Example 3 Based on the same inventive concept, the third embodiment of the present invention also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of any of the methods in the road recognition method based on a large visual model of embodiment two.
[0179] Example 4 Based on the same inventive concept, the fourth embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods of the road recognition method based on a large visual model described in the second embodiment above.
[0180] Those skilled in the art will understand that although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the invention.
[0181] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A road recognition method based on a large visual model, characterized in that, Applied to the edge, the method includes: Each road surface image in the road surface image stream is acquired, and the road surface images are processed by a target detection model to obtain a target set and target information for each target in the target set. The target information includes: target number, coordinates of the target bounding box, prediction category, original confidence score, original Logits vector, target image patch, context image, timestamp, location information, and road scene auxiliary information. In each target, based on the target information, it is determined whether the target is a calibration sample; If the target is a calibration sample, then the target information of the target is used as the sample information of the calibration sample, and the calibration sample and the sample information of the calibration sample are sent to the cloud, so that the cloud can perform calibration processing on the calibration sample through a large visual model to obtain the calibration information of the calibration sample, and then output the structured road report of the calibration sample. The calibration information includes: target category, target confidence and supplementary attributes. If the target is an uncalibrated sample, a structured road report for the target is obtained and output based on the target information of the target.
2. The road recognition method based on a large visual model as described in claim 1, characterized in that, The step of determining whether the target is a calibration sample based on the target information of the target includes: Based on the original confidence level of the target, the confidence uncertainty score of the target is obtained; Based on the predicted category of the target, the category confusion weakness score of the target is obtained; Based on the target bounding box of the target, the morphological anomaly score of the target is obtained; Based on the confidence uncertainty score, category confusion weakness score, and morphological anomaly score of the target, the distribution decision value of the target is obtained, as shown in the following formula: ; Wherein, Uconf(i) is the confidence uncertainty score of target i, Uweak(i) is the category confusion weakness score of target i, Umorph(i) is the morphological anomaly score of target i, α is the weight coefficient corresponding to the confidence uncertainty score, β is the weight coefficient corresponding to the category confusion weakness score, and γ is the weight coefficient corresponding to the morphological anomaly score. If the distribution decision value of the target is not less than the distribution threshold, then the target is determined as the calibration sample; If the distribution decision value of the target is less than the distribution threshold, then the target is determined as the uncalibrated sample.
3. The road recognition method based on a large visual model as described in claim 2, characterized in that, After obtaining the confidence uncertainty score, the category confusion weakness score, and the morphological anomaly score of the target, the following is also included: If the confidence uncertainty score is within the confidence threshold range and the predicted category belongs to the confused category pair set, then the target is determined as the calibration sample; Alternatively, if the score of the morphological anomaly item is within the morphological anomaly threshold range, then the target is determined as the calibration sample.
4. The road recognition method based on a large visual model as described in claim 1, characterized in that, Also includes: Under the specified network environment, receive the upgrade package sent by the cloud; After the upgrade package passes verification, an updated version of the target detection model is obtained, and the updated version is written to a spare partition or a spare running directory. After the updated version passes the preset test, the current version of the target detection model is updated to the updated version to realize the update and upgrade of the target detection model.
5. A road recognition method based on a large visual model, characterized in that, Applied to the cloud, the method includes: Obtain the calibration sample sent from the edge and the sample information of the calibration sample; Based on the sample information, the VQA cue words of the calibration sample are obtained. Then, through the visual big model, the calibration information of the calibration sample is obtained according to the sample information and the VQA cue words. The calibration information includes: target category, target confidence and supplementary attributes. The calibration information and sample information of the calibration sample are fused together to obtain and output the structured road report of the calibration sample.
6. The road recognition method based on a large visual model as described in claim 5, characterized in that, Obtaining the target confidence level of the calibration sample includes: The original Logits vector in the sample information is scaled according to the optimal temperature parameter to obtain the post-calibration confidence of the calibration sample belonging to each candidate category, and then the target confidence is obtained, as shown in the following formula: ; Where K is the total set of categories, k represents the index of the target category, m represents the category traversal index in the total set of categories, and z i =(z i,1 , z i,2 , ..., z i,K ) represents the original Logits vector of the calibration sample; z i,K z is the original Logits value corresponding to the Kth category in the original Logits vector. i,m p is the original Logits value corresponding to the m-th category in the original Logits vector. cali,i (k) represents the post-calibration confidence that the calibration sample belongs to the k-th category, T The optimal temperature parameter is defined as follows; If the target category is the predicted category c in the sample information i Then, based on the predicted category, the calibrated confidence level p is obtained. cali,i (c i ), and the post-calibration confidence level p cali,i (c i The target confidence level is determined as follows; If the target category is a calibrated category, then the calibrated confidence level p is obtained based on the calibrated category. cali,i (c new ), and the post-calibration confidence level p cali,i (c new The target confidence level is determined as follows: the calibrated category is the corrected category c after semantic calibration using the large visual model. new ; When the predicted category belongs to a set of confused category pairs, a comprehensive confidence score is obtained based on the calibrated confidence score of the target category and the category probability of the target category obtained through the visual big model, and the comprehensive confidence score is determined as the target confidence score.
7. The road recognition method based on a large visual model as described in claim 5, characterized in that, The process of fusing the calibration information and the sample information of the calibration sample to obtain and output the structured road report of the calibration sample includes: Based on the target number, timestamp difference, intersection-over-union ratio of the target bounding box, distance and location information of the center point of the target bounding box, the sample information and the calibration information are associated and fused to obtain the fused information of the calibration sample, wherein the category of the fused information is the target category, and the confidence level of the fused information is the target confidence level; If the target confidence level is not less than the first target confidence level threshold, then the structured road report is generated directly based on the fusion information, and the structured road report is output. If the target confidence level is less than the second target confidence level threshold, the fused information is output to the manual inspection queue for staff to review and inspect, wherein the second target confidence level threshold is less than the first target confidence level threshold. If the target confidence level is not less than the second target confidence level threshold, and the target confidence level of the calibration information is less than the first target confidence level threshold, then the supplementary attribute is retained, and the fusion score of the candidate category is obtained according to the candidate category in the fusion information. The candidate category corresponding to the maximum fusion score is determined as the final category of the calibration sample, and the maximum fusion score is determined as the final confidence level of the final category. Based on the supplementary attribute, the final category, the final confidence level, and the added manual review mark, the structured road report is generated and the structured road report is output.
8. The road recognition method based on a large visual model as described in claim 5, characterized in that, Also includes: Obtain the set of difficult examples; The student model is trained using a knowledge distillation method based on the teacher model to obtain updated version parameters of the student model. The student model is the target detection model at the edge. The first total loss function used in the distillation training is: + + ; in, This is the first total loss function value. This represents the distillation loss value. For classification loss value, L represents the bounding box regression loss or the bounding box intersection-union ratio loss. attr L represents the attribute supervision loss value. sem The semantic consistency loss value corresponding to the text interpretation. The weighting coefficient for the distillation loss value. The weighting coefficients for the classification loss values are... The weighting coefficients for the bounding box regression loss value or the weighting coefficients for the bounding box intersection-union ratio loss value. The weighting coefficients for the attribute supervision loss value are... The weighting coefficients for the corresponding semantic consistency loss values; Alternatively, the second total loss function used in the distillation training is: ; in, This is the second total loss function value. This represents the distillation loss value. For classification loss value, L represents the bounding box regression loss or the bounding box intersection-union ratio loss. attr The attribute supervision loss value, Alignment loss values with semantic features For salient regions or segmentation mask alignment loss values, , , , , and These are the weighting coefficients for the corresponding losses; Alternatively, the third total loss function used in the distillation training is: ; in, This is the third total loss function value. For training set D based on high confidence distillation distill The resulting fractional distillation losses, For the high-confidence distillation training set D distill The obtained classification loss, For the high-confidence distillation training set D distill The obtained bounding box regression loss, , , These are the weighting coefficients for the corresponding loss terms; The version parameters of the updated version are encapsulated to obtain an upgrade package, and the upgrade package is sent to the edge device.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the road recognition method based on a large visual model as described in any one of claims 6-8.
10. A computer-readable storage medium storing a computer program thereon, characterized in that, When executed by a processor, the program implements the steps of the road recognition method based on a large visual model as described in any one of claims 6-8.