Control system and method based on unattended road patrol wrecker

Through the combination of multimodal perception module and deep neural network, the problem of insufficient identification accuracy and data processing intelligence in complex environments is solved, and accurate assessment of vehicle damage and intelligent selection of the cleaning mode is realized, which significantly improves the efficiency and safety of cleaning operations.

CN120220116AActive Publication Date: 2025-06-27FUJIAN PINGTAN RUIQIAN INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510381147.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

The existing unmanned clearance system has shortcomings in identification accuracy and data processing intelligence, and it is difficult to accurately evaluate vehicle damage in complex environments, resulting in low efficiency of clearance operations and frequent misjudgment.

Method used

The multimodal perception module is adopted, combining image sensors, laser sensors and depth sensors, and the feature extraction and fusion of multiple modal data is carried out through the deep neural network and self-attention mechanism of the dual-current architecture, and a confidence score is generated to select the barrier mode.

Benefits of technology

It improves the accuracy and comprehensiveness of damage to the faulty vehicle, reduces the probability of misjudgment, avoids the risk of secondary damage, and improves the efficiency and safety of the obstacle cleaning operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220116A_ABST
    Figure CN120220116A_ABST
Patent Text Reader

Abstract

The invention provides a control system and method based on an unmanned road guarding patrol wrecker, and relates to the technical field of automatic obstacle clearance, the control system comprises a multi-mode sensing module composed of an image sensor, a laser radar and a depth sensor, and the multi-mode sensing module obtains two-dimensional image data, three-dimensional point cloud data and three-dimensional depth data; a hierarchical detection module combining a double-flow deep neural network and a self-attention mechanism is used for carrying out data feature fusion and outputting multi-level scores, accurate evaluation of vehicle damage and intelligent selection of an obstacle clearance mode are realized, and the system is continuously optimized through adaptive learning. Through multi-sensor cooperation, double-flow deep neural network feature extraction and a self-attention dynamic fusion mechanism, high-precision vehicle state analysis and intelligent obstacle clearance decision making are realized, intelligent and accurate judgment of road fault vehicle damage conditions is realized, secondary damage risks caused by damage assessment errors are reduced, and the road fault vehicle damage assessment accuracy is improved. The frequency of manual intervention is reduced, and the efficiency of road obstacle clearing operation is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of automatic obstacle removal, and particularly to a control system and method based on an unmanned road patrol and obstacle removal vehicle. Background Art

[0002] With the rapid development of artificial intelligence technology, autonomous driving technology, and intelligent transportation technology, an intelligent and automated road management system centered around unmanned vehicles has received increasing attention. Among them, an unmanned road patrol and obstacle removal vehicle, as a device that can autonomously patrol, promptly detect road breakdown vehicles, and quickly implement effective disposal, has become one of the important development directions of intelligent transportation and smart cities. The obstacle removal vehicle usually selects the obstacle removal method according to the damage condition of the vehicle. For example, for a breakdown vehicle with less damage, the obstacle removal vehicle towes it to the support part by connecting the towing hook of the vehicle; for a breakdown vehicle with severe damage, the obstacle removal vehicle connects the four wheels of the breakdown vehicle through a hoisting mechanism, hoists it to the support part, or by means of lifting, raises and fixes one end of the vehicle and directly drags it for obstacle removal.

[0003] In recent years, researchers have proposed some automated or semi-automated obstacle removal technical solutions, relying on cameras or lidar to identify by optically visual recognition of vehicle appearance features. However, the current unmanned obstacle removal systems generally have problems such as low recognition accuracy and single data processing methods, and it is difficult to accurately and effectively evaluate the damage condition of vehicles in complex environments, resulting in the inevitable need for manual intervention in the actual obstacle removal operation process. At the same time, due to the lack of intelligence in data processing and decision-making of traditional unmanned obstacle removal vehicles, they cannot accurately distinguish the obstacle removal methods corresponding to different damage degrees (such as towing, lifting, or hoisting), often making misjudgments, thus reducing the efficiency of the obstacle removal work and even causing secondary damage accidents.

[0004] Therefore, in order to meet the higher requirements of modern traffic management for the high efficiency, intelligence, and reliability of obstacle removal operations, it is necessary to develop an unmanned obstacle removal control system and method with more accurate damage assessment capabilities, more flexible operation decision-making methods, and more intelligent data fusion processing mechanisms. Summary of the Invention

[0005] In order to overcome the deficiencies of the prior art, the technical problem to be solved by the present invention is to propose a control system and method based on an unmanned road patrol and obstacle removal vehicle, adopting the following technical solutions:

[0006] On the one hand, the present invention provides a control system based on an unmanned road patrol and obstacle removal vehicle, including:

[0007] A multimodal perception module, at least including:

[0008] An image sensor for acquiring two-dimensional image data of a faulty vehicle and extracting body identification information through OCR;

[0009] A laser sensor for generating three-dimensional point cloud data of the appearance of a faulty vehicle;

[0010] A depth sensor for generating three-dimensional depth data of the local appearance of a faulty vehicle;

[0011] A hierarchical detection module, including a depth neural network computing model with a dual-stream architecture:

[0012] The first stream network uses a convolutional neural network to extract the image feature vector F of the above two-dimensional image data img , and outputs the first confidence score C1;

[0013] The second stream network uses a point cloud processing algorithm to extract the point cloud feature vector F of the above three-dimensional point cloud data pc and the depth feature vector F of the three-dimensional depth data dep ; superimpose and fuse the above two-dimensional image data and three-dimensional point cloud data, and output the second confidence score C2; if the above second confidence score C2 ≤ 85, then update and fuse the above three-dimensional point cloud data and three-dimensional depth data, and output the third confidence score C3;

[0014] The feature fusion unit uses a self-attention mechanism to weight-fuse the above image feature vector F img , point cloud feature vector F pc and depth feature vector F dep , and generates a fused feature vector F fusion ;

[0015] The control processing unit is connected to the above hierarchical detection module and selects a breakdown clearing mode according to a preset threshold rule. When the above third confidence score C3 ≤ 85, it requests manual intervention.

[0016] For further improvement, the above image sensor scans the front and rear of the faulty vehicle, and the calculation formula of the above first confidence score C1 is:

[0017] C1 = 100·σ(W1·F img +b1)

[0018] where F img represents the image feature vector output by the first-stream convolutional neural network;

[0019] W1 is the weight matrix and b1 is the bias term;

[0020] σ(·) is the Sigmoid activation function, which maps the above first confidence score C1 to the range of 0 to 100.

[0021] Scan the front and rear of the faulty vehicle. Usually, the front and rear are marked with information such as the vehicle brand logo and vehicle model. On the one hand, the image sensor judges the damage condition of the front and rear of the vehicle, and on the other hand, it collects the vehicle information of the faulty vehicle through OCR.

[0022] For further improvement, the above laser sensor scans both sides and the roof of the faulty vehicle. The calculation formula for the above second confidence score C2 is:

[0023] C2 = 100·σ(W2·F fusion(img,pc) +b2)

[0024] where, W2 is the weight matrix and b2 is the bias term;

[0025] σ(·) is the Sigmoid activation function, which maps the above second confidence score C2 to the range of 0 - 100;

[0026] The fused feature vector F fusion(img,pc) is generated by fusing the above two-dimensional image data and three-dimensional point cloud data through a self-attention mechanism. The specific formula is:

[0027] F fusion(img,pc) = γ1·Attention(Q img ,K img ,V img ) + γ2·Arrention(Q pc ,K pc ,V pc )

[0028] The above weight coefficients γ1 and γ2 satisfy γ1 + γ2 = 1, and when the first confidence score C1 > 75, that is, when the front and rear of the faulty vehicle are less damaged, the weight coefficient γ1 is reduced and the weight coefficient γ2 is increased.

[0029] For further improvement, the above depth sensor scans the wheels and axle areas of the faulty vehicle. The calculation formula for the above third confidence score C3 is:

[0030] C3 = 100·σ(W3·F fusion(pc,dep) +b3)

[0031] where, W3 is the weight matrix and b3 is the bias term;

[0032] σ(·) is the Sigmoid activation function, which maps the above third confidence score C3 to the range of 0 - 100;

[0033] The fused feature vector F fusion(pc,dep) is generated by fusing the above two-dimensional image data and three-dimensional point cloud data through a self-attention mechanism. The specific formula is:

[0034] F fusion(pc,dep) = (1 - M) ⊙ [γ2' · Attention(Q pc , K pc , V pc )] + M ⊙ [γ3 · Attention(Q dep , K dep , V dep )]

[0035] Wherein, M is a mask matrix, and the overlapping part of the above three-dimensional point cloud data and the above three-dimensional depth data is updated to the above three-dimensional depth data through the mask matrix M;

[0036] The above weight coefficients γ2' and γ3 satisfy γ2' + γ3 = 1, and the weight coefficient γ3 > γ2'.

[0037] For further improvement, it further includes an adaptive learning module, which is connected to the above control processing unit, and optimizes the weight coefficients in the above feature fusion unit through offline learning and incremental learning.

[0038] For further improvement, the offline learning of the above weight coefficients adopts a loss function that minimizes the mean square error:

[0039]

[0040] Wherein, F target,i is the standard feature vector of the vehicle state annotated after manual intervention.

[0041] For further improvement, the incremental learning method of the above weight coefficients is:

[0042] When the recent system detection error shows a continuous increasing trend, the weight coefficients are updated using the following rules:

[0043]

[0044] Wherein, Accuracy is the proportion of tasks with an error within ±5 between the system and the measurement score and the manual verification score in the last 5 tasks;

[0045] Δ is the learning rate, initially limited between 0.01 and 0.05;

[0046] When Accuracy ≥ 95%, the update of the weight coefficients is stopped to maintain the parameter stability of the system.

[0047] For further improvement, the preset rule for the above control processing unit to select the obstacle clearing mode is:

[0048] When the above first confidence score C1 > 75 and the second confidence score C2 > 85, control the execution of the towing mode to tow the faulty vehicle at a speed not exceeding 20 km / h;

[0049] When the first confidence score C1 > 75 and the second confidence score C2 ≤ 85, if the above third confidence score C3 > 85, then control the execution of the lifting mode, and if the third confidence score C3 ≤ 85, then request manual intervention;

[0050] When the first confidence score C1 ≤ 75 and the second confidence score C2 > 85, control the execution of the hoisting mode;

[0051] When the first confidence score C1 ≤ 75 and the second confidence score C2 ≤ 85, if the above third confidence score C3 > 85, then control the execution of the lifting mode, and if the third confidence score C3 ≤ 85, then request manual intervention.

[0052] On the other hand, the present invention provides a control method based on an unmanned road patrol and breakdown vehicle, which is applied to the control system proposed in any one of the above, and includes the following steps:

[0053] S10: Control the above image sensor to collect two-dimensional image data of the front and rear of the faulty vehicle, and extract the image feature vector F of the two-dimensional image data through the first flow network img and generate the first confidence score C1;

[0054] S20: Control the above laser sensor to obtain three-dimensional point cloud data of both sides and the roof area of the faulty vehicle body, and extract the point cloud feature vector F of the three-dimensional point cloud data by using the second flow network pc ;

[0055] S21: Superpose and fuse the above two-dimensional image data and three-dimensional point cloud data through the self-attention mechanism, and generate the second confidence score C2;

[0056] S30: Select the corresponding breakdown mode according to the scoring results:

[0057] When the first confidence score C1 > 75 and the second confidence score C2 > 85, select the towing mode and perform breakdown operations by connecting the tow hook of the faulty vehicle;

[0058] When the first confidence score C1 ≤ 75 and the second confidence score C2 > 85, select the hoisting mode and lift the faulty vehicle by connecting the four wheels of the faulty vehicle to perform breakdown operations;

[0059] S40: If the second confidence score C2 ≤ 85, control the above depth sensor to acquire three-dimensional depth data of the wheels and axles of the faulty vehicle, and use the second stream network to extract the depth feature vector F of the above three-dimensional depth data dep ;

[0060] S41: Update and fuse the above three-dimensional point cloud data and three-dimensional depth data through the self-attention mechanism, and generate the third confidence score C3;

[0061] S50: Select the corresponding breakdown clearing mode according to the scoring result:

[0062] When the third confidence score C3 > 85, select the lifting mode, lift and fix one end of the faulty vehicle, and tow the faulty vehicle for breakdown clearing operations;

[0063] When the third confidence score C3 ≤ 85, request manual intervention;

[0064] S60: Record the scoring results and actual breakdown clearing operation feedback of each task in real time, and optimize the model weight coefficients through online incremental learning and offline batch learning to improve the accuracy of subsequent detection tasks.

[0065] Compared with the prior art, the beneficial effects of the present invention are:

[0066] First, the technical solution of the present invention uses a multi-modal perception module composed of an image sensor, a laser sensor, and a depth sensor to comprehensively perceive and collect data from a faulty vehicle, and jointly constructs a multi-dimensional representation of vehicle damage with two-dimensional images, three-dimensional point clouds, and three-dimensional depth data. The multi-modal data acquisition method enables the system to overcome the problem of one-sided identification information of a single sensor. Specifically, the front and rear damage conditions of the vehicle body and vehicle identification information are obtained through the image sensor, and it is judged whether the towing condition is met through the first confidence score C1. Further, the three-dimensional point cloud data of the vehicle body except the front and rear ends is obtained through the laser sensor and superimposed with the two-dimensional image data to form a second confidence score C2 representing the damage state of the entire vehicle body. When the towing condition is not met, if the second confidence score C2 is higher than the preset value, that is, the vehicle body is slightly damaged, the four wheels of the faulty vehicle are connected by a hoisting mechanism, and the vehicle is lifted for obstacle clearance. Further, when the second confidence score C2 is lower than the preset value, it indicates that the vehicle body is severely damaged. At this time, if hoisting is used, the vehicle body may overturn due to imbalance. Therefore, the wheel and axle area is scanned by the depth sensor, and the third confidence score C3 is output to judge whether the wheel can rotate. If it can, the vehicle is cleared by the lifting mode, that is, one end of the faulty vehicle is lifted and fixed, and then towed and transported. The above technical solution effectively improves the accuracy and comprehensiveness of the assessment of the damage condition of the faulty vehicle. By judging the damage condition of the vehicle and selecting an appropriate obstacle clearance method, the misjudgment probability is significantly reduced, and at the same time, the risk of secondary damage caused by wrong decisions is effectively avoided, improving the overall operation efficiency and safety of traffic management, and having obvious practical application advantages.

[0067] Second, the technical solution of the present invention includes a hierarchical detection module based on a two-stream architecture, which realizes an efficient feature extraction and feature fusion process by separately processing multi-modal data through a deep neural network. When the second confidence score C2 is not sufficient to provide a clear judgment, the system updates and fuses the three-dimensional depth data, replaces the three-dimensional depth data representing the wheels and axles with the original three-dimensional point cloud data, and generates a more refined third confidence score C3 to enhance the damage recognition accuracy of the wheel area. The misjudgment of the operation mode caused by insufficient damage assessment is effectively avoided through the gradually refined evaluation method. When facing a small damage situation, the depth sensor is not called, saving system resources, greatly improving the adaptability of the system to complex scenarios, and enhancing the accuracy of decision-making and the reliability of obstacle clearance operation selection.

[0068] Thirdly, the technical solution of the present invention further includes an adaptive learning module to achieve dynamic optimization of the weight coefficients. After manual intervention, the processing results are entered into the system to generate standard feature vectors, and the weight coefficients are corrected through offline learning. On the other hand, when the deviation between the recent 5 scoring results and the manual verification is too large, the weight coefficients are corrected by means of incremental learning. The adaptive learning module effectively enhances the robustness and environmental adaptability of the system, ensuring the high precision and stability of the control system of the unmanned wrecking truck during long-term operation under actual complex road conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0070] Figure 1 is the framework structure diagram of the system of the present invention;

[0071] Figure 2 is the step flow chart of the method of the present invention;

[0072] Figure 3 is the flow chart of the wrecking mode selection rule in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] For the convenience of those skilled in the art to understand, the following further describes the structure of the present invention in detail by combining the embodiments with the drawings:

[0074] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. The terms "part", "side", "end", etc. indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the indicated device or element must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation of the present invention.

[0075] As Figure 1 shown, the present application provides a control system based on an unmanned road patrol wrecking truck, including a multi-modal perception module, a hierarchical detection module, and a control processing unit.

[0076] As Figure 1As shown in the figure, the multi-modal perception module at least includes an image sensor, a laser sensor, and a depth sensor. Among them, the image sensor is used to obtain the two-dimensional image data of the faulty vehicle and extract the vehicle body identification information through OCR. Specifically, it is an optical shooting device such as a high-resolution camera, which can clearly capture the license plate number, vehicle model information, and the details of the damaged surface of the vehicle body of the faulty vehicle. The laser sensor is used to generate the three-dimensional point cloud data of the appearance of the faulty vehicle. Specifically, it is a device such as a 3D lidar. By actively emitting laser pulses and measuring the reflection time, it can obtain the three-dimensional point cloud data of the surface of the target object. The lidar has high distance and angle measurement accuracy, and at the same time can be unaffected by changes in ambient light, and can efficiently and stably obtain the three-dimensional structure data of the body shape and damaged parts of the faulty vehicle. The depth sensor is specifically a depth imaging device such as a structured light scanner or a TOF camera, which is used to generate the three-dimensional depth data of the local appearance of the faulty vehicle. By calculating the deformation of light, it can obtain the depth information of the target, which is suitable for accurately obtaining the detailed information of the damage in the vehicle chassis area, the wheel and axle area, facilitating further judgment of the vehicle's moving ability, and avoiding the failure of the breakdown operation caused by missed inspection of the vehicle bottom damage. The above-mentioned image sensor, laser sensor, and depth sensor scan the areas such as the front of the faulty vehicle, the rear of the vehicle, the left and right sides of the vehicle body, and the roof of the vehicle through the movement of the breakdown vehicle body and the control of the body manipulator, and align the data through spatio-temporal calibration, and the data representation error ≤ 2 cm, so as to ensure accurate subsequent data fusion.

[0077] Further, as Figure 1 shown, the hierarchical detection module includes a deep neural network computing model with a two-stream architecture. Specifically: The first-stream network uses a convolutional neural network, such as a ResNet model or a DenseNet model, and gradually extracts the image feature vector F img from the input two-dimensional image data through multiple convolutional operations, forming the first confidence score C1. It is used to determine the damage degree of the front and rear of the faulty vehicle, and is also used to judge whether the faulty vehicle meets the towing conditions.

[0078] The second-stream network uses a point cloud processing neural network model such as PointNet or PointNet++, fuses and extracts the three-dimensional point cloud data obtained by the laser sensor and the three-dimensional depth data obtained by the depth sensor, and performs feature mapping, clustering, and dimensionality reduction on each point cloud data point through the MLP structure inside the neural network, obtaining the point cloud feature vector F pc of the three-dimensional point cloud data and the depth feature vector F depOverlay and fuse the two-dimensional image data and the three-dimensional point cloud data, and output the second confidence score C2. In one embodiment, the two-dimensional image data represents the damage information of the front and rear of the vehicle body, and the three-dimensional point cloud data represents the damage information of both sides and the roof of the vehicle body. The superposition and fusion of the two represent the damage information of the vehicle body appearance. If the second confidence score C2 ≤ 85, update and fuse the three-dimensional point cloud data and the three-dimensional depth data, and output the third confidence score C3, which is used to further judge the damage degree of the faulty vehicle and the applicable breakdown clearing method. Specifically, the three-dimensional depth data represents the local information of the wheels and axles, which overlaps with part of the representation information of the three-dimensional point cloud data. Therefore, the three-dimensional depth data is updated into the three-dimensional point cloud data, that is, the two are updated and fused to form new three-dimensional data, which is used to further judge the damage degree of the wheels of the faulty vehicle.

[0079] Further, the hierarchical detection module further includes a feature fusion unit. The feature fusion unit adopts the self-attention mechanism based on the Transformer architecture. When this mechanism fuses the feature vectors of multiple data modalities, it can automatically generate attention weight coefficients according to the importance of each modality feature to highlight the key damage information. Through the feature fusion process, the negative impact caused by the missing or misjudgment of single-modal information can be further reduced, and the accuracy and robustness of the overall vehicle damage assessment can be improved. Specifically, the self-attention mechanism is used to weightedly fuse the image feature vector F img , the point cloud feature vector F pc , and the depth feature vector F dep to generate the fused feature vector F fusion .

[0080] As Figure 1 shown, the control and processing unit is connected to the hierarchical detection module and selects the breakdown clearing mode according to the preset threshold rule. When the third confidence score C3 ≤ 85, it requests manual intervention.

[0081] As Figure 3 shown, the preset rule for the control and processing unit to select the breakdown clearing mode is specifically:

[0082] When the first confidence score C1 > 75 and the second confidence score C2 > 85, control the execution of the towing mode to tow the faulty vehicle at a speed not exceeding 20 km / h.

[0083] In this rule, when the first confidence score C1 > 75, it is judged that the front and rear of the vehicle head are less damaged and meet the condition for installing the tow hook. Therefore, the damage situation of the vehicle body is further judged. If the second confidence score C2 > 85, it is judged that the towing mode can be executed.

[0084] When the first confidence score C1 ≤ 75 and the second confidence score C2 > 85, control the execution of the lifting mode.

[0085] In this rule, when the first confidence score C1 ≤ 75, it indicates that the front and rear of the vehicle are severely damaged and do not meet the installation conditions of the tow hook. Therefore, the damage condition of the vehicle body is further judged. If the second confidence score C2 > 85, it means that the vehicle body is slightly damaged, and the hoisting mode is executed. The vehicle is fixed in the wheel area through the hoisting mechanism, and the faulty vehicle is lifted and carried on the carrying part of the breakdown truck for transportation.

[0086] When the first confidence score C1 > 75 and the second confidence score C2 ≤ 85, if the third confidence score C3 > 85, the control executes the lifting mode. If the third confidence score C3 ≤ 85, a manual intervention is requested.

[0087] When the first confidence score C1 ≤ 75 and the second confidence score C2 ≤ 85, if the third confidence score C3 > 85, the control executes the lifting mode. If the third confidence score C3 ≤ 85, a manual intervention is requested.

[0088] In this rule, when the second confidence score C2 ≤ 85, it means that the vehicle body is severely damaged. The damaged part may be the vehicle body or the wheels. If the vehicle body is severely damaged, the faulty vehicle may overturn during hoisting due to the imbalance of the vehicle body, causing secondary damage. If the wheels are severely damaged, unpredictable risks may occur during towing due to the inability of the wheels to turn or the deviation of the direction. Therefore, it is necessary to further introduce the third confidence score C3 for judgment. If the third confidence score C3 > 85, it means that the wheel damage is relatively light (at this time, the first confidence score C1 and the second confidence score C2 do not need to be considered), and the breakdown can be carried out by the lifting method, that is, one end of the faulty vehicle is lifted, and the other end relies on the wheels to support on the ground, and the breakdown truck drags the faulty vehicle for breakdown. If the third confidence score C3 ≤ 85, it means that the damage in the wheel area is relatively heavy, and a manual intervention is requested to adopt a more complex breakdown mode, such as hoisting the vehicle chassis, manually fixing the faulty vehicle for transportation, etc.

[0089] In an embodiment, the breakdown truck identifies information such as the vehicle brand, vehicle model, and vehicle identification code through an image sensor, and transports the faulty vehicle to a repair station or a temporary storage point on a preset path to achieve an automated breakdown task.

[0090] The present invention proposes a control system for an unmanned breakdown truck based on multi-modal perception and hierarchical confidence evaluation. Through multi-sensor collaboration, two-stream deep neural network feature extraction, and self-attention dynamic fusion mechanism, it realizes high-precision vehicle state analysis and intelligent breakdown decision-making, realizes intelligent and accurate judgment of the damage condition of road faulty vehicles, can effectively reduce the risk of secondary damage caused by incorrect damage assessment, reduce the frequency of manual intervention, has the advantages of high automation, strong stability, and strong environmental adaptability, can significantly improve the efficiency, safety, and intelligent level of road breakdown operations, and has broad practical application value and good market prospects.

[0091] The working principles of the hierarchical detection module and the adaptive learning module are described below through a specific embodiment:

[0092] The image sensor scans the front and rear of the faulty vehicle. The calculation formula for the first confidence score C1 is:

[0093] C1 = 100·σ(W1·F img +b1)

[0094] where F img represents the image feature vector output by the first-stream convolutional neural network, W1 is the weight matrix, which is a vector of the same dimension as F img and is used to map F img to the scoring space, b1 is the bias term with an initial value of 0; σ(·) is the Sigmoid activation function, and its original linear output mapping interval is [0,1]. It is magnified 100 times so that the first confidence score C1 is mapped to the range of 0 to 100.

[0095] The calculation formula for the image feature vector F img is:

[0096] F img = Attention(Q img ,K img ,V img )

[0097] where Q img ,K img ,V img are the query matrix Q, key matrix K, and value matrix V in the self-attention mechanism, which are generated from the two-dimensional image data.

[0098] The laser sensor scans both sides and the roof of the faulty vehicle. The calculation formula for the second confidence score C2 is:

[0099] C2 = 100·σ(W2·F fusion(img,pc) +b2)

[0100] where W2 is the weight matrix, b2 is the bias term; σ(·) is the Sigmoid activation function, which is magnified 100 times so that the second confidence score C2 is mapped to the range of 0 to 100.

[0101] The fused feature vector F fusion(img,pc) is generated by fusing the two-dimensional image data and the three-dimensional point cloud data through the self-attention mechanism. The specific formula is:

[0102] F fusin(iPmg,pc) = γ1·Attention(Q img ,K img ,Vimg ) + γ2·Attention(Q pc , K pc , V pc )

[0103] where Q img , K img , V img are the query matrix Q, key matrix K, and value matrix V in the self-attention mechanism, which are generated from two-dimensional image data; Q pc , K pc , V pc are the query matrix Q, key matrix K, and value matrix V generated from three-dimensional point cloud data.

[0104] The weight coefficients γ1 and γ2 satisfy γ1 + γ2 = 1, and when the first confidence score C1 > 75, that is, when the front and rear of the faulty vehicle are less damaged, the weight coefficient γ1 is reduced and the weight coefficient γ2 is increased. In a preferred embodiment, the initial value of γ1 is 0.5, the initial value of γ2 is 0.5, and if C1 > 75, then γ1 changes to 0.3 and γ2 changes to 0.7.

[0105] The depth sensor scans the wheel and axle areas of the faulty vehicle, and the calculation formula for the third confidence score C3 is:

[0106] C3 = 100·σ(W3·F fusion(pc,dep) + b3)

[0107] where W3 is the weight matrix, b3 is the bias term; σ(·) is the Sigmoid activation function, amplified by 100 times, so that the third confidence score C3 is mapped to the range of 0 - 100.

[0108] The fused feature vector F fusion(pc,dep) is generated by fusing two-dimensional image data and three-dimensional point cloud data through the self-attention mechanism. The specific formula is:

[0109] F fusion(pc,dep) = (1 - M) ⊙ [γ2′·Attention(Q pc , K pc , V pc )] + M ⊙ [γ3·Attention(Q dep , K dep , V dep )]

[0110] Among them, M is the mask matrix. The overlapping part of the three-dimensional point cloud data and the three-dimensional depth data is updated to the three-dimensional depth data through the mask matrix M. In this embodiment, the overlapping part is the side surfaces of the wheel and the axle. In another embodiment, the depth sensor is also used to detect the vehicle chassis, and the three-dimensional depth data of the vehicle chassis is not used as the updated content of the three-dimensional point cloud data.

[0111] The weight coefficients γ2′ and γ3 satisfy γ2′ + γ3 = 1, and the weight coefficient γ3 > γ2′. Preferably, γ2′ = 0.3 and γ3 = 0.7.

[0112] The system further includes an adaptive learning module, which is connected to the control processing unit and optimizes the weight coefficients in the feature fusion unit through offline learning and incremental learning.

[0113] The offline learning of the weight coefficients uses a loss function that minimizes the mean square error:

[0114]

[0115] Among them, F target,i is the standard feature vector of the vehicle state labeled after manual intervention, which can be the standard feature vector converted from historical average data or the standard feature vector manually calibrated by professionals. Through N groups of data, the fused feature vector F fusion(·) is calculated forward, and the learnable variables such as the above weight matrix and weight parameters are updated.

[0116] The incremental learning method of the weight coefficients is as follows:

[0117] When the recent system detection error shows a continuous increasing trend, the following rules are used to update the weight coefficients:

[0118]

[0119] Among them, Accuracy is the proportion of tasks in the last 5 tasks where the error between the system's measured score and the manual verification score is within ±5;

[0120] Δ is the learning rate, initially limited between 0.01 and 0.05;

[0121] When Accuracy ≥ 95%, the update of the weight coefficients is stopped to maintain the parameter stability of the system.

[0122] In some embodiments, the learning rate can be dynamically adjusted according to Accuracy. For example:

[0123] If Accuracy = 93%, Δ is reduced to 0.01 to stabilize the system;

[0124] If Accuracy = 70%, Δ is increased to 0.05 to accelerate convergence.

[0125] As Figure 2 shown, on the other hand, the present invention provides a control method based on an unmanned road patrol and breakdown vehicle, which is applied to the above control system, and includes the following steps:

[0126] S10: Control the image sensor to collect two-dimensional image data of the front and rear of the breakdown vehicle, extract the image feature vector F of the two-dimensional image data through the first flow network img , and generate the first confidence score C1;

[0127] S20: Control the laser sensor to obtain three-dimensional point cloud data of both sides and the roof area of the breakdown vehicle, and extract the point cloud feature vector F of the three-dimensional point cloud data by using the second flow network pc ;

[0128] S21: Superimpose and fuse the two-dimensional image data and the three-dimensional point cloud data through the self-attention mechanism, and generate the second confidence score C2;

[0129] S30: Select the corresponding breakdown mode according to the scoring result:

[0130] When the first confidence score C1 > 75 and the second confidence score C2 > 85, select the towing mode and perform breakdown operations by connecting the tow hook of the breakdown vehicle;

[0131] When the first confidence score C1 ≤ 75 and the second confidence score C2 > 85, select the hoisting mode, lift the breakdown vehicle by connecting the four wheels of the breakdown vehicle, and perform breakdown operations;

[0132] S40: If the second confidence score C2 ≤ 85, control the depth sensor to obtain three-dimensional depth data of the wheels and axles of the breakdown vehicle, and extract the depth feature vector F of the three-dimensional depth data by using the second flow network dep ;

[0133] S41: Update and fuse the three-dimensional point cloud data and the three-dimensional depth data through the self-attention mechanism, and generate the third confidence score C3;

[0134] S50: Select the corresponding breakdown mode according to the scoring result:

[0135] When the third confidence score C3 > 85, select the lifting mode, lift one end of the breakdown vehicle and fix it, and tow the breakdown vehicle for breakdown operations;

[0136] When the third confidence score C3 ≤ 85, request manual intervention;

[0137] S60: Record the scoring results and actual obstacle clearing operation feedback of each task in real time, and optimize the model weight coefficients through online incremental learning and offline batch learning to improve the accuracy of subsequent detection tasks.

[0138] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A control system based on an unmanned road patrol tow truck, characterized in that: include: The multimodal perception module includes at least: Image sensor, used to obtain two-dimensional image data of the faulty vehicle and extract vehicle body identification information through OCR; A laser sensor for generating three-dimensional point cloud data of the appearance of the faulty vehicle; A depth sensor for generating three-dimensional depth data of the partial appearance of the faulty vehicle; Hierarchical detection module, including a deep neural network computing model with a two-stream architecture: The first stream network uses a convolutional neural network to extract the image feature vector F of the two-dimensional image data. img , output the first confidence score C1; The second stream network uses a point cloud processing algorithm to extract the point cloud feature vector F of the three-dimensional point cloud data. pc and the depth feature vector F of the three-dimensional depth data dep ; Superimpose and fuse the two-dimensional image data and the three-dimensional point cloud data, and output a second confidence score C2; ​​if the second confidence score C2≤85, update and fuse the three-dimensional point cloud data and the three-dimensional depth data, and output a third confidence score C3; The feature fusion unit uses a self-attention mechanism to weightedly fuse the image feature vector F img , point cloud feature vector F pc And the deep feature vector F dep , generate the fused feature vector F fusion ; The control processing unit is connected to the hierarchical detection module and selects an obstacle removal mode according to a preset threshold rule, and requests manual intervention when the third confidence score C3≤85.

2. A control system based on an unmanned road patrol tow truck as claimed in claim 1, characterized in that: The image sensor scans the front and rear of the faulty vehicle, and the calculation formula of the first confidence score C1 is: C1=100·σ(W1·F img +b1) Among them, F img The image feature vector representing the output of the first-stream convolutional neural network; W1 is the weight matrix, b1 is the bias term; σ(·) is a Sigmoid activation function, which maps the first confidence score C1 to a range of 0 to 100.

3. A control system based on an unmanned road patrol tow truck as claimed in claim 2, characterized in that: The laser sensor scans both sides of the body and the roof of the faulty vehicle, and the calculation formula of the second confidence score C2 is: C2=100·σ(W2·F fusion(img,pc) +b2) Among them, W2 is the weight matrix and b2 is the bias term; σ(·) is a Sigmoid activation function, which maps the second confidence score C2 to a range of 0 to 100; Fusion feature vector F fusion(img,pc) The two-dimensional image data and the three-dimensional point cloud data are fused and generated through the self-attention mechanism. The specific formula is: F fusion(img,pc) =γ1·Attention(Q img ,K img ,V img )+γ2·Attention(Q pc ,K pc ,V pc ) The weight coefficients γ1 and γ2 satisfy γ1+γ2=1, and when the first confidence score C1>75, that is, when the front and rear of the faulty vehicle are less damaged, the weight coefficient γ1 is reduced and the weight coefficient γ2 is increased.

4. A control system based on an unmanned road patrol tow truck as claimed in claim 3, characterized in that: The depth sensor scans the wheel and axle area of ​​the faulty vehicle, and the calculation formula of the third confidence score C3 is: C3=100·σ(W3·F fusion(pc,dep) +b3) Among them, W3 is the weight matrix and b3 is the bias term; σ(·) is a Sigmoid activation function, which maps the third confidence score C3 to a range of 0 to 100; Fusion feature vector F fusion(pc,dep) The two-dimensional image data and the three-dimensional point cloud data are fused and generated through the self-attention mechanism. The specific formula is: F fusion(pc,dep) =(1-M)⊙[γ2′·Attention(Q pc ,K pc ,V pc )]+M⊙[γ3·Attention(Q dep ,K dep ,V dep )] Among them, M is a mask matrix, and the overlapping part of the three-dimensional point cloud data and the three-dimensional depth data is updated to the three-dimensional depth data through the mask matrix M. Specifically, in one embodiment, since the laser sensor scans the two sides and the top of the vehicle body, and the depth sensor scans the wheels and the axle, there is a certain overlap in the scanning areas of the two. Therefore, a mask matrix M with the same dimension as the above-mentioned data feature vector is introduced, and the elements of the mask matrix M are 0 or 1. It is specifically defined as: when an element of the mask matrix M is 1, it indicates that the corresponding position is scanned by the laser sensor and the depth sensor at the same time, and the position is replaced by the three-dimensional depth data of the depth sensor. The weight coefficients γ2′ and ′3 satisfy γ2′+γ3=1, and the weight coefficient γ3>γ2′.

5. A control system based on an unmanned road patrol tow truck as claimed in claim 1, characterized in that: It also includes an adaptive learning module, which is connected to the control processing unit and optimizes the weight coefficient in the feature fusion unit through offline learning and incremental learning.

6. A control system based on an unmanned road patrol tow truck as claimed in claim 5, characterized in that: The offline learning of the weight coefficients adopts the loss function of minimizing the mean square error: Among them, F target,i It is the standard feature vector of the vehicle state annotated after manual intervention.

7. A control system based on an unmanned road patrol tow truck as claimed in claim 6, characterized in that: The incremental learning method of the weight coefficient is: When the recent system detection error shows a continuous increasing trend, the following rules are used to update the weight coefficient: Among them, Accuracy is the percentage of tasks in the last five tasks where the error between the system and the test score and the manual verification score is within ±5; Δ is the learning rate, initially limited to between 0.01 and 0.05; When Accuracy ≥ 95%, stop updating the weight coefficients to maintain the parameter stability of the system.

8. A control system based on an unmanned road patrol tow truck as claimed in claim 1, characterized in that: The preset rule for the control processing unit to select the obstacle removal mode is: When the first confidence score C1>75 and the second confidence score C2>85, the control executes the towing mode to tow the faulty vehicle at a speed not exceeding 20 km / h; When the first confidence score C1>75 and the second confidence score C2≤85, if the third confidence score C3>85, the lifting mode is executed, and if the third confidence score C3≤85, manual intervention is requested; When the first confidence score C1≤75 and the second confidence score C2>85, the control executes the hoisting mode; When the first confidence score C1≤75 and the second confidence score C2≤85, if the third confidence score C3>85, the lifting mode is executed, and if the third confidence score C3≤85, manual intervention is requested.

9. A control method based on an unmanned road patrol tow truck, applied to the control system according to any one of claims 1 to 8, characterized in that: The steps include: S10: Control the image sensor to collect two-dimensional image data of the front and rear of the faulty vehicle, and extract the image feature vector F of the two-dimensional image data through the first-stream network img , and generate a first confidence score C1; S20: Control the laser sensor to obtain three-dimensional point cloud data of the two sides and the roof area of ​​the faulty vehicle, and use the second stream network to extract the point cloud feature vector F of the three-dimensional point cloud data pc ; S21: superimposing and fusing the two-dimensional image data and the three-dimensional point cloud data through a self-attention mechanism, and generating a second confidence score C2; S30: Select the corresponding obstacle removal mode according to the scoring result: When the first confidence score C1>75 and the second confidence score C2>85, the towing mode is selected to perform the towing operation by connecting the tow hook of the faulty vehicle; When the first confidence score C1≤75 and the second confidence score C2>85, the hoisting mode is selected, and the faulty vehicle is hoisted by connecting the four wheels of the faulty vehicle to perform the towing operation; S40: If the second confidence score C2≤85, control the depth sensor to obtain the three-dimensional depth data of the wheel and axle area of ​​the faulty vehicle, and use the second flow network to extract the depth feature vector F of the three-dimensional depth data dep ; S41: updating and fusing the three-dimensional point cloud data and the three-dimensional depth data through a self-attention mechanism, and generating a third confidence score C3; S50: Select the corresponding obstacle removal mode according to the scoring result: When the third confidence score C3>85, the lifting mode is selected to lift and fix one end of the faulty vehicle and tow the faulty vehicle for towing; When the third confidence score C3≤85, request manual intervention; S60: Record the scoring results of each task and the actual obstacle clearance feedback in real time, and optimize the model weight coefficients through online incremental learning and offline batch learning to improve the accuracy of subsequent detection tasks.

Citation Information

Patent Citations

  • Multi-modal target detection method and device and multi-modal identification system

    CN119625279A

  • Unmanned aerial vehicle routing inspection line adaptive obstacle detection method and system based on monocular camera

    CN119672577A

  • Navigation obstacle avoidance method and system in low-confidence and feature similar environment

    CN119687918A

  • Multi-modal data fusion for enhanced 3D perception for platforms

    US20200184718A1

  • Systems and methods for object detection using stereovision information

    US20220188578A1