Multi-modular heavy-duty vehicle collaborative transportation risk assessment method and system based on high-altitude visual angle

By employing a multi-modal heavy-duty vehicle collaborative transportation risk assessment method from a high-altitude perspective, and utilizing a high-altitude imaging device and a multi-task detection model to generate global spatial perception results, combined with Kalman filtering and Hungarian matching, the problem of insufficient perception of SPMT in unstructured environments is solved. This enables accurate risk assessment and early warning of vehicles and the environment, improving the safety and operability of heavy-duty platooning operations.

CN121961367APending Publication Date: 2026-05-01WUHAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN UNIV OF TECH
Filing Date
2025-12-31
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing multi-modular heavy-duty transport vehicles (SPMTs) have limited perception range and incomplete outline recognition in unstructured environments, resulting in large errors in human judgment and thus low safety and operability of heavy-duty platooning operations.

Method used

A high-altitude imaging device is used to acquire video of vehicle transportation. A multi-task detection model is used to perform instance segmentation, rotation detection, and semantic segmentation to generate global spatial perception results. Kalman filtering and Hungarian matching are used for trajectory smoothing and short-term prediction. Combined with sweep set and environmental boundary detection, minimum distance and contact time are calculated to achieve risk level assessment.

Benefits of technology

It achieves globally consistent assessment under strong occlusion and complex space conditions, reduces reliance on human experience, and significantly improves the safety and operability of heavy-load formation operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961367A_ABST
    Figure CN121961367A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modular heavy-load vehicle collaborative transportation risk assessment method and system based on a high-altitude visual angle, and the method comprises the steps: collecting a vehicle transportation video through a high-altitude imaging device, inputting a multi-task detection model for instance segmentation, rotation detection and semantic segmentation, and outputting a global space perception result; performing Kalman prediction and Hungary matching on the global spatial perception result to obtain a standardized result and a contour matching result; the observable state of the vehicle platform is judged, and a system safety outline is generated; trajectory derivation and short-term prediction are carried out according to a standardization result and a system safety contour, and a future pose and a future contour sequence of the vehicle are obtained; constructing a scanning set, and obtaining the minimum spacing and the contact time according to the scanning set and the global space sensing result; and determining a collaborative transportation risk level of the multi-modular heavy-load vehicle. The method can remarkably improve the safety and operability of heavy-load formation operation, and can be widely applied to the technical field of vehicle transportation safety.
Need to check novelty before this filing date? Find Prior Art

Description

A Risk Assessment Method and System for Multi-Modular Heavy-Duty Vehicle Collaborative Transportation Based on High-Altitude Perspective Technical Field

[0001] This application relates to the field of vehicle transportation safety technology, and in particular to a method and system for risk assessment of multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective. Background Technology

[0002] With the rapid development of large equipment manufacturing, wind power transportation, and engineering construction industries, multi-modal heavy-duty transport vehicles (SPMTs) are widely used in the loading, unloading, and transfer of oversized and irregularly shaped cargo. SPMTs, characterized by self-drive, multi-module assembly, and multi-axle steering, can flexibly handle complex sites and narrow road sections, making them crucial equipment for heavy-duty transport. However, existing SPMT safety monitoring and driving control methods still primarily rely on human experience and limited onboard sensing systems, making it difficult to meet the demands for high-precision, global perception and real-time risk assessment. In unstructured environments, existing collaborative transport of heavy-duty SPMTs suffers from limited sensing range, incomplete outline recognition, and large errors in human judgment, resulting in low safety and operability of heavy-duty platooning operations. Summary of the Invention

[0003] The main objective of this application is to propose a method and system for risk assessment of multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective, which can significantly improve the safety and operability of heavy-duty platooning operations.

[0004] To achieve the above objectives, one aspect of this application proposes a risk assessment method for multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective, comprising the following steps: deploying a high-altitude imaging device above the target road segment, calibrating the high-altitude imaging device, and obtaining calibration results; acquiring vehicle transportation videos through the high-altitude imaging device, inputting the vehicle transportation videos into a pre-trained multi-task detection model for instance segmentation, rotation detection, and semantic segmentation, and outputting global spatial perception results of the target vehicle and cargo; performing Kalman prediction and Hungarian matching on the global spatial perception results to obtain the target vehicle ID and temporal outline. The system generates a system safety profile by determining the observable state of the vehicle platform based on the calibration results, the standardization results, and the profile matching results; performing trajectory derivation and short-term prediction based on the standardization results and the system safety profile to obtain the vehicle's future pose and future profile sequence; constructing a sweep set based on the vehicle's future pose and the future profile sequence; and obtaining the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception results; and determining the cooperative transportation risk level of the multi-modal heavy-duty vehicles based on the minimum distance and the contact time.

[0005] In some embodiments, the method further includes a step of pre-training the multi-task detection model, wherein the pre-training of the multi-task detection model specifically includes: acquiring a high-altitude image training dataset through the high-altitude imaging device; annotating the vehicle-marked bounding box set, vehicle rotation bounding rectangle set, instance mask set, outline polygon set, scene semantic annotation set, safety boundary set, and dynamic obstacle set in the high-altitude image training dataset to obtain an annotation set; constructing a deep learning network model, wherein the deep learning network model includes a shared encoder and a three-task decoder; inputting the high-altitude image training dataset into the shared encoder for encoding and outputting encoded features; inputting the encoded features into the three-task decoder for instance segmentation training, rotation detection training, and semantic segmentation training, and adjusting the parameters of the shared encoder and the three-task decoder according to the annotation set to obtain the multi-task detection model.

[0006] In some embodiments, the global spatial perception result includes a set of identified bounding boxes. The step of performing Kalman prediction and Hungarian matching on the global spatial perception result to obtain a standardized result and a contour matching result for the target vehicle ID and temporal contour specifically includes: constructing a trajectory set of the target vehicle from the previous frame based on the global spatial perception result, the trajectory set including the target vehicle ID and the state vector of the target vehicle from the previous frame; constructing a state transition matrix and an observation matrix; obtaining a prior state estimate and a prior covariance matrix for the current frame based on the state transition matrix and the state vector; obtaining an estimated bounding box set based on the prior state estimate; performing Hungarian matching on the identified bounding box set and the estimated bounding box set to obtain the contour matching result; updating the prior state estimate based on the observation matrix and the prior covariance matrix to obtain a posterior state estimate for the current frame; and obtaining the standardized result of the target vehicle ID and temporal contour based on the posterior state estimate and the contour matching result.

[0007] In some embodiments, the calibration result includes a geography homography matrix. The step of determining the observable state of the vehicle platform based on the calibration result, the standardization result, and the outline matching result, and generating a system safety outline, specifically includes: obtaining a cargo outline polygon and a vehicle platform outline polygon based on the geography homography matrix and the standardization result; determining whether the vehicle platform outline polygon is observable based on the outline matching result; if the vehicle platform outline polygon is observable, calculating the union convex hull of the cargo outline polygon and the vehicle platform outline polygon to obtain a first system outline; if the vehicle platform outline polygon is unobservable, performing Minkowski dilation on the cargo outline polygon to obtain a second system outline; and performing Minkowski dilation on either the first system outline or the second system outline to obtain the system safety outline.

[0008] In some embodiments, the step of deriving and short-term predicting the trajectory based on the standardized result and the system safety profile to obtain the vehicle's future pose and future profile sequence specifically includes: defining a discrete time step and a prediction time domain; calculating the linear velocity based on the trajectory arc length in the standardized result; calculating the curvature based on the heading angle in the standardized result; updating the heading angle and the vehicle center position in the prediction time domain based on the curvature, the discrete time step, and the linear velocity to obtain the vehicle's future pose; and transforming the system safety profile based on the vehicle's future pose to obtain the future profile sequence.

[0009] In some embodiments, the global spatial perception result includes a set of safety boundaries, and the contact time includes a first contact time or a second contact time. The step of obtaining the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception result specifically includes: determining whether the sweep set intersects with the set of safety boundaries; if the sweep set intersects with the set of safety boundaries, calculating the minimum distance between the target vehicle and the environment based on the sweep set and the set of safety boundaries, determining an index, and then performing linear interpolation on the index to obtain the first contact time; if the sweep set does not intersect with the set of safety boundaries, calculating the minimum distance between the target vehicle and the environment based on the sweep set and the set of safety boundaries, and then obtaining the second contact time based on the minimum distance.

[0010] In some embodiments, determining the cooperative transportation risk level of multi-modal heavy-duty vehicles based on the minimum distance and the contact time specifically includes: setting a safety distance upper limit threshold, a collision critical distance threshold, a safety time upper limit threshold, a warning time threshold, and an emergency response time threshold; if the minimum distance is greater than the safety distance upper limit threshold and the contact time is greater than the safety time upper limit threshold, the cooperative transportation risk level is determined to be a first risk level; if the minimum distance is greater than the collision critical distance threshold and less than or equal to the safety distance upper limit threshold, or the contact time is greater than the warning time threshold and less than or equal to the safety time upper limit threshold, the cooperative transportation risk level is determined to be a second risk level; if the minimum distance is less than or equal to the collision critical distance threshold, or the contact time is less than or equal to the warning time threshold, the cooperative transportation risk level is determined to be a third risk level; if the sweep set intersects with the safety boundary set, or the contact time is less than or equal to the emergency response time threshold, the cooperative transportation risk level is determined to be a fourth risk level.

[0011] To achieve the above objectives, another aspect of this application proposes a multi-modal heavy-duty vehicle collaborative transportation risk assessment system based on a high-altitude perspective, comprising: a first module for deploying a high-altitude imaging device above the target road segment, calibrating the high-altitude imaging device, and obtaining calibration results; a second module for acquiring vehicle transportation videos through the high-altitude imaging device, inputting the vehicle transportation videos into a pre-trained multi-task detection model for instance segmentation, rotation detection, and semantic segmentation, and outputting global spatial perception results of the target vehicle and cargo; and a third module for performing Kalman prediction and Hungarian matching on the global spatial perception results to obtain standardized results of the target vehicle ID and temporal outline. The system comprises seven modules: a fourth module, a fifth module, and a sixth module. The sixth module is used to determine the observable state of the vehicle platform based on the calibration results, the standardization results, and the outline matching results, and to generate a system safety outline. The seventh module is used to perform trajectory derivation and short-term prediction based on the standardization results and the system safety outline to obtain the vehicle's future pose and future outline sequence. The seventh module is used to construct a sweep set based on the vehicle's future pose and the future outline sequence, and then obtain the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception results. The seventh module is used to determine the cooperative transportation risk level of multi-modular heavy-duty vehicles based on the minimum distance and the contact time.

[0012] To achieve the above objectives, another aspect of this application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the method described above.

[0013] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer program product, including a computer program that, when executed by a processor, implements the aforementioned method.

[0014] The embodiments of this application include at least the following beneficial effects: The multi-modal heavy-duty vehicle collaborative transportation risk assessment method and system based on a high-altitude perspective of this application acquires vehicle transportation videos from a high-altitude perspective using a high-altitude imaging device above the target road section. It then uses a multi-task detection model for instance segmentation, rotation detection, and semantic segmentation to generate a global spatial perception result that accurately reflects the vehicle's occupied space and posture changes, thus solving the problem of limited perception range. Furthermore, through temporal tracking, trajectory smoothing, and short-term prediction based on Kalman filtering and Hungarian matching, it generates the vehicle's future pose and future outline sequence. It also introduces sweep set and environmental boundary intersection detection to calculate quantitative indicators such as the minimum distance and contact time between the target vehicle and the environment, thereby providing a basis for real-time risk classification and early warning. Under conditions of strong occlusion, confined space, and complex load configurations, it achieves a globally consistent assessment of the overall geometric relationship between the vehicle and the environment, reducing reliance on human experience and significantly improving the safety and operability of heavy-duty platooning operations. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the embodiments of this application are described below. It should be understood that the drawings described below are only for the purpose of clearly illustrating some embodiments of the technical solutions in this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0016] Figure 1 is a flowchart of the steps of a multi-modal heavy-duty vehicle collaborative transportation risk assessment method based on a high-altitude perspective provided in an embodiment of this application; Figure 2 is a schematic diagram of the system safety profile generation process provided in an embodiment of this application; Figure 3 is a schematic diagram of the risk level determination process provided in an embodiment of this application; Figure 4 is a schematic diagram of the process of a multi-modal heavy-duty vehicle collaborative transportation risk assessment method based on a high-altitude perspective provided in an embodiment of this application; Figure 5 is a schematic diagram of the structure of a multi-modal heavy-duty vehicle collaborative transportation risk assessment system based on a high-altitude perspective provided in an embodiment of this application; Figure 6 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0019] With the rapid development of large-scale equipment manufacturing, wind power transportation, and engineering construction industries, multi-modal heavy-duty transport vehicles (SPMTs) are widely used in the loading, unloading, and transfer of oversized and irregularly shaped cargo. SPMTs, characterized by self-drive, multi-module assembly, and multi-axle steering, can flexibly handle complex sites and narrow road sections, making them crucial equipment for heavy-duty transport. However, existing SPMT safety monitoring and driving control still primarily rely on human experience and limited onboard perception systems, failing to meet the demands for high-precision, global perception, and real-time risk assessment. In unstructured environments, existing SPMT collaborative transport for heavy-duty cargo suffers from the following shortcomings: 1. Onboard vision and laser sensors are limited by installation height and obstruction, failing to obtain complete vehicle outline and environmental boundary information; 2. Precise modeling of the geometric features of oversized and irregularly shaped cargo is lacking, making it difficult to accurately reflect the system's space occupancy; 3. When vehicles travel on special road sections, reliance on human experience for judgment introduces risks of delays and blind spots.

[0020] Therefore, how to obtain complete scene information by constructing a high-altitude overhead view, accurately reconstruct the system outline of SPMT and cargo, and realize the quantitative assessment and early warning of driving risks has become a key technical problem that urgently needs to be solved in the field of intelligent heavy-duty transportation.

[0021] In view of this, this application proposes a risk assessment method for multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective. It acquires vehicle transportation videos from a high-altitude perspective using a high-altitude imaging device above the target road segment, and performs instance segmentation, rotation detection, and semantic segmentation using a multi-task detection model to generate a global spatial perception result that accurately reflects the vehicle's occupied space and posture changes, thus solving the problem of limited perception range. Furthermore, through temporal tracking, trajectory smoothing, and short-term prediction based on Kalman filtering and Hungarian matching, it generates the vehicle's future pose and future outline sequence, and introduces sweep set and environmental boundary intersection detection to calculate quantitative indicators such as the minimum distance and contact time between the target vehicle and the environment, thereby providing a basis for real-time risk classification and early warning. Under conditions of strong occlusion, confined space, and complex load configurations, it achieves a globally consistent assessment of the overall geometric relationship between the vehicle and the environment, reducing reliance on human experience and significantly improving the safety and operability of heavy-duty platooning operations.

[0022] Referring to Figure 1, which is a flowchart of the steps of a multi-modal heavy-duty vehicle collaborative transportation risk assessment method based on a high-altitude perspective according to an embodiment of this application, this embodiment proposes a multi-modal heavy-duty vehicle collaborative transportation risk assessment method based on a high-altitude perspective. This method may include, but is not limited to, the following steps S101 to S107: Step S101: Deploy a high-altitude imaging device above the target road segment, calibrate the high-altitude imaging device, and obtain calibration results; Specifically, in this embodiment, a high-altitude imaging device is deployed above the target road segment or at a slightly oblique position to establish a stable video acquisition link, and a global observation of the SPMT and its transported goods is performed using a high-altitude overhead perspective. Ground control points are deployed at each fixed camera position or a known geometric structure is selected to solve for the camera intrinsic parameters K, distortion parameters D, and ground homography matrix H, obtaining calibration results.

[0023] It should be noted that the embodiments of this application use a high-altitude imaging device to conduct a global observation of the SPMT and its transported goods from a high-altitude top-down perspective. This can simultaneously cover the vehicle body, the shape of the goods, and the boundaries of the surrounding environment, fundamentally eliminating the blind spots caused by occlusion and height limitations of the vehicle's vision.

[0024] Step S102: Acquire vehicle transportation video using an aerial imaging device, input the video into a pre-trained multi-task detection model for instance segmentation, rotation detection, and semantic segmentation, and output global spatial perception results of the target vehicle and cargo. In some optional embodiments, the computing host in this application establishes a video data link with the fixed-position aerial imaging device, and forms a steady-state reference sequence using feature matching and homography geometric alignment to provide stable input for subsequent synchronous detection and temporal processing. On the steady-state pixel sequence, a pre-trained multi-task detection model (AT-MUSE model) is loaded, and for each frame... A single forward pass simultaneously performs instance segmentation, rotated target detection, and scene semantic parsing, generating a unified pixel-domain result at the same timestamp—that is, a global spatial perception result for the target vehicle and cargo. This global spatial perception result includes a set of recognized bounding boxes. Instance mask set Set of outer polygons Rotated circumscribed rectangle set Dynamic obstacle set Scene semantic graph and security boundary set .

[0025] Specifically, the vehicle transportation video is input into a pre-trained multi-task detection model, which first outputs a set of "SPMT + cargo" instance masks. And confidence level, extract the outer boundary of each mask and use RDP at a threshold Simplified to polygons Record instance identifier and category, vertex format , For the first Frame number The first instance of the outline polygon The vertex pixel coordinates are then output. Finally, the set of rotated bounding rectangles is output. ,in For pixel domain parameters, As a category, Confidence level; synchronously output dynamic obstacle observations of conventional vehicles and pedestrians; for the set of rotated circumscribed rectangles With the corresponding set of outline polygons Perform a geometric consistency check (the four corners of the rotated bounding box should fall within the instance polygon or its neighborhood). If there is inconsistency, backtrack or re-score based on the confidence level to improve geometric stability. Finally, output a semantic graph. Covering categories such as roads, curbs / guardrails, islands / water edges, fixed obstacles / restricted areas, contours are extracted and simplified through opening / closing operations and hole filling to form a set of pixel-domain safety boundaries. .

[0026] It should be noted that, in response to the problem that the "SPMT + cargo" shape is irregular and difficult to represent by traditional rectangular boxes when transporting oversized or irregularly shaped goods, this application embodiment constructs a multi-task detection model based on instance segmentation, rotation detection, and semantic segmentation. By extracting the polygon outline, the rotated circumscribed rectangle, and the semantic boundary, the outline of the combined system of platform and cargo is dynamically generated, which can truly reflect the space occupied by the vehicle and the characteristics of its posture changes.

[0027] As an optional implementation, the risk assessment method for multi-modular heavy-duty vehicle collaborative transportation based on a high-altitude perspective also includes the step of pre-training a multi-task detection model, which can be further divided into the following steps S1021 to S1025: Step S1021: Obtain a high-altitude image training dataset through a high-altitude imaging device; Specifically, the computing host controls the high-altitude imaging device to be deployed directly above or slightly tilted down from the target road section via wireless means. Under three loading modes—empty, standard cargo, and oversized cargo—video image data is continuously collected in typical scenarios such as narrow corridors, access control / door frames, ramp / roundabout entrances, dock ramps, bridge decks / edges, and intersections / merging points, forming a high-altitude image training dataset of multi-modular distributed driving heavy-duty vehicles and road environment.

[0028] Step S1022: Label the vehicle bounding box set, vehicle rotation bounding rectangle set, instance mask set, outline polygon set, scene semantic annotation set, safety boundary set, and dynamic obstacle set in the high-altitude image training dataset to obtain an annotation set; specifically, manually annotate the high-altitude image training dataset, label the bounding rectangle bounding box of the target vehicle and the target type to form the high-altitude image training vehicle bounding box set, and simultaneously complete the joint annotation of the "SPMT+cargo" instance outline, scene semantics, and dynamic obstacles to obtain an annotation set for use in the training of the integrated multi-task model.

[0029] High-altitude image training vehicle marker bounding box set for: ;in, The first vehicle in the training vehicle bounding box set of the high-altitude image represents the... The first frame of the image The x-coordinate of the top-left corner of the rectangle marking the vehicle target. Represents the y-coordinate of the top left corner. Indicates the x-coordinate of the bottom right corner. Indicates the y-coordinate of the bottom right corner; Indicates the first The first frame of the image The tagging category of each vehicle target (including but not limited to: SPMT platform, cargo, and conventional vehicle).

[0030] High-altitude image training vehicle rotation circumscribed rectangle set for: ;in, and The center pixel coordinates of the rotated rectangle. and Width and height (in pixels). For the orientation angle, For the target category.

[0031] High-altitude image training instance mask set for: ; Set of outer polygons for: ;in, For "SPMT+cargo" pixel-level mask, This is the vertex sequence obtained by extracting from the outer boundary of the mask and simplifying it using the RDP algorithm.

[0032] Scene semantic annotation set for: ;in, Indicates the total number of semantic categories. This represents the semantic label value of a pixel in frame e, with categories including road, shoulder / edge, guardrail / wall, island / water edge, fixed obstacle / no-entry zone, etc.

[0033] Safety boundary set for:

[0034] Dynamic obstacle set for: ;in, Depend on Obtained through morphological and contour extraction; Indicates the first Frame number One dynamic obstacle target; , These are the pixel coordinates of the top left and bottom right corners of a dynamic obstacle (such as a regular vehicle or pedestrian). Represented as the first Frame number Category labels for each obstacle; This is for axis alignment annotations of dynamic obstacles such as regular vehicles and pedestrians.

[0035] Step S1023: Construct a deep learning network model, which includes a shared encoder and a three-task decoder. Specifically, a forward synchronization structure of a shared encoder + three-task decoder is constructed, i.e., the deep learning network model. In some optional embodiments, the shared encoder uses CSPDarknet as the backbone to solve the gradient copying problem during the optimization process. It supports feature propagation and feature reuse, reducing parameters and computational cost. Therefore, it is beneficial to ensure the real-time performance of the network. A multi-scale neck (SPP+FPN+PAN) is used to output multi-level features; the instance segmentation head is based on upsampling and multi-scale fusion to output a "SPMT+cargo" instance mask and supervises the vertices of the outline polygon during training; the rotation detection head uses single-stage gridded prediction and regression. and category and confidence level ( The center pixel coordinates, (Width and height), while also considering dynamic obstacles; semantic segmentation head, outputting scene semantics such as roads, curbs / guardrails, etc., used to construct a set of safety boundaries. .

[0036] Step S1024: Input the high-altitude image training dataset into the shared encoder for encoding and output the encoded features; Step S1025: Input the encoded features into the three-task decoder for instance segmentation training, rotation detection training and semantic segmentation training, and adjust the parameters of the shared encoder and the three-task decoder according to the annotation set to obtain the multi-task detection model.

[0037] Specifically, based on the aforementioned annotation set { Each frame of the high-altitude image training dataset and its annotation set are sequentially input into the deep learning network model for training. A multi-task joint loss function model is constructed using the GIOU method, combining the outputs of instance segmentation, rotation detection, and semantic segmentation. The loss function value is iteratively optimized using the Adam optimization algorithm. Training stops when the number of iterations reaches the target number or the model accuracy reaches the target model accuracy, resulting in a well-trained multi-task detection model.

[0038] Step S103: Perform Kalman prediction and Hungarian matching on the global spatial perception results to obtain the standardized results of the target vehicle ID and temporal outline and the outline matching results; specifically, Kalman prediction + Hungarian matching is used to maintain the stable update of the target vehicle ID and temporal outline, and the standardized and smoothed multi-frame results are output.

[0039] As an optional implementation, the global spatial perception result includes a set of identified bounding boxes. Step S103 can be further divided into the following steps S1031 to S1036: Step S1031: Based on the global spatial perception result, construct a trajectory set of the target vehicle in the previous frame. The trajectory set includes the target vehicle ID and the state vector of the target vehicle in the previous frame; specifically, establish the trajectory set of the target vehicle in the previous frame. for: ; where the motion vector of the vehicle target bounding box for: ;in, The coordinates of the bounding box center are For area, The aspect ratio (usually considered a constant). These are the rates of change of the quantities mentioned above; The system error covariance matrix is... For the target vehicle ID, For vehicle categories.

[0040] Step S1032: Construct the state transition matrix and observation matrix. Based on the state transition matrix and state vector, obtain the prior state estimate and prior covariance matrix of the current frame. Step S1033: Based on the prior state estimate, obtain the estimated bounding box set. Specifically, a linear uniform velocity model is used for prediction, and the state transition matrix... Used to predict the current state, initialized as follows:

[0041] covariance matrix Empirical parameters; system noise covariance matrix Assume it follows a normal distribution; observation matrix Initialize to:

[0042] Observation noise covariance matrix Assume that it follows a normal distribution.

[0043] Then, from the state transition matrix The optimal estimate of the vehicle target state vector in the previous frame The prior state estimate is obtained. and prior covariance matrix : ; Then, the prior state is estimated. The estimated corner points of the bounding box are obtained:

[0044] ;gather ;in, This represents the estimated width of the vehicle target bounding box obtained through Kalman prediction. These are the pixel coordinates of the top left and bottom right corners of the prediction box. Indicates the first Frame number A set of predicted bounding boxes This represents the set of estimated bounding boxes on the trajectory from the previous frame.

[0045] Step S1034: Perform Hungarian matching between the identified bounding box set and the estimated bounding box set to obtain the outline matching result; specifically, detect the identified bounding box set in the output global spatial perception result. for: ;in, As a category, , where is the confidence level.

[0046] First, calculate the set of recognized bounding boxes. With the estimated bounding box set IOU crossover ratio This is used to measure the degree of overlap between the recognized bounding box and the detected bounding box. Furthermore, based on the IOU intersection-union ratio Calculate the cost matrix : For those with inconsistent categories, set a cost matrix. This indicates that when the categories are inconsistent or When the trajectory is below the set threshold, With detection The matching cost is set to positive infinity to prohibit the matching. In the cost matrix... The Hungarian algorithm is used for matching to obtain the outer contour matching set. .

[0047] Step S1035: Update the prior state estimate based on the observation matrix and the prior covariance matrix to obtain the posterior state estimate of the current frame; Step S1036: Obtain the standardized result of the target vehicle ID and temporal contour based on the posterior state estimate and the contour matching result.

[0048] Specifically, firstly based on the observation matrix Observation noise covariance matrix and the prior covariance matrix Calculate Kalman gain :

[0049] Furthermore, based on Kalman gain Perform state and covariance updates to obtain the posterior state estimate of the current frame. and posterior covariance matrix : ;

[0050] Among them, the observed values , , , , .

[0051] Next, from the posterior state estimation Extract the estimated bounding box of the current frame ;

[0052] .

[0053] Finally, the output is standardized, in frames. Output the normalized result indexed by the target vehicle ID: Furthermore, regarding Temporal filtering is applied to the polygon vertices to stabilize the heading and outline. Among these, The coordinates of the target center in the ground plane coordinate system. The heading angle (in radians) and For equivalent width and height.

[0054] Step S104: Determine the observable state of the vehicle platform based on the calibration results, standardization results, and outline matching results, and generate the system safety outline; specifically, based on the observable state of the vehicle platform, take the union convex hull or perform Minkowski expansion on the cargo outline, and then superimpose the safety margin to obtain the system safety outline.

[0055] As an optional implementation, the calibration result includes a geography homography matrix. Step S104 can be further divided into the following steps S1041 to S1045: Step S1041: Obtain the cargo outline polygon and the vehicle platform outline polygon based on the geography homography matrix and the standardization result; Step S1042: Determine whether the vehicle platform outline polygon is observable based on the outline matching result; Step S1043: If the vehicle platform outline polygon is observable, calculate the convex hull of the union of the cargo outline polygon and the vehicle platform outline polygon to obtain the first system outline; Step S1044: If the vehicle platform outline polygon is unobservable, perform Minkowski dilation on the cargo outline polygon to obtain the second system outline; Step S1045: Perform Minkowski dilation on the first system outline or the second system outline to obtain the system safety outline.

[0056] Specifically, Figure 2 shows a flowchart of the system safety outline generation process. Based on the smoothed outline and pairing results aggregated by target vehicle ID obtained in step S103, the geography homography matrix obtained in step S101 is first used as a reference. By projecting the center point of the border and the vertices of the polygon from the image coordinate system to the ground plane coordinate system, the outline polygon of the cargo is obtained. The vehicle platform has a polygonal outline. .in, For the first The frame contains a smooth polygonal outline of cargo. For the first The frame platform has a smooth outer polygonal outline.

[0057] Next, based on the outline matching result obtained in step S103, it is determined whether the vehicle platform outline polygon is in an observable state (visible state). When the vehicle platform outline polygon is in an observable state, the union convex hull of the cargo and the vehicle platform outline is calculated to obtain the first system outline. : ;in, This is the standard convex hull operator; if the union is already convex, then... It equals the union itself.

[0058] When the vehicle platform's outer polygon is unobservable or unreliable, a conservative expansion of the cargo's outer contour is used to cover the uncertainty of the platform's outline, resulting in the second system outline. : ;in, For Minkowski and, It is a circular structural element; in some alternative embodiments, the expansion radius is... The value ranges from 0.2 to 2.0 m in physical space; in this embodiment, the expansion radius is taken as... The value is 1.0m, which is automatically set after the pixel scale is calculated from the camera calibration matrix.

[0059] Furthermore, to introduce safety clearance and positioning error margin, in the outer contour of the first system Or the outline of the second system On top of that, add another layer of expansion to obtain the system's safety outline. : ;in, This is a safety margin. In some optional embodiments, the safety margin... The value range is 0.1–2.5m, and the embodiment in this application uses a safety margin. It is 0.5m.

[0060] Step S105: Based on the standardization results and the system safety profile, perform trajectory derivation and short-term prediction to obtain the vehicle's future pose and future profile sequence; specifically, based on the smooth trajectory and heading, use a constant curvature model to extrapolate the vehicle's future pose and generate the corresponding future profile sequence.

[0061] As an optional implementation, step S105 can be further divided into the following steps S1051 to S1055: Step S1051: Define the discrete time step and the prediction time domain; Step S1052: Calculate the linear velocity based on the trajectory arc length in the normalization result; Step S1053: Calculate the curvature based on the heading angle in the normalization result; Step S1054: Based on the curvature, update the heading angle and vehicle center position in the prediction time domain according to the discrete time step and the linear velocity to obtain the future pose of the vehicle; Step S1055: Transform the system safety profile according to the future pose of the vehicle to obtain the future profile sequence.

[0062] Specifically, calculated from the smooth trajectory Where s is the arc length along the trajectory, For linear velocity, For curvature, The heading angle is used; the future vehicle pose is predicted in the 1–3s roll time domain using constant curvature. With future outline sequence First, the output of step S103... For a time series, define the discrete time step. With prediction time domain s represents the predicted vehicle pose within the next 1–3 seconds. This is determined by the trajectory arc length. With heading angle To obtain the linear velocity With curvature : ; ;in, and The center of the vehicle is respectively , The instantaneous velocity component in the direction. Get the current value Speed .

[0063] Next, update the heading angle and vehicle center position based on the curvature. Then the heading angle and vehicle center position are updated using the following formula: ; ;like If the speed changes, it degenerates into uniform linear motion, and the vehicle's center position is updated using the following formula, meaning the vehicle moves at a uniform linear speed along the current heading: Furthermore, based on the obtained future vehicle pose (heading angle and vehicle center position), combined with smoothed width and height... With system security outline The future outline sequence is obtained by rigid body transformation with pose. .

[0064] Step S106: Based on the vehicle's future pose and future outline sequence, construct a sweep set, and then, based on the sweep set and the global spatial perception results, obtain the minimum distance and contact time between the target vehicle and the environment; specifically, based on the vehicle's future pose obtained in step S105... With future outline sequence Construct sweep set Furthermore, based on the sweep set and the safety boundary set in the global spatial perception result... Calculate the minimum distance between the target vehicle and the environment. Contact time (TTP).

[0065] As a further optional implementation, the global spatial perception result includes a set of safety boundaries, and the contact time includes a first contact time or a second contact time. The step of obtaining the minimum distance between the target vehicle and the environment and the contact time based on the sweep set and the global spatial perception result can be further divided into the following steps S1061 to S1063: Step S1061: Determine whether the sweep set and the safety boundary set intersect; Step S1062: If the sweep set and the safety boundary set intersect, calculate the minimum distance between the target vehicle and the environment based on the sweep set and the safety boundary set, determine the index, and then perform linear interpolation on the index to obtain the first contact time; Step S1063: If the sweep set and the safety boundary set do not intersect, calculate the minimum distance between the target vehicle and the environment based on the sweep set and the safety boundary set, and then obtain the second contact time based on the minimum distance.

[0066] Specifically, the sweep collection With security boundary set The comparison is performed to determine if a collision / sweep is imminent. The system determines that "sweeping will occur" and calculates the minimum distance between the target vehicle and the environment using the following formula. And on the way to find the earliest... index And thus to The first contact time is obtained by linear interpolation. : ; ;like If it is determined that there are no intersections in the entire region, then the second contact time is approximated by the equivalent closing rate along the normal direction of the nearest distance. : ;in, And set a minimum speed threshold. .

[0067] Step S107: Determine the risk level of collaborative transportation of multi-modal heavy-duty vehicles based on the minimum spacing and contact time.

[0068] Specifically, based on the minimum spacing obtained above The risk level of collaborative transportation of multi-modal heavy-duty vehicles is determined by comprehensively considering factors such as contact time (TTP) and scenario category.

[0069] As an optional implementation, step S107 can be further divided into the following steps S1071 to S1075: Step S1071: Set the upper limit threshold for safe distance, the critical distance threshold for collision, the upper limit threshold for safe time, the warning time threshold, and the emergency response time threshold; Step S1072: If the minimum distance is greater than the upper limit threshold for safe distance and the contact time is greater than the upper limit threshold for safe time, determine the cooperative transportation risk level as the first risk level; Step S1073: If the minimum distance is greater than the critical distance threshold for collision and less than or equal to the upper limit threshold for safe distance, or the contact time is greater than the warning time threshold and less than or equal to the upper limit threshold for safe time, determine the cooperative transportation risk level as the second risk level; Step S1074: If the minimum distance is less than or equal to the critical distance threshold for collision, or the contact time is less than or equal to the warning time threshold, determine the cooperative transportation risk level as the third risk level; Step S1075: If the sweep set intersects with the safe boundary set, or the contact time is less than or equal to the emergency response time threshold, determine the cooperative transportation risk level as the fourth risk level.

[0070] Specifically, Figure 3 shows a flowchart of the risk level determination process. First, a safety distance upper limit threshold is set. (Low-risk boundary), collision critical distance threshold (High-risk boundary), upper limit threshold of safe time (Low-risk boundary), alert time threshold (High-risk boundary) and emergency response time threshold : and then based on the set threshold ( , , , , Minimum Spacing The risk level is determined based on the Time to Contact (TTP) and the following criteria:

[0071] The meanings of each risk level are as follows: First risk level L1 (low risk): The vehicle maintains a safe distance from the environmental boundary and its operating status is stable; Second risk level L2 (Medium Risk): There is an approaching trend, so pay attention to slowing down or warning prompts; L3 (High Risk): The vehicle outline is approaching or about to enter the danger boundary, triggering active braking or path correction; L4 (Emergency): Geometric contact occurs or TTP is below the emergency threshold, so stop immediately or manual intervention is required.

[0072] The final output includes the risk level. The corresponding triggering criteria and early warning signal parameters can be directly used in the vehicle monitoring and dispatch control module to realize automatic early warning and intervention decisions.

[0073] In some optional embodiments, the upper limit threshold for safe distance is... Collision critical distance threshold Safe time upper limit threshold Alert time threshold and emergency response time threshold The system can adaptively adjust based on vehicle loading status (empty, standard, overloaded), speed level, and scenario type. For example, this application embodiment uses a safe distance upper limit threshold. Collision critical distance threshold Safe time upper limit threshold Alert time threshold Emergency response time threshold .

[0074] In summary, the process of the risk assessment method for multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective proposed in this application is shown in Figure 4: Step 1: Collect multi-scene videos from a fixed high-altitude camera position and complete the calibration; Step 2: Train and deploy a shared coding-multi-task decoding detection model to output the "SPMT+cargo" instance mask and outline polygon, rotated circumscribed rectangle and dynamic obstacles, scene semantics and safety boundaries in a single forward synchronous output; Steps 3 to 5: Perform steady-state initialization, Kalman filtering + Hungarian correlation temporal tracking and ID stabilization; Step 6: Generate the system outline (union convex hull / conservative expansion of cargo outline) according to the observable state of the platform, and construct the system safety outline; Steps 7 to 9: Predict the future pose and outline in a short time based on the constant curvature model, construct the sweep set and compare it with the safety boundary set, and calculate the minimum distance and contact time; Step 10: Perform L1–L4 risk classification according to the threshold structure and output the triggering basis.

[0075] The above describes the risk assessment method for multi-modal heavy-duty vehicle collaborative transportation based on a high-altitude perspective according to the embodiments of this application. It can be recognized that the embodiments of this application have the following advantages: First, by combining a high-altitude perspective with a multi-task detection model (instance segmentation, rotation detection, semantic segmentation), the precise geometric outline and environmental boundary of SPMT and cargo can be obtained in real time without relying on on-board sensors, thereby achieving complete spatial perception of the transportation system.

[0076] Second, based on traditional outline recognition, trajectory smoothing, short-term pose prediction and sweep ensemble analysis are introduced to realize dynamic distance determination and contact time (TTP) calculation between vehicles and the environment, thereby enabling early warning before vehicles enter dangerous areas, significantly improving the timeliness and reliability of safety response.

[0077] Third, it supports three loading modes: empty, standard cargo, and oversized cargo. It is suitable for various road sections such as narrow corridors, doorway passages, dock ramps, bridge face edges, and intersection merging, and has strong environmental adaptability and scalability.

[0078] Fourth, by replacing manual visual observation with algorithms, the SPMT operating status can be automatically identified, its outline updated, and risk classified. This reduces reliance on manual labor and can effectively reduce errors and delays caused by operator subjective judgment, thereby improving the safety and decision-making efficiency of transportation tasks.

[0079] Fifth, the output trajectory, outline, and risk level information can be directly connected to the upper-level monitoring and dispatching system for automatic path planning, speed control, and early warning decision-making, laying the foundation for intelligent management and unmanned operation of multi-module vehicles.

[0080] Referring to Figure 5, this application embodiment also provides a multi-modal heavy-duty vehicle collaborative transportation risk assessment system based on a high-altitude perspective, including: a first module for deploying a high-altitude imaging device above the target road segment, calibrating the high-altitude imaging device, and obtaining calibration results; a second module for acquiring vehicle transportation videos through the high-altitude imaging device, inputting the vehicle transportation videos into a pre-trained multi-task detection model for instance segmentation, rotation detection, and semantic segmentation, and outputting global spatial perception results of the target vehicle and cargo; and a third module for performing Kalman prediction and Hungarian matching on the global spatial perception results to obtain the target vehicle ID and temporal outline calibration. The system consists of seven modules: a standardization module and a contour matching module; a fourth module, which determines the observable state of the vehicle platform based on the calibration, standardization, and contour matching results, and generates a system safety contour; a fifth module, which performs trajectory derivation and short-term prediction based on the standardization results and the system safety contour, to obtain the vehicle's future pose and future contour sequence; a sixth module, which constructs a sweep set based on the vehicle's future pose and future contour sequence, and then obtains the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception results; and a seventh module, which determines the cooperative transportation risk level of multi-modular heavy-duty vehicles based on the minimum distance and contact time.

[0081] It is understood that the content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0082] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0083] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0084] Please refer to Figure 6, which illustrates the hardware structure of an electronic device according to another embodiment. The electronic device includes: a processor 1001, which can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, for executing related programs to implement the technical solutions provided in the embodiments of this application; and a memory 1002, which can be implemented using a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM), etc. The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001. The input / output interface 1003 is used to implement information input and output. The communication interface 1004 is used to realize communication interaction between this device and other devices. Communication can be realized through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.). The bus 1005 transmits information between various components of the device (such as processor 1001, memory 1002, input / output interface 1003 and communication interface 1004). The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device through the bus 1005.

[0085] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0086] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0087] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0088] It is understood that the content of the above method embodiments is applicable to the embodiments of this program product. The specific functions implemented by the embodiments of this program product are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0089] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0090] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0091] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0092] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0093] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0094] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0095] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0097] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0099] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0100] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A risk assessment method for collaborative transportation of multi-modal heavy-duty vehicles based on a high-altitude perspective, characterized in that, Includes the following steps: A high-altitude imaging device is deployed above the target road section, and the high-altitude imaging device is calibrated to obtain calibration results; The high-altitude imaging device acquires vehicle transportation videos, which are then input into a pre-trained multi-task detection model for instance segmentation, rotation detection, and semantic segmentation. The model outputs global spatial perception results for the target vehicle and cargo. Kalman prediction and Hungarian matching are applied to these global spatial perception results to obtain standardized results of the target vehicle ID and temporal outline, as well as outline matching results. Based on the calibration results, the standardized results, and the outline matching results, the observable state of the vehicle platform is determined, and a system safety outline is generated. Based on the standardized results and the system safety profile, trajectory derivation and short-term prediction are performed to obtain the vehicle's future pose and future profile sequence. Based on the vehicle's future pose and the future profile sequence, a sweep set is constructed. Then, based on the sweep set and the global spatial perception results, the minimum distance and contact time between the target vehicle and the environment are obtained. Based on the minimum distance and the contact time, the cooperative transportation risk level of multi-modal heavy-duty vehicles is determined.

2. The method according to claim 1, characterized in that, The method further includes a step of pre-training the multi-task detection model, which specifically includes: acquiring a high-altitude image training dataset through the high-altitude imaging device; annotating the vehicle-marked bounding box set, vehicle rotation bounding rectangle set, instance mask set, outline polygon set, scene semantic annotation set, safety boundary set, and dynamic obstacle set in the high-altitude image training dataset to obtain an annotation set; constructing a deep learning network model, which includes a shared encoder and a three-task decoder; inputting the high-altitude image training dataset into the shared encoder for encoding and outputting encoded features; inputting the encoded features into the three-task decoder for instance segmentation training, rotation detection training, and semantic segmentation training, and adjusting the parameters of the shared encoder and the three-task decoder according to the annotation set to obtain the multi-task detection model.

3. The method according to claim 1, characterized in that, The global spatial perception result includes a set of identified bounding boxes. The process of performing Kalman prediction and Hungarian matching on the global spatial perception result to obtain the standardized result of the target vehicle ID and temporal outline, and the outline matching result, specifically includes: constructing a trajectory set of the target vehicle from the previous frame based on the global spatial perception result; constructing a state transition matrix and an observation matrix; obtaining the prior state estimate and prior covariance matrix of the current frame based on the state transition matrix and the state vector; obtaining an estimated bounding box set based on the prior state estimate; performing Hungarian matching between the identified bounding box set and the estimated bounding box set to obtain the outline matching result; updating the prior state estimate based on the observation matrix and the prior covariance matrix to obtain the posterior state estimate of the current frame; and obtaining the standardized result of the target vehicle ID and temporal outline based on the posterior state estimate and the outline matching result.

4. The method according to claim 1, characterized in that, The calibration result includes a geography homography matrix. The step of determining the observable state of the vehicle platform based on the calibration result, the standardization result, and the outline matching result, and generating a system safety outline, specifically includes: obtaining cargo outline polygons and vehicle platform outline polygons based on the geography homography matrix and the standardization result; determining whether the vehicle platform outline polygon is observable based on the outline matching result; if the vehicle platform outline polygon is observable, calculating the union convex hull of the cargo outline polygon and the vehicle platform outline polygon to obtain a first system outline; if the vehicle platform outline polygon is unobservable, performing Minkowski dilation on the cargo outline polygon to obtain a second system outline; and performing Minkowski dilation on either the first or second system outline to obtain the system safety outline.

5. The method according to claim 1, characterized in that, The process of deriving and predicting the trajectory based on the standardized result and the system safety profile to obtain the vehicle's future pose and future profile sequence specifically includes: defining a discrete time step and a prediction time domain; calculating the linear velocity based on the trajectory arc length in the standardized result; calculating the curvature based on the heading angle in the standardized result; updating the heading angle and the vehicle center position in the prediction time domain based on the curvature, the discrete time step, and the linear velocity to obtain the vehicle's future pose; and transforming the system safety profile based on the vehicle's future pose to obtain the future profile sequence.

6. The method according to claim 1, characterized in that, The global spatial perception result includes a set of safety boundaries, and the contact time includes a first contact time or a second contact time. The step of obtaining the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception result specifically includes: determining whether the sweep set intersects with the set of safety boundaries; if the sweep set intersects with the set of safety boundaries, calculating the minimum distance between the target vehicle and the environment based on the sweep set and the set of safety boundaries, determining an index, and then performing linear interpolation on the index to obtain the first contact time; if the sweep set does not intersect with the set of safety boundaries, calculating the minimum distance between the target vehicle and the environment based on the sweep set and the set of safety boundaries, and then obtaining the second contact time based on the minimum distance.

7. The method according to claim 6, characterized in that, The step of determining the collaborative transportation risk level of multi-modal heavy-duty vehicles based on the minimum distance and the contact time specifically includes: setting a safety distance upper limit threshold, a collision critical distance threshold, a safety time upper limit threshold, a warning time threshold, and an emergency response time threshold; if the minimum distance is greater than the safety distance upper limit threshold and the contact time is greater than the safety time upper limit threshold, the collaborative transportation risk level is determined to be a first risk level; if the minimum distance is greater than the collision critical distance threshold and less than or equal to the safety distance upper limit threshold, or the contact time is greater than the warning time threshold and less than or equal to the safety time upper limit threshold, the collaborative transportation risk level is determined to be a second risk level; if the minimum distance is less than or equal to the collision critical distance threshold, or the contact time is less than or equal to the warning time threshold, the collaborative transportation risk level is determined to be a third risk level; if the sweep set intersects with the safety boundary set, or the contact time is less than or equal to the emergency response time threshold, the collaborative transportation risk level is determined to be a fourth risk level.

8. A risk assessment system for collaborative transportation of multi-modal heavy-duty vehicles based on a high-altitude perspective, characterized in that, include: The first module is used to deploy a high-altitude imaging device above the target road section, calibrate the high-altitude imaging device, and obtain calibration results. The second module is used to acquire vehicle transportation video through the high-altitude imaging device, input the vehicle transportation video into a pre-trained multi-task detection model for instance segmentation, rotation detection and semantic segmentation, and output the global spatial perception results of the target vehicle and cargo; the third module is used to perform Kalman prediction and Hungarian matching on the global spatial perception results to obtain the standardized results of the target vehicle ID and temporal outline and the outline matching results. The fourth module is used to determine the observable state of the vehicle platform based on the calibration results, the standardization results, and the outline matching results, and to generate the system safety outline. The fifth module is used to perform trajectory derivation and short-term prediction based on the standardization results and the system safety profile to obtain the vehicle's future pose and future profile sequence. The sixth module is used to construct a sweep set based on the future pose of the vehicle and the future outline sequence, and then obtain the minimum distance and contact time between the target vehicle and the environment based on the sweep set and the global spatial perception result. The seventh module is used to determine the risk level of collaborative transportation of multi-modular heavy-duty vehicles based on the minimum spacing and the contact time.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method of any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.