Road collapse treatment method, device, chip, vehicle, medium and product
Through the road collapse identification model, multi-layer feature extraction and classification of the spatial data collected by the vehicle is used to identify the type and location of the road collapse, and solve the problem that the vehicle cannot effectively detect and respond to road collapse during driving, achieving efficient and accurate collapse recognition and timely response, and improving the safety of vehicles and other vehicles.
Patent Information
- Application Number
- CN202510807505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the prior art, vehicles cannot effectively detect and respond to road collapse during driving, resulting in greater safety hazards.
Through the road collapse identification model, multi-layer feature extraction and classification processing is performed on the spatial data of the front road collected by the vehicle, identify the collapse type and location, and output it in real time for the vehicle to respond.
It improves the accuracy and timeliness of vehicles identifying road collapses, enhances the safety of vehicles, and improves the driving safety of other vehicles through Internet of vehicles and cloud backup.
Smart Images

Figure CN120318794B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of automotive technology, and in particular to methods, devices, chips, vehicles, media, and products for handling road collapse. Background Art
[0002] With the increasing number of road construction, road safety hazards are becoming more and more prominent. Among them, road collapse causing vehicle driving safety problems is also relatively common. How to solve the problems caused by road collapse during vehicle driving is particularly important. Summary of the Invention
[0003] One of the purposes of this application is to provide a method, device, chip, vehicle, medium and product for handling road collapse. This solution can autonomously detect whether there is a collapse in the road ahead during vehicle driving, and can obtain the collapse type and collapse location, so that the vehicle can respond in a timely manner based on the collapse type and collapse location of the road ahead, thereby improving the safety of vehicle driving.
[0004] In order to achieve the above objectives, the technical solutions adopted in this application are as follows:
[0005] In a first aspect, the present application provides a method for processing road collapse, the method comprising: obtaining spatial data of the road ahead collected by a vehicle; performing identification processing on the spatial data through a road collapse recognition model to determine an identification result of the road ahead; the identification result comprises a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; wherein, the identification processing of the road collapse recognition model comprises: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain a multi-layer target feature map; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a prediction detection frame of each layer; the size of the prediction detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the prediction detection frame of each layer through a classification model to obtain a classification result of each layer; determining an identification result of the road ahead based on the classification result of each layer; when the first identifier indicates that there is a collapse in the road ahead, outputting the collapse type and collapse position in the identification result so that the vehicle performs response processing based on the collapse type and collapse position; the collapse types include: road rupture, lateral road collapse, and longitudinal road collapse.
[0006] Based on the above technical means, while a vehicle is driving, it can obtain real-time road surface recognition results from the collected spatial data of the road ahead. This allows it to determine whether there is a collapse ahead, as well as the type and location of the collapse, and thus respond accordingly. On the one hand, responding based on the collapse type and location can improve the accuracy of the response process. On the other hand, in this solution, the vehicle can autonomously detect whether there is a collapse ahead, allowing it to respond promptly based on the collapse type and location, improving the timeliness of the response and the safety of the vehicle. Furthermore, when determining the recognition result, a feature extraction model is first used to perform multi-layer feature extraction on the spatial data. Then, for each layer, a prediction reference frame is determined using the target reference frame of that layer. The classification result for that layer is determined based on the prediction reference frame of that layer, and the final recognition result is determined based on the classification results of each layer. As can be seen, when determining the recognition result, the resolution of the features corresponding to each layer matches the collapse type, and the size of the prediction detection frame corresponding to each layer matches the collapse type. This allows each layer to recognize collapse types of different sizes. For example, one layer can be used to identify road cracks, while another layer can be used to identify road collapse. This allows multiple types of road collapse problems to be identified simultaneously, improving recognition efficiency and accuracy.
[0007] In one possible implementation, the method further includes: determining the vehicle's own handling method based on the collapse type and collapse location; sending a first prompt message, the collapse type, and the collapse location to the vehicle behind the vehicle through vehicle network communication; the first prompt message is used to remind that there is a collapse ahead; and sending the collapse type and collapse location to the cloud.
[0008] Based on these technologies, a vehicle's response to a collapse based on its type and location includes not only its own response but also alerts to vehicles behind it and backs up information in the cloud. This comprehensive approach improves not only the vehicle's own safety but also the safety of other vehicles.
[0009] In one possible implementation, multi-layer feature extraction is performed on spatial data through a feature extraction model to obtain a multi-layer target feature map, including: performing multi-scale convolution processing on the spatial data through a fully convolutional network, splicing the outputs of multiple stages to obtain a multi-layer first feature map; the resolution of the first feature map of each layer in the multi-layer first feature map matches the collapse type; detailed feature extraction is performed on the first feature map of each layer in the multi-layer first feature map based on the attention mechanism to obtain a multi-layer second feature map; the multi-layer second feature map is fused to obtain a multi-layer third feature map; and it is determined that the multi-layer target feature map includes a multi-layer first feature map, a multi-layer second feature map, or a multi-layer third feature map.
[0010] Based on the above technical means, when determining the target feature map, the target feature map can be determined to be a multi-layer first feature map output by a fully convolutional network. The resolution of the multi-layer first feature map matches the collapse type, which can meet the layer identification requirements. The target feature map is determined by the fully convolutional network, which has the characteristics of simple implementation logic and high processing efficiency. When determining the target feature map, the target feature map can be determined to be a multi-layer second feature map. The resolution of the multi-layer second feature map matches the collapse type, which can meet the layer identification requirements. The multi-layer second feature map can also extract detail features based on the attention mechanism. The target feature map obtained in this way includes detail features, the obtained features are more comprehensive, and the collapse identification is more accurate. When determining the target feature map, the target feature map can be determined to be a multi-layer third feature map. The resolution of the multi-layer third feature map matches the collapse type, which can meet the layer identification requirements. The multi-layer third feature map is obtained by fusing the second feature map. Therefore, it not only has the detail features in the second feature map, but also the fusion of each layer allows each layer to retain the features of other layers. The target feature map obtained in this way is more comprehensive and the collapse identification is more accurate.
[0011] In one possible embodiment, when the multi-layer includes three layers, multi-scale convolution processing is performed on the spatial data through a fully convolutional network, and the outputs of multiple stages are spliced to obtain the first feature maps of the multi-layer, including: convolution processing on the spatial data based on a first downsampling rate to obtain the first feature map of the first layer; convolution processing on the spatial data based on a second downsampling rate to obtain the first feature map of the second layer; convolution processing on the spatial data based on a third downsampling rate to obtain the first feature map of the third layer; based on the first feature map of the first layer, the first feature map of the second layer and the first feature map of the third layer, determining the first feature maps of the multi-layer; wherein the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.
[0012] Based on the above technical means, when determining the first feature map, convolution processing is performed through different downsampling rates to obtain the first feature maps at different resolutions. This has the characteristics of simple implementation and reliability.
[0013] In a possible embodiment, when the first feature maps of the multiple layers include the first feature maps of the first layer, the first feature maps of the second layer, and the first feature maps of the third layer, detail features are extracted from the first feature maps of each layer in the multiple layers based on the attention mechanism to obtain the second feature maps of the multiple layers, including: extracting detail features from the first feature map of the first layer based on the first receptive field range by the first attention module to obtain the second feature map of the first layer; extracting detail features from the first feature map of the second layer based on the second receptive field range by the second attention module to obtain the second feature map of the second layer; extracting detail features from the first feature map of the third layer based on the third receptive field range by the third attention module to obtain the second feature map of the third layer; determining the second feature maps of the multiple layers based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer; wherein the first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.
[0014] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, each of which corresponds to a different receptive field, second feature maps of different layers can be obtained. This is characterized by simple implementation, reliability, and accuracy.
[0015] In a possible implementation, when the second feature maps of the multiple layers include the second feature maps of the first layer, the second feature maps of the second layer, and the second feature maps of the third layer, the second feature maps of the multiple layers are fused to obtain the third feature maps of the multiple layers, including: performing cascade convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to a first path to obtain the intermediate feature maps of the three layers; performing cascade convolution processing on the intermediate feature maps of the three layers according to a second path to obtain the third feature maps of the multiple layers; the second path is different from the first path.
[0016] Based on the above technical means, when fusing the second feature maps of multiple layers, the fusion is first performed based on the first path, and then based on the second path. Through the fusion of the two paths, the target feature map of each layer after fusion can retain the features of all upper and lower layers. The feature map of each layer in the obtained target feature map has comprehensive information, which improves the accuracy of identifying the collapse type of each layer.
[0017] In one possible implementation, the target features corresponding to the predicted detection frame of each layer are classified by a classification model to obtain a classification result for each layer, including: obtaining a target reference frame for each layer; determining the matching degree between the predicted detection frame and the target reference frame of each layer respectively to obtain multiple matching degrees; the matching degree is used to characterize the degree of overlap between the predicted detection frame and the target reference frame; determining the collapse type based on multiple matching degrees; determining the collapse position based on the position of the predicted detection frame; and determining that the classification result for each layer includes the collapse type and the collapse position.
[0018] Based on the above technical means, when determining the collapse type, it is determined based on the matching degree between the predicted detection frame and the target reference frame of each layer. In this way, since the target reference frame of a layer can represent the characteristics of a type of collapse, the collapse type is determined more accurately based on the matching degree. When determining the collapse location, since the predicted detection frame is used to indicate the collapse location, it is more accurate to determine the collapse location based on the position of the predicted detection frame.
[0019] In a possible embodiment, the method also includes: obtaining a sample data set; the sample data set includes multiple sample space data, each sample space data is marked with a real box; the real box is used to point to the collapse position in the sample space data; the multiple real boxes in the sample data set are distributed to multiple layers according to size, so that each layer includes multiple real boxes; for each layer in the multiple layers, based on the multiple real boxes included in the layer, the center of the target reference box of the layer is determined; based on the center of the target reference box of the layer, the target reference box of the layer is determined.
[0020] Based on the above technical means, the sample space data in the sample dataset can be divided into different layers according to the size of the ground truth box. The center of the target reference box of each layer is determined based on the size of the ground truth box of each layer, and then the target reference box of each layer is determined. In this way, the target reference box determined for each layer can meet the size requirements of a certain type of collapse, improving the recognition accuracy of that layer for that collapse type.
[0021] In a possible implementation, determining the center of a target reference frame of a layer based on multiple real frames included in the layer includes: determining center points of multiple real frames based on the multiple real frames of the layer; respectively determining the probability density of the center point of each real frame of the layer; and determining the center point with the highest probability density as the center of the target reference frame.
[0022] Based on the above technical means, when determining the center of the target reference frame of a layer, it is determined based on the probability density method, which can meet the needs of most real frames and is more accurate.
[0023] In a possible implementation, a target reference frame of a layer is determined based on the center of the target reference frame of the layer, including: generating a candidate reference frame with the center point of the target reference frame as the center; determining the dynamic distance between each real frame of the layer and the candidate reference frame respectively; the dynamic distance is negatively correlated with the overlap, and the dynamic distance is positively correlated with the aspect ratio similarity; adjusting the candidate reference frame based on the dynamic distance until the dynamic distance meets a preset condition; and determining the candidate reference frame that meets the preset condition as the target reference frame of the layer.
[0024] Based on the above technical means, when determining the target reference frame, a candidate reference frame is first generated based on the center point, and then the shape and size of the candidate reference frame are continuously adjusted based on the dynamic distance, so that the candidate reference frame and the real frame of the layer meet the preset conditions. In this way, the target reference frame can have a large overlap with the real frame and a close aspect ratio with the real frame, that is, the position and shape are close to the real frame, thereby meeting actual needs and improving accuracy.
[0025] In a possible implementation, the method further includes: determining, for each true box of the layer, a degree of matching between the true box and the target reference box of the layer; adjusting the layer to which the true box belongs based on the degree of matching; and updating the target reference box of each layer based on the adjusted true box included in each layer.
[0026] Based on the above technical means, after each ground-truth frame is stratified, the allocation method can be adjusted based on the matching degree. This can optimize the stratification, improve the accuracy of the segmentation, and also improve the accuracy of the collapse type identification.
[0027] In a second aspect, the present application provides a device for handling road collapse, the device comprising:
[0028] An acquisition unit, used to acquire spatial data of the road ahead collected by the vehicle;
[0029] An identification unit is used to identify and process spatial data through a road collapse identification model to determine an identification result of the road ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; wherein the identification processing of the road collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain a multi-layer target feature map; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame of each layer; the size of the predicted detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection frame of each layer through a classification model to obtain a classification result of each layer; and determining an identification result of the road ahead based on the classification result of each layer.
[0030] The output unit is used to output the collapse type and collapse location in the identification result when the first identifier indicates that there is a collapse in the road ahead, so that the vehicle can respond based on the collapse type and collapse location; the collapse types include: road surface fracture, road surface lateral collapse, and road surface longitudinal collapse.
[0031] In a third aspect, the present application provides a system-on-chip, wherein the system-on-chip is connected to a controller for a road collapse;
[0032] The system-level chip is used to: obtain spatial data of the road ahead collected by the vehicle; identify and process the spatial data through a road collapse recognition model to determine the recognition result of the road ahead; the recognition result includes a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; wherein, the recognition processing of the road collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining the target reference frame of each layer, and predicting and processing the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame of each layer; the size of the predicted detection frame of each layer matches the collapse type; classifying and processing the target features corresponding to the predicted detection frame of each layer through a classification model to obtain the classification result of each layer; and determining the recognition result of the road ahead based on the classification result of each layer;
[0033] The system-level chip is also used to: when the first identifier indicates that there is a collapse in the road ahead, send the collapse type and collapse location in the identification result to the controller of the road collapse; so that the controller can respond based on the collapse type and collapse location; the collapse types include: road rupture, road lateral collapse, and road longitudinal collapse.
[0034] Because collapses are long-tail scenarios—highly destructive but low-probability events—this solution implements these low-probability events through the system-on-chip (SoC), rather than the vehicle's main controller. This prevents these low-probability events from impacting the vehicle's normal control efficiency. When a collapse is detected, the type and location of the collapse are transmitted to the main controller in real time, allowing it to respond and process, while also improving safety.
[0035] In a fourth aspect, the present application further provides a vehicle, comprising a processor and a memory, wherein the memory stores a computer program or instructions, and when the computer program or instructions are executed by the processor, the method provided in the first aspect is implemented.
[0036] In a fifth aspect, the present application also provides a storage medium on which a computer program or instruction is stored. When the computer program or instruction is executed by a processor, the method provided in the first aspect is implemented.
[0037] In a sixth aspect, the present application also provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the method provided in the first aspect is implemented.
[0038] It should be noted that the technical effects of the second to sixth aspects can refer to the detailed description of the first aspect above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 A schematic diagram of a first optional flow chart of a method for handling road collapse provided in an embodiment of the present application;
[0040] Figure 2 A second optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0041] Figure 3 A schematic diagram of a third optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0042] Figure 4 A fourth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0043] Figure 5 A fifth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0044] Figure 6 A sixth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0045] Figure 7 A seventh optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0046] Figure 8 This is a schematic diagram of an eighth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0047] Figure 9 A ninth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0048] Figure 10 A tenth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0049] Figure 11 This is a schematic diagram of an eleventh optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0050] Figure 12A twelfth optional flow chart of the method for handling road collapse provided in an embodiment of the present application;
[0051] Figure 13 An optional flow chart of a road collapse treatment process provided in an embodiment of the present application;
[0052] Figure 14 An optional structural diagram of the FPGA hardware chip system provided in the embodiment of the present application;
[0053] Figure 15 An optional structural diagram of the FPGA hardware accelerator provided in an embodiment of the present application;
[0054] Figure 16 A schematic diagram of an optional structure of a device for handling road collapse provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the specific technical solutions of the application will be further described in detail below in conjunction with the drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.
[0056] In the following description, reference is made to “some embodiments”, which describes a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0057] In the following description, the terms "first, second, and third" are used merely as examples to distinguish between different objects and do not represent a specific order or precedence for the objects. It is understood that the specific order or precedence of "first, second, and third" can be interchanged where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0058] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.
[0059] The present application provides embodiments of a method, device, chip, vehicle, medium, and product for handling road collapse. The method for handling road collapse is performed by a device for handling road collapse, which can be deployed in a vehicle. For example, the device for handling road collapse can be a controller in the vehicle. The following describes various embodiments of the method, device, chip, vehicle, medium, and product for handling road collapse provided in the present application.
[0060] In a first aspect, embodiments of the present application provide a method for handling road collapse. The method is described below using a vehicle as an example.
[0061] refer to Figure 1 As shown in the content, the process may include but is not limited to S101 to S103.
[0062] S101: Acquire spatial data of the road ahead collected by the vehicle.
[0063] The embodiments of this application do not limit the type of vehicle and can be configured according to actual needs. The vehicles here can include but are not limited to sedans, commercial vehicles, sports cars, etc.; they can also include but are not limited to: gasoline vehicles, electric vehicles, hydrogen energy vehicles, etc.
[0064] Spatial data refers to the collected image data or point cloud data of the road ahead, or the data obtained by fusing image data and point cloud data.
[0065] During the driving process, the vehicle collects information of the road ahead through cameras and radar sensors to obtain spatial data. S101 can be implemented as: obtaining the spatial data of the road ahead collected by the vehicle sensors.
[0066] S102: Perform recognition processing on the spatial data using a road collapse recognition model to determine a recognition result of the road ahead.
[0067] The recognition result includes a first identifier.
[0068] The first indicator is used to indicate whether there is a collapse in the road ahead. The embodiment of the present application does not limit the value of the first indicator, and can be configured according to actual needs. For example, a first indicator of 1 indicates that there is a collapse in the road ahead, and a first indicator of 0 indicates that there is no collapse in the road ahead.
[0069] When the first sign indicates that there is a collapse on the road ahead, the recognition result may include the collapse type and collapse location. The collapse type may include but is not limited to: road rupture, road lateral collapse, road longitudinal collapse, etc.
[0070] The identification result may include one collapse or multiple collapses. In the case where the identification result includes multiple collapses, the types corresponding to the multiple collapses may be consistent or inconsistent.
[0071] Among them, the recognition processing of the road collapse recognition model includes: performing multi-layer feature extraction on spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining the target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain the prediction detection frame of each layer; the size of the prediction detection frame of each layer matches the collapse type; classifying the target features corresponding to the prediction detection frame of each layer through a classification model to obtain the classification results of each layer; and determining the recognition results of the road ahead based on the classification results of each layer.
[0072] For example, one collapse type corresponds to a target feature map of a layer with a certain resolution, and one collapse type corresponds to a predicted detection box of a certain size.
[0073] S102 can be implemented as follows: inputting the spatial data of the road ahead into a road collapse recognition model, and performing recognition processing on the spatial data through the road collapse recognition model. The road collapse recognition model recognizes whether there is a collapse in the road ahead. If there is a collapse, the road collapse recognition model recognizes the type and location of the collapse, thereby obtaining the recognition result of the road ahead.
[0074] The road collapse recognition model can include a convolutional network model, a prediction detection box model, and a classification model. The convolutional network model is used to extract multi-layer target feature maps, the prediction detection box model is used to determine multi-layer prediction detection boxes, and the classification model is used to determine multi-layer classification results.
[0075] S103: When the first indicator indicates that there is a collapse on the road ahead, output the collapse type and collapse location in the recognition result, so that the vehicle performs a response process based on the collapse type and collapse location.
[0076] The embodiment of the present application does not limit the method of outputting the recognition results, which can be configured according to actual needs. For example, outputting the recognition results can include but is not limited to: sending to other devices, displaying on the vehicle's central control screen, outputting via voice, etc.
[0077] S103 may be implemented as follows: determining the content of the first indicator in the recognition result; if the first indicator indicates that there is a collapse in the road ahead, determining the type and location of the collapse, and outputting the type and location of the collapse in the recognition result so that the vehicle can respond based on the type and location of the collapse. If the first indicator indicates that there is no collapse in the road ahead, no response is taken, and the sensor continues to collect spatial data of the road ahead, and a new round of determination is performed.
[0078] Steps S101 to S103 can be performed by the vehicle's main controller. This eliminates the need for an additional controller, simplifying implementation. Alternatively, a dedicated controller dedicated to collapse detection can be added to the vehicle. This dedicated controller performs the detection without impacting the main controller's control functions or processing efficiency.
[0079] In this embodiment, the method includes: obtaining spatial data of the road ahead collected by the vehicle; identifying and processing the spatial data through a road collapse recognition model to determine an identification result of the road ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; when the first identifier indicates that there is a collapse in the road ahead, outputting the collapse type and collapse location in the identification result so that the vehicle can respond based on the collapse type and collapse location.
[0080] Based on the above technical means, while a vehicle is driving, it can obtain real-time road surface recognition results by collecting spatial data from the road ahead. This allows it to determine whether a collapse is occurring ahead, as well as the type and location of the collapse, and then respond accordingly. On the one hand, responding based on the collapse type and location can improve the accuracy of response processing. On the other hand, in this solution, the vehicle can autonomously detect whether a collapse is occurring ahead, allowing it to respond promptly based on the collapse type and location, improving both timeliness and driving safety. Furthermore, the road collapse recognition model utilizes multiple layers of recognition. Because the resolution of each layer's features matches the collapse type, and the size of each layer's predicted detection box matches the collapse type, each layer can identify collapses of varying sizes. For example, one layer can identify road fractures, while another layer can identify road collapses. This allows for simultaneous identification of multiple road collapse types, improving recognition efficiency and accuracy.
[0081] The processing method provided in the embodiment of the present application may further include a processing process based on the collapse type and collapse location after determining the collapse type and collapse location.
[0082] In one possible embodiment, referring to Figure 2 The process may include but is not limited to the following S201 to S203.
[0083] S201. Determine the vehicle's own handling method based on the collapse type and collapse location.
[0084] In one possible implementation, if the vehicle is in autonomous driving mode, an avoidance strategy or a driving stop strategy is determined based on the type and location of the road collapse ahead. For example, if the collapse type includes a lateral collapse, and the collapse is large and extends across the entire width of the road, the vehicle's response is to stop. If the collapse type includes a road crack, and the crack is small, the vehicle's response is to circumvent the road.
[0085] In another possible implementation, if the vehicle is in manual driving mode, the vehicle's handling method is output via a display or voice prompt to remind the driver of the driving plan for the road ahead. This allows the driver to promptly understand the handling method for the road ahead and take timely action, preventing the driver from being unable to arrive at a correct solution in a timely manner due to thinking about how to handle the situation, which can cause a malfunction.
[0086] S202: Sending a first prompt message, a collapse type, and a collapse location to a vehicle behind the vehicle via Internet of Vehicles communication.
[0087] The first prompt information is used to remind that there is a collapse ahead.
[0088] Rear vehicles refer to vehicles within a certain range behind your vehicle. Rear vehicles here are not limited to vehicles behind your lane, but can also include vehicles in other lanes traveling in the same direction.
[0089] Since a communication connection is established between the vehicle and surrounding vehicles in the IoV, S202 can be implemented as follows: first prompt information, collapse type, and collapse location are sent to other vehicles within a certain range behind the vehicle via IoV communication, so that other vehicles can respond immediately after receiving the first prompt information, collapse type, and collapse location, thereby increasing response time and reducing accident rates.
[0090] S203: Send the collapse type and collapse location to the cloud.
[0091] S203 may be implemented as follows: the vehicle sends the collapse type and collapse location of the road ahead detected by the vehicle to the cloud through the communication connection between the vehicle and the cloud device, so that the cloud records the collapse location and collapse type.
[0092] The cloud can continuously update the map based on the collapse locations and types of the road surface reported by many vehicles, thereby obtaining a continuously updated map including collapse information.
[0093] Based on these technologies, a vehicle's response to a collapse based on its type and location includes not only its own response but also alerts to vehicles behind it and backs up information in the cloud. This comprehensive approach improves not only the vehicle's own safety but also the safety of other vehicles.
[0094] During the execution of S201 to S203, the execution subject of S201 to S203 can be the main controller in the vehicle. In the case that the execution subject of S101 to S103 is a newly added controller, the new controller sends the collapse information to the vehicle's controller, and the main controller responds based on the collapse information. Since collapse belongs to a long-tail scenario, that is, an event with high damage but low probability, in this solution, low-probability events are implemented through a new controller rather than the vehicle's main controller, so that low-probability events do not affect the efficiency of normal control of the vehicle. When a collapse is identified, the collapse type and location are sent to the main control in real time, allowing the main control to respond and process, while improving safety.
[0095] In other processing methods, one or more of the above S201 to S203 can also be configured according to actual needs. For example, only the self-processing process in S201 can be included, or the self-processing process and the following vehicle reminder can be included. Other situations are not listed one by one.
[0096] Next, the process of performing recognition processing on the spatial data by using the road collapse recognition model and determining the recognition result of the road ahead in S102 will be described.
[0097] refer to Figure 3 The process may include but is not limited to the following S301 to S304.
[0098] S301. Perform multi-layer feature extraction on spatial data through a feature extraction model to obtain a multi-layer target feature map.
[0099] The resolution of each target feature map in the multi-layer target feature map matches the collapse type.
[0100] The embodiment of the present application does not limit the number of layers, and can be configured according to actual needs. In a possible implementation, the number of layers can be consistent with the number of collapse types, so that one layer corresponds to one type of collapse.
[0101] For example, when the collapse types include pavement rupture, pavement lateral collapse, and pavement longitudinal collapse, the number of layers can be configured as 3.
[0102] S301 can be implemented as follows: determine the number of layers of the feature extraction model based on the collapse type, configure the size of the feature extraction of each layer, and then perform feature extraction on each layer of the spatial data through the feature extraction model. The resolution of each layer obtained matches the collapse type, so that multi-layer target feature maps of different resolutions can be obtained. The implementation of this application does not limit the type of feature extraction model and can be configured according to actual needs. For example, the feature extraction model can be a full convolutional network model. Or it can be a combination of a full convolutional network model and an attention model.
[0103] S302: Obtain the target reference frame of each layer, and perform prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain the predicted detection frame of each layer.
[0104] The predicted detection box refers to the box corresponding to the predicted collapse content of that layer. Each layer corresponds to a different collapse type, so a predicted detection box must be generated for each layer. A layer can have one or more predicted detection boxes (corresponding to multiple collapses of the same type), or no predicted detection box (corresponding to no collapse of a certain type).
[0105] The size of the predicted detection box at each layer matches the collapse type.
[0106] Target reference frame: Each layer corresponds to a target reference frame, corresponding to a collapse type. For example, if the collapse type is road fracture, a slender target reference frame is configured; if the collapse type is horizontal collapse, a horizontal rectangular reference frame is configured; if the collapse type is vertical collapse, a vertical rectangular reference frame is configured.
[0107] S302 can be implemented as follows: for each of the multiple layers, perform the following processing respectively: obtain the target reference frame of each layer, perform corresponding prediction processing on the target feature map of each layer based on the target reference frame of each layer through the detection frame prediction model, and obtain the predicted detection frame of each layer.
[0108] S303: Classify the target features corresponding to the predicted detection boxes of each layer through the classification model to obtain the classification results of each layer.
[0109] S303 can be implemented as follows: inputting the target features in the predicted detection frame into the classification model, and the target classification model classifies the target features in the predicted detection frame to obtain a classification result for the input. Then, the predicted detection frame of each layer is traversed to obtain a classification result for each layer.
[0110] Example 1: For the first layer, one fracture is identified at position 1; for the second layer, two transverse collapses are identified at positions 2 and 3; and for the third layer, one longitudinal collapse is identified at position 4.
[0111] S304: Determine the recognition result of the road ahead based on the classification result of each layer.
[0112] S304 can be implemented by summarizing the classification results of all layers to obtain an identification result of the road ahead. For example, based on Example 1, the identification result may include: one fracture, two lateral collapses, and one longitudinal collapse. The identification result also includes the location information of the one fracture, two lateral collapses, and one longitudinal collapse.
[0113] Based on the above technical means, when determining the recognition result, the spatial data is first subjected to multi-layer feature extraction through the feature extraction model. Then, for each layer, the predicted reference frame is determined through the target reference frame of the layer, the classification result of the layer is determined based on the predicted reference frame of the layer, and the final recognition result is determined based on the classification result of each layer. It can be seen that when determining the recognition result, it is achieved through multiple layers. Since the resolution corresponding to the features of each layer matches the collapse type, and the size corresponding to the predicted detection frame of each layer matches the collapse type, each layer can realize the recognition of collapse types of different sizes. For example, road surface fractures are identified through one layer, and road surface collapses are identified through another layer. In this way, multiple types of road collapse problems can be identified at the same time, which improves the recognition efficiency and has a high recognition accuracy.
[0114] Next, the process of performing multi-layer feature extraction on spatial data by using a feature extraction model in S301 to obtain a multi-layer target feature map will be described.
[0115] refer to Figure 4 The process may include but is not limited to the following S401 to S404.
[0116] S401. Perform multi-scale convolution processing on the spatial data through a fully convolutional network, and splice the outputs of multiple stages to obtain a multi-layer first feature map.
[0117] The resolution of the first feature map of each layer in the multi-layer first feature map matches the collapse type.
[0118] S401 can be implemented as follows: configuring the convolution scale of each layer, then sequentially performing convolution processing on the spatial data of each layer at the corresponding scale through a fully convolutional network, and concatenating the outputs of multiple stages to obtain a multi-layer first feature map. The first feature map can be a matrix feature map. The concatenation of multiple layers, i.e., the concatenation of multiple stages, corresponds to the concatenation of matrices.
[0119] S402: Perform detail feature extraction on the first feature map of each layer in the multi-layer first feature map based on the attention mechanism to obtain a multi-layer second feature map.
[0120] S402 can be implemented as follows: for each layer's output first feature map, configure an attention module respectively, perform detail feature extraction on the first feature map of the layer through the attention module, traverse the first feature map of each layer, obtain a second feature map of the corresponding layer based on the first feature map of each layer, and splice the second feature maps of multiple layers to obtain the second feature maps of multiple layers.
[0121] S403: Perform fusion processing on the second feature maps of the multiple layers to obtain the third feature maps of the multiple layers.
[0122] S403 may be implemented as follows: fusing the multi-layer second feature maps in the order of the first layer to the last layer and / or the last layer to the first layer to obtain a multi-layer third feature map. The fusion may be performed once or multiple times.
[0123] S404: Determine that the target feature map of the multiple layers includes the first feature map of the multiple layers, the second feature map of the multiple layers, or the third feature map of the multiple layers.
[0124] S404 can be implemented as follows: determining, according to actual needs, that the multi-layer target feature map includes a multi-layer first feature map, a multi-layer second feature map, or a multi-layer third feature map.
[0125] Based on the above technical means, when determining the target feature map, the target feature map can be determined to be a multi-layer first feature map output by a fully convolutional network. The resolution of the multi-layer first feature map matches the collapse type, which can meet the layer identification requirements. The target feature map is determined by the fully convolutional network, which has the characteristics of simple implementation logic and high processing efficiency. When determining the target feature map, the target feature map can be determined to be a multi-layer second feature map. The resolution of the multi-layer second feature map matches the collapse type, which can meet the layer identification requirements. The multi-layer second feature map can also extract detail features based on the attention mechanism. The target feature map obtained in this way includes detail features, the obtained features are more comprehensive, and the collapse identification is more accurate. When determining the target feature map, the target feature map can be determined to be a multi-layer third feature map. The resolution of the multi-layer third feature map matches the collapse type, which can meet the layer identification requirements. The multi-layer third feature map is obtained by fusing the second feature map. Therefore, it not only has the detail features in the second feature map, but also the fusion of each layer allows each layer to retain the features of other layers. The target feature map obtained in this way is more comprehensive and the collapse identification is more accurate.
[0126] Next, the process of performing multi-scale convolution processing on spatial data through a fully convolutional network in S401 and splicing the outputs of multiple stages to obtain a multi-layer first feature map is described.
[0127] In the case of a multi-layer comprising three layers, reference Figure 5The process may include but is not limited to the following S501 to S504. For other numbers of layers, the three-layer processing process can be referred to and will not be described in detail here.
[0128] S501: Perform convolution processing on spatial data based on a first downsampling rate to obtain a first feature map of a first layer.
[0129] Among them, the first downsampling rate is smaller than the second downsampling rate, and the second downsampling rate is smaller than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.
[0130] The embodiment of the present application does not limit the value of the first downsampling ratio and can be configured according to actual needs. For example, assuming the input image size is L×L×3, the first downsampling ratio can be 2, and the size of the first feature map of the first layer can be L / 2×L / 2×16.
[0131] S501 can be implemented as follows: inputting the spatial data into a fully convolutional network, performing convolution processing on the spatial data based on a first downsampling rate, and obtaining a first feature map of the first layer.
[0132] S502: Perform convolution processing on the spatial data based on the second downsampling rate to obtain a first feature map of the second layer.
[0133] The embodiment of the present application does not limit the value of the second downsampling ratio and can be configured according to actual needs. For example, assuming the input image size is L×L×3, the second downsampling ratio can be 4, and the size of the first feature map of the second layer can be L / 4×L / 4×24.
[0134] The implementation of S502 can refer to the description of performing convolution processing on the spatial data based on the first downsampling rate in S501 to obtain the first feature map of the first layer, which will not be repeated here.
[0135] S503 : Perform convolution processing on the spatial data based on the third downsampling rate to obtain a first feature map of the third layer.
[0136] The embodiment of the present application does not limit the value of the third downsampling ratio and can be configured according to actual needs. For example, assuming the input image size is L×L×3, the third downsampling ratio can be 8, and the size of the first feature map of the third layer can be L / 8×L / 8×40.
[0137] The implementation of S503 may refer to the description of performing convolution processing on the spatial data based on the first downsampling rate in S501 to obtain the first feature map of the first layer, which will not be described in detail here.
[0138] S504: Determine multiple layers of first feature maps based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer.
[0139] S504 performs matrix splicing on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, thereby obtaining multiple layers of first feature maps.
[0140] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, each of which corresponds to a different receptive field, second feature maps of different layers can be obtained. This is characterized by simple implementation, reliability, and accuracy.
[0141] Next, the process of performing detail feature extraction on the first feature map of each layer in the multi-layer first feature map based on the attention mechanism in S402 to obtain the multi-layer second feature map is described.
[0142] In the case where the first feature map of the multiple layers includes the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, reference Figure 6 The process may include but is not limited to the following S601 to S604.
[0143] S601. Perform detail feature extraction on a first feature map of a first layer based on a first receptive field range through a first attention module to obtain a second feature map of the first layer.
[0144] The first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.
[0145] The embodiment of the present application does not limit the size of the first receptive field, and can be configured according to actual needs. For example, the first receptive field can be 3×3 in size.
[0146] S601 can be implemented as follows: inputting the first image feature of the first layer into the first attention module, the first attention module is configured with the size of the first receptive field range, and the first feature map of the first layer is progressively traversed in sequence based on the first receptive field range by the first attention module, extracting the detail features within the receptive field each time, and then moving the receptive field backward or downward, step by step, and then extracting the detail features, thereby obtaining the second feature map of the first layer.
[0147] S602. Perform detail feature extraction on the first feature map of the second layer based on the second receptive field range through the second attention module to obtain a second feature map of the second layer.
[0148] The embodiment of the present application does not limit the size of the second receptive field, and can be configured according to actual needs. For example, the second receptive field can be 7×7 in size.
[0149] The implementation of S602 can refer to S601 in which the first attention module extracts detail features from the first feature map of the first layer based on the first receptive field range to obtain a detailed description of the second feature map of the first layer, which will not be repeated here.
[0150] S603. Perform detail feature extraction on the first feature map of the third layer based on the third receptive field range through the third attention module to obtain a second feature map of the third layer.
[0151] The embodiment of the present application does not limit the size of the third receptive field, and can be configured according to actual needs. For example, the third receptive field can be 15×15 in size.
[0152] The implementation of S603 can refer to S601 in which the first attention module extracts detail features from the first feature map of the first layer based on the first receptive field range to obtain a detailed description of the second feature map of the first layer, which will not be repeated here.
[0153] S604: Determine multiple layers of second feature maps based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer.
[0154] S604 may be implemented as: performing matrix splicing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, thereby obtaining multiple layers of second feature maps.
[0155] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, each of which corresponds to a different receptive field, second feature maps of different layers can be obtained. This is characterized by simple implementation, reliability, and accuracy.
[0156] Next, the process of fusing the multi-layer second feature maps in S403 to obtain the multi-layer third feature maps is described. Figure 7 The process may include but is not limited to the following S701 and S702.
[0157] S701. Perform cascade convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to the first path to obtain intermediate feature maps of the three layers.
[0158] The first path may be from the first layer to the last layer. Of course, the first path may also be from the last layer to the first layer.
[0159] During the cascade convolution process, different weights can be configured for the second feature map of each layer. The weight configuration can be determined according to actual needs.
[0160] For example, the weight may be gradually increased or gradually decreased.
[0161] The intermediate feature map of the three layers refers to the feature map of the three layers after the first path cascade convolution.
[0162] S701 can be implemented as follows: perform cascade convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer in the order of the first path, and complete the cascade convolution of all layers to obtain the intermediate feature maps of the three layers.
[0163] S702: Perform cascade convolution processing on the intermediate feature maps of the three layers according to the second path to obtain a multi-layer third feature map.
[0164] The second path is different from the first path.
[0165] For example, in the case where the first path is from the first layer to the last layer, the second path may be from the last layer to the first layer.
[0166] The implementation of S702 can refer to S701, which performs cascade convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to the first path to obtain a description of the intermediate feature maps of the three layers. The difference is that in S702, the cascade convolution order needs to be adjusted from the first path to the second path.
[0167] Based on the above technical means, when fusing the second feature maps of multiple layers, the fusion is first performed based on the first path, and then based on the second path. Through the fusion of the two paths, the target feature map of each layer after fusion can retain the features of all upper and lower layers. The feature map of each layer in the obtained target feature map has comprehensive information, which improves the accuracy of identifying the collapse type of each layer.
[0168] Next, the process of classifying the target features corresponding to the predicted detection boxes of each layer using the classification model in S303 to obtain the classification results of each layer will be described.
[0169] The processing process of the predicted detection frame of each layer is similar. The following is an example of the processing of the predicted detection frame of one layer. The predicted detection frames of other layers can refer to the implementation process of this layer, which will not be repeated here.
[0170] refer to Figure 8 The process may include but is not limited to the following S801 to S805.
[0171] S801: Obtain the target reference frame of each layer.
[0172] In advance, a target reference frame is set for each layer. S801 can be implemented as: directly reading the set target reference frame for each layer. The target reference frame for each layer is different.
[0173] S802: Determine the matching degree between the predicted detection box and the target reference box of each layer respectively to obtain multiple matching degrees.
[0174] The matching degree is used to represent the degree of overlap between the predicted detection box and the target reference box. The greater the overlap, the higher the matching degree. In other words, the more similar they are, the higher the matching degree.
[0175] The embodiment of the present application does not limit the method for determining the matching degree, and can be configured according to actual needs. For example, the matching degree can be determined based on the intersection-and-union ratio of the areas of two boxes, or based on the coordinates of the areas.
[0176] S802 can be implemented as follows: for the predicted detection frame, based on the matching degree determination algorithm, the matching degree between the predicted detection frame and the target reference frame of each layer is determined respectively, and the target reference frames of all layers are traversed, so that the matching degree between the target reference frames of all layers can be obtained, that is, multiple matching degrees.
[0177] S803: Determine the collapse type based on multiple matching degrees.
[0178] S803 may be implemented as follows: determining the collapse type corresponding to the target reference frame with the highest matching degree as the collapse type of the layer.
[0179] S804: Determine the collapse position based on the predicted position of the detection frame.
[0180] S804 may be implemented as: reading the position of the predicted detection frame, and determining the position of the predicted detection frame as the determined collapse position.
[0181] S805: Determine the classification result of each layer including the collapse type and collapse location.
[0182] S805 can be implemented as follows: determining the classification results of each layer, including the collapse type and collapse location corresponding to the predicted detection frame of the layer identified. Here, each layer may include one or more or no predicted detection frames, corresponding to one, multiple, or no collapse of a certain type. However, all layers have at least one predicted detection frame. If all layers do not have predicted detection frames, it is considered that there is no potential collapse, and the following processing is urgently performed to continue detecting spatial data through sensors and continue monitoring.
[0183] Based on the above technical means, when determining the collapse type, it is determined based on the matching degree between the predicted detection frame and the target reference frame of each layer. In this way, since the target reference frame of a layer can represent the characteristics of a type of collapse, the collapse type is determined more accurately based on the matching degree. When determining the collapse location, since the predicted detection frame is used to indicate the collapse location, it is more accurate to determine the collapse location based on the position of the predicted detection frame.
[0184] The processing method provided in the embodiment of the present application may include but is not limited to the process of determining the target reference frame of each layer.
[0185] refer to Figure 9 The process may include but is not limited to the following S901 to S904.
[0186] S901: Obtain a sample data set.
[0187] The sample data set includes multiple sample space data, each of which is marked with a true box; the true box is used to point to the collapse position in the sample space data.
[0188] The sample space data can also be labeled with the collapse type. The labeling here can be manual labeling, or reliable labeling that has been manually verified after machine labeling.
[0189] The embodiment of the present application does not limit the amount of sample space data included in the sample data set, and can be configured according to actual needs. The larger the amount of sample space data, that is, the richer the sample data set, the more accurate the result obtained.
[0190] After labeling a plurality of sample space data in the sample data set, S901 may be implemented as reading each sample space data in the labeled sample data set.
[0191] S902 : Distribute multiple ground-truth boxes in the sample data set into multiple layers according to their sizes, so that each layer includes multiple ground-truth boxes.
[0192] For example, a fracture type collapse is configured for the first layer, a horizontal collapse type collapse is configured for the second layer, and a vertical collapse type collapse is configured for the third layer. Each layer has a size range.
[0193] In one possible implementation, S902 may be implemented as follows: determining a collapse type corresponding to each layer and a size range corresponding to the collapse type, and then allocating multiple ground truth boxes in the sample dataset to multiple layers according to the size range, so that each layer includes multiple ground truth boxes.
[0194] Of course, multiple ground truth boxes can also be assigned to different layers according to the collapse type.
[0195] S903 : For each layer in the multiple layers, determine the center of the target reference frame of the layer based on multiple real frames included in the layer.
[0196] The embodiment of the present application does not limit the method for determining the center of the target reference frame, and can be configured according to actual needs. For example, the average value or the median position of the centers of all real frames can be determined as the center of the target reference frame.
[0197] S903 may be implemented as follows: for each of the multiple layers, based on the multiple real frames included in the layer, calculating the centers of the multiple real frames by a center determination algorithm to obtain the center of the target reference frame of the layer.
[0198] S904: Determine the target reference frame of the layer based on the center of the target reference frame of the layer.
[0199] In a possible implementation, S904 may be implemented as: generating a target reference frame for the layer based on the center of the target reference frame of the layer.
[0200] In another possible implementation, S904 may be implemented as follows: generating a candidate reference frame for the layer based on the center of the target reference frame of the layer, and then continuously optimizing and adjusting the target reference frame measured for the layer.
[0201] Based on the above technical means, the sample space data in the sample dataset can be divided into different layers according to the size of the ground truth box. The center of the target reference box of each layer is determined based on the size of the ground truth box of each layer, and then the target reference box of each layer is determined. In this way, the target reference box determined for each layer can meet the size requirements of a certain type of collapse, improving the recognition accuracy of that layer for that collapse type.
[0202] Next, the process of determining the center of the target reference frame of the layer based on the multiple real frames included in the layer in S903 is described. Figure 10 The process may include but is not limited to the following S1001 to S1003.
[0203] S1001: Determine the center points of multiple real boxes based on the multiple real boxes of the layer.
[0204] S1001 can be implemented as follows: based on multiple real frames of the layer, first determine the coordinates of each real frame, and then based on the real coordinates of each real frame, determine the collective center of each real frame through a collective center point determination algorithm, thereby obtaining the coordinates of the center point of each real frame.
[0205] S1002: Determine the probability density of the center point of each ground truth box of the layer.
[0206] The concept density is used to cluster the ground truth boxes of this layer.
[0207] S1002 can be implemented as follows: for each real frame of the layer, based on the coordinates of each feature point in the real frame, determine the probability density of the center point of each real frame in the real frame. This embodiment of the application does not limit the method for determining the probability density, and can be configured according to actual needs.
[0208] S1003: Determine the center point with the highest probability density as the center of the target reference frame.
[0209] S1003 can be implemented as follows: the center point with the highest probability density is determined as the center of the target reference frame. If there are multiple points with the highest probability density, one of them is selected as the center of the target reference frame. If the first multiple probability density values are all high and relatively close, the center points of multiple high probability densities can also be used as the center of the target reference frame.
[0210] Based on the above technical means, when determining the center of the target reference frame of a layer, it is determined based on the probability density method, which can meet the needs of most real frames and is more accurate.
[0211] Next, the process of determining the target reference frame of the layer based on the center of the target reference frame of the layer in S904 is described. Figure 11 The process may include but is not limited to the following S1101 to S1104.
[0212] S1101. Generate a candidate reference frame with the center point of the target reference frame as the center.
[0213] S1101 can be implemented as follows: generating a candidate reference frame with the center point of the target reference frame as the center. The candidate reference frame generated here can be determined based on any real frame. For example, in the case of a fracture corresponding to the first layer, when the center point is determined, a long strip shape is generated with the center point as the center as the candidate reference frame.
[0214] Here, for each layer, the shapes of the generated candidate reference boxes can be the same or different. Even if the shapes of the candidate reference boxes for each layer are initially the same, the shapes of the final target reference boxes for each layer will be different with subsequent adjustments and optimizations.
[0215] S1102: Determine the dynamic distance between each ground truth box of the layer and the candidate reference box respectively.
[0216] The embodiment of the present application does not limit the method for determining the dynamic distance, and can be configured according to actual needs. For example, the dynamic distance is negatively correlated with the overlap, and the dynamic distance is positively correlated with the aspect ratio similarity.
[0217] S1102 may be implemented as follows: determining the dynamic distance between each real frame of the layer and the candidate reference frame according to a dynamic distance determination algorithm or formula, thereby obtaining multiple dynamic distances.
[0218] Then, it is determined whether the multiple dynamic distances meet the preset conditions.
[0219] S1103: Adjust the candidate reference frame based on the dynamic distance until the dynamic distance meets a preset condition.
[0220] The embodiments of the present application do not limit the preset conditions and can be configured according to actual needs.
[0221] For example, the preset conditions may include but are not limited to any of the following: the dynamic distance between each real frame is less than or equal to the first threshold; the dynamic distance between most real frames is less than or equal to the first threshold, that is, the number greater than or equal to the first threshold is greater than the first number threshold, that is, individual interference is allowed; or, the adjustment of all sample space data has been completed.
[0222] S1103 can be implemented as follows: based on the dynamic distance, the size and shape of the candidate reference frame are adjusted, and then the adjusted candidate reference frame is used as a new candidate reference frame to recalculate the dynamic distance, and this cycle is repeated until the dynamic distance meets the preset conditions.
[0223] S1104: Determine the candidate reference frame that meets the preset conditions as the target reference frame of the layer.
[0224] S1104 may be implemented as follows: when it is determined that the preset condition is satisfied, determining the candidate reference frame satisfying the preset condition as the target reference frame of the layer.
[0225] Based on the above technical means, when determining the target reference frame, a candidate reference frame is first generated based on the center point, and then the shape and size of the candidate reference frame are continuously adjusted based on the dynamic distance, so that the candidate reference frame and the real frame of the layer meet the preset conditions. In this way, the target reference frame can have a large overlap with the real frame and a close aspect ratio with the real frame, that is, the position and shape are close to the real frame, thereby meeting actual needs and improving accuracy.
[0226] The processing method provided in the embodiment of the present application can also update the target reference frame of each layer.
[0227] refer to Figure 12 The content shown, the process may include but is not limited to the following S1201 to S1203.
[0228] S1201 : For each ground-truth frame of a layer, determine the matching degree between the ground-truth frame and the target reference frame of the layer.
[0229] S1201 may be implemented as follows: for each real frame of the layer, determining the matching degree between the real frame and the target reference frame of the layer based on a matching degree algorithm, thereby obtaining multiple matching degrees.
[0230] S1202: Adjust the layer to which the real frame belongs based on the matching degree.
[0231] S1202 may be implemented as follows: determining the real frames whose matching degree is less than a threshold as the real frames to be adjusted, and then deleting the real frames to be adjusted from the layer. In this way, the real frames included in each layer will be redistributed and adjusted based on the matching degree.
[0232] S1203: Update the target reference frame of each layer based on the adjusted real frame included in each layer.
[0233] The process of updating the target reference frame of each layer is similar to the process of determining the target reference frame of each layer. The implementation of S1203 can refer to the description of S903 to S904, which will not be repeated here.
[0234] Based on the above technical means, after each ground-truth frame is stratified, the allocation method can be adjusted based on the matching degree. This can optimize the stratification, improve the accuracy of the segmentation, and also improve the accuracy of the collapse type identification.
[0235] Below, taking an autonomous driving vehicle as an example, the process of handling road collapse is explained through an embodiment.
[0236] With the increasing number of road construction projects, road safety hazards are becoming increasingly prominent. Pavement collapse is a critical and common problem, posing a significant threat to people's lives and property. In 2020, a serious road collapse occurred in a certain city, resulting in numerous casualties and drawing significant public attention. In 2024, a road collapse disaster occurred on a certain expressway, resulting in a particularly serious accident with 48 fatalities and over 100 million yuan in direct economic losses. Statistics show that the incidence of road collapse accidents in my country is increasing year by year, with an average annual growth rate of 81%. To effectively protect the lives and property of the people, it is necessary to analyze the causes of road collapse and implement preventive measures. Furthermore, it is necessary to identify and detect natural disasters such as road collapses through vehicle-side automated assisted driving, allowing for proactive intervention.
[0237] The mainstream target recognition and detection technology in the current autonomous driving field uses cameras, lidar, and sensors to obtain real-world views of the surrounding environment, and performs real-time target detection and analysis on these real-world views. These real-world views mainly include instances of people, vehicles, and objects for tracking, calibration, and decision-making. However, there is no detection and recognition method for special long-tail scenarios in the surrounding environment (such as disasters such as road collapse, road faults, and bridge deck fractures). From the current perspective, most of the detection methods and maintenance plans are centered around the analysis and prevention of the causes of traffic road collapse.
[0238] For example, related technology 1 provides a method and system for detecting abnormal conditions on a highway, which includes: constructing a detection management platform, setting up detection devices in sections along the highway, and connecting each detection device to the detection management platform to form a monitoring network; using the detection device to monitor whether there are abnormal conditions on the highway in the corresponding section; if an abnormal condition is found, issuing traffic warning information to each section through the monitoring network.
[0239] Related technology 2 provides a road collapse identification and detection method, which aims to provide a solution to the problem of real-time identification of underground collapse and reporting the detected road section at the fastest speed, so as to report hidden dangers in the first time and prevent them from happening.
[0240] Related Technology 3 provides a method and device for remotely controlling an autonomous vehicle operating without an in-vehicle driver. This method aims to detect an emergency and transmit emergency information to a remote control platform. The remote control platform then responds to the emergency information and executes the response measures. This method aims to address the issue of traditional remote systems and personnel being unable to respond to emergencies in long-tail scenarios, which can lead to operational risks for autonomous vehicles and hinder their development and application.
[0241] By detecting abnormal road conditions, such as emergencies, remote systems are used to provide feedback and control the control platform to intervene in the autonomous vehicle, without identifying and intervening within the autonomous vehicle itself. This shows that current technical solutions for predicting road collapses rely solely on reliability monitoring of the underlying road surface, without identifying and determining road safety reliability within the autonomous vehicle domain. Relying solely on reliability detection of road surface defects cannot fully protect against the potential harm to life and property caused by such traffic accidents.
[0242] This embodiment primarily utilizes vehicle-side autonomous driving object detection algorithms to identify and detect long-tail scenarios such as road surface collapse and provide decision-making. To ensure the safe implementation of autonomous driving in real-world traffic environments, continuous optimization of autonomous driving algorithms is crucial to ensure effective emergency response measures for the 10% of these scenarios. This can significantly reduce or mitigate personal and property damage and ensure vehicle safety.
[0243] To address this technical issue, the technical solution and working principle of this embodiment are as follows: A road collapse identification and detection solution based on an autonomous driving system is proposed, aiming to continuously optimize the autonomous driving algorithm to ensure effective emergency response measures in the 10% long-tail scenario.
[0244] To this end, the first purpose of this embodiment is to propose an algorithm for identifying and detecting road collapse based on an autonomous driving system, including but not limited to the following steps 1 to 3.
[0245] Step 1: Perform environmental perception based on the vehicle's visual perception equipment (cameras and lidar) to obtain environmental perception point cloud data and image data; collect road condition images of several different roads, including various road diseases (road collapse, road faults, block cracks, bridge deck fractures, and other disasters) to form a sample image set. Use the deep learning architecture PyTorch to fuse the image and point cloud data to establish an atlas that clearly defines the location of road diseases (equivalent to the above-mentioned collapse location) and type (equivalent to the above-mentioned collapse type) as a standard road disease dataset (equivalent to the above-mentioned sample dataset).
[0246] Step 2: Select the road surface condition images annotated in the sample annotation step as the training dataset for the network model. Set the loss function, optimizer, and hyperparameters to build a deep learning model based on the Mobile NetV3 convolutional neural network. Complete model training in Python 3.6 using the deep learning framework TensorFlow and the neural network library Keras.
[0247] Step 3: Perform quantitative feature extraction of road collapse on the images in the pavement disease segmentation image dataset and obtain the quantitative feature extraction results of the collapse area. The preprocessed data of the road collapse images are input into the trained deep learning model based on Mobile NetV3. After inference, the prediction output is obtained and the model performance is verified.
[0248] refer to Figure 13 As shown in the content, the road collapse processing process may include but is not limited to the following S1301 to S1315.
[0249] S1301. Collect road subsidence disease image data.
[0250] S1302: Establish a pavement disease image dataset.
[0251] S1303: Preprocess the pavement damage image data.
[0252] S1304. Set the network model loss function and hyperparameters.
[0253] S1305: Train a road collapse environment model based on MobileNetV3.
[0254] S1306: Radar and camera input road image dataset.
[0255] S1307. Apply a road subsidence environment model for identification.
[0256] S1308. Extraction of quantitative parameter features of pavement collapse disease.
[0257] S1309, using the classic crack disease classification algorithm.
[0258] S1310. Classify and save road subsidence disease information.
[0259] S1311. Road environment model based on autonomous driving.
[0260] S1312. Determine whether there is an emergency event (equivalent to a collapse event).
[0261] If so, that is, there is an emergency, execute the following S1313 to S1315; if not, that is, there is no emergency, execute the above S1307.
[0262] The sudden incident here is equivalent to a collapse event.
[0263] S1313. The vehicle makes a decision and executes corresponding measures.
[0264] S1314. Use vehicle-to-everything (V2X) wireless communication technology to provide warning information to nearby vehicles.
[0265] S1315, upload to the vehicle-road-cloud integrated system through V2X wireless communication technology.
[0266] The implementation of step 1 is described in detail below, including but not limited to steps 1.1 to 1.3.
[0267] Step 1.1: In order to facilitate the unified preprocessing of the acquired environmental perception point cloud data and image data, it is necessary to use the deep learning architecture pytorch to fuse the image and point cloud data to produce a richer dataset.
[0268] Step 1.2: Use the labeling software labelimg to label the images of the pavement disease standard dataset, obtain the range coordinates of the diseased area, and annotate the samples of the diseased area according to the classification category and segmentation label. Label the specific conditions such as the location and type of road diseases in the samples and divide them into training set and validation set.
[0269] Step 1.3: Preprocess the images in the training and validation sets. For the road condition images of different roads and various road damages included in the sample image set, perform cropping, geometric switching, and brightness / contrast / hue conversion. Add Gaussian noise and salt and pepper noise, resize the images, and adjust the pixel values.
[0270] The implementation of step 2 is described in detail below.
[0271] Step 2.1: Select the road surface condition images annotated in the sample annotation step as the training dataset for the network model. Set the loss function, optimizer, and hyperparameters to build a deep learning model based on the Mobile NetV3 convolutional neural network. In the Python 3.6 environment, first calculate the difference between the predicted value and the true label by defining the loss function. Then use backpropagation gradient descent to update the network model weights so that the network model output prediction value is close to the true label value. In this way, you can obtain an object detection model based on the deep learning framework TensorFlow and the neural network library Keras.
[0272] Step 2.1.1: Based on the U-shaped network architecture, use MobileNetV3-Large as the backbone feature network to extract features from the dataset images. The size of the feature maps at each stage is defined primarily based on the actual downsampling step. The downsampling rate refers to the reduction factor relative to the input image (e.g., A1 is 1 / 2 the input size). The output feature maps at different stages are denoted as A1 (equivalent to the first feature map of the first layer mentioned above), A2 (equivalent to the first feature map of the second layer mentioned above), A3 (equivalent to the first feature map of the third layer mentioned above), A4, and A5.
[0273] Assuming the input image size is L×L×3, the size of A1 is L / 2×L / 2×16. After multi-stage convolution and pooling, the size of A2 is L / 4×L / 4×24, the size of A3 is L / 8×L / 8×40, the size of A4 is L / 16×L / 16×112, and the size of A5 is L / 16×L / 16×160. The processing parameters of the convolution stage can be found in Table 1.
[0274] Table 1
[0275] Example of processing parameters for the convolution stage
[0276]
[0277] Next, for each feature map Ai, a dynamic multi-kernel collaborative receptive field boosting (DMKC-RFB) module (RFB) is designed. Taking Ai as input (i=1-5), DMKC-RFB is used to enhance the feature maps extracted in the previous step. The output enhanced feature maps are B1 (equivalent to the second feature map of the first layer), B2 (equivalent to the second feature map of the second layer), B3 (equivalent to the second feature map of the third layer), B4, and B5, with the same size as the input. Feature fusion is then used to fuse the enhanced feature maps into C1 (equivalent to the third feature map of the first layer), C2 (equivalent to the third feature map of the second layer), C3 (equivalent to the third feature map of the third layer), C4, and C5. Finally, the fused feature maps are used to predict the target.
[0278] Step 2.1.2: Feature Enhancement The receptive field uses the DMKC-RFB module to reduce the model calculation amount and improve the model feature extraction capability.
[0279] The feature enhancement process includes: inserting the DMKC-RFB module (equivalent to the above-mentioned attention mechanism) after each level Ai to generate enhanced feature maps B1, B2, B3, B4 and B5.
[0280] Input branch division: Each Ai is processed in 3 parallel branches, and the branch parameters are adaptively adjusted according to the size of the feature map: Branch 1: 3×3 deformable convolution (expansion rate 1); Branch 2: 3×3 deformable convolution (expansion rate 3); Branch 3: 5×5 deformable convolution (expansion rate 5).
[0281] In DMKC-RFB, the design of three parallel branches aims to adaptively capture multi-scale feature information through deformable convolutions of different scales while balancing computational efficiency and model performance.
[0282] The parameters of each branch in DMKC-RFB can be found in Table 2.
[0283] Table 2
[0284] Parameter examples for each branch in DMKC-RFB
[0285]
[0286] The collaborative process can include: Input feature map partitioning: The feature map (Ai) of the same level is input into three branches for parallel processing. Multi-scale feature extraction: Branch 1 focuses on local details with a small receptive field; Branch 2 captures surrounding context with a medium receptive field; Branch 3 covers the global area with a large receptive field.
[0287] Dynamic weighted fusion and residual connections: A lightweight attention module (Squeeze-and-Excitation, SE) dynamically assigns weights based on channel importance. The fusion result is added to the input feature map, preserving the original information.
[0288] The weight of each branch is dynamically adjusted through SE. Please refer to the following formula (1).
[0289] Formula (1);
[0290] In formula (1), The channel attention weights generated by the SE module, Represents the features of each layer after convolution processing, Represents the enhanced features of each layer; Represents a variable convolution operation.
[0291] Step 2.1.3: Feature fusion adopts a multi-scale feature fusion expression, based on a bidirectional feature pyramid, and adds a cross-layer weight adaptation mechanism. The fusion of the top-down path (equivalent to the first path mentioned above) can refer to the following formulas (2-1) to (2-5), and the fusion of the bottom-up path (equivalent to the second path mentioned above) can refer to the following formulas (3-1) to (3-5).
[0292] Formula (2-1);
[0293] Formula (2-2);
[0294] Formula (2-3);
[0295] Formula (2-4);
[0296] Formula (2-5);
[0297] In formulas (2-1) to (2-5), 、 、 、 Fusion weights, 、 、 、 Initialization can be 1.0. 、 、 、 、 are the fusion features after fusion of each two layers (equivalent to the above intermediate feature map). : Bilinear interpolation upsampling by a factor of 2. Indicates size 1 1 convolution operation.
[0298] Formula (3-1);
[0299] Formula (3-2);
[0300] Formula (3-2);
[0301] Formula (3-4);
[0302] Formula (3-5);
[0303] In formulas (3-1) to (3-5), 、 、 、 Fusion weights, 、 、 、 Initialization can be 1.0. 、 、 、 、 are the fusion features after fusion of each two layers. : Max pooling downsamples by a factor of two. Indicates size 3 3 convolution operations.
[0304] in, , The weight parameters are learnable and automatically updated through back-propagation without manual design, and are used to dynamically balance the importance of features at different resolutions.
[0305] In the object detection model, C1 to C5 are multi-scale fused feature maps, corresponding to different downsampling ratios (e.g., C1 downsampling ratio 2, C5 downsampling ratio 16). The layered density-aware K-means++ algorithm is used to generate an adaptive anchor box size for each level.
[0306] 1. Target layer allocation (preprocessing stage).
[0307] Function: Assign the real target boxes in the dataset to the feature maps of the corresponding levels (C1-C5) according to their sizes.
[0308] For example:
[0309] Small objects (such as cracks, size 20×20 pixels): matched to C1 (downsampling rate 2, original image anchor box size 40×40).
[0310] Large objects (e.g. collapsed, 160×160 pixels): matched to C5 (downsampling rate 16, original image anchor box size 160×160).
[0311] 2. Run K-means++ independently layer by layer.
[0312] The assigned target boxes of each level (C1-C5) are clustered independently to generate the anchor box size of that level.
[0313] 1. Input data preparation: The target box set of level i: only contains the ground-truth target boxes assigned to this level. For example, the C3 layer processes all medium-sized target boxes matched to this level.
[0314] According to the matching degree between the target box size and the anchor box size of each layer, the target is assigned to the corresponding layer.
[0315] The matching degree can refer to the following formula (4).
[0316] P= Formula (4);
[0317] In formula (4), P represents the matching degree, Indicates the predicted anchor box size (width × height). Indicates the size (width × height) of the true target box in the dataset.
[0318] IoU stands for Intersection over Union (IoU), which measures the degree of overlap between two boxes. max(Sanchor,Starget) represents the maximum area between the anchor box and the target box (used for normalization).
[0319] 2. Predicted Box
[0320] Definition: The target bounding box output by the model indicates the location and size of the target detected by the algorithm.
[0321] Generation process: Based on the anchor box: the model (V3) generates a prediction box through the anchor box (Anchor), and the coordinates of the prediction box are the offset adjustment results of the anchor box. Regression parameters: The model predicts the center point offset and width and height scaling factors .
[0322] For example, if the anchor box size is 64x64 and the model predicts the offset (Δx = 0.1, Δy = 0.2, Δw = 1.2, Δh = 0.8), the predicted box size is 64x1.2 = 76.8 (width) and 64x0.8 = 51.2 (height).
[0323] The receptive field (RF) is the size of the area of the input image corresponding to a point on the feature map, indicating the field of view of the input image that the point can "see".
[0324] The parameters corresponding to the receptive fields of each layer can be found in Table 3.
[0325] Table 3
[0326] Example of parameters corresponding to the receptive field of each layer
[0327]
[0328] Among them, the low-level layers (such as C1) have a small receptive field and are suitable for detecting details (such as crack edges), while the high-level layers (such as C5) have a large receptive field and are suitable for integrating global information (such as collapsed areas).
[0329] Calculate the density map of the target distribution and preferentially select the center point of the high-density area as the initial cluster center.
[0330] The role of the density map is to quantify the distribution density of objects in different areas of the dataset, guide the anchor box generation process to focus on high-density areas, and improve the adaptability of the anchor box to dense objects.
[0331] Layered processing: Each layer (C1-C5) independently computes its corresponding density map, considering only the object boxes assigned to that layer. For example, layer C1 (downsampling rate 2) processes small objects (such as cracks), and its density map reflects the dense areas of cracks in the image. Layer C5 (downsampling rate 16) processes large objects (such as collapses), and its density map reflects the distribution of large-scale collapses.
[0332] The calculation of the hierarchical density map can include: Input data: the center coordinates of all target boxes assigned to this level. Hierarchical adaptability: Small target layer (C1): The target distribution may be concentrated in local areas (such as dense cracks on the edge of the road). Large target layer (C5): The target distribution may be more dispersed or cover a wide area (such as a collapsed area). Parameter adjustment: Density maps of different levels use different Gaussian kernel bandwidths σ to adapt to the target scale: Small target layer (C1): σ is small (such as σ=2) to capture local density. Large target layer (C5): σ is large (such as σ=8) to smooth the global distribution. Hierarchical division of labor: C1-C5 are responsible for the detection of targets of different scales. The density map ensures that the anchor box generation of each layer adapts to the distribution characteristics of the targets in that layer.
[0333] The density map guides cluster initialization, making the anchor boxes more consistent with the target shape and size in the actual dense area.
[0334] The necessity of layered processing: Feature maps at different levels have different receptive fields and need to match targets of different scales. For example, the high resolution of C1 is suitable for detecting fine cracks, while the low resolution of C5 is suitable for detecting large-scale collapses.
[0335] Advantages of dynamic distance metrics: It simultaneously optimizes position overlap (IoU) and shape similarity (aspect ratio), avoiding the scale sensitivity of traditional Euclidean distance. Value of density-sensitive initialization: It generates anchor boxes for densely populated areas, improving the model's recall for small or densely populated objects.
[0336] The dataset contains 1,000 road surface images annotated with three types of objects: cracks, potholes, and collapses. Objects are assigned layer by layer: Cracks (20×5 pixels) → Layer C1 (s=2). Potholes (80×80 pixels) → Layer C3 (s=8). Collapses (160×160 pixels) → Layer C5 (s=16). Layer-by-layer clustering: Layer C1: K-means++ is run on cracks to generate elongated anchor boxes (e.g., 24×6). Layer C3: K-means++ is run on potholes to generate square anchor boxes (e.g., 64×64). Layer C5: K-means++ is run on collapses to generate large anchor boxes (e.g., 160×160).
[0337] Anchor box application: During training, the model performs target regression and classification based on anchor boxes at each level. During inference, the anchor boxes and predicted boxes calculate the Intersection over Union (IoU) and output the detection results.
[0338] Dynamic distance metric: Dynamic distance metric calculates the distance between the ground truth box and the candidate anchor box (predicted box), used to measure the degree of match between the two. Ground truth box: The manually annotated object bounding box in the dataset, including location (center point coordinates) and size (width and height). Candidate anchor box: The anchor box to be optimized generated during the clustering process, its size is iteratively adjusted to match the distribution of the ground truth box.
[0339] The determination of the dynamic distance can refer to the following formula (5).
[0340] Formula (5);
[0341] In formula (5), represents the dynamic distance between b and c, γ is an adjustable parameter that controls the aspect ratio weight; b represents the real target box (width and height are ); c represents the candidate anchor box (width and height are ).
[0342] The dynamic distance formula optimizes anchor box generation by combining two metrics: The Intersection over Union (IoU) term measures the degree of overlap between the ground-truth box and the candidate anchor box. A larger IoU indicates a smaller distance. This ensures that the anchor box covers the ground-truth object area as much as possible. The aspect ratio similarity term measures shape similarity by using the arctangent difference of the aspect ratio, avoiding reliance solely on area. This optimizes the anchor box shape to more closely match the object's aspect ratio distribution (e.g., elongated cracks or square holes).
[0343] In the hierarchical density-aware K-means++ algorithm, dynamic distance measurement is the core step in the clustering process. Its logic includes the following: Pre-processing step: Target layer assignment: assigning ground-truth target boxes to corresponding layers (C1-C5) based on size. Density-sensitive initialization: selecting initial cluster centers based on the density map, focusing on high-density areas.
[0344] Dynamic distance measurement step: Input: The ground-truth target box assigned to this level + the initial cluster center (candidate anchor box). Calculation: For each ground-truth target box, calculate its dynamic distance with all candidate anchor boxes. Assignment: Assign the ground-truth target box to the cluster center with the smallest distance.
[0345] Subsequent steps: Update cluster centers: Adjust the anchor box size based on the mean and shape of the target boxes within the cluster. Iterative optimization: Repeat distance calculation, allocation, and update until convergence. Anchor box restoration: Restore the clustering results to the original image scale using the hierarchical downsampling ratio. Match the receptive field of the feature map.
[0346] Scenario: Detect cracks (small objects) and collapses (large objects) in images. Target layer assignment: cracks (20×5 pixels) are assigned to layer C1 (downsampling rate 2), and collapses (160×160 pixels) are assigned to layer C5 (downsampling rate 16).
[0347] Hierarchical clustering output: C1 layer: generates anchor boxes (16×4, 24×6) using dynamic distance metric, with the original image scale being (32×8, 48×12). C5 layer: generates anchor boxes (160×160, 192×192).
[0348] Detection process: C1 layer prediction: Based on the 32×8 anchor box, regress the crack position and output the predicted box (30, 80, 38, 88). C5 layer prediction: Based on the 160×160 anchor box, regress the collapse position and output the predicted box (200, 300, 360, 460).
[0349] Set the loss function to cross entropy loss function.
[0350] The cross entropy loss function can refer to the following formula (6).
[0351] Formula (6);
[0352] In formula (6), represents the cross entropy, is the input vector sample, is the true label of the sample, N represents the number of samples, is the network weight parameter, by adjusting the weight parameter , the loss function can be minimized, thereby improving the recognition accuracy of the model.
[0353] Set the optimizer and hyperparameters: Use the Adam optimizer, set the batch size to 4, the learning rate to 1e-3, the number of iterations to 300, the adjustment factor to 0.75, and the default values for weight decay and initial momentum.
[0354] During model training, the training set is first used as input, and average pooling operations and regularization techniques are introduced. The deep learning model is trained according to the hyperparameter settings. The label tensor and the network output tensor are used to calculate the training error according to the loss function. The optimizer uses the training error to update the model parameters and verifies the recognition effect of the model on the validation set. When the model performance reaches a stable state, that is, when the loss function value and the recognition accuracy of the validation set tend to converge, the MobileNetV3 model parameters that perform best on the validation set are saved.
[0355] The implementation process of step 3 is described below.
[0356] The model uses an encoder-decoder structure (e.g., MobileNetV3 as the encoder and UNet as the decoder) to generate pixel-level segmentation masks. Specifically, the model takes the following form: A probability map outputs a two-dimensional matrix (e.g., H×W) of the same size as the input image, with each pixel value representing the probability of that location being a collapsed region (in the range [0, 1]). By setting a threshold (e.g., 0.5), the probability map is converted into a binary marker for collapsed regions. 1 (white pixel): collapsed region. 0 (black pixel): non-collapsed region.
[0357] Input: Preprocessed road surface image (sized to fit the model input, e.g., 224×224×3). Model inference: Output probability map (224×224×1). Post-processing: Generate a binary mask using thresholding to extract the collapsed areas. Performance verification: Calculate metrics such as Intersection over Union (IoU) (e.g., IoU = 0.85). Feature extraction: Quantify parameters such as the collapsed area, location, and shape (e.g., area = 1500 pixels, center of mass = (112, 89)).
[0358] The connected domain separation function is used to extract the connected domains in each image from the predicted output results of road collapse. The crack disease identification results with a connected domain smaller than a certain number of pixels are removed. The positional relationship of the connected domains of road collapse diseases is determined, and the bounding box of the collapsed area is annotated.
[0359] The classic crack disease classification algorithm is used to classify the pavement collapse diseases in the processed images into longitudinal pavement collapse, transverse pavement collapse, and pavement cracks. Geometric information quantitative evaluation is performed for different types of pavement collapse diseases. The geometric information quantitative indicators mainly include:
[0360] mAP@0.5 represents the mean average precision (mAP) when the IoU threshold is 0.5, which comprehensively reflects the detection accuracy.
[0361] Recall@0.5 represents the recall rate of collapsed targets and measures missed detections.
[0362] FPS represents the inference speed of the model on the deployed hardware (such as NVIDIA Jetson AGX Xavier).
[0363] Finally, a road environment model based on autonomous driving is obtained. When the vehicle is in motion, it processes the environmental perception data and determines whether there is an emergency of road collapse based on the detection and identification of the road collapse environment model. When it is determined that an emergency occurs, the autonomous driving algorithm will make a decision and execute corresponding measures, including emergency braking, lane changing, and defensive driving.
[0364] To improve the inference efficiency of the road collapse target detection and recognition network on FPGAs, a fully convolutional network is adopted. This means that the network model does not include fully connected layers. This significantly reduces the number of network model parameters, further reducing data exchange. Furthermore, a convolutional attention module is added after each convolutional layer to enhance feature extraction capabilities and compress unimportant feature information, enabling faster target identification and effectively improving target recognition accuracy.
[0365] A second objective of this embodiment is to provide a field-programmable gate array (FPGA) chip system and an early warning mechanism based on an autonomous driving system's road collapse recognition and detection algorithm, including the following steps:
[0366] An FPGA chip system for detecting road collapse in an autonomous driving system utilizes a top-down modular design based on an FPGA platform, using the hardware description language (Verilog HDL) for logic circuit design. After the top-level module definition is completed, the logical function definitions for each submodule are further broken down.
[0367] refer to Figure 14 As shown, the FPGA hardware chip system includes an image acquisition module 1401 , a data communication module 1402 , a cache module 1403 , a display control module 1404 , and an FPGA hardware accelerator 1405 .
[0368] The logical relationship between each module is as follows: after the system is powered on and reset, the FPGA uses the register lookup table and configures the internal registers of the image acquisition module; the on-board laser radar and camera output the video image data to the FPGA through clock signals and control signals; the central processing unit (CPU) receives and parses the trained network weight parameters, and then sends them to the FPGA through the data communication module; then the communication data packet is parsed. Since the weight parameters need to be sorted and spliced before the convolution operation, it is convenient to transmit and calculate the custom bit width data, so the parsed weight parameters are cached in the cache center to avoid interaction between the FPGA and external storage, reduce data transmission time, and thus improve data calculation efficiency; when all network parameters are transmitted, the frame read and write control module begins to cache the video images in sequence to the double data rate (DDR) 3, a ping-pong operation is implemented in the DDR3 cache to improve read and write efficiency. Video images read from the DDR3 are displayed in real time on the display screen via the display control module. Simultaneously, the video images are transmitted to the FPGA hardware accelerator. Because the network model's input feature maps need to be compressed, the video images are preprocessed and temporarily stored in the Block Random Access Memory (BRAM) cache. The FPGA hardware accelerator reads the image data and network parameters to extract features for road collapse identification, marking the detection results in real time on the display screen.
[0369] refer to Figure 15 As shown in the figure, the FPGA hardware accelerator includes a main control module 1501, a data cache module 1502, a convolution calculation module 1503, a zero filling module 1504 and an average pooling module 1505.
[0370] The main control module includes an enable signal generator and an address controller, which is responsible for re-sorting and splicing the unsorted and unspliced weight parameters and bias parameters of the BRAM cache center according to the network calculation process, and then writing them into other corresponding BRAM data cache modules to realize the control of the enable signals and address signals of each module.
[0371] The data cache module is used to cache input image data, intermediate feature map data, weight parameters, and bias parameter data. When the number of input feature map channels N in the input image BRAM block is greater than 8, the main control module reads image data and parameters from the intermediate feature map BRAM block, weight parameter BRAM block, and bias parameter BRAM block in batches according to the progress of depthwise convolution.
[0372] The convolution computation module primarily extracts image features through convolution operations. It includes a general-purpose convolution computation engine, a standard convolution controller, a depthwise convolution controller, and a pointwise convolution controller. The general-purpose convolution computation engine is applicable to all convolutional layers in the network model, utilizing the 48 resources of 24 digital signal processors (DSPs) for convolution operations. A three-stage pipeline based on shift registers completes the multiplication and accumulation of 24 input data points within a single clock cycle. For standard convolution, an 8-channel 5×5 convolution window can be calculated in three clock cycles, outputting an 8-channel result. As the network progresses, the number of input and output channels for depthwise and pointwise convolutions increases. By reusing this convolution computation engine until all input feature maps within a layer are calculated, resource utilization is greatly improved and system power consumption is reduced.
[0373] The zero-padding module uses the timing of the input image to perform padding operations. Based on the enable signal of the read-in image, zero padding is performed on the edges of the input feature map. Because the padded pixel values need to be multiplied by the weight parameters in the convolution kernel, to reduce data interaction time and computational latency, only the weight parameters of some edge data are multiplied by the pixel value 0 during the convolution calculation, further improving network computation speed.
[0374] The average pooling module calculates the average value of the local area after all convolutional layers are calculated and replaces it to reduce the amount of network calculation.
[0375] After the initialization of the FPGA hardware accelerator is completed, the enable signal of the main control module is pulled high. First, the preprocessed video point cloud fusion image data is read from the cache module and entered into the convolution calculation module to perform standard convolution calculation, depth convolution calculation and point-by-point convolution operation. The feature map output each time is repeatedly written into the intermediate feature map BRAM block; finally, after the depth-separable convolution operation is completed, the result is output through the average pooling layer.
[0376] This embodiment, an FPGA chip system and warning mechanism based on an autonomous driving system's road collapse detection algorithm, has at least the following beneficial effects: When the recognition algorithm identifies a road collapse, fault, or bridge deck fracture, the autonomous driving system will make decisions and implement corresponding measures, including assisting the driver with warnings, emergency braking, lane changes, and defensive driving. In emergencies such as road collapses and bridge deck fractures, the vehicle's geographic location and driving data generate vehicle warning information. This warning information is displayed to a second vehicle through vehicle-to-vehicle (V2X) communication technology and commercial high-precision maps. This solves the problem of rear-end drivers being unable to promptly understand the current vehicle's driver's driving intentions in emergency situations with poor visibility, such as at night or in rain, snow, frost, or fog. The rear-end driver can promptly understand the current vehicle's driver's driving intentions and be prompted to take emergency braking measures, thereby reducing rear-end collisions, scrapes, and the occurrence of further serious accidents. In the event of a road collapse emergency, information can be fed back to the highway remote control platform through integrated vehicle-road-cloud collaboration, enabling timely feedback and enabling highway maintenance personnel to take timely measures. This target recognition and detection algorithm further reduces data exchange by performing layer fusion design on normalization. Normalization and activation functions are added after each convolution operation to accelerate the convergence of the network model, greatly improving the parallel computing efficiency of the FPGA chip system.
[0377] In a second aspect, an embodiment of the present application provides a device for processing road collapse. The road collapse processing device can be deployed in vehicle equipment. Figure 16 As shown in the content, the road collapse processing device 160 may include but is not limited to: an acquisition unit 1601, an identification unit 1602 and an output unit 1603.
[0378] in:
[0379] An acquisition unit 1601 is used to acquire spatial data of the road ahead collected by the vehicle;
[0380] Identification unit 1602 is used to perform identification processing on spatial data through a road collapse identification model to determine an identification result of the road ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse on the road ahead; wherein the identification processing of the road collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain a multi-layer target feature map; matching the resolution of the target feature map of each layer in the multi-layer target feature map with the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame of each layer; matching the size of the predicted detection frame of each layer with the collapse type; performing classification processing on the target features corresponding to the predicted detection frame of each layer through a classification model to obtain a classification result of each layer; and determining an identification result of the road ahead based on the classification result of each layer.
[0381] Output unit 1603 is used to output the collapse type and collapse location in the identification result when the first identifier indicates that there is a collapse in the road ahead, so that the vehicle can respond based on the collapse type and collapse location; the collapse types include: road rupture, road lateral collapse, and road longitudinal collapse.
[0382] In some embodiments, the road collapse processing 160 may also include a processing unit, which is used to: determine the vehicle's own processing method based on the collapse type and collapse location; send a first prompt message, collapse type and collapse location to the vehicle behind the vehicle through vehicle network communication; the first prompt message is used to remind that there is a collapse ahead; send the collapse type and collapse location to the cloud.
[0383] In some embodiments, the identification unit 1602 is further configured to perform:
[0384] Multi-scale convolution processing is performed on spatial data through a fully convolutional network, and the outputs of multiple stages are spliced to obtain multi-layer first feature maps; the resolution of the first feature map of each layer in the multi-layer first feature map matches the collapse type; the detail features of the first feature map of each layer in the multi-layer first feature map are extracted based on the attention mechanism to obtain a multi-layer second feature map; the multi-layer second feature map is fused to obtain a multi-layer third feature map; it is determined that the multi-layer target feature map includes the multi-layer first feature map, the multi-layer second feature map or the multi-layer third feature map.
[0385] In some embodiments, when the multiple layers include three layers, the identification unit 1602 is further used to perform: convolution processing on the spatial data based on the first downsampling rate to obtain the first feature map of the first layer; convolution processing on the spatial data based on the second downsampling rate to obtain the first feature map of the second layer; convolution processing on the spatial data based on the third downsampling rate to obtain the first feature map of the third layer; determine the first feature maps of the multiple layers based on the first feature map of the first layer, the first feature map of the second layer and the first feature map of the third layer; wherein the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.
[0386] In some embodiments, when the multi-layer first feature map includes the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, the identifying unit 1602 is further configured to perform:
[0387] The first attention module is used to extract detail features from the first feature map of the first layer based on the first receptive field range to obtain the second feature map of the first layer; the second attention module is used to extract detail features from the first feature map of the second layer based on the second receptive field range to obtain the second feature map of the second layer; the third attention module is used to extract detail features from the first feature map of the third layer based on the third receptive field range to obtain the second feature map of the third layer; based on the second feature map of the first layer, the second feature map of the second layer and the second feature map of the third layer, the second feature maps of the multiple layers are determined; wherein the first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.
[0388] In some embodiments, when the second feature maps of the multiple layers include the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, the identifying unit 1602 is further configured to perform:
[0389] The second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer are subjected to cascade convolution processing according to the first path to obtain intermediate feature maps of the three layers; the intermediate feature maps of the three layers are subjected to cascade convolution processing according to the second path to obtain third feature maps of the multiple layers; the second path is different from the first path.
[0390] In some embodiments, the identification unit 1602 is further configured to perform:
[0391] Obtain a target reference frame for each layer; determine the degree of match between the predicted detection frame and the target reference frame of each layer respectively, and obtain multiple matching degrees; the matching degree is used to characterize the degree of overlap between the predicted detection frame and the target reference frame; determine the collapse type based on the multiple matching degrees; determine the collapse location based on the position of the predicted detection frame; determine the classification result of each layer including the collapse type and collapse location.
[0392] In some embodiments, the road collapse processing 160 may further include a pre-processing unit configured to perform:
[0393] Acquire a sample data set; the sample data set includes multiple sample space data, each sample space data is marked with a true box; the true box is used to point to the collapse position in the sample space data; the multiple true boxes in the sample data set are distributed to multiple layers according to size, so that each layer includes multiple true boxes; for each layer in the multiple layers, based on the multiple true boxes included in the layer, determine the center of the target reference box of the layer; based on the center of the target reference box of the layer, determine the target reference box of the layer.
[0394] In some embodiments, the pre-processing unit is further configured to perform:
[0395] Based on multiple true frames of the layer, the center points of the multiple true frames are determined; the probability density of the center point of each true frame of the layer is determined respectively; and the center point with the highest probability density is determined as the center of the target reference frame.
[0396] In some embodiments, the preprocessing unit is further used to perform: generating a candidate reference frame with the center point of the target reference frame as the center; determining the dynamic distance between each real frame of the layer and the candidate reference frame respectively; the dynamic distance is negatively correlated with the overlap, and the dynamic distance is positively correlated with the aspect ratio similarity; adjusting the candidate reference frame based on the dynamic distance until the dynamic distance meets the preset conditions; determining the candidate reference frame that meets the preset conditions as the target reference frame of the layer.
[0397] In some embodiments, the preprocessing unit is further used to perform: determining, for each real box of the layer, a degree of matching between the real box and the target reference box of the layer; adjusting the layer to which the real box belongs based on the degree of matching; and updating the target reference box of each layer based on the real boxes included in each layer after adjustment.
[0398] In a third aspect, the present application provides a system-on-chip, wherein the system-on-chip is connected to a controller of a vehicle;
[0399] The system-level chip is used to: obtain spatial data of the road ahead collected by the vehicle; identify and process the spatial data through a road collapse recognition model to determine the recognition result of the road ahead; the recognition result includes a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; wherein, the recognition processing of the road collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining the target reference frame of each layer, and predicting and processing the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame of each layer; the size of the predicted detection frame of each layer matches the collapse type; classifying and processing the target features corresponding to the predicted detection frame of each layer through a classification model to obtain the classification result of each layer; and determining the recognition result of the road ahead based on the classification result of each layer;
[0400] The system-level chip is also used to: when the first identifier indicates that there is a collapse in the road ahead, send the collapse type and collapse location in the identification results to the vehicle controller; so that the controller can respond based on the collapse type and collapse location; the collapse types include: road surface fracture, road surface lateral collapse, and road surface longitudinal collapse.
[0401] Because collapses are long-tail scenarios—highly destructive but low-probability events—this solution implements these low-probability events through the system-on-chip (SoC), rather than the vehicle's main controller. This prevents these low-probability events from impacting the vehicle's normal control efficiency. When a collapse is detected, the type and location of the collapse are transmitted to the main controller in real time, allowing it to respond and process, while also improving safety.
[0402] In a fourth aspect, the present application further provides a vehicle, comprising a processor and a memory, wherein the memory stores a computer program or instructions, and when the computer program or instructions are executed by the processor, the method provided in the first aspect is implemented.
[0403] In a fifth aspect, an embodiment of the present application provides a storage medium, that is, a computer-readable storage medium, which stores a computer program or instructions. When the computer program or instructions are executed by a processor, the method provided in the first aspect above is implemented.
[0404] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the method provided in the first aspect above is implemented.
[0405] It should be noted that the descriptions of the above storage medium, device, and program product embodiments are similar to the descriptions of the above method embodiments and have similar beneficial effects as the method embodiments. For technical details not disclosed in the storage medium, device, apparatus, and program product embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0406] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, "in one embodiment" or "in some embodiments" appearing throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application. The above-mentioned serial numbers of the embodiments of the present application are for description only and do not represent the advantages and disadvantages of the embodiments.
[0407] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0408] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0409] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units; they may be located in one place or distributed across multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0410] In addition, all functional units in the embodiments of the present application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0411] Those skilled in the art will understand that all or part of the steps of the above-mentioned method embodiment can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above-mentioned method embodiment; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM), magnetic disks or optical disks, and other media that can store program codes.
[0412] Alternatively, if the above-mentioned integrated unit of the present application is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the methods of each embodiment of the present application. The aforementioned storage medium includes: various media that can store program code, such as mobile storage devices, ROMs, magnetic disks, or optical disks.
[0413] The above description is only an implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.
Claims
1. A method for treating road collapse, characterized in that: The method comprises: Acquire spatial data of the road ahead collected by the vehicle; The spatial data is identified and processed by a road collapse recognition model to determine an identification result of the road ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse in the road ahead; wherein the identification processing of the road collapse recognition model includes: performing multi-layer feature extraction on the spatial data by a feature extraction model to obtain multi-layer target feature maps; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a prediction detection frame of each layer; the size of the prediction detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the prediction detection frame of each layer by a classification model to obtain a classification result of each layer; and determining the identification result of the road ahead based on the classification result of each layer; When the first identifier indicates that there is a collapse in the road ahead, the collapse type and collapse location in the identification result are output so that the vehicle can respond based on the collapse type and collapse location; the collapse types include: road surface fracture, road surface lateral collapse, and road surface longitudinal collapse.
2. The method according to claim 1, characterized in that The method further comprises: Determining a self-handling method of the vehicle based on the collapse type and the collapse location; Sending a first prompt message, the collapse type, and the collapse location to a vehicle behind the vehicle via vehicle network communication; the first prompt message is used to remind the vehicle that a collapse exists ahead; The collapse type and collapse location are sent to the cloud.
3. The method according to claim 1, characterized in that The method of performing multi-layer feature extraction on the spatial data by using a feature extraction model to obtain a multi-layer target feature map includes: Performing multi-scale convolution processing on the spatial data through a fully convolutional network, and splicing outputs of multiple stages to obtain a multi-layer first feature map; the resolution of each layer of the multi-layer first feature map matches the collapse type; Performing detail feature extraction on the first feature map of each layer in the multiple layers of the first feature map based on the attention mechanism to obtain the second feature map of the multiple layers; Performing a fusion process on the second feature maps of the multiple layers to obtain a third feature map of the multiple layers; Determining the target feature map of multiple layers includes the first feature map of multiple layers, the second feature map of multiple layers, or the third feature map of multiple layers.
4. The method according to claim 3, characterized in that In the case where the multi-layer includes three layers, the multi-scale convolution processing is performed on the spatial data through the full convolution network, and the outputs of multiple stages are spliced to obtain the first feature map of the multi-layer, including: Performing convolution processing on the spatial data based on a first downsampling rate to obtain a first feature map of a first layer; Performing convolution processing on the spatial data based on a second downsampling rate to obtain a first feature map of a second layer; Performing convolution processing on the spatial data based on a third downsampling rate to obtain a first feature map of a third layer; Determining the first feature maps of multiple layers based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer; Among them, the first downsampling rate is smaller than the second downsampling rate, and the second downsampling rate is smaller than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.
5. The method according to claim 3, characterized in that In a case where the multi-layer first feature maps include a first feature map of a first layer, a first feature map of a second layer, and a first feature map of a third layer, performing detail feature extraction on the first feature map of each layer in the multi-layer first feature maps based on an attention mechanism to obtain a multi-layer second feature map, including: Extracting detail features from the first feature map of the first layer based on a first receptive field using a first attention module to obtain a second feature map of the first layer; Performing detail feature extraction on the first feature map of the second layer based on a second receptive field using a second attention module to obtain a second feature map of the second layer; Performing detail feature extraction on the first feature map of the third layer based on a third receptive field using a third attention module to obtain a second feature map of the third layer; Determining multiple layers of the second feature map based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer; The first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.
6. The method according to claim 3, characterized in that In a case where the second feature maps of the multiple layers include the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, the fusing processing of the second feature maps of the multiple layers to obtain the third feature map of the multiple layers includes: Performing cascade convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to the first path to obtain intermediate feature maps of the three layers; The intermediate feature maps of the three layers are subjected to cascade convolution processing according to a second path to obtain the third feature maps of the multiple layers; the second path is different from the first path.
7. The method according to claim 1, characterized in that The classification model is used to classify the target features corresponding to the predicted detection box of each layer to obtain the classification results of each layer, including: Get the target reference frame of each layer; Determining the matching degree between the predicted detection frame and the target reference frame of each layer respectively to obtain multiple matching degrees; the matching degree is used to represent the degree of overlap between the predicted detection frame and the target reference frame; determining the collapse type based on the multiple matching degrees; Determining the collapse position based on the position of the predicted detection frame; The classification result of each layer is determined to include the collapse type and the collapse location.
8. The method according to claim 1, characterized in that The method further comprises: Acquire a sample data set; the sample data set includes a plurality of sample space data, each of the sample space data is marked with a true box; the true box is used to point to the collapse position in the sample space data; Distributing multiple ground-truth boxes in the sample data set into multiple layers according to their sizes, so that each layer includes multiple ground-truth boxes; For each of the plurality of layers, determining a center of a target reference frame of the layer based on a plurality of ground truth frames included in the layer; A target reference frame of the layer is determined based on a center of the target reference frame of the layer.
9. The method according to claim 8, characterized in that The determining, based on a plurality of real frames included in the layer, a center of a target reference frame of the layer, comprises: Based on the multiple ground truth boxes of the layer, determining the center points of the multiple ground truth boxes; Determine the probability density of the center point of each ground truth box of the layer respectively; The center point with the highest probability density is determined as the center of the target reference frame.
10. The method according to claim 8, characterized in that The determining the target reference frame of the layer based on the center of the target reference frame of the layer includes: Generate a candidate reference frame with the center point of the target reference frame as the center; Determine a dynamic distance between each ground truth box of the layer and the candidate reference box respectively; the dynamic distance is negatively correlated with the overlap, and the dynamic distance is positively correlated with the aspect ratio similarity; Adjusting the candidate reference frame based on the dynamic distance until the dynamic distance meets a preset condition; The candidate reference frame that meets the preset condition is determined as the target reference frame of the layer.
11. The method according to claim 8, characterized in that The method further comprises: For each ground-truth frame of the layer, determining a degree of matching between the ground-truth frame and a target reference frame of the layer; Adjusting the layer to which the real frame belongs based on the matching degree; Based on the adjusted true frame included in each of the layers, the target reference frame of each of the layers is updated.
12. A device for handling road collapse, characterized in that: The device comprises: An acquisition unit, used to acquire spatial data of the road ahead collected by the vehicle; An identification unit is used to perform identification processing on the spatial data through a road collapse identification model to determine an identification result of the road ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse on the road ahead; wherein the identification processing of the road collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of the target feature map of each layer in the multi-layer target feature map matches the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame of each layer; the size of the predicted detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection frame of each layer through a classification model to obtain a classification result of each layer; and determining the identification result of the road ahead based on the classification result of each layer. An output unit is used to output the collapse type and collapse location in the identification result when the first identifier indicates that there is a collapse of the road ahead, so that the vehicle can respond based on the collapse type and collapse location; the collapse type includes: road surface fracture, road surface lateral collapse, and road surface longitudinal collapse.
13. A system-on-chip, characterized in that: The system-on-chip is connected to a controller for road collapse; The system-level chip is used to: obtain spatial data of the road ahead collected by the vehicle; identify and process the spatial data using a road collapse identification model to determine an identification result of the road ahead; The recognition result includes a first identifier; The first identifier is used to characterize whether there is a collapse on the road ahead; wherein, the recognition processing of the road collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; matching the resolution of the target feature maps of each layer in the multi-layer target feature maps with the collapse type; obtaining a target reference frame of each layer, and performing prediction processing on the target feature maps of each layer based on the target reference frame of each layer to obtain a prediction detection frame of each layer; matching the size of the prediction detection frame of each layer with the collapse type; performing classification processing on the target features corresponding to the prediction detection frame of each layer through a classification model to obtain a classification result of each layer; and determining the recognition result of the road ahead based on the classification result of each layer. The system-level chip is also used to: when the first identification indicates that there is a collapse of the road ahead, send the collapse type and collapse location in the identification result to the controller of the road collapse; so that the controller can respond based on the collapse type and the collapse location; the collapse types include: road rupture, road lateral collapse, and road longitudinal collapse.
14. A vehicle, characterized in that: The vehicle includes a processor and a memory, wherein a computer program or instruction is stored in the memory, and when the computer program or instruction is executed by the processor, the method according to any one of claims 1 to 11 is implemented.
15. A computer-readable storage medium, characterized in that The storage medium stores a computer program or instructions, and when the computer program or instructions are executed by the processor, the method according to any one of claims 1 to 11 is implemented.
16. A computer program product, characterized in that The computer program product comprises a computer program or instructions, and when the computer program or instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Double-branch low-illumination image enhancement method based on Retinex theory
CN117994155A
Neural network for object detection in images
US20180121762A1