Road surface collapse processing method and device, chip, vehicle, medium and product

The system enhances vehicle safety by accurately detecting and responding to road collapses using a road collapse identification model for vehicles, improving response accuracy and timeliness and providing warnings.

CN120318794AActive Publication Date: 2025-07-15CHONGQING CHANGAN AUTOMOBILE CO LTD

Patent Information

Application Number
CN202510807505.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

The prior art has failed to effectively solve the safety problems caused by road collapse during driving, especially in the autonomous driving environment, which are insufficiently identified and responded to the types and locations of road collapses.

Method used

The road collapse recognition model is used to identify the collapse type and location of the road ahead through feature extraction, prediction detection frame and classification model, and multi-layer feature extraction and fusion is used to combine vehicle network communication and cloud backup to achieve autonomous response and timely processing of vehicles.

Benefits of technology

It improves the accuracy and timeliness of the vehicle's response in the situation of road collapse, ensures the safety of the vehicle and the rear vehicles, and reduces accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318794A_ABST
    Figure CN120318794A_ABST
Patent Text Reader

Abstract

The invention relates to a road surface collapse processing method and device, a chip, a vehicle, a medium and a product. The method comprises the steps that spatial data, collected by the vehicle, of a front road surface are acquired; performing identification processing on the spatial data through a road surface collapse identification model, and determining an identification result of a front road surface; the identification result comprises a first identifier; and under the condition that the first identifier represents that the front road surface collapses, the collapse type and the collapse position in the recognition result are output, so that the vehicle performs response processing based on the collapse type and the collapse position. According to the scheme, whether the front road surface collapses or not can be automatically detected in the vehicle driving process, the collapse type and the collapse position can be obtained, and therefore the vehicle can respond in time based on the collapse type and the collapse position of the front road surface, and the vehicle driving safety is improved. And various types of collapse problems of the road surface can be identified at the same time, so that the identification efficiency is improved, and the identification accuracy is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of automotive technologies, and particularly to a method, device, chip, vehicle, medium, and product for dealing with road surface collapses. Background Art

[0002] With the increasing construction of roads, road safety hazards have become more and more prominent. Among them, the safety problems of vehicle driving caused by road surface collapses are also relatively common. Therefore, it is particularly important to solve the problems caused by road surface collapses during vehicle driving. Summary of the Invention

[0003] One of the purposes of this application is to provide a method, device, chip, vehicle, medium, and product for dealing with road surface collapses. This solution can autonomously detect whether there is a collapse on the road surface ahead during vehicle driving, and can obtain the collapse type and collapse location, so that the vehicle can respond in a timely manner based on the collapse type and collapse location of the road surface ahead, improving the safety of vehicle driving.

[0004] To achieve the above purpose, the technical solutions adopted in this application are as follows: In a first aspect, this application provides a method for dealing with road surface collapses. The method includes: obtaining spatial data of the road surface ahead collected by the vehicle; performing identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the road surface ahead; the identification result includes a first identifier; the first identifier is used to represent whether there is a collapse on the road surface ahead; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; obtaining the target reference frame for each layer, and performing prediction processing on the target feature map for each layer based on the target reference frame for each layer to obtain the predicted detection frame for each layer; the size of the predicted detection frame for each layer matches the collapse type; performing classification processing on the target feature corresponding to the predicted detection frame for each layer through a classification model to obtain the classification result for each layer; determining the identification result of the road surface ahead based on the classification result for each layer; in the case where the first identifier represents that there is a collapse on the road surface ahead, outputting the collapse type and collapse location in the identification result so that the vehicle can perform response processing based on the collapse type and collapse location; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

[0005] Based on the above technical means, during the driving process of the vehicle, the recognition result of the road surface ahead can be obtained in real time through the collected spatial data of the road surface ahead, so as to determine whether there is a collapse ahead, and the type and location of the collapse can also be determined, and then the response can be made according to the type and location of the collapse. On the one hand, making a response based on the type and location of the collapse can improve the accuracy of the response processing. On the other hand, in this solution, during the driving process of the vehicle at the vehicle end, the vehicle can autonomously detect whether there is a collapse on the road surface ahead, and then make a timely response according to the type and location of the collapse on the road surface ahead, improving the timeliness of the processing and the safety of the vehicle driving. On the other hand, when determining the recognition result, first perform multi-layer feature extraction on the spatial data through the feature extraction model, and then for each layer, determine the prediction reference box through the target reference box of this layer, determine the classification result of this layer based on the prediction reference box of this layer, and determine the final recognition result based on the classification results of each layer. It can be seen that when determining the recognition result, it is achieved through multiple layers. Since the resolution corresponding to the features of each layer matches the collapse type, and the size of the prediction detection box corresponding to each layer matches the collapse type, each layer can realize the recognition of collapse types of different sizes. For example, identify road fractures through one layer and road collapses through another layer. In this way, various types of collapse problems on the road surface can be recognized simultaneously, improving the recognition efficiency and having a relatively high recognition accuracy.

[0006] In a possible implementation manner, the method further includes: determining the vehicle's own processing method based on the collapse type and location; sending a first prompt message, the collapse type, and the location of the collapse to the vehicle behind the vehicle through vehicle-to-internet communication; the first prompt message is used to remind that there is a collapse ahead; sending the collapse type and the location of the collapse to the cloud.

[0007] Based on the above technical means, when the vehicle makes a response based on the collapse type and location, it not only includes its own response processing, but also includes reminding the vehicle behind and backing up in the cloud. The processing method is comprehensive, which not only improves the safety of the vehicle itself, but also improves the driving safety of other vehicles.

[0008] In a possible implementation, the feature extraction model performs multi-layer feature extraction on the spatial data to obtain multi-layer target feature maps, including: performing multi-scale convolution processing on the spatial data through a fully convolutional network, and splicing the outputs of multiple stages to obtain multi-layer first feature maps; the resolution of each first feature map in the multi-layer first feature maps matches the collapse type; performing detail feature extraction on each first feature map in the multi-layer first feature maps respectively based on the attention mechanism to obtain multi-layer second feature maps; performing fusion processing on the multi-layer second feature maps to obtain multi-layer third feature maps; determining that the multi-layer target feature maps include the multi-layer first feature maps, the multi-layer second feature maps or the multi-layer third feature maps.

[0009] Based on the above technical means, when determining the target feature map, it can be determined that the target feature map is the multi-layer first feature maps output by the fully convolutional network. The resolution of the multi-layer first feature maps matches the collapse type, which can meet the layer recognition requirements. Moreover, using the fully convolutional network to determine the target feature map has the characteristics of simple implementation logic and high processing efficiency. When determining the target feature map, it can be determined that the target feature map is the multi-layer second feature maps. The resolution of the multi-layer second feature maps matches the collapse type, which can meet the layer recognition requirements. Additionally, the multi-layer second feature maps can also extract detail features based on the attention mechanism. In this way, the obtained target feature map includes detail features, and the obtained features are more comprehensive and the collapse recognition is more accurate. When determining the target feature map, it can be determined that the target feature map is the multi-layer third feature maps. The resolution of the multi-layer third feature maps matches the collapse type, which can meet the layer recognition requirements. Moreover, the multi-layer third feature maps are obtained by fusing the second feature maps. Therefore, it not only has the detail features in the second feature maps, but also fusing each layer can make each layer retain the features of other layers. In this way, the obtained target feature map is more comprehensive and the collapse recognition is more accurate.

[0010] In a possible implementation, when the multi-layer includes three layers, performing multi-scale convolution processing on the spatial data through a fully convolutional network, and splicing the outputs of multiple stages to obtain multi-layer first feature maps, including: performing convolution processing on the spatial data based on the first downsampling rate to obtain the first feature map of the first layer; performing convolution processing on the spatial data based on the second downsampling rate to obtain the first feature map of the second layer; performing convolution processing on the spatial data based on the third downsampling rate to obtain the first feature map of the third layer; determining the multi-layer first feature maps based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer; where the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the resolution of the second feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.

[0011] Based on the above technical means, when determining the first feature map, convolution processing is performed through different downsampling rates, so as to obtain the first feature maps at different resolutions. It has the characteristics of simple and reliable implementation.

[0012] In a possible implementation manner, when the multi-layer first feature maps include the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, detail feature extraction is respectively performed on each first feature map in the multi-layer first feature maps based on the attention mechanism to obtain multi-layer second feature maps, including: performing detail feature extraction on the first feature map of the first layer through the first attention module based on the first receptive field range to obtain the second feature map of the first layer; performing detail feature extraction on the first feature map of the second layer through the second attention module based on the second receptive field range to obtain the second feature map of the second layer; performing detail feature extraction on the first feature map of the third layer through the third attention module based on the third receptive field range to obtain the second feature map of the third layer; determining the multi-layer second feature maps based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer; wherein, the first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.

[0013] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, and then each attention module corresponds to a different receptive field, so that the second feature maps of different layers can be obtained. It has the characteristics of simple, reliable, and accurate implementation.

[0014] In a possible implementation manner, when the multi-layer second feature maps include the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, fusion processing is performed on the multi-layer second feature maps to obtain multi-layer third feature maps, including: performing concatenated convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to the first path to obtain an intermediate feature map of three layers; performing concatenated convolution processing on the intermediate feature map of three layers according to the second path to obtain multi-layer third feature maps; the second path is different from the first path.

[0015] Based on the above technical means, when fusing the multi-layer second feature maps, first fuse based on the first path, and then fuse based on the second path. Through the fusion of the two paths, each target feature map after fusion can retain the features of all upper and lower layers, and the feature maps of each layer in the obtained target feature maps have comprehensive information, improving the accuracy of identifying the collapse type of each layer.

[0016] In a possible implementation, the target features corresponding to the predicted detection boxes of each layer are classified by a classification model to obtain the classification result of each layer, including: obtaining the target reference box of each layer; respectively determining the matching degree between the predicted detection box and the target reference box of each layer to obtain a plurality of matching degrees; the matching degree is used to characterize the overlapping degree between the predicted detection box and the target reference box; determining the collapse type based on the plurality of matching degrees; determining the collapse position based on the position of the predicted detection box; determining that the classification result of each layer includes the collapse type and the collapse position.

[0017] Based on the above technical means, when determining the collapse type, it is determined based on the matching degree between the predicted detection box and the target reference box of each layer. In this way, since the target reference box of one layer can represent the characteristics of a type of collapse type, the determination method based on the matching degree is more accurate for determining the collapse type. When determining the collapse position, since the predicted detection box is used to indicate the collapse position, it is more accurate to determine the collapse position through the position of the predicted detection box.

[0018] In a possible implementation, the method further includes: obtaining a sample data set; the sample data set includes a plurality of sample space data, and each sample space data is marked with a ground truth box; the ground truth box is used to point to the collapse position in the sample space data; the plurality of ground truth boxes in the sample data set are allocated to a plurality of layers according to the size, so that each layer includes a plurality of ground truth boxes; for each layer in the plurality of layers, based on the plurality of ground truth boxes included in the layer, determining the center of the target reference box of the layer; based on the center of the target reference box of the layer, determining the target reference box of the layer.

[0019] Based on the above technical means, the sample space data in the sample data set can be divided into different layers according to the size of the ground truth box, so as to determine the center of the target reference box of each layer according to the size of the ground truth box of different layers, and then determine the target reference box of each layer. In this way, the determined target reference box of each layer can meet the size requirements of a type of collapse type, and improve the recognition accuracy of the layer for this collapse type.

[0020] In a possible implementation, based on the plurality of ground truth boxes included in the layer, determining the center of the target reference box of the layer includes: based on the plurality of ground truth boxes of the layer, determining the center points of the plurality of ground truth boxes; respectively determining the probability density of the center point of each ground truth box of the layer; determining the center point with the highest probability density as the center of the target reference box.

[0021] Based on the above technical means, when determining the center of the target reference box of a layer, it is determined based on the probability density, which can meet the needs of most ground truth boxes and is also more accurate.

[0022] In a possible implementation manner, determining the target reference box of a layer based on the center of the target reference box of the layer includes: generating a candidate reference box centered on the center point of the target reference box; respectively determining the dynamic distance between each ground truth box of the layer and the candidate reference box; the dynamic distance is negatively correlated with the overlap degree and positively correlated with the aspect ratio similarity; adjusting the candidate reference box based on the dynamic distance until the dynamic distance meets a preset condition; determining the candidate reference box that meets the preset condition as the target reference box of the layer.

[0023] Based on the above technical means, when determining the target reference box, first generate a candidate reference box based on the center point, and then continuously adjust the shape and size of the candidate reference box based on the dynamic distance, so that the candidate reference box meets the preset condition with the ground truth box of the layer. In this way, the determined target reference box can have a large overlap degree with the ground truth box and an aspect ratio close to that of the ground truth box, that is, it is close to the position and shape of the ground truth box, thus meeting the actual requirements and improving the accuracy.

[0024] In a possible implementation manner, the method further includes: for each ground truth box of the layer, respectively determining the matching degree between the ground truth box and the target reference box of the layer; adjusting the layer to which the ground truth box belongs based on the matching degree; updating the target reference box of each layer based on the ground truth boxes included in each adjusted layer.

[0025] Based on the above technical means, after stratifying each ground truth box, the allocation method can also be adjusted based on the matching degree. This can optimize the stratification, improve the accuracy of the division, and also improve the accuracy of collapse type recognition.

[0026] In a second aspect, the present application provides a device for processing road surface collapse, and the device includes: An acquisition unit, configured to acquire spatial data of the road surface ahead collected by a vehicle; An identification unit, configured to perform identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the road surface ahead; the identification result includes a first identifier; the first identifier is used to represent whether there is a collapse on the road surface ahead; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; acquiring the target reference box of each layer, and performing prediction processing on each layer of the target feature map based on the target reference box of each layer to obtain the predicted detection box of each layer; the size of the predicted detection box of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection box of each layer through a classification model to obtain the classification result of each layer; determining the identification result of the road surface ahead based on the classification result of each layer; An output unit, configured to output the collapse type and the collapse location in the recognition result when the first identifier indicates that there is a collapse on the front road surface, so that the vehicle performs response processing based on the collapse type and the collapse location; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

[0027] In a third aspect, the present application provides a system-on-chip, which is connected to the controller of the road surface collapse; The system-on-chip is configured to: obtain the spatial data of the front road surface collected by the vehicle; perform recognition processing on the spatial data through a road surface collapse recognition model to determine the recognition result of the front road surface; the recognition result includes a first identifier; the first identifier is used to indicate whether there is a collapse on the front road surface; wherein, the recognition processing of the road surface collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each layer of the target feature maps in the multi-layer target feature maps matches the collapse type; obtaining the target reference frame of each layer, and performing prediction processing on each layer of the target feature maps based on the target reference frame of each layer to obtain the predicted detection frame of each layer; the size of the predicted detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection frame of each layer through a classification model to obtain the classification result of each layer; determining the recognition result of the front road surface based on the classification result of each layer; The system-on-chip is further configured to: when the first identifier indicates that there is a collapse on the front road surface, send the collapse type and the collapse location in the recognition result to the controller of the road surface collapse; so that the controller performs response processing based on the collapse type and the collapse location; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

[0028] Since the collapse belongs to a long-tail scenario, that is, an event with large damage but low probability, in this solution, the low-probability event is implemented by the system-on-chip instead of the main controller of the vehicle, so that the low-probability event does not affect the normal control efficiency of the vehicle. When a collapse is recognized, the collapse type and the collapse location are sent to the main control in real time, and the main control performs response processing, while improving safety.

[0029] In a fourth aspect, the present application further provides a vehicle, which includes a processor and a memory, and a computer program or instruction is stored on the memory. When the computer program or instruction is executed by the processor, the method provided in the first aspect is implemented.

[0030] In a fifth aspect, the present application further provides a storage medium, and a computer program or instruction is stored on the storage medium. When the computer program or instruction is executed by the processor, the method provided in the first aspect is implemented.

[0031] In a sixth aspect, the present application further provides a computer program product, which includes a computer program or instructions. When the computer program or instructions are executed by a processor, the method provided in the first aspect above is implemented.

[0032] It should be noted that the technical effects of the second aspect to the sixth aspect can be referred to the detailed description of the first aspect above, and will not be elaborated here one by one. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 FIG. 9 is a first alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 2 FIG. 12 is a second alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 3 FIG. 15 is a third alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 4 FIG. 18 is a fourth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 5 FIG. 21 is a fifth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 6 FIG. 24 is a sixth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 7 FIG. 27 is a seventh alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 8 FIG. 30 is an eighth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 9 FIG. 33 is a ninth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 10 FIG. 36 is a tenth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 11 FIG. 39 is an eleventh alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 12 FIG. 42 is a twelfth alternative flowchart of the method for treating road surface collapse provided by the embodiment of the present application; Figure 13 FIG. 45 is an alternative flowchart of the process for treating road surface collapse provided by the embodiment of the present application; Figure 14 FIG. 48 is an alternative structural diagram of the FPGA hardware chip system provided by the embodiment of the present application; Figure 15 An optional schematic structural diagram of the FPGA hardware accelerator provided by an embodiment of the present application; Figure 16 An optional schematic structural diagram of the device for treating road surface collapse provided by an embodiment of the present application. Detailed implementation manners

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.

[0035] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0036] In the following description, the terms "first / second / third" are only used to distinguish different objects, and do not represent a specific order for the objects, and there is no limitation on the order of precedence. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0037] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0038] The embodiments of the present application provide a method, device, chip, vehicle, medium, and product for treating road surface collapse. The method for treating road surface collapse is executed by the device for treating road surface collapse, and the device for treating road surface collapse can be deployed in a vehicle. For example, the device for treating road surface collapse can be a controller in the vehicle. Next, the embodiments of the method, device, chip, vehicle, medium, and product for treating road surface collapse provided by the embodiments of the present application will be described.

[0039] In a first aspect, the embodiments of the present application provide a method for treating road surface collapse. The following will illustrate this method by taking the vehicle as the execution subject.

[0040] Referring to Figure 1 the content shown, this process may include but is not limited to S101 to S103.

[0041] S101. Obtain the spatial data of the road surface ahead collected by the vehicle.

[0042] The embodiments of the present application do not limit the type of vehicle, which can be configured according to actual needs. The vehicles here may include but are not limited to sedans, commercial vehicles, sports cars, etc.; they may also include but are not limited to gasoline vehicles, electric vehicles, hydrogen energy vehicles, etc.

[0043] Spatial data refers to the image data or point cloud data of the road surface ahead collected, or the data after the fusion of image data and point cloud.

[0044] During the driving process of the vehicle, the vehicle collects the information of the road surface ahead through cameras, radar sensors, etc. to obtain spatial data. S101 can be implemented as: obtaining the spatial data of the road surface ahead collected by vehicle sensors.

[0045] S102. Perform recognition processing on the spatial data through a road surface collapse recognition model to determine the recognition result of the road surface ahead.

[0046] The recognition result includes a first identifier.

[0047] The first identifier is used to characterize whether there is a collapse on the road surface ahead. The embodiments of the present application do not limit the value of the first identifier, which can be configured according to actual needs. For example, the first identifier being 1 indicates that there is a collapse ahead, and the first identifier being 0 indicates that there is no collapse ahead.

[0048] In the case where the first identifier characterizes that there is a collapse on the road surface ahead, the recognition result may include the collapse type and the collapse location. The collapse type may include but is not limited to: road surface fracture, lateral collapse of the road surface, longitudinal collapse of the road surface, etc.

[0049] The recognition result may include one collapse or multiple collapses. In the case where the recognition result includes multiple collapses, the types corresponding to the multiple collapses may be the same or different.

[0050] Among them, the recognition processing of the road surface collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each layer of the target feature maps in the multi-layer target feature maps matches the collapse type; obtaining the target reference frame of each layer, and performing prediction processing on each layer of the target feature maps based on the target reference frame of each layer to obtain the prediction detection frame of each layer; the size of the prediction detection frame of each layer matches the collapse type; performing classification processing on the target features corresponding to the prediction detection frame of each layer through a classification model to obtain the classification result of each layer; determining the recognition result of the road surface ahead based on the classification result of each layer.

[0051] For example, one collapse type corresponds to the target feature map of a layer with a certain resolution, and one collapse type corresponds to a prediction detection frame with a certain size.

[0052] S102 can be implemented as follows: input the spatial data of the road surface ahead into the road surface collapse recognition model, and perform recognition processing on the spatial data through the road surface collapse recognition model. The road surface collapse recognition model identifies whether there is a collapse on the road surface ahead. If there is a collapse, the road surface collapse recognition model identifies the type and location of the collapse, so as to obtain the recognition result of the road surface ahead.

[0053] The road surface collapse recognition model can include a convolutional network model, a model for predicting detection boxes, a classification model, and so on. Among them, the convolutional network model is used for extracting multi-layer target feature maps, the model for predicting detection boxes is used for determining multi-layer prediction detection boxes, and the classification model is used for determining multi-layer classification results.

[0054] S103. When the first identifier indicates that there is a collapse on the road surface ahead, output the collapse type and location in the recognition result, so that the vehicle can perform response processing based on the collapse type and location.

[0055] The embodiments of the present application do not limit the manner of outputting the recognition result, which can be configured according to actual needs. For example, the output recognition result can include but is not limited to: sending to other devices, displaying on the vehicle center control screen, outputting by voice, and so on.

[0056] S103 can be implemented as follows: determine the content of the first identifier in the recognition result. When the first identifier indicates that there is a collapse on the road surface ahead, determine the collapse type and location, and output the collapse type and location in the recognition result, so that the vehicle can perform response processing based on the collapse type and location. When the first identifier indicates that there is no collapse on the road surface ahead, no response processing is performed, and continue to collect the spatial data of the road surface ahead through the sensor, and continue a new round of judgment.

[0057] For the execution subject of S101 to S103, it can be implemented by the main controller in the vehicle. In this way, no additional controller needs to be added, and the implementation is simple. It can also be a newly added dedicated controller for collapse recognition in the vehicle. In this way, through the dedicated controller for recognition, it will not affect the main control function of the main controller and will not affect the processing efficiency of the main control.

[0058] In this embodiment, the method includes: obtaining the spatial data of the road surface ahead collected by the vehicle; performing recognition processing on the spatial data through the road surface collapse recognition model to determine the recognition result of the road surface ahead; the recognition result includes a first identifier; the first identifier is used to indicate whether there is a collapse on the road surface ahead; when the first identifier indicates that there is a collapse on the road surface ahead, output the collapse type and location in the recognition result, so that the vehicle can perform response processing based on the collapse type and location.

[0059] Based on the above technical means, during the driving process of the vehicle, the recognition result of the road surface ahead can be obtained in real time through the collected spatial data of the road surface ahead, so as to determine whether there is a collapse ahead, and the type and location of the collapse can also be determined, and then a response can be made according to the type and location of the collapse. On the one hand, making a response based on the type and location of the collapse can improve the accuracy of the response processing. On the other hand, in this solution, during the driving process of the vehicle at the vehicle end, the vehicle can independently detect whether there is a collapse on the road surface ahead, and then make a timely response processing according to the type and location of the collapse on the road surface ahead, which improves the timeliness of the processing and the safety of the vehicle driving. On the other hand, when the road surface collapse recognition model performs recognition processing, it is implemented through multiple layers. Since the resolution corresponding to the features of each layer matches the type of collapse, and the size of the predicted detection box corresponding to each layer matches the type of collapse, each layer can realize the recognition of different sizes of collapse types. For example, a layer is used to recognize road fractures, and another layer is used to recognize road collapses. In this way, multiple types of collapse problems on the road surface can be recognized simultaneously, which improves the recognition efficiency and the recognition accuracy is also relatively high.

[0060] The processing method provided by the embodiment of the present application may further include a processing process based on the type and location of the collapse after determining the type and location of the collapse.

[0061] In a possible embodiment, referring to Figure 2 the content shown, this process may include but is not limited to the following S201 to S203.

[0062] S201. Determine the vehicle's own processing method based on the type and location of the collapse.

[0063] In a possible implementation, if the vehicle is in an autonomous driving state, directly determine an avoidance plan or a plan to stop driving based on the type and location of the collapse on the road surface ahead. For example, when the type of collapse includes a lateral collapse and the collapse location is large, running through the entire road width, the vehicle's own processing method is to stop advancing. When the type of collapse includes a road fracture and the fracture location is small, determine the vehicle's own processing method as detouring to avoid.

[0064] In another possible implementation, if the vehicle is in a manual driving state, output the vehicle's own processing method through display or voice, etc., to remind the driver of the driving plan for the road surface ahead. In this way, the driver can timely perceive the processing method of the road surface ahead, and the processing is relatively timely, which can prevent the driver from thinking about how to process and resulting in an inability to obtain a correct processing plan in time, thus causing a failure.

[0065] S202. Send a first prompt message, the type of collapse, and the location of the collapse to the vehicle behind the vehicle through vehicle-to-internet communication.

[0066] The first prompt message is used to remind that there is a collapse ahead.

[0067] Rear vehicles refer to vehicles within a certain range behind the vehicle. The rear vehicles here are not limited to the vehicles behind in the same lane, but can also include the vehicles traveling in other lanes in the same driving direction.

[0068] Since a communication connection is established between the vehicle and surrounding vehicles in the vehicle network. Therefore, S202 can be implemented as: sending the first prompt message, the type of collapse, and the location of the collapse to other vehicles within a certain range behind the vehicle through vehicle network communication, so that other vehicles can perform response processing immediately after receiving the first prompt message, the type of collapse, and the location of the collapse, increasing the response processing time, thereby reducing the accident rate.

[0069] S203: Send the type of collapse and the location of the collapse to the cloud.

[0070] S203 can be implemented as: The vehicle sends the type of collapse and the location of the collapse on the road surface in front detected by the vehicle itself to the cloud through the communication connection between the vehicle and the cloud device, so that the cloud records the location and type of the collapse.

[0071] The cloud can continuously update the map according to the collapse locations and types on the road surface reported by many vehicles, so as to obtain a continuously updated map including collapse information.

[0072] Based on the above technical means, when the vehicle performs response processing based on the type of collapse and the location of the collapse, it includes not only its own response processing, but also the reminder for rear vehicles and the backup in the cloud. The processing method is comprehensive, which not only improves the safety of the vehicle itself, but also improves the driving safety of other vehicles.

[0073] During the execution of S201 to S203, the execution subject of S201 to S203 can be the main controller in the vehicle. When the execution subject of S101 to S103 is the newly added controller, the new controller sends the collapse information to the vehicle's controller, and the main controller responds based on the collapse information. Since collapse belongs to a long-tail scenario, that is, an event with great damage but low probability, in this solution, the low-probability event is implemented by the new controller instead of the vehicle's main controller, so that the low-probability event does not affect the normal control efficiency of the vehicle. When a collapse is recognized, the type of collapse and the location of the collapse are sent to the main control in real time for the main control to perform response processing, while improving safety.

[0074] In other processing methods, one or more of the above S201 to S203 can also be configured according to actual needs. For example, it can only include the self-processing process in S201, or it can include the self-processing process and the reminder for the vehicle behind. Other situations will not be listed one by one Next, the process of using the road surface collapse recognition model to identify and process the spatial data in S102 to determine the recognition result of the road surface ahead will be described.

[0075] Reference Figure 3 As shown in the content, this process can include but is not limited to the following S301 to S304.

[0076] S301. Perform multi-layer feature extraction on the spatial data through the feature extraction model to obtain multi-layer target feature maps.

[0077] The resolution of each target feature map in the multi-layer target feature maps matches the collapse type.

[0078] In the embodiments of the present application, the number of layers in the multi-layer is not limited and can be configured according to actual needs. In a possible implementation manner, the number of layers here can be the same as the number of collapse types, so as to implement one layer corresponding to one type of collapse.

[0079] For example, in the case where the collapse types include road surface fracture, road surface lateral collapse, and road surface longitudinal collapse, the number of layers can be configured as 3 layers.

[0080] S301 can be implemented as follows: Determine the number of layers of the feature extraction model based on the collapse type, configure the size of each layer of feature extraction, and then perform feature extraction on each layer of the spatial data through the feature extraction model. The resolution of each obtained layer matches the collapse type, so that multi-layer target feature maps with different resolutions can be obtained. In the embodiments of the present application, the type of the feature extraction model is not limited and can be configured according to actual needs. For example, the feature extraction model can be a fully convolutional network model. Or it can be a combination of a fully convolutional network model and an attention model.

[0081] S302. Obtain the target reference frame for each layer, and perform prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain the predicted detection frame for each layer.

[0082] The predicted detection frame refers to the frame corresponding to the collapse content predicted for this layer. Each layer corresponds to a different collapse type. Therefore, a predicted detection frame needs to be generated for each layer. One layer can have one or more predicted detection frames (corresponding to multiple collapses of the same type), or it can also have no predicted detection frame (corresponding to no collapse of a certain type).

[0083] The size of the predicted detection frame for each layer matches the collapse type.

[0084] The target reference box, one layer corresponds to one target reference box, corresponding to one type of collapse. For example, for the case where the collapse type is road surface fracture, a slender target reference box is configured; for the case where the collapse type is lateral collapse, a horizontal rectangular reference box is configured; for the case where the collapse type is longitudinal collapse, a vertical rectangular reference box is configured.

[0085] S302 can be implemented as: for each layer in multiple layers, the following processing is respectively performed: obtain the target reference box of each layer, and respectively perform corresponding prediction processing on the target feature map of each layer based on the target reference box of each layer through the detection box prediction model to obtain the predicted detection box of each layer.

[0086] S303. Classify the target features corresponding to the predicted detection box of each layer through the classification model to obtain the classification result of each layer.

[0087] S303 can be implemented as: input the target features in the predicted detection box into the classification model, and the target classification model classifies the target features in the predicted detection box to obtain the classification result of the input. Then traverse the predicted detection boxes of each layer to obtain the classification result of each layer.

[0088] Example 1, for the first layer, one fracture is identified at position 1; for the second layer, two lateral collapses are identified at positions 2 and 3; for the third layer, one longitudinal collapse is identified at position 4.

[0089] S304. Determine the recognition result of the road surface ahead based on the classification result of each layer.

[0090] S304 can be implemented as: summarize the classification results of all layers to obtain the recognition result of the road surface ahead. For example, based on Example 1, the obtained recognition result may include: one fracture, two lateral collapses, and one longitudinal collapse. The recognition result also includes the position information of one fracture, two lateral collapses, and one longitudinal collapse.

[0091] Based on the above technical means, when determining the recognition result, first perform multi-layer feature extraction on the spatial data through the feature extraction model, and then for each layer, determine the prediction reference frame through the target reference frame of this layer, determine the classification result of this layer based on the prediction reference frame of this layer, and determine the final recognition result based on the classification results of each layer. It can be seen that when determining the recognition result, it is achieved through multiple layers. Since the resolution of the features corresponding to each layer matches the collapse type, and the size of the prediction detection frame corresponding to each layer matches the collapse type, each layer can achieve the recognition of different sizes of collapse types. For example, identify pavement fractures through one layer and identify pavement collapses through another layer. In this way, multiple types of collapse problems on the road surface can be identified simultaneously, improving the recognition efficiency and having a relatively high recognition accuracy.

[0092] Next, the process of performing multi-layer feature extraction on the spatial data through the feature extraction model in S301 to obtain multi-layer target feature maps will be described.

[0093] Refer to Figure 4 the content shown. This process may include but is not limited to the following S401 to S404.

[0094] S401: Perform multi-scale convolution processing on the spatial data through a fully convolutional network, and splice the outputs of multiple stages to obtain multi-layer first feature maps.

[0095] The resolution of each first feature map in the multi-layer first feature maps matches the collapse type.

[0096] S401 can be implemented as: Configure the convolution scale of each layer, and then sequentially perform convolution processing on the spatial data of each layer through the fully convolutional network at the corresponding scale, splice the outputs of multiple stages, and obtain multi-layer first feature maps. The first feature map can be a matrix feature map, and the splicing of multiple layers, that is, the splicing of multiple stages, corresponds to the splicing of matrices.

[0097] S402: Respectively perform detailed feature extraction on each first feature map in the multi-layer first feature maps based on the attention mechanism to obtain multi-layer second feature maps.

[0098] S402 can be implemented as: For each first feature map output by each layer, configure an attention module, perform detailed feature extraction on the first feature map of this layer through this attention module, traverse the first feature maps of each layer, and a second feature map corresponding to the corresponding layer can be obtained based on the first feature map of each layer. Splice the multi-layer second feature maps to obtain multi-layer second feature maps.

[0099] S403: Perform fusion processing on the multi-layer second feature maps to obtain multi-layer third feature maps.

[0100] S403 can be implemented as follows: performing a fusion process on the multi-layer second feature maps in the order from the first layer to the last layer and / or from the last layer to the first layer to obtain multi-layer third feature maps. The fusion here can be performed once or multiple times.

[0101] S404. Determine that the multi-layer target feature maps include the multi-layer first feature maps, the multi-layer second feature maps, or the multi-layer third feature maps.

[0102] S404 can be implemented as follows: It can be determined according to actual requirements that the multi-layer target feature maps include the multi-layer first feature maps, the multi-layer second feature maps, or the multi-layer third feature maps.

[0103] Based on the above technical means, when determining the target feature maps, it can be determined that the target feature maps are the multi-layer first feature maps output by the fully convolutional network. The resolution of the multi-layer first feature maps matches the collapse type, which can meet the layer recognition requirements. And determining the target feature maps through the fully convolutional network has the characteristics of simple implementation logic and high processing efficiency. When determining the target feature maps, it can be determined that the target feature maps are the multi-layer second feature maps. The resolution of the multi-layer second feature maps matches the collapse type, which can meet the layer recognition requirements. Moreover, the multi-layer second feature maps can also extract detailed features based on the attention mechanism. In this way, the obtained target feature maps include detailed features, and the obtained features are more comprehensive and the collapse recognition is more accurate. When determining the target feature maps, it can be determined that the target feature maps are the multi-layer third feature maps. The resolution of the multi-layer third feature maps matches the collapse type, which can meet the layer recognition requirements. And the multi-layer third feature maps are obtained by fusing the second feature maps. Therefore, not only do they have the detailed features in the second feature maps, but also fusing each layer can enable each layer to retain the features of other layers. In this way, the obtained target feature maps are more comprehensive and the collapse recognition is more accurate.

[0104] Next, the process of performing multi-scale convolutional processing on the spatial data through the fully convolutional network in S401, splicing the outputs of multiple stages, and obtaining the multi-layer first feature maps will be described.

[0105] In the case where the multi-layer includes three layers, referring to Figure 5 the content shown, this process may include but is not limited to the following S501 to S504. For the case of other numbers of layers, the processing process of three layers can be referred to, and details will not be elaborated here.

[0106] S501. Perform convolutional processing on the spatial data based on the first downsampling rate to obtain the first feature map of the first layer.

[0107] Among them, the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.

[0108] In the embodiment of the present application, the value of the first downsampling rate is not limited and can be configured according to actual requirements. For example, assuming that the input image size is L×L×3, the first downsampling rate can be 2, and the size of the first feature map of the first layer obtained can be L / 2×L / 2×16.

[0109] S501 can be implemented as: inputting the spatial data into the fully convolutional network, and performing convolutional processing on the spatial data based on the first downsampling rate to obtain the first feature map of the first layer.

[0110] S502. Perform convolutional processing on the spatial data based on the second downsampling rate to obtain the first feature map of the second layer.

[0111] In the embodiment of the present application, the value of the second downsampling rate is not limited and can be configured according to actual requirements. For example, assuming that the input image size is L×L×3, the second downsampling rate can be 4, and the size of the first feature map of the second layer obtained can be L / 4×L / 4×24.

[0112] The implementation of S502 can refer to the description of performing convolutional processing on the spatial data based on the first downsampling rate in S501 to obtain the first feature map of the first layer, and will not be elaborated here one by one.

[0113] S503. Perform convolutional processing on the spatial data based on the third downsampling rate to obtain the first feature map of the third layer.

[0114] In the embodiment of the present application, the value of the third downsampling rate is not limited and can be configured according to actual requirements. For example, assuming that the input image size is L×L×3, the third downsampling rate can be 8, and the size of the first feature map of the third layer obtained can be L / 8×L / 8×40.

[0115] The implementation of S503 can refer to the description of performing convolutional processing on the spatial data based on the first downsampling rate in S501 to obtain the first feature map of the first layer, and will not be elaborated here one by one.

[0116] S504. Determine the first feature maps of multiple layers based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer.

[0117] S504 will perform matrix splicing based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, so as to obtain the first feature maps of multiple layers.

[0118] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, and each attention module corresponds to a different receptive field, different layers of the second feature map can be obtained. It has the characteristics of simple, reliable, and accurate implementation.

[0119] Next, the process of performing detail feature extraction on each layer of the first feature maps in multiple layers based on the attention mechanism in S402 to obtain the second feature maps in multiple layers will be described.

[0120] When the first feature maps in multiple layers include the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, referring to Figure 6 the content shown, this process may include but is not limited to the following S601 to S604.

[0121] S601: Perform detail feature extraction on the first feature map of the first layer based on the first receptive field range through the first attention module to obtain the second feature map of the first layer.

[0122] Among them, the first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.

[0123] The embodiments of the present application do not limit the size of the first receptive field range, which can be configured according to actual needs. For example, the first receptive field range can all be of the size of 3×3.

[0124] S601 can be implemented as: input the first image feature of the first layer into the first attention module, the first attention module is configured with the size of the first receptive field range, and through the first attention module, perform progressive traversal on the first feature map of the first layer based on the first receptive field range, extract the detail features within each receptive field, then move the receptive field backward or downward, move step by step, and then perform the extraction of detail features, so as to obtain the second feature map of the first layer.

[0125] S602: Perform detail feature extraction on the first feature map of the second layer based on the second receptive field range through the second attention module to obtain the second feature map of the second layer.

[0126] The embodiments of the present application do not limit the size of the second receptive field range, which can be configured according to actual needs. For example, the second receptive field range can all be of the size of 7×7.

[0127] The implementation of S602 can refer to the detailed description of performing detail feature extraction on the first feature map of the first layer based on the first receptive field range through the first attention module in S601 to obtain the second feature map of the first layer, and will not be elaborated here one by one.

[0128] S603. The third attention module performs detailed feature extraction on the first feature map of the third layer based on the third receptive field range to obtain the second feature map of the third layer.

[0129] In the embodiments of the present application, the size of the third receptive field range is not limited and can be configured according to actual needs. For example, the third receptive field range can all be of the size of 15×15.

[0130] For the implementation of S603, reference can be made to the detailed description of S601 in which the first attention module performs detailed feature extraction on the first feature map of the first layer based on the first receptive field range to obtain the second feature map of the first layer, and details will not be repeated here.

[0131] S604. Based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, determine the second feature map of multiple layers.

[0132] S604 can be implemented as: splicing the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer in a matrix diagram to obtain the second feature map of multiple layers.

[0133] Based on the above technical means, when determining the second feature map, by configuring multiple attention modules, and then each attention module corresponds to a different receptive field, the second feature map of different layers can be obtained. It has the characteristics of simple, reliable, and accurate implementation.

[0134] Next, the process of performing fusion processing on the second feature map of multiple layers in S403 to obtain the third feature map of multiple layers will be described. Refer to Figure 7 the content shown. This process may include but is not limited to the following S701 and S702.

[0135] S701. Perform cascaded convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer along the first path to obtain the intermediate feature map of three layers.

[0136] The first path can be: from the first layer to the last layer. Of course, the first path can also be from the last layer to the first layer.

[0137] During the process of performing cascaded convolution processing, different weights can also be configured for the second feature map of each layer. The weight configuration can be determined according to actual needs.

[0138] For example, the weights can gradually increase or gradually decrease.

[0139] The intermediate feature map of three layers refers to the feature map of three layers after cascaded convolution along the first path.

[0140] S701 can be implemented as follows: for the second feature maps of the first layer, the second layer, and the third layer, perform concatenated convolution processing on each layer in sequence according to the order in the first path. After performing concatenated convolution on all layers, an intermediate feature map of three layers is obtained.

[0141] S702. Perform concatenated convolution processing on the intermediate feature maps of three layers according to the second path to obtain a third feature map of multiple layers.

[0142] The second path is different from the first path.

[0143] For example, when the first path is from the first layer to the last layer, the second path can be from the last layer to the first layer.

[0144] The implementation of S702 can refer to the description of performing concatenated convolution processing on the second feature maps of the first layer, the second layer, and the third layer according to the first path in S701 to obtain an intermediate feature map of three layers. The difference is that in S702, the concatenated convolution order needs to be adjusted from the first path to the second path.

[0145] Based on the above technical means, when fusing the second feature maps of multiple layers, first fuse based on the first path, and then fuse based on the second path. Through the fusion of the two paths, the target feature map of each layer after fusion can retain the features of all upper and lower layers, and the feature map of each layer in the obtained target feature map has comprehensive information, improving the accuracy of identifying the collapse type of each layer.

[0146] Next, the process of classifying the target features corresponding to the predicted detection boxes of each layer through the classification model in S303 to obtain the classification result of each layer will be described.

[0147] The processing process of the predicted detection box of each layer here is similar. Below, the processing of the predicted detection box of one layer will be used as an example for description. The implementation process of the predicted detection boxes of other layers can refer to that of this layer and will not be elaborated here one by one.

[0148] Refer to Figure 8 As shown in the content, this process may include but is not limited to the following S801 to S805.

[0149] S801. Obtain the target reference box of each layer.

[0150] Previously, a target reference box is set for each layer. S801 can be implemented as: directly read the set target reference box of each layer. Among them, the target reference boxes of each layer are different.

[0151] S802. Respectively determine the matching degree between the predicted detection box and the target reference box of each layer to obtain multiple matching degrees.

[0152] The matching degree is used to characterize the overlapping degree between the predicted detection box and the target reference box. The greater the overlapping degree, the higher the matching degree. That is, the more similar, the higher the matching degree.

[0153] The embodiments of the present application do not limit the method for determining the matching degree, which can be configured according to actual requirements. For example, the matching degree can be determined according to the intersection over union of the areas of two boxes, or the matching degree can be determined according to the coordinates of the area.

[0154] S802 can be implemented as: for the predicted detection box, based on the determination algorithm of the matching degree, determine the matching degree between the predicted detection box and the target reference box of each layer respectively, and traverse the target reference boxes of all layers, so that the matching degrees with the target reference boxes of all layers can be obtained, that is, multiple matching degrees.

[0155] S803. Determine the collapse type based on multiple matching degrees.

[0156] S803 can be implemented as: determine the collapse type corresponding to the target reference box with the highest matching degree as the collapse type of this layer.

[0157] S804. Determine the collapse position based on the position of the predicted detection box.

[0158] S804 can be implemented as: read the position of the predicted detection box, and determine the position of the predicted detection box as the determined collapse position.

[0159] S805. Determine that the classification result of each layer includes the collapse type and the collapse position.

[0160] S805 can be implemented as: determine that the classification result of each layer includes the collapse type and the collapse position corresponding to the predicted detection box of this layer identified. Each layer here can include one or more or no predicted detection boxes, corresponding to one place, multiple places or no collapse of a certain type. However, all layers have at least one predicted detection box. If there are no predicted detection boxes in all layers, it is considered that there is no potential collapse, and the following processing is urgently needed, continue to detect spatial data through sensors and continue to monitor.

[0161] Based on the above technical means, when determining the collapse type, it is determined based on the matching degree between the predicted detection box and the target reference box of each layer. In this way, since the target reference box of one layer can represent the characteristics of a certain type of collapse type, the determination method based on the matching degree is more accurate in determining the collapse type. When determining the collapse position, since the predicted detection box is used to indicate the collapse position, it is more accurate to determine the collapse position through the position of the predicted detection box.

[0162] The processing method provided by the embodiments of the present application may include but is not limited to the process of determining the target reference box of each layer.

[0163] Reference Figure 9 Referring to the content shown, the process may include but is not limited to the following S901 to S904.

[0164] S901. Obtain a sample data set.

[0165] The sample data set includes multiple sample space data, and each sample space data is annotated with a ground truth box; the ground truth box is used to point to the collapse position in the sample space data.

[0166] The collapse type may also be annotated in the sample space data. The annotation here may be manual annotation, or a trustworthy annotation that has been machine-annotated and then verified manually.

[0167] The embodiments of the present application do not limit the number of sample space data included in the sample data set, which can be configured according to actual needs. Among them, the larger the number of sample space data, that is, the richer the sample data set, the more accurate the result obtained.

[0168] After annotating the multiple sample space data of the sample data set, S901 can be implemented as reading each sample space data in the annotated sample data set.

[0169] S902. Allocate the multiple ground truth boxes in the sample data set to multiple layers according to the size, so that each layer includes multiple ground truth boxes.

[0170] For example, configure the collapse corresponding to the crack type to the first layer, the collapse corresponding to the transverse collapse type to the second layer, and the collapse corresponding to the longitudinal collapse type to the third layer. Each layer has a size range.

[0171] In a possible implementation manner, S902 can be implemented as: determining the collapse type corresponding to each layer, and the size range corresponding to the collapse type. Then, according to the size range, allocate the multiple ground truth boxes in the sample data set to multiple layers, so that each layer includes multiple ground truth boxes.

[0172] Of course, the multiple ground truth boxes can also be allocated to different layers according to the collapse type.

[0173] S903. For each layer among the multiple layers, based on the multiple ground truth boxes included in the layer, determine the center of the target reference box of the layer.

[0174] The embodiments of the present application do not limit the method for determining the center of the target reference box, which can be configured according to actual needs. For example, the average value or the median of the centers of all ground truth boxes can be determined as the center of the target reference box.

[0175] S903 can be implemented as follows: for each of multiple layers, based on the multiple ground truth boxes included in that layer, the centers of the multiple ground truth boxes are calculated through an algorithm for determining the center, and the center of the target reference box for that layer is obtained.

[0176] S904. Determine the target reference box for that layer based on the center of the target reference box for that layer.

[0177] In a possible implementation manner, S904 can be implemented as: based on the center of the target reference box of the layer, generate a target reference box for that layer.

[0178] In another possible implementation manner, S904 can be implemented as: based on the center of the target reference box of the layer, generate a candidate reference box for that layer, and then continuously optimize and adjust to obtain the target reference box for that layer.

[0179] Based on the above technical means, the sample space data in the sample dataset can be divided into different layers according to the sizes of the ground truth boxes, so as to determine the centers of the target reference boxes for each layer according to the sizes of the ground truth boxes in different layers, and then determine the target reference boxes for each layer. In this way, the determined target reference box for each layer can meet the size requirements of a type of collapse type, and the recognition accuracy of that layer for that collapse type is improved.

[0180] Next, the process of determining the center of the target reference box for that layer based on the multiple ground truth boxes included in that layer in S903 will be described. Refer to Figure 10 the content shown. This process may include but is not limited to the following S1001 to S1003.

[0181] S1001. Determine the center points of the multiple ground truth boxes based on the multiple ground truth boxes of that layer.

[0182] S1001 can be implemented as: based on the multiple ground truth boxes of that layer, first determine the coordinates of each ground truth box, and then based on the true coordinates of each ground truth box, determine the set center of each ground truth box through the set center point determination algorithm, so as to obtain the coordinates of the center point of each ground truth box.

[0183] S1002. Respectively determine the probability density of the center point of each ground truth box of that layer.

[0184] The concept density is used to cluster the ground truth boxes of that layer.

[0185] S1002 can be implemented as: for each ground truth box of that layer, based on the coordinates of each feature point in that ground truth box, determine the probability density of the center point of each ground truth box in that ground truth box. The embodiments of the present application do not limit the determination method of the probability density, and it can be configured according to actual needs.

[0186] S1003. Determine the center point with the highest probability density as the center of the target reference box.

[0187] S1003 can be implemented as: determine the center point with the highest probability density as the center of the target reference box. If there are multiple center points with the highest probability density, any one of them can be selected as the center of the target reference box. If the probability density values of the first several ones are all relatively high and similar, the center points with multiple high probability densities can also be used as the center of the target reference box.

[0188] Based on the above technical means, when determining the center of the target reference box of a layer, it is determined based on the probability density, which can meet the requirements of most real boxes and is more accurate.

[0189] Next, the process of determining the target reference box of the layer based on the center of the target reference box of the layer in S904 will be described. Refer to Figure 11 the content shown. This process may include but is not limited to the following S1101 to S1104.

[0190] S1101. Generate a candidate reference box with the center point of the target reference box as the center.

[0191] S1101 can be implemented as: generate a candidate reference box with the center point of the target reference box as the center. The candidate reference box generated here can be determined based on any real box. For example, in the case of corresponding fractures in the first layer, after determining the center point, a long strip shape is generated with the center point as the center as the candidate reference box.

[0192] Here, for each layer, the shapes of the generated candidate reference boxes can be the same or different. Even if the shapes of the initially determined candidate reference boxes for each layer are the same, with subsequent continuous adjustment and optimization, the shapes of the final target reference boxes for each layer are also different.

[0193] S1102. Respectively determine the dynamic distance between each real box of the layer and the candidate reference box.

[0194] The embodiments of the present application do not limit the method for determining the dynamic distance, which can be configured according to actual needs. For example, the dynamic distance is negatively correlated with the overlap degree and positively correlated with the similarity of the aspect ratio.

[0195] S1102 can be implemented as: respectively determine the dynamic distance between each real box of the layer and the candidate reference box according to the determination algorithm or formula of the dynamic distance, so as to obtain multiple dynamic distances.

[0196] Then determine whether the multiple dynamic distances meet the preset conditions.

[0197] S1103. Adjust the candidate reference box based on the dynamic distance until the dynamic distance meets the preset condition.

[0198] The embodiments of the present application do not limit the preset condition, which can be configured according to actual needs.

[0199] For example, the preset condition may include but is not limited to any of the following: the dynamic distance to each ground truth box is less than or equal to the first threshold; the dynamic distance to most ground truth boxes is less than or equal to the first threshold, that is, the number greater than or equal to the first threshold is greater than the first quantity threshold, that is, the existence of individual interferences is allowed; or, the adjustment of all sample space data has been completed.

[0200] S1103 can be implemented as: based on the dynamic distance, adjust the size and shape dimensions in the candidate reference box, and then use the adjusted candidate reference box as the new candidate reference box to recalculate the dynamic distance, and so on in a loop until the dynamic distance meets the preset condition.

[0201] S1104. Determine the candidate reference box that meets the preset condition as the target reference box of this layer.

[0202] S1104 can be implemented as: when it is determined that the preset condition is met, determine the candidate reference box that meets the preset condition as the target reference box of this layer.

[0203] Based on the above technical means, when determining the target reference box, first generate a candidate reference box based on the center point, and then continuously adjust the shape and size of the candidate reference box based on the dynamic distance, so that the preset condition is met between the candidate reference box and the ground truth box of this layer. In this way, the determined target reference box can have a large overlap degree with the ground truth box, and the aspect ratio between the target reference box and the ground truth box is close, that is, the position and shape of the target reference box are close to the ground truth box, thus meeting the actual needs and improving the accuracy.

[0204] The processing method provided by the embodiments of the present application can also update the target reference box of each layer.

[0205] Refer to Figure 12 the content shown in

[0206] S1201. For each ground truth box of the layer, respectively determine the matching degree between the ground truth box and the target reference box of the layer.

[0207] S1201 can be implemented as: for each ground truth box of the layer, respectively determine the matching degree between the ground truth box and the target reference box of this layer based on the matching degree algorithm, so as to obtain multiple matching degrees.

[0208] S1202. Adjust the layer to which the ground truth box belongs based on the matching degree.

[0209] S1202 can be implemented as: determining the true bounding boxes with a matching degree less than the threshold as the true bounding boxes to be adjusted, and then deleting the true bounding boxes to be adjusted from this layer. In this way, the true bounding boxes included in each layer will be redistributed and adjusted based on the matching degree.

[0210] S1203. Update the target reference box of each layer based on the true bounding boxes included in each adjusted layer.

[0211] The process of updating the target reference box of each layer is similar to the process of determining the target reference box of each layer. The implementation of S1203 can refer to the descriptions of S903 to S904, and will not be elaborated here one by one.

[0212] Based on the above technical means, after stratifying each true bounding box, the allocation method can also be adjusted based on the matching degree. This can optimize the stratification, improve the accuracy of the division, and also improve the accuracy of collapse type recognition.

[0213] Next, taking an autonomous vehicle as an example, the processing process of road surface collapse will be described through an embodiment.

[0214] With the increasing construction of roads, road safety hazards have become more and more prominent. Among them, road surface collapse is a very critical and common problem, bringing great harm to people's lives and property. In 2020, a serious road surface collapse accident occurred in a certain city, causing multiple casualties and attracting strong social attention. In 2024, a road surface collapse disaster occurred on a certain highway, resulting in a particularly serious accident with 48 deaths and direct economic losses of more than 100 million yuan. According to the statistical data, the road surface collapse accidents in our country are on the rise year by year, occurring frequently, with an average annual growth rate of 81%. In order to effectively protect people's lives and property, on the one hand, analyze the causes of road surface collapse and take prevention and control measures for traffic roads, and on the other hand, more importantly, identify and detect natural disasters such as road surface collapse from the vehicle-end automatic driving aspect and make intervention decisions in advance.

[0215] The current mainstream target recognition and detection technology in the field of autonomous driving uses cameras, lidar, and sensors to obtain the real scenes of the surrounding environment, and conducts real-time target detection and analysis on these real scenes. These real scenes mainly include instances such as people, vehicles, and objects for tracking, calibration, and decision-making, but there is no detection and recognition method for special long-tail scenes in the surrounding environment (such as disasters like road surface collapse, road surface fault, and bridge fracture). At present, basically all are detection methods and maintenance plans around the analysis of the causes of preventing traffic road surface collapse.

[0216] For example, in Related Technology 1, a method and system for detecting abnormal conditions on expressways are provided. The method includes: constructing a detection management platform, setting detection devices segment by segment along the expressway, and connecting each detection device to the detection management platform to form a monitoring network; using the detection devices to monitor whether there are abnormal conditions on the corresponding sections of the expressway; if an abnormal condition is found, traffic warning information is sent to each section through the monitoring network.

[0217] Related Technology 2 provides a method for detecting and identifying road surface collapses, aiming to solve the problem of being able to report the detected section of underground collapse in the fastest speed in real time, so as to report potential hazards in the first time and prevent problems before they occur.

[0218] Related Technology 3 provides a method and device for remotely controlling an unattended autonomous driving operating vehicle. The aim is to transmit the emergency event information to a remote operation control platform when an emergency event occurs; receive the countermeasures feedback by the remote operation control platform according to the emergency event information, and execute the countermeasures. The aim is to solve the problem that when long-tail scenarios occur, traditional remote systems and personnel cannot conduct emergency response to emergency events, thus triggering the driving risks of autonomous driving vehicles and hindering the development and application of unattended autonomous driving vehicles.

[0219] Through the detection of abnormal conditions on the road, in case of an emergency event, the remote system is used for information feedback, and the control platform is manipulated to intervene in the autonomous driving vehicle, without identifying and intervening from the autonomous driving itself. It can be seen from this that the current technical solutions based on road surface collapse prediction are all for reliability monitoring of the basic traffic road surface itself, and do not conduct road surface safety and reliability identification and decision-making from the vehicle-end autonomous driving field. Simply relying on the reliability detection of road surface diseases based on the traffic road surface cannot fully guarantee the harm of such traffic accidents to the safety of people's lives and property.

[0220] This embodiment mainly uses the object detection algorithm in the vehicle-end autonomous driving field to identify and detect long-tail scenarios such as road surface collapses and provide decisions. To achieve the safe application of autonomous driving in the actual traffic environment, it is urgent to continuously optimize the autonomous driving algorithm to ensure that effective emergency response measures are taken in the case of 10% long-tail scenarios, which can largely avoid or reduce personal and property losses and ensure the safe driving of vehicles.

[0221] To address this technical problem, the technical solution and working principle of this embodiment are as follows. A road surface collapse identification and detection scheme based on an autonomous driving system is proposed, aiming to address one of the technical problems of continuously optimizing the autonomous driving algorithm to ensure effective emergency response measures in the case of 10% long-tail scenarios to a certain extent.

[0222] To this end, the first objective of this embodiment is to propose a road surface collapse recognition and detection algorithm based on an autonomous driving system, including but not limited to the following steps 1 to 3.

[0223] Step 1: Perform environmental perception based on vehicle vision perception devices (cameras and lidar) to obtain environmental perception point cloud data and image data; collect images of road surface conditions (such as road surface collapses, road surface faults, block cracks, bridge fractures, etc.) on several different roads containing various road diseases to form a sample image set, and use the deep learning framework PyTorch to fuse the images and point cloud data to establish an atlas that clarifies the specific conditions such as the location of road diseases (equivalent to the above-mentioned collapse location) and type (equivalent to the above-mentioned collapse type) as the road disease standard data set (equivalent to the above-mentioned sample data set).

[0224] Step 2: Select the images of road surface lesions marked through the sample annotation step as the training data set of the network model, set the loss function, optimizer, and hyperparameters, construct a deep learning model based on the convolutional neural network of the third version of the lightweight neural network (Mobile NetV3), and complete the model training based on the deep learning framework TensorFlow and the neural network library Keras in the Python 3.6 environment.

[0225] Step 3: Extract the quantitative features of road surface collapses from the images in the road disease segmentation image data set and obtain the quantitative feature extraction results of the collapse area. Input the preprocessed data of the road surface collapse images into the trained deep learning model based on Mobile NetV3, obtain the prediction output through inference, and perform model performance verification.

[0226] Reference Figure 13 As shown in the content, the road surface collapse treatment process can include but not limited to the following S1301 to S1315.

[0227] S1301. Collect image data of road surface collapse diseases.

[0228] S1302. Establish a road disease image data set.

[0229] S1303. Preprocess the road disease image data.

[0230] S1304. Set the loss function and hyperparameters of the network model.

[0231] S1305. Train the road surface collapse environment model based on MobileNetV3.

[0232] S1306. Input the road surface image data set through radar and camera.

[0233] S1307. Identify using the road surface collapse environment model.

[0234] S1308. Extract the quantitative parameter characteristics of road surface collapse diseases.

[0235] S1309. Adopt the classical crack - type disease classification algorithm.

[0236] S1310. Classify and save the information of road surface collapse diseases.

[0237] S1311. Road environment model based on autonomous driving.

[0238] S1312. Determine whether there is an emergency event (equivalent to a collapse event).

[0239] If yes, that is, there is an emergency event, execute the following S1313 to S1315; if no, that is, there is no emergency event, execute the above S1307.

[0240] The emergency event here is equivalent to a collapse event.

[0241] S1313. The vehicle makes a decision and executes corresponding measures.

[0242] S1314. Through vehicle - to - everything (V2X) wireless communication technology, prompt early warning information to nearby vehicles.

[0243] S1315. Upload to the vehicle - road - cloud integrated system through V2X wireless communication technology.

[0244] The implementation of step 1 is described in detail below, including but not limited to steps 1.1 to 1.3.

[0245] Step 1.1: To facilitate the unified pre - processing of the acquired environmental perception point cloud data and image data, it is necessary to use the deep - learning framework pytorch to fuse the image and point cloud data to generate a more abundant data set.

[0246] Step 1.2: Use the annotation software labelimg to mark the images of the road disease standard data set to obtain the range coordinates of the disease area, and perform sample annotation on the disease area according to the classification category and segmentation label. Label the specific situations such as the location and type of the clear road diseases in the sample and divide them into the training set and the validation set.

[0247] Step 1.3: Preprocess the images in the training set and the validation set. Crop, geometrically transform, and perform brightness / contrast / hue transformation on the pavement condition images containing different roads and various road diseases in the sample image set. Add Gaussian noise and salt-and-pepper noise, reconstruct the image size, and adjust the pixel values of the pixel points.

[0248] Next, the implementation of Step 2 will be described in detail.

[0249] Step 2.1: Select the pavement lesion images labeled in the sample annotation step as the training data set of the network model. Set the loss function, optimizer, and hyperparameters, and construct a deep learning model based on the Mobile NetV3 convolutional neural network. In the Python 3.6 environment, first calculate the difference between the predicted value and the true label by defining the loss function, and then use backpropagation gradient descent to update the weights of the network model, so that the predicted value output by the network model is close to the true label value. In this way, an object detection model based on the deep learning framework Tensor flow and the neural network library Keras can be obtained.

[0250] Step 2.1.1: Based on the U-shaped network architecture, use MobileNetV3-Large as the backbone feature network, which is responsible for extracting features from the dataset images. Define the size of the feature maps in each stage mainly according to its actual downsampling steps. The downsampling rate refers to the reduction multiple relative to the input image (e.g., A1 is 1 / 2 of the input size). The output feature maps at different stages are denoted as A1 (equivalent to the first feature map of the first layer above), A2 (equivalent to the first feature map of the second layer above), A3 (equivalent to the first feature map of the third layer above), A4, and A5.

[0251] Assume the input image size is L×L×3. Then the size of A1 is L / 2×L / 2×16, and after multi-stage convolutional pooling operations, the size of A2 is L / 4×L / 4×24, the size of A3 is L / 8×L / 8×40, the size of A4 is L / 16×L / 16×112, and the size of A5 is L / 16×L / 16×160. The processing parameters in the convolutional stage can refer to the content of Table 1.

[0252] Table 1 Example of processing parameters in the convolutional stage

[0253] Then, for each feature map Ai, a Dynamic Multi-Kernel Collaborative Receptive Field Boosting (DMKC-RFB) module is designed. Input Ai (i = 1~5), and then use DMKC-RFB to enhance the feature map extracted in the previous step to output enhanced feature maps B1 (equivalent to the second feature map of the first layer above), B2 (equivalent to the second feature map of the second layer above), B3 (equivalent to the second feature map of the third layer above), B4, and B5, with the same size as the input. Then, use feature fusion to fuse the enhanced feature maps from the previous step into C1 (equivalent to the third feature map of the first layer above), C2 (equivalent to the third feature map of the second layer above), C3 (equivalent to the third feature map of the third layer above), C4, and C5; finally, use the fused feature maps to predict the target.

[0254] Step 2.1.2: The feature-enhanced receptive field uses the DMKC-RFB module to reduce the model's computational load and improve the model's feature extraction ability.

[0255] The process of feature enhancement includes: inserting the DMKC-RFB module (equivalent to the attention mechanism above) after each level Ai to generate enhanced feature maps B1, B2, B3, B4, and B5.

[0256] Input branch division: Each Ai is processed in 3 parallel branches, and the branch parameters are adaptively adjusted according to the size of the feature map: Branch 1: 3×3 deformable convolution (dilation rate 1); Branch 2: 3×3 deformable convolution (dilation rate 3); Branch 3: 5×5 deformable convolution (dilation rate 5).

[0257] In DMKC-RFB, the design of the three parallel branches aims to adaptively capture multi-scale feature information through deformable convolutions of different scales, while balancing computational efficiency and model performance.

[0258] The parameters of each branch in DMKC-RFB can refer to the content of Table 2.

[0259] Table 2 Parameter examples of each branch in DMKC-RFB

[0260] The collaboration process can include: Input feature map division: Input the feature maps (Ai) of the same level into three branches for parallel processing. Multi-scale feature extraction: Branch 1 focuses on local details through a small receptive field; Branch 2 captures surrounding context through a medium receptive field; Branch 3 covers the global area through a large receptive field.

[0261] Dynamic weighted fusion and residual connection: The lightweight attention module (Squeeze-and-Excitation, SE) dynamically assigns weights according to channel importance. The fusion result is added to the input feature map to retain the original information.

[0262] Dynamically adjust the weights of each branch through SE. The following formula (1) can be referred to.

[0263] Formula (1); In formula (1), is the channel attention weight generated by the SE module, represents the features of each layer after convolution processing, represents the enhanced features of each layer; represents the deformable convolution operation.

[0264] Step 2.1.3: Feature fusion adopts a multi-scale feature fusion expression. Based on the bidirectional feature pyramid, an adaptive cross-layer weight mechanism is added. For the fusion of the top-down path (equivalent to the first path above), the following formulas (2-1) to (2-5) can be referred to. For the fusion of the bottom-up path (equivalent to the second path above), the following formulas (3-1) to (3-5) can be referred to.

[0265] Formula (2-1); Formula (2-2); Formula (2-3); Formula (2-4); Formula (2-5); In formulas (2-1) to (2-5), , , , Fusion weights, , , , The initialization can be 1.0. , , , , Are the fusion features after fusing every two layers (equivalent to the intermediate feature maps above). : Bilinear interpolation upsampling by 2 times. Represents a convolution operation with a size of 1 1.

[0266] Formula (3-1); Formula (3-2); Formula (3-2); Formula (3-4); Formula (3-5); In Formulas (3-1) to (3-5), 、 、 、 are the fusion weights, 、 、 、 and the initial value can be 1.0. 、 、 、 、 are the fusion features after fusion of every two layers respectively. : Max-pooling downsamples by a factor of two. represents a convolutional operation with a size of 3 ×3.

[0267] Among them, , are learnable weight parameters, and the weight parameters are automatically updated through backpropagation without manual design, which is used to dynamically balance the importance of features with different resolutions.

[0268] In the object detection model, C1 to C5 are the feature maps after multi-scale fusion, corresponding to different levels of downsampling rates (e.g., C1 has a downsampling rate of 2, and C5 has a downsampling rate of 16). Through the hierarchical density-aware K-means++ algorithm, the appropriate anchor box sizes are generated for each level.

[0269] I. Object Hierarchical Assignment (Preprocessing Stage).

[0270] Function: Assign the real object boxes in the dataset to the corresponding level feature maps (C1 - C5) according to their sizes.

[0271] For example: Small objects (such as cracks, size 20×20 pixels): Matched to C1 (downsampling rate 2, original image anchor box size 40×40).

[0272] Large objects (such as collapses, size 160×160 pixels): Matched to C5 (downsampling rate 16, original image anchor box size 160×160).

[0273] II. Run K-means++ Independently for Each Level

[0274] For each level (C1 - C5), perform clustering on the assigned target boxes independently to generate the anchor box sizes for that level.

[0275] 1. Input data preparation: The set of target boxes for level i: Only contains the ground truth target boxes assigned to that level. For example: The C3 layer processes all medium - sized target boxes that match this layer.

[0276] Assign the targets to the corresponding levels according to the matching degree between the target box sizes and the anchor box sizes of each layer.

[0277] The matching degree can refer to the following formula (4).

[0278] P = Formula (4); In formula (4), P represents the matching degree, represents the predicted anchor box size (width × height). represents the size of the ground truth target box in the dataset (width × height).

[0279] IoU represents the Intersection over Union, which measures the overlapping degree of two boxes. max(Sanchor, Starget) represents the maximum area between the anchor box and the target box (for normalization).

[0280] 2. Predicted Box Definition: The target bounding box output by the model, representing the detected target position and size by the algorithm.

[0281] Generation process: Based on the anchor box: The model (V3) generates the predicted box through the anchor box. The coordinates of the predicted box are the adjusted results of the offsets of the anchor box. Regression parameters: The model predicts the offset of the center point and the scaling factors of width and height .

[0282] For example, if the anchor box size is 64x64 and the model predicts the offsets (Δx = 0.1, Δy = 0.2, Δw = 1.2, Δh = 0.8), then the predicted box size is 64x1.2 = 76.8 (width) and 64x0.8 = 51.2 (height).

[0283] The receptive field (RF) is the area of the input image corresponding to a point on the feature map, representing the field of view range that this point can "see" in the input image.

[0284] The parameters corresponding to the receptive fields of each layer can refer to Table 3.

[0285] Table 3 Parameter examples corresponding to the receptive fields of each layer

[0286] Among them, the receptive field of the low-level layer (such as C1) is small and suitable for detecting details (such as crack edges). The receptive field of the high-level layer (such as C5) is large and suitable for integrating global information (such as the collapsed area).

[0287] Calculate the density map of the target distribution, and preferentially select the center point of the high-density area as the initial clustering center.

[0288] The role of the density map: Quantify the distribution density of the targets in the dataset in different regions, guide the anchor box generation process to focus on the high-density areas, and improve the adaptability of the anchor boxes to the dense targets.

[0289] Hierarchical processing: Each layer (C1 - C5) independently calculates its corresponding density map, only considering the target boxes assigned to that layer. For example, the C1 layer (downsampling rate 2) processes small targets (such as cracks), and its density map reflects the dense areas of cracks in the image. The C5 layer (downsampling rate 16) processes large targets (such as collapses), and its density map reflects the distribution area of large-scale collapses.

[0290] The calculation of the hierarchical density map can include: Input data: The center coordinates of all target boxes assigned to that layer. Hierarchical adaptability: Small target layer (C1): The target distribution may be concentrated in local areas (such as dense cracks at the road edge). Large target layer (C5): The target distribution may be more dispersed or cover a wide area (such as the collapsed area). Parameter adjustment: Different Gaussian kernel bandwidths σ are used for the density maps of different layers to adapt to the target scale: Small target layer (C1): σ is smaller (such as σ = 2) to capture local density. Large target layer (C5): σ is larger (such as σ = 8) to smooth the global distribution. Hierarchical division of labor: C1 - C5 are responsible for detecting targets of different scales, and the density map ensures that the anchor box generation of each layer adapts to the distribution characteristics of the targets in that layer.

[0291] The density map guides the clustering initialization, making the anchor boxes more conform to the target shapes and sizes in the actual dense areas.

[0292] The necessity of hierarchical processing: The receptive fields of different hierarchical feature maps are different, and different scale targets need to be matched. Example: The high resolution of C1 is suitable for detecting thin cracks, and the low resolution of C5 is suitable for large-scale collapses.

[0293] The advantages of dynamic distance metric: Optimize both the position overlap (IoU) and shape similarity (aspect ratio) simultaneously, and avoid the scale sensitivity problem of the traditional Euclidean distance. The value of density-sensitive initialization: Generate anchor boxes for disease-dense areas, and improve the recall rate of the model for small targets or dense targets.

[0294] Dataset: It contains 1000 road surface images, with three types of targets annotated: cracks, potholes, and collapses. Target hierarchical allocation: Cracks (20×5 pixels) → Layer C1 (s = 2). Potholes (80×80 pixels) → Layer C3 (s = 8). Collapses (160×160 pixels) → Layer C5 (s = 16). Hierarchical clustering: Layer C1: Run K-means++ on cracks to generate slender anchor boxes (such as 24×6). Layer C3: Run K-means++ on potholes to generate square anchor boxes (such as 64×64). Layer C5: Run K-means++ on collapses to generate large-sized anchor boxes (such as 160×160).

[0295] Application of anchor boxes: During training, the model performs object regression and classification based on the anchor boxes at each level. During inference, the IoU is calculated between the anchor boxes and the predicted boxes, and the detection results are output.

[0296] Dynamic distance metric: The dynamic distance metric calculates the distance between the ground truth box and the candidate anchor box (predicted box), which is used to measure the matching degree between the two. Ground truth box: The object bounding box manually annotated in the dataset, which contains the position (center point coordinates) and size (width and height). Candidate anchor box: The anchor box to be optimized generated during the clustering process, and its size is iteratively adjusted to match the distribution of the ground truth box.

[0297] The determination of the dynamic distance can refer to the following formula (5).

[0298] Formula (5); In formula (5), represents the dynamic distance between b and c, γ is an adjustable parameter that controls the aspect ratio weight; b represents the ground truth box (with width and height of ); c represents the candidate anchor box (with width and height of ).

[0299] The dynamic distance formula optimizes the generation of anchor boxes by combining two metrics: The IoU term measures the degree of position overlap between the ground truth box and the candidate anchor box. The larger the IoU, the smaller the distance. Ensure that the anchor box covers the area of the ground truth object as much as possible. Aspect ratio similarity term: Measure the shape similarity through the arctangent difference of the aspect ratio, avoiding relying solely on the area. Optimize the shape of the anchor box to make it more conform to the aspect ratio distribution of the object (such as slender cracks or square potholes).

[0300] In the hierarchical density-aware K-means++ algorithm, the dynamic distance metric is the core step of the clustering process, and its front and back logic includes: Preceding steps: Target hierarchical allocation: Allocate the ground truth boxes to the corresponding levels (C1 - C5) according to their sizes. Density-sensitive initialization: Select the initial clustering centers based on the density map, focusing on high-density regions.

[0301] Dynamic distance metric steps: Input: True target boxes assigned to this level + initial clustering centers (candidate anchor boxes). Calculation: For each true target box, calculate its dynamic distance to all candidate anchor boxes. Assignment: Assign the true target box to the cluster to which the clustering center with the minimum distance belongs.

[0302] Subsequent steps: Update clustering centers: Adjust the anchor box sizes according to the mean and shape of the target boxes within the cluster. Iterative optimization: Repeat distance calculation, assignment, and update until convergence. Anchor box restoration: Restore the clustering results to the original image scale according to the level downsampling rate. Match with the receptive field of the feature map.

[0303] Scenario: Detect cracks (small targets) and collapses (large targets) in the image. Target hierarchical assignment: Cracks (20×5 pixels) are assigned to layer C1 (downsampling rate 2), and collapses (160×160 pixels) are assigned to layer C5 (downsampling rate 16).

[0304] Hierarchical clustering output: Layer C1: Generate anchor boxes (16×4, 24×6) through dynamic distance metric, and the original image scale is (32×8, 48×12). Layer C5: Generate anchor boxes (160×160, 192×192).

[0305] Detection process: Prediction for layer C1: Based on the anchor box 32×8, regress the crack position, and output the prediction box (30, 80, 38, 88). Prediction for layer C5: Based on the anchor box 160×160, regress the collapse position, and output the prediction box (200, 300, 360, 460).

[0306] Set the loss function as the cross-entropy loss function.

[0307] The cross-entropy loss function can refer to the following formula (6).

[0308] Formula (6); In formula (6), represents the cross-entropy, is the input vector sample, is the true label of the sample, N represents the number of samples, is the network weight parameter, by adjusting the weight parameter , the loss function can be minimized, thereby improving the recognition accuracy of the model.

[0309] Set the optimizer and hyperparameters: Adopt the Adam optimizer, set the hyperparameter batch size to 4, learning rate to 1e-3, number of iterations to 300, adjustment factor 0.75, and take the default values for weight decay and initial momentum.

[0310] During the model training process, first, the training set is used as the input, and average pooling operation and regularization techniques are introduced. The deep learning model is trained according to the settings of hyperparameters. The training error is calculated using the label tensor and the network output tensor according to the loss function. The optimizer updates the model parameters using the training error, and the recognition effect of the model is verified on the validation set. When the performance of the model reaches a stable state, that is, both the loss function value and the recognition accuracy of the validation set tend to converge, the parameters of the MobileNetV3 model with the best performance on the validation set are saved.

[0311] The implementation process of step 3 is described below.

[0312] The model generates pixel-level segmentation masks through an encoder-decoder structure (such as MobileNetV3 as the encoder and UNet as the decoder). The specific form is as follows: Probability Map: Output a two-dimensional matrix (e.g., H×W) with the same size as the input image. Each pixel value represents the probability that the position belongs to the collapsed area (ranging from [0,1]). By setting a threshold (e.g., 0.5), the probability map is converted into a binary label of the collapsed area. 1 (white pixel): Collapsed area. 0 (black pixel): Non-collapsed area.

[0313] Input: Preprocessed road surface image (size adapted to the model input, such as 224×224×3). Model Inference: Output probability map (224×224×1). Post-processing: Generate a binary mask through the threshold to extract the collapsed area. Performance Verification: Calculate metrics such as IoU (e.g., IoU = 0.85). Feature Extraction: Quantify parameters such as the collapsed area, position, and shape (e.g., area = 1500 pixels, centroid = (112,89)).

[0314] The predicted output results for road surface collapse are used to extract the connected components in each image using the connected component separation function, remove the crack disease recognition results where the connected components are less than a certain number of pixel points, judge the positional relationship of the connected components of the road surface collapse disease, and perform bounding box annotation on the collapsed area.

[0315] The classic crack disease classification algorithm is used to classify the road surface collapse diseases in the processed images into longitudinal road surface collapse, transverse road surface collapse, and road surface cracks. Geometric information quantification evaluation is carried out for different types of road surface collapse diseases. The geometric information quantification indicators mainly include: mAP@0.5 represents the mean average precision when the IoU threshold is 0.5, comprehensively reflecting the detection accuracy.

[0316] Recall@0.5 represents the recall rate of the collapsed targets, measuring the missed detection situation.

[0317] The FPS represents the inference speed of the model on the deployed hardware (such as NVIDIA Jetson AGX Xavier).

[0318] Finally, a road environment model based on autonomous driving is obtained. When a moving vehicle processes environmental perception data, it determines whether there is an emergency of road collapse according to the detection and recognition of the road collapse environment model. When it is determined that there is an emergency, the autonomous driving algorithm will make decisions and execute corresponding measures, including emergency braking, lane change, and defensive driving, etc.

[0319] To improve the inference efficiency of the road collapse target detection and recognition network on the FPGA, a fully convolutional layer network is adopted, that is, the network model does not include fully connected layers, which can greatly reduce the number of network model parameters and further reduce data exchange. At the same time, a convolutional attention module is added after each convolutional layer, which can enhance the feature extraction ability and compress unimportant feature information, so as to quickly identify target information and effectively improve the accuracy of target recognition.

[0320] The second objective of this embodiment is to propose a field-programmable gate array (FPGA) chip system and an early warning mechanism for the road collapse recognition and detection algorithm based on an autonomous driving system, including the following steps: An FPGA chip system for the road collapse recognition and detection method based on an autonomous driving system is modularly designed from top to bottom based on the FPGA platform, and a hardware description language (Verilog HDL) is used for logic circuit design. After the definition of the top-level module design is completed, the logical function definitions of each sub-module are further divided.

[0321] Reference Figure 14 As shown in the content, the FPGA hardware chip system includes an image acquisition module 1401, a data communication module 1402, a cache module 1403, a display control module 1404, and an FPGA hardware accelerator 1405.

[0322] The logical relationships of each module are as follows: After the system power-on reset is completed, the FPGA searches for the internal registers of the image acquisition module through the register lookup table and configuration. The vehicle-mounted lidar and camera output video image data to the FPGA through the clock signal and control signal. The central processing unit (CPU) receives and parses the trained network weight parameters, and then sends them to the FPGA through the data communication module. Then, the communication data packet is parsed. Since the weight parameters need to be sorted and spliced before the convolution operation to facilitate the transmission and calculation of custom bit-width data, the parsed weight parameters are cached in the cache center, which also avoids the interaction between the FPGA and external storage and reduces the data transmission time, thereby improving the data calculation efficiency. When all network parameters are transmitted, the frame read-write control module starts to cache the video images into two different storage spaces of double data rate (DDR) 3 in sequence. During the DDR3 caching process, a ping-pong operation of the storage space is designed to improve the read-write efficiency. The video images read from the DDR3 are displayed on the display screen in real time through the display control module, and at the same time, the video images are transmitted to the FPGA hardware accelerator. Since the input feature map of the network model needs to be compressed, the video images are preprocessed, and the reduced images are temporarily stored in the block random access memory (BRAM) cache center. The FPGA hardware accelerator reads the image data and network parameters to extract the features for road collapse recognition, and marks the detection results on the display screen in real time.

[0323] Reference Figure 15 As shown in the content, the FPGA hardware accelerator includes a main control module 1501, a data cache module 1502, a convolution calculation module 1503, a zero-padding module 1504, and an average pooling module 1505.

[0324] The main control module includes an enable signal generator and an address controller, which are responsible for re-sorting and splicing the unsorted and unspliced weight parameters and bias parameters in the BRAM cache center according to the network calculation process, and then writing them into other corresponding BRAM data cache modules to realize the control of the enable signal and address signal for each module.

[0325] The data cache module is used to cache the input image data, intermediate feature map data, weight parameters, and bias parameter data respectively. When the number of input image BRAM block input feature map channels N is greater than 8, the main control module reads the image data and parameters from the intermediate feature map BRAM block, weight parameter BRAM block, and bias parameter BRAM block in batches according to the process of depth convolution.

[0326] The convolution calculation module mainly extracts image features through convolution operations, including a general-purpose convolution calculation engine, a standard convolution controller, a depth convolution controller, and a pointwise convolution controller. The general-purpose convolution calculation engine is applicable to the operations of all convolutional layers in the network model. It uses 48 resources of 24 Digital Signal Processors (DSPs) to perform convolution operations. In one clock cycle, it completes the multiply-accumulate calculation of 24 input data based on a shift register using a 3-stage pipeline. For standard convolution, it can complete the calculation of an 8-channel 5×5 convolution window in 3 clock cycles and output the calculation results of 8 channels. As the network process progresses, the number of input and output channels of depth convolution and pointwise convolution continues to increase. By circularly reusing this convolution calculation engine for processing until all input feature maps within the layer are calculated, the resource utilization rate is greatly improved, and the system power consumption is reduced.

[0327] The zero-padding module performs padding operations using the timing of the input image. According to the enable signal of the read-in image, zero-padding operations are respectively performed on the edges of the input feature map. Since the padded pixel values need to be multiplied by the weight parameters in the convolution kernel, in order to reduce the data interaction time and calculation latency, during convolution calculation, only the multiplication operations of the weight parameters of partial edge data and the pixel value 0 are performed, thereby further improving the network calculation speed.

[0328] The average pooling module calculates the local region average value for replacement after all convolutional layers are calculated to reduce the network calculation amount.

[0329] After the initialization of the FPGA hardware accelerator is completed, the enable signal of the main control module is pulled high. First, the preprocessed video point cloud fusion image data is read from the cache module into the convolution calculation module for standard convolution calculation, depth convolution calculation, and pointwise convolution operations. Each output feature map is repeatedly written into the intermediate feature map BRAM block; finally, after the depthwise separable convolution operation ends, the result is output through the average pooling layer.

[0330] An FPGA chip system and early warning mechanism based on an automatic driving system for a road surface collapse recognition and detection algorithm in this embodiment have at least the following beneficial effects: When the recognition algorithm recognizes a road surface collapse, fault, or bridge fracture, the automatic driving system will make a decision and execute corresponding measures, including assisting in reminding the driver, as well as emergency braking, lane change, and defensive driving, etc.; When an emergency such as a road surface collapse or bridge fracture occurs, the geographical location and driving data of the vehicle generate vehicle early warning information; The vehicle early warning information is displayed to the second vehicle through vehicle-to-vehicle V2X communication technology, commercial high-precision maps, etc. It solves the problem that the driver of the rear vehicle cannot timely know the driving intention of the driver of the current vehicle in an emergency with poor visibility conditions such as at night, rain, snow, frost, and fog; It achieves the effect that the driver of the rear vehicle can timely know the driving intention of the driver of the current vehicle, remind the driver to take emergency braking measures, and reduce rear-end collisions, scratches, and the recurrence of serious accidents. When an emergency of road surface collapse occurs, it can also be coordinated and fed back to the highway remote control platform through vehicle-road-cloud integration to timely feedback information and help highway maintenance personnel take measures in time. The target recognition and detection algorithm further reduces data exchange through layer fusion design of normalization, adds normalization and activation functions after each convolution operation to accelerate the convergence of the network model, and greatly improves the parallel computing efficiency of the FPGA chip system.

[0331] In a second aspect, an embodiment of the present application provides a device for dealing with road surface collapse. The device for dealing with road surface collapse can be deployed on vehicle equipment. Refer to Figure 16 As shown in the content, the device 160 for dealing with road surface collapse may include, but is not limited to: an acquisition unit 1601, an identification unit 1602, and an output unit 1603.

[0332] Among them: The acquisition unit 1601 is used to acquire the spatial data of the road surface in front collected by the vehicle; The identification unit 1602 is used to identify and process the spatial data through a road surface collapse recognition model to determine the recognition result of the road surface in front; The recognition result includes a first identifier; The first identifier is used to represent whether there is a collapse on the road surface in front; Among them, the recognition and processing of the road surface collapse recognition model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; The resolution of each target feature map in the multi-layer target feature maps matches the collapse type; Obtain the target reference frame for each layer, and perform prediction processing on each layer of the target feature map based on the target reference frame for each layer to obtain the predicted detection frame for each layer; The size of the predicted detection frame for each layer matches the collapse type; Classify the target features corresponding to the predicted detection frame for each layer through a classification model to obtain the classification result for each layer; Determine the recognition result of the road surface in front based on the classification result for each layer; An output unit 1603, configured to output the type and location of the collapse in the recognition result when the first identifier indicates that there is a collapse on the road surface ahead, so that the vehicle can perform response processing based on the type and location of the collapse; the types of collapse include: road surface fracture, lateral collapse of the road surface, and longitudinal collapse of the road surface.

[0333] In some embodiments, the processing of the road surface collapse 160 may further include a processing unit, and the processing unit is configured to: determine the vehicle's own processing method based on the type and location of the collapse; send a first prompt message, the type of collapse, and the location of the collapse to the vehicle behind the vehicle through vehicle-to-everything communication; the first prompt message is used to remind that there is a collapse ahead; send the type of collapse and the location of the collapse to the cloud.

[0334] In some embodiments, the recognition unit 1602 is further configured to perform: Perform multi-scale convolutional processing on the spatial data through a fully convolutional network, splice the outputs of multiple stages to obtain a multi-layer first feature map; the resolution of each first feature map in the multi-layer first feature map matches the type of collapse; perform detail feature extraction on each first feature map in the multi-layer first feature map based on the attention mechanism respectively to obtain a multi-layer second feature map; perform fusion processing on the multi-layer second feature maps to obtain a multi-layer third feature map; determine that the multi-layer target feature map includes the multi-layer first feature map, the multi-layer second feature map, or the multi-layer third feature map.

[0335] In some embodiments, when the multi-layer includes three layers, the recognition unit 1602 is further configured to perform: perform convolutional processing on the spatial data based on a first downsampling rate to obtain a first feature map of the first layer; perform convolutional processing on the spatial data based on a second downsampling rate to obtain a first feature map of the second layer; perform convolutional processing on the spatial data based on a third downsampling rate to obtain a first feature map of the third layer; determine the multi-layer first feature map based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer; wherein, the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.

[0336] In some embodiments, when the multi-layer first feature map includes the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, the recognition unit 1602 is further configured to perform: The first attention module extracts detailed features from the first feature map of the first layer based on the first receptive field range to obtain the second feature map of the first layer; the second attention module extracts detailed features from the first feature map of the second layer based on the second receptive field range to obtain the second feature map of the second layer; the third attention module extracts detailed features from the first feature map of the third layer based on the third receptive field range to obtain the second feature map of the third layer; based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, the second feature maps of multiple layers are determined; wherein, the first receptive field range is smaller than the second receptive field range, and the second receptive field range is smaller than the third receptive field range.

[0337] In some embodiments, when the second feature maps of multiple layers include the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, the recognition unit 1602 is further configured to perform: Perform concatenated convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer along the first path to obtain the intermediate feature map of three layers; perform concatenated convolution processing on the intermediate feature map of three layers along the second path to obtain the third feature map of multiple layers; the second path is different from the first path.

[0338] In some embodiments, the recognition unit 1602 is further configured to perform: Obtain the target reference box for each layer; respectively determine the matching degree between the predicted detection box and the target reference box of each layer to obtain multiple matching degrees; the matching degree is used to characterize the overlapping degree between the predicted detection box and the target reference box; determine the collapse type based on the multiple matching degrees; determine the collapse position based on the position of the predicted detection box; determine that the classification result of each layer includes the collapse type and the collapse position.

[0339] In some embodiments, the processing 160 of road surface collapse may further include a preprocessing unit, and the preprocessing unit is configured to perform: Obtain a sample data set; the sample data set includes multiple sample space data, and each sample space data is labeled with a ground truth box; the ground truth box is used to point to the collapse position in the sample space data; allocate the multiple ground truth boxes in the sample data set to multiple layers according to the size, so that each layer includes multiple ground truth boxes; for each layer in the multiple layers, based on the multiple ground truth boxes included in the layer, determine the center of the target reference box of the layer; based on the center of the target reference box of the layer, determine the target reference box of the layer.

[0340] In some embodiments, the preprocessing unit is further configured to perform: Based on the multiple ground truth boxes of the layer, determine the center points of the multiple ground truth boxes; respectively determine the probability density of the center point of each ground truth box of the layer; determine the center point with the highest probability density as the center of the target reference box.

[0341] In some embodiments, the preprocessing unit is further configured to perform: generating a candidate reference box centered at the center point of the target reference box; respectively determining the dynamic distances between each ground truth box of the layer and the candidate reference box; the dynamic distance is negatively correlated with the overlap degree and positively correlated with the aspect ratio similarity; adjusting the candidate reference box based on the dynamic distance until the dynamic distance meets a preset condition; and determining the candidate reference box that meets the preset condition as the target reference box of the layer.

[0342] In some embodiments, the preprocessing unit is further configured to perform: for each ground truth box of the layer, respectively determining the matching degree between the ground truth box and the target reference box of the layer; adjusting the layer to which the ground truth box belongs based on the matching degree; and updating the target reference box of each layer based on the ground truth boxes included in each adjusted layer.

[0343] In a third aspect, the present application provides a system-on-chip, which is connected to the controller of the vehicle; The system-on-chip is configured to: obtain the spatial data of the road surface ahead collected by the vehicle; perform identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the road surface ahead; the identification result includes a first identifier; the first identifier is used to represent whether there is a collapse on the road surface ahead; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; obtaining the target reference box of each layer, and performing prediction processing on each layer of the target feature map based on the target reference box of each layer to obtain the predicted detection box of each layer; the size of the predicted detection box of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection box of each layer through a classification model to obtain the classification result of each layer; and determining the identification result of the road surface ahead based on the classification result of each layer; The system-on-chip is further configured to: when the first identifier represents that there is a collapse on the road surface ahead, send the collapse type and the collapse position in the identification result to the controller of the vehicle; so that the controller performs response processing based on the collapse type and the collapse position; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

[0344] Since collapse belongs to a long-tail scenario, that is, an event with large damage but low probability, in this solution, the low-probability event is implemented by the system-on-chip instead of the main controller of the vehicle, so that the low-probability event does not affect the normal control efficiency of the vehicle. When a collapse is identified, the collapse type and the collapse position are sent to the main control in real time, allowing the main control to perform response processing, while improving safety.

[0345] Fourthly, the present application also provides a vehicle, which includes a processor and a memory. A computer program or instruction is stored on the memory, and when the computer program or instruction is executed by the processor, it implements the method provided in the first aspect above.

[0346] Fifthly, an embodiment of the present application provides a storage medium, that is, a computer-readable storage medium. A computer program or instruction is stored on the storage medium, and when the computer program or instruction is executed by the processor, it implements the method provided in the first aspect above.

[0347] Sixthly, an embodiment of the present application provides a computer program product, which includes a computer program or instruction. When the computer program or instruction is executed by the processor, it implements the method provided in the first aspect above.

[0348] It should be noted here that the descriptions of the embodiments of the above storage medium, device, and program product are similar to the descriptions of the above method embodiments and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the storage medium, device, apparatus, and program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0349] It should be understood that the term "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The serial numbers of the embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.

[0350] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such a process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, method, article, or device including the element.

[0351] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed. In addition, the couplings, direct couplings, or communication connections between the various components shown or discussed can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0352] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0353] In addition, each functional unit in the embodiments of this application can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0354] Those of ordinary skill in the art can understand that all or part of the steps to implement the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage media include various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0355] Alternatively, if the above-mentioned integrated units of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of this application, in essence, or the parts that contribute to the related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of this application. The foregoing storage media include various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.

[0356] The above description is only an implementation mode of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A method for dealing with road surface collapse, characterized in that, The method includes: Obtaining spatial data of the road surface ahead collected by the vehicle; Performing identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the road surface ahead; the identification result includes a first identifier; the first identifier is used to characterize whether there is a collapse on the road surface ahead; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; obtaining the target reference frame for each layer, and performing prediction processing on the target feature map for each layer based on the target reference frame for each layer to obtain the predicted detection frame for each layer; the size of the predicted detection frame for each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection frame for each layer through a classification model to obtain the classification result for each layer; determining the identification result of the road surface ahead based on the classification result for each layer; When the first identifier characterizes that there is a collapse on the road surface ahead, outputting the collapse type and collapse location in the identification result, so that the vehicle performs response processing based on the collapse type and collapse location; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

2. The method according to claim 1, wherein The method further includes: Determining the vehicle's own processing method based on the collapse type and the collapse location; Sending a first prompt message, the collapse type, and the collapse location to the vehicle behind the vehicle through vehicle-to-infrastructure communication; the first prompt message is used to remind that there is a collapse ahead; Sending the collapse type and the collapse location to the cloud.

3. The method according to claim 1, characterized in that, The performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps includes: Performing multi-scale convolution processing on the spatial data through a fully convolutional network, and splicing the outputs of multiple stages to obtain multi-layer first feature maps; the resolution of each first feature map in the multi-layer first feature maps matches the collapse type; Performing detailed feature extraction on each first feature map in the multi-layer first feature maps respectively based on an attention mechanism to obtain multi-layer second feature maps; Performing fusion processing on the multi-layer second feature maps to obtain multi-layer third feature maps; Determining that the multi-layer target feature maps include the multi-layer first feature maps, the multi-layer second feature maps, or the multi-layer third feature maps.

4. The method according to claim 3, wherein When the multi-layer includes three layers, the performing multi-scale convolution processing on the spatial data through a fully convolutional network and splicing the outputs of multiple stages to obtain multi-layer first feature maps includes: Performing convolution processing on the spatial data based on a first downsampling rate to obtain the first feature map of the first layer; Performing convolution processing on the spatial data based on a second downsampling rate to obtain the first feature map of the second layer; Performing convolution processing on the spatial data based on a third downsampling rate to obtain the first feature map of the third layer; Determine the first feature maps of multiple layers based on the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer; Wherein, the first downsampling rate is less than the second downsampling rate, and the second downsampling rate is less than the third downsampling rate; the resolution of the first feature map of the first layer is greater than the second resolution of the first feature map of the second layer, and the resolution of the first feature map of the second layer is greater than the resolution of the first feature map of the third layer.

5. The method according to claim 3, characterized in that, When the first feature maps of multiple layers include the first feature map of the first layer, the first feature map of the second layer, and the first feature map of the third layer, the detailed feature extraction is respectively performed on the first feature maps of each layer in the first feature maps of multiple layers based on the attention mechanism to obtain the second feature maps of multiple layers, including: Performing detailed feature extraction on the first feature map of the first layer through the first attention module based on the first receptive field range to obtain the second feature map of the first layer; Performing detailed feature extraction on the first feature map of the second layer through the second attention module based on the second receptive field range to obtain the second feature map of the second layer; Performing detailed feature extraction on the first feature map of the third layer through the third attention module based on the third receptive field range to obtain the second feature map of the third layer; Determine the second feature maps of multiple layers based on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer; Wherein, the first receptive field range is less than the second receptive field range, and the second receptive field range is less than the third receptive field range.

6. The method according to claim 3, wherein When the second feature maps of multiple layers include the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer, the fusion processing is performed on the second feature maps of multiple layers to obtain the third feature maps of multiple layers, including: Performing concatenated convolution processing on the second feature map of the first layer, the second feature map of the second layer, and the second feature map of the third layer according to the first path to obtain the intermediate feature maps of three layers; Performing concatenated convolution processing on the intermediate feature maps of three layers according to the second path to obtain the third feature maps of multiple layers; the second path is different from the first path.

7. The method according to claim 1, wherein The classification processing of the target features corresponding to the predicted detection boxes of each layer through the classification model to obtain the classification results of each layer includes: Obtain the target reference box of each layer; Respectively determine the matching degrees between the predicted detection box and the target reference box of each layer to obtain multiple matching degrees; the matching degree is used to characterize the overlapping degree between the predicted detection box and the target reference box; Determine the collapse type based on the multiple matching degrees; Determine the collapse position based on the position of the predicted detection box; Determine that the classification result of each layer includes the collapse type and the collapse position.

8. The method according to claim 1, characterized in that, The method further includes: Obtain a sample data set; the sample data set includes multiple sample space data, and each sample space data is labeled with a ground truth box; the ground truth box is used to point to the collapse position in the sample space data; Distribute multiple ground truth boxes in the sample dataset into multiple layers according to their sizes, so that each layer contains multiple ground truth boxes; For each of the multiple layers, based on the multiple ground truth boxes included in the layer, determine the center of the target reference box of the layer; Based on the center of the target reference box of the layer, determine the target reference box of the layer.

9. The method according to claim 8, wherein The step of determining the center of the target reference box of the layer based on the multiple ground truth boxes included in the layer includes: Based on the multiple ground truth boxes of the layer, determine the center points of the multiple ground truth boxes; Respectively determine the probability density of the center point of each ground truth box of the layer; Determine the center point with the highest probability density as the center of the target reference box.

10. The method according to claim 8, wherein The step of determining the target reference box of the layer based on the center of the target reference box of the layer includes: Generate a candidate reference box with the center point of the target reference box as the center; Respectively determine the dynamic distance between each ground truth box of the layer and the candidate reference box; the dynamic distance is negatively correlated with the overlap degree and positively correlated with the aspect ratio similarity; Adjust the candidate reference box based on the dynamic distance until the dynamic distance meets the preset condition; Determine the candidate reference box that meets the preset condition as the target reference box of the layer.

11. The method according to claim 8, characterized in that The method further includes: For each ground truth box of the layer, respectively determine the matching degree between the ground truth box and the target reference box of the layer; Adjust the layer to which the ground truth box belongs based on the matching degree; Based on the ground truth boxes included in each adjusted layer, update the target reference box of each layer.

12. A treatment device for road surface collapse, characterized in that, The device includes: An acquisition unit, configured to acquire the spatial data of the front road surface collected by the vehicle; An identification unit, configured to perform identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the front road surface; the identification result includes a first identifier; the first identifier is used to represent whether there is a collapse on the front road surface; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; acquiring the target reference box of each layer, and performing prediction processing on the target feature map of each layer based on the target reference box of each layer to obtain the predicted detection box of each layer; the size of the predicted detection box of each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection box of each layer through a classification model to obtain the classification result of each layer; determining the identification result of the front road surface based on the classification result of each layer; An output unit, configured to output the collapse type and collapse location in the identification result when the first identifier represents that there is a collapse on the front road surface, so that the vehicle performs response processing based on the collapse type and collapse location; the collapse types include: road surface fracture, road surface lateral collapse, road surface longitudinal collapse.

13. A system-on-chip, characterized in that, The system-on-chip is connected to the controller of the road surface collapse; The system-on-chip is used for: obtaining spatial data of the road surface ahead collected by the vehicle; performing identification processing on the spatial data through a road surface collapse identification model to determine the identification result of the road surface ahead; The identification result includes a first identifier; The first identifier is used to characterize whether there is a collapse on the road surface ahead; wherein, the identification processing of the road surface collapse identification model includes: performing multi-layer feature extraction on the spatial data through a feature extraction model to obtain multi-layer target feature maps; the resolution of each target feature map in the multi-layer target feature maps matches the collapse type; obtaining a target reference frame for each layer, and performing prediction processing on the target feature map of each layer based on the target reference frame of each layer to obtain a predicted detection frame for each layer; the size of the predicted detection frame for each layer matches the collapse type; performing classification processing on the target features corresponding to the predicted detection frame for each layer through a classification model to obtain a classification result for each layer; determining the identification result of the road surface ahead based on the classification result for each layer; The system-on-chip is further used for: when the first identifier characterizes that there is a collapse on the road surface ahead, sending the collapse type and the collapse position in the identification result to the controller of the road surface collapse; so that the controller performs response processing based on the collapse type and the collapse position; the collapse types include: road surface fracture, lateral road surface collapse, and longitudinal road surface collapse.

14. A vehicle, characterized in that, The vehicle includes a processor and a memory, and a computer program or instruction is stored on the memory. When the computer program or instruction is executed by the processor, the method according to any one of claims 1-11 is implemented.

15. A computer-readable storage medium, characterized in that, A computer program or instruction is stored on the storage medium. When the computer program or instruction is executed by the processor, the method according to any one of claims 1-11 is implemented.

16. A computer program product, characterized in that, The computer program product includes a computer program or instruction. When the computer program or instruction is executed by the processor, the method according to any one of claims 1-11 is implemented.

Citation Information

Patent Citations

  • Double-branch low-illumination image enhancement method based on Retinex theory

    CN117994155A

  • Neural network for object detection in images

    US20180121762A1

Cited By

  • Road settlement automatic detection method and system

    CN120564155A