Road safety risk identification and early warning system and method based on image semantic segmentation
By adopting multi-scale feature extraction and cross-modal relationship enhancement methods in the road image semantic segmentation system, the problems of poor identification of small-scale obstacles and insufficient multi-level assessment capabilities in the prior art are solved, and high-precision identification and hierarchical early warning of road safety risks are achieved.
Patent Information
- Application Number
- CN202510631426.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The existing road image semantic segmentation method performs well when identifying large-scale obstacles, but ignores small-scale obstacles and cannot effectively process confusing objects that are spatial information and color texture features, resulting in low recognition accuracy and inability to conduct road safety assessments of multi-level and multi-tasks.
The road safety risk identification and early warning system based on image semantic segmentation is adopted. Through the image acquisition module, feature extraction module, cross-modal relationship enhancement module, semantic segmentation module and safety early warning module, multi-scale color texture features and spatial features are extracted, cross-modal feature relationship matrix is constructed, and multi-scale features are dynamically fusion, so as to realize multi-level detection and hierarchical early warning of road safety.
It improves the high-precision recognition ability of multi-scale objects on the road, enhances the recognition ability of objects that are easily confused by space or color textures, ensures accurate identification of road obstacles, and improves the accuracy and early warning efficiency of road safety detection.
Smart Images

Figure CN120147996A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of road safety warning, and more specifically, to a road safety risk identification and warning system and method based on image semantic segmentation. Background Art
[0002] With the continuous growth of motor vehicles and the development of intelligent transportation systems and autonomous driving technologies, people's demand for road safety is increasing day by day. Conducting road anomaly prediction and warning can effectively reduce traffic accidents and improve driving safety. In a complex traffic environment, timely identification and warning of potential road anomalies can greatly avoid possible dangers.
[0003] Environmental perception is the first link to achieve intelligent driving. Intelligent vehicles obtain information around the vehicle through various sensors such as cameras, millimeter-wave radars, ultrasonic radars, lidars, etc. Detection and recognition of obstacles and road surface structured data are important research contents of environmental perception, which can be achieved by semantic segmentation of structured roads based on deep learning.
[0004] Existing road image semantic segmentation methods only perform well in large-scale obstacle categories, often ignoring small-scale obstacles (such as road nails, cracks, etc.). Moreover, existing methods generally pay more attention to the spatial information of data and cannot achieve adaptive feature enhancement for different regions, resulting in easy confusion of the category segmentation results of objects with highly similar spatial information and relatively low segmentation accuracy. At the same time, existing methods often directly output the types of potential safety hazards of the current road conditions and cannot perform multi-level and multi-task road safety assessments. Therefore, existing deep learning models cannot take into account multi-scale objects on the road surface, cannot distinguish road surface obstacles that are easy to confuse in terms of space or color texture, and cannot effectively respond to road anomalies caused by road construction, traffic accidents or other emergencies, and their detection accuracy needs to be improved urgently. Summary of the Invention
[0005] In view of this, the present invention provides a road safety risk identification and warning system and method based on image semantic segmentation, which can take into account high-precision identification of multi-scale and multi-category objects on the road and improve the detection accuracy of road anomalies during driving.
[0006] To achieve the above object, the present invention adopts the following technical solutions:
[0007] In the first aspect, the present invention provides a road safety risk identification and warning system based on image semantic segmentation, including: an image acquisition module, a feature extraction module, a cross-modal relationship enhancement module, a semantic segmentation module, and a safety warning module;
[0008] The image acquisition module is used to acquire a real-time panoramic image within a preset range around the driving vehicle;
[0009] The feature extraction module is used to extract multi-scale color texture features and multi-scale spatial features from the real-time panoramic image;
[0010] The cross-modal relationship enhancement module is used to calculate the cosine similarity between the color texture features and the spatial features at multiple scales, construct a cross-modal feature relationship matrix, and perform dynamic fusion of multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features;
[0011] The semantic segmentation module hierarchically analyzes the multi-level detection problem of road safety from the perspective of task decomposition, and detects the multi-modal fusion features to obtain road safety detection results;
[0012] The safety warning module is used to classify and warn road safety according to the road safety detection results.
[0013] Furthermore, the feature extraction module includes a first feature extraction branch and a second feature extraction branch. The first feature extraction branch captures semantic information, context information, and local detail information at different scales of the real-time panoramic image through dense convolution and residual connection as multi-scale spatial features;
[0014] The second feature extraction branch captures color information and texture information at different scales of the real-time panoramic image through the combination of color histogram and gray-level co-occurrence matrix as multi-scale color texture features.
[0015] Furthermore, the first feature extraction branch includes a residual network unit, a dense convolution block, a multi-scale spatial fusion unit, and an attention guidance unit;
[0016] The residual network unit is used to downsample the real-time panoramic image and extract intermediate feature maps at different scales;
[0017] The dense convolution block consists of multiple depthwise separable convolution layers with the same dilation rate and residual connections, and is used to extract and transform features of the real-time panoramic image to obtain low-level features of the real-time panoramic image;
[0018] The multi-scale spatial fusion unit is used to add the features output by the residual network unit and the features output by the dense convolution block at the same dimension element by element;
[0019] The attention guidance unit is used to perform multi-angle screening and aggregation on the features output by the multi-scale spatial fusion unit from the channel level, pixel level, and global level to obtain the finally extracted multi-scale spatial features.
[0020] Further, the second feature extraction branch includes a scale division unit, a color extraction unit, a texture extraction unit, and a multi-scale color texture fusion unit;
[0021] The scale division unit is used to divide the real-time panoramic image into different scale images;
[0022] The color extraction unit is used to convert different scale images from the RGB space to the HSV space, calculate the color distribution of the images, and constitute the color features of the images;
[0023] The texture extraction unit is used to convert different scale images into grayscale images, calculate the local pixel point relationship values of the grayscale images, and statistically analyze the histograms of the grayscale relationships of the grayscale images to obtain the LBP texture features of the images;
[0024] The multi-scale color texture fusion unit is used to fuse the color features and LBP texture features of each scale image to obtain multi-scale color texture features.
[0025] Further, the cross-modal relationship enhancement module includes a relationship matrix construction unit and a dynamic fusion unit;
[0026] The relationship matrix construction unit is used to splice the spatial features and color texture features of the real-time panoramic image at multiple scales respectively to obtain the spatial feature splicing result and the color texture feature splicing result, and then construct a cross-modal feature relationship matrix according to the cosine similarity between the spatial feature splicing result and the color texture feature splicing result;
[0027] The dynamic fusion unit is used to identify high-similarity regions, medium-similarity regions, and low-similarity regions according to the cross-modal feature relationship matrix; if the cosine similarity between the spatial feature splicing result and the color texture feature splicing result of a certain region in the real-time panoramic image is between 0.8 and 1.0, the current region is regarded as a high-similarity region. For high-similarity regions, it is suitable for the recognition of color-sensitive targets, and the color texture feature splicing result is used as the dominant;
[0028] If the cosine similarity between the spatial feature splicing result and the color texture feature splicing result of a certain region in the real-time panoramic image is between 0.3 and 0.8, the current region is regarded as a medium-similarity region. For medium-similarity regions, the two modal feature splicing results are spliced or weighted and fused;
[0029] If the cosine similarity between the spatial feature splicing result and the color texture feature splicing result of a certain region in the real-time panoramic image is less than 0.3, the current region is regarded as a low-similarity region. For low-similarity regions, it is suitable for the recognition of geometric feature targets, and the spatial feature splicing result is used as the dominant.
[0030] Furthermore, the expression of the cross-modal feature relationship matrix is as follows:
[0031] ;
[0032] ;
[0033] where , representing the cross-modal feature relationship matrix, whose dimension is ; ; represents the inner product, represents the L2 norm; represents the i-th row vector of the cross-modal feature relationship matrix; represents the j-th row vector of the cross-modal feature relationship matrix; represents the result of spatial feature splicing; represents the result of color texture feature splicing; represents the dimension of the feature splicing result, represents the tensor length, represents the tensor width, and the tensor length is equivalent to the number of sampling points, that is each of the points has a vector of length
[0034] Furthermore, the semantic segmentation module divides the multi-level detection problem of road safety into three subtasks, and each subtask corresponds to a prediction network;
[0035] Among them, the prediction network corresponding to the first subtask identifies and segments the object categories contained in the road in the panoramic image around the vehicle;
[0036] The prediction network corresponding to the second subtask identifies whether the identified object categories pose a safety hazard to road safety;
[0037] The prediction network corresponding to the third subtask detects the road safety hazard level.
[0038] Furthermore, the loss functions of the three prediction networks are respectively expressed as:
[0039] ;
[0040] ;
[0041] ;
[0042] where represents the total number of pixels in the panoramic image, , , Labels representing three subtasks respectively, , , represent the parameters of three prediction networks respectively, represents the i-th pixel point in the input multi-modal fusion feature; , , represent the prediction results of three prediction networks respectively;
[0043] Weighted sum the losses of three prediction networks as the final loss of the semantic segmentation module:
[0044] ;
[0045] wherein, , , represent the weights of three losses respectively.
[0046] Furthermore, the object categories recognized by the prediction network corresponding to the first subtask include road potholes, road surface water accumulation, road surface reflection, overpass reflection, low-altitude floating objects, traffic signs, road surface cracks, road surface micro sharp objects, dynamic obstacles, static obstacles, accident areas and construction areas; the prediction network corresponding to the second subtask analyzes the positions of the objects recognized by the first prediction network to predict whether there are potential safety hazards; the prediction network corresponding to the third subtask comprehensively analyzes the safety hazard judgment results of each object category output by the second prediction network and outputs the safety hazard level, including normal passage, mild safety hazard, moderate safety hazard and severe safety hazard.
[0047] In a second aspect, the present invention provides a road safety risk identification and early warning method based on image semantic segmentation, applicable to the system as described above, including the following steps:
[0048] Obtain a real-time panoramic image within a preset range around the driving vehicle;
[0049] Extract multi-scale color texture features and multi-scale spatial features from the real-time panoramic image;
[0050] Calculate the cosine similarity between the color texture features and spatial features at multiple scales, construct a cross-modal feature relationship matrix, and perform dynamic fusion of multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features;
[0051] Conduct hierarchical analysis on the multi-level detection problem of road safety from the perspective of task decomposition, and detect the multi-modal fusion features to obtain road safety detection results;
[0052] Perform hierarchical early warning on road safety according to the road safety detection results.
[0053] As can be seen from the above technical solutions, compared with the prior art, the present invention has the following beneficial effects:
[0054] 1. The present invention makes full use of the spatial features and color texture features of road images, constructs a cross-modal feature relationship matrix, considers the spatial and color texture relationships of each region at different scales, enables the model to distinguish objects that are easily confused in space or color texture, and adaptively fuses the features of different regions, enhancing the dominant degree of spatial or color texture features, and ensuring accurate recognition of road obstacles.
[0055] 2. By extracting multi-scale features of the panoramic image around the vehicle, the present invention can take into account multi-scale obstacles or road surface defects in the positioning area, without ignoring small-scale objects, making the recognition of regional objects more complete, ensuring more accurate recognition of small-scale obstacles or sharp objects on the road, and guaranteeing driving safety.
[0056] 3. By performing multi-task detection on road safety issues, the present invention effectively reduces the deviation caused by a single network for multi-task detection, simplifies the learning objectives, and can ensure the prediction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0058] Figure 1 It is a schematic structural diagram of a road safety risk identification and warning system based on image semantic segmentation provided by the present invention;
[0059] Figure 2 It is a schematic structural diagram of the first feature extraction branch provided by the present invention;
[0060] Figure 3 It is a schematic structural diagram of the second feature extraction branch provided by the present invention;
[0061] Figure 4 It is a schematic structural diagram of the cross-modal relationship enhancement module provided by the present invention;
[0062] Figure 5 It is a schematic structural diagram of the semantic segmentation module provided by the present invention;
[0063] Figure 6 It is a flowchart of a road safety risk identification and warning method based on image semantic segmentation provided by the present invention. Detailed implementation manners
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] As Figure 1 shown, an embodiment of the present invention discloses a road safety risk identification and early warning system based on image semantic segmentation, including: an image acquisition module, a feature extraction module, a cross-modal relationship enhancement module, a semantic segmentation module, and a safety early warning module;
[0066] The image acquisition module is used to acquire a real-time panoramic image within a preset range around the driving vehicle;
[0067] The feature extraction module is used to extract multi-scale color texture features and multi-scale spatial features in the real-time panoramic image;
[0068] The cross-modal relationship enhancement module is used to calculate the cosine similarity between the color texture features and spatial features at multiple scales, construct a cross-modal feature relationship matrix, and perform dynamic fusion of multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features;
[0069] The semantic segmentation module hierarchically analyzes the multi-level detection problem of road safety from the perspective of task decomposition, detects the multi-modal fusion features, and obtains the road safety detection result; the semantic segmentation module includes a total of three prediction networks, namely prediction network 1, prediction network 2, and prediction network 3. Among them, prediction network 1 is used to identify the object categories contained in the road, prediction network 2 is used to judge whether each object category contained in the road will cause safety hazards, and prediction network 3 is used to output the safety hazard level.
[0070] The safety early warning module is used to perform hierarchical early warning on road safety according to the road safety detection result.
[0071] Next, the implementation details of the present invention will be further described.
[0072] 1) The image acquisition module is used to acquire a panoramic image within a preset range around the currently driving vehicle, including road surface image data and surrounding environment image data.
[0073] 2) The feature extraction module includes a first feature extraction branch and a second feature extraction branch. The first feature extraction branch captures semantic information, context information, and local detail information at different scales of the real-time panoramic image through dense convolution and residual connection as multi-scale spatial features;
[0074] The second feature extraction branch captures color information and texture information at different scales of the real-time panoramic image through the combination of color histogram and gray-level co-occurrence matrix as multi-scale color texture features.
[0075] 2.1) As Figure 2 shown, the first feature extraction branch includes a residual network unit, a dense convolutional block, a multi-scale spatial fusion unit, and an attention guidance unit.
[0076] The residual network unit is used to downsample the panoramic image and extract intermediate feature maps at different scales; specifically, it downsamples the panoramic image through four convolutions with a stride of 1 to extract multi-scale features of the panoramic image, obtaining four intermediate feature maps r1~r4.
[0077] The dense convolutional block performs preliminary feature extraction and transformation on the panoramic image to capture low-level features of the panoramic image, thereby obtaining intermediate features and encoding them to obtain low-level features d1~d4 of the panoramic image. The specific structure of the dense convolutional block consists of four depthwise separable convolutional layers with the same dilation rate and residual connections, aiming to expand the receptive field with fewer parameters.
[0078] The multi-scale spatial fusion unit is used to add the matrix elements of the features r1~r4 output by the residual network unit and the features d1~d4 output by the dense convolutional block under the same dimension.
[0079] The dense flow convolution branch focuses on context information and detail features, and the residual network unit provides high-level semantic features. By associating these two branches, the performance and accuracy of the model in image processing tasks can be improved, and the expression ability and performance of the model can be enhanced by fusing the features of different branches.
[0080] The attention guidance unit performs progressive fusion of adjacent scales on the multi-scale features output by the multi-scale spatial fusion unit. Specifically, it can fuse features of three scales, and the number of scales for fusion can be selected according to the requirements of the task and the model. The use of three-scale feature fusion is to balance low-level detail information, middle-level semantic information, and high-level global information. Such a fusion strategy can balance between details, context, and global vision to meet the requirements of different tasks.
[0081] Specifically, low-level features have a small receptive field and can capture local details of the image. Therefore, depthwise separable convolution can be used for downsampling to reduce the image size and expand the receptive field range of each output pixel point to capture broader context information.
[0082] High-level features contain more semantic information and can be dimensionally reduced to reduce the number of feature channels, thereby reducing the number of parameters and computational complexity. Through appropriate convolution operations, feature channels with similar semantics can be fused to enhance the semantic expression ability of high-level features. This can help the model better understand the objects, scenes, and semantic information in the image.
[0083] After performing the dimensionality reduction operation, deconvolution or transposed convolution is carried out to restore the size of the feature map to the original size.
[0084] Finally, through concatenation operations and convolution operations, low-level features with spatial information and high-level features with semantic information are fused to obtain multi-scale features. 。
[0085] The attention guidance unit is used to perform multi-angle screening and aggregation on the multi-scale features output by the multi-scale fusion unit from the channel level, pixel level, and global level to obtain an enhanced representation of the multi-scale features as the finally extracted multi-scale spatial features.
[0086] Specifically, the convolutional block-based attention mechanism, per-pixel attention mechanism, and global attention mechanism can be used to screen and aggregate the multi-scale features respectively. Each attention mechanism extracts information at different levels of the multi-scale features, and then the features extracted by each attention mechanism are aggregated to obtain the final multi-scale spatial features, which are finally represented as:
[0087] ;
[0088] Among them, represents the Sigmoid activation function, represents the output of the convolutional block-based attention mechanism, represents the pointwise attention feature with the same shape as the input , that is, the output of the per-pixel attention mechanism, represents the global attention output matrix.
[0089] 2.2) As Figure 3 shown, the second feature extraction branch includes a scale division unit, a color extraction unit, a texture extraction unit, and a multi-scale color-texture fusion unit;
[0090] The scale division unit is used to divide the real-time panoramic image into different scale images;
[0091] The color extraction unit is used to convert different scale images from the RGB space to the HSV space, calculate the color distribution of the images, and constitute the color features of the images;
[0092] The texture extraction unit is used to convert images of different scales into grayscale images, calculate the local pixel point relationship values of the grayscale images, and count the histograms of the grayscale relationships of the grayscale images to obtain the LBP texture features of the images. Specifically, let the grayscale values of P adjacent pixel points with a radius of R be g 0 ,g 1 ,...,g P-1 , and compare them with the grayscale value g c of the current pixel point in turn. If it is larger than g c , the pixel point is set to 1, otherwise it is set to 0.
[0093] The multi-scale color texture fusion unit is used to fuse the color features and LBP texture features of the images at each scale to obtain multi-scale color texture features 。
[0094] 3) As Figure 4 shown, the cross-modal relationship enhancement module includes a relationship matrix construction unit and a dynamic fusion unit;
[0095] The relationship matrix construction unit is used to splice the spatial features and color texture features at multiple scales of the real-time panoramic image respectively to obtain the spatial feature splicing result and the color texture feature splicing result, and then construct a cross-modal feature relationship matrix according to the cosine similarity between the spatial feature splicing result and the color texture feature splicing result;
[0096] The dynamic fusion unit is used to identify the high-similarity region, medium-similarity region, and low-similarity region according to the cross-modal feature relationship matrix; for the high-similarity region, the color texture feature splicing result is used as the dominant; for the medium-similarity region, the two modal feature splicing results are spliced or weighted and fused; for the low-similarity region, the spatial feature splicing result is used as the dominant.
[0097] Specifically, the expression of the cross-modal feature relationship matrix is:
[0098] ;
[0099] ;
[0100] Among them, , represents the cross-modal feature relationship matrix, and its dimension is ; ; represents the inner product, represents the L2 norm; represents the i-th row vector of the cross-modal feature relationship matrix; represents the j-th row vector of the cross-modal feature relationship matrix; represents the spatial feature splicing result; represents the color texture feature splicing result; Indicates the dimension of the feature splicing result, Indicates the tensor length, Indicates the tensor width. The tensor length is equivalent to the number of sampling points, that is, each of the points has a vector of length
[0101] Since the features for calculating the cross-modal feature relationship matrix respectively emphasize spatial information and color texture information, the segmentation result of the present invention takes into account the relationship between cross-modal points, enabling the model to more accurately identify small-scale and easily confused objects.
[0102] When the cosine similarity between the spatial feature splicing result and the color texture feature splicing result in a certain area of the real-time panoramic image is between 0.8 and 1.0, the current area is regarded as a high similarity area, which is suitable for the recognition of color-sensitive targets; when the cosine similarity between the spatial feature splicing result and the color texture feature splicing result in a certain area of the real-time panoramic image is between 0.3 and 0.8, the current area is regarded as a medium similarity area; when the cosine similarity between the spatial feature splicing result and the color texture feature splicing result in a certain area of the real-time panoramic image is less than 0.3, the current area is regarded as a low similarity area, which is suitable for the recognition of geometric feature targets.
[0103] For example, for a reflective road surface and an overpass reflection, the calculation result is that the area belongs to a low similarity area. At this time, the spatial feature is dominant and is input to the subsequent module for prediction, avoiding misidentifying the reflection in the reflective road surface as a real object and misidentifying the overpass reflection as a crack.
[0104] For partially occluded pedestrians, low-altitude drones, inflatable advertising balls, etc., the calculated result is that the area is a high similarity area. At this time, the color texture feature is dominant and is input to the subsequent module for prediction, avoiding blurring the boundary contour of the pedestrian, misidentifying the drone as a falling object, and misjudging the inflatable advertising balloon as a pedestrian, so as to avoid category confusion and reduce the false alarm rate.
[0105] For the medium similarity area, at this time, the spatial feature and the color texture feature are weighted and fused to dynamically balance the color and shape information. For example, for a night traffic sign, the weight of the color texture feature can be enhanced to avoid color feature distortion. For a water accumulation area, to avoid misexpanding the road area, the weight of the spatial feature can be enhanced.
[0106] For tiny sharp objects or obstacles on the road surface, such as nails and stones, at this time, the spatial feature and the color texture feature are spliced to achieve accurate recognition of sharp objects or stones, avoiding flat tire or collision accidents of vehicles caused by misidentification.
[0107] In addition, the image features can be globally adjusted according to specific driving scenarios. For night scenes, the spatial features are enhanced to avoid reflection interference. For rainy and foggy weather, the color texture features are enhanced to compensate for insufficient visibility.
[0108] 4) As Figure 5 shown, the semantic segmentation module divides the multi-level detection problem of road safety into three subtasks, and each subtask corresponds to a prediction network;
[0109] Among them, the first subtask corresponds to prediction network 1, and prediction network 1 identifies and segments the object categories contained in the road in the panoramic image around the vehicle;
[0110] The second subtask corresponds to prediction network 2, and prediction network 2 identifies whether the identified object categories pose potential safety hazards to road safety;
[0111] The third subtask corresponds to prediction network 3, and prediction network 3 detects the level of road safety hazards.
[0112] The first subtask is "object category identification and segmentation", which is a multi-classification task. The prediction network identifies each object category in the panoramic image of the current vehicle's surrounding environment, and the output is category 1 - category n. The identified object categories include one or more of potholes on the road surface, water accumulation on the road surface, road surface reflection, overpass reflection, low-altitude floating objects, traffic signs, road surface cracks, tiny sharp objects on the road surface, dynamic obstacles, static obstacles, accident areas, and construction areas.
[0113] The second subtask is "whether it poses a safety hazard", which includes multiple output heads. Each output head corresponds to a safety hazard category and is a binary classification task. Based on the object categories output by the first prediction network, it predicts "whether there is a safety hazard", and outputs whether there is a safety hazard for category 1 - category n. The second prediction network analyzes the positions of the objects identified by the first prediction network and predicts whether each category of object will pose a safety hazard.
[0114] The prediction network corresponding to the third subtask comprehensively analyzes the safety hazard judgment results of each object category output by the second prediction network and outputs the safety hazard level. This is a four-classification problem, and the output safety hazard level includes normal passage, minor safety hazard, moderate safety hazard, and severe safety hazard.
[0115] The present invention trains three prediction networks for the three subtasks respectively. Each prediction network only focuses on its own small part of the task, which can effectively reduce the deviation caused by the multi-task coupling part, simplifies the learning objective of the model, and enables the model to better learn the key points of the task.
[0116] Specifically, the three prediction networks share underlying features at relatively shallow layers and learn their own exclusive parameters at relatively deep layers. Their loss functions are respectively expressed as:
[0117] ;
[0118] ;
[0119] ;
[0120] where represents the total number of pixels in the panoramic image, , , respectively represent the labels of the three subtasks, , , respectively represent the parameters of the three prediction networks, represents the i-th pixel point in the input multi-modal fusion feature; , , respectively represent the prediction results of the three prediction networks;
[0121] The losses of the three prediction networks are weighted and summed to obtain the final loss of the semantic segmentation module:
[0122] ;
[0123] where , , respectively represent the weights of the three losses. The weights of the three losses are set according to the difficulty level between tasks, or according to the importance of tasks, or a set of weights are learned by the model using the attention method.
[0124] 5) The warning module can broadcast the warning information in voice form to remind the driver to drive safely and synchronize the warning information to the user's handheld terminal or in-vehicle terminal.
[0125] As Figure 6 shown, an embodiment of the present invention further provides a road safety risk identification and warning method based on image semantic segmentation, which is applicable to the above system and includes the following steps:
[0126] S1. Obtain a real-time panoramic image within a preset range around the driving vehicle;
[0127] S2. Extract multi-scale color texture features and multi-scale spatial features from the real-time panoramic image;
[0128] S3. Calculate the cosine similarity between the color texture features and the spatial features at multiple scales, construct a cross-modal feature relationship matrix, and perform dynamic fusion of the multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features;
[0129] S4. Conduct a hierarchical analysis of the multi-level detection problem of road safety from the perspective of task decomposition, and detect the multi-modal fusion features to obtain the road safety detection results;
[0130] S5. According to the road safety detection results, conduct hierarchical early warning for road safety.
[0131] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part.
[0132] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A road safety risk identification and early warning system based on image semantic segmentation, characterized in that: include: Image acquisition module, feature extraction module, cross-modal relationship enhancement module, semantic segmentation module and safety warning module; The image acquisition module is used to acquire a real-time panoramic image of a preset range around the driving vehicle; The feature extraction module is used to extract multi-scale color texture features and multi-scale spatial features in the real-time panoramic image; The cross-modal relationship enhancement module is used to calculate the cosine similarity between color texture features and spatial features at multiple scales, construct a cross-modal feature relationship matrix, and dynamically fuse multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features; The semantic segmentation module performs hierarchical analysis on the multi-level detection problem of road safety from the perspective of task decomposition, detects the multimodal fusion features, and obtains road safety detection results; The safety warning module is used to provide graded warnings on road safety according to the road safety detection results.
2. The road safety risk identification and early warning system based on image semantic segmentation according to claim 1 is characterized in that: The feature extraction module includes a first feature extraction branch and a second feature extraction branch, wherein the first feature extraction branch captures semantic information, context information and local detail information of the real-time panoramic image at different scales as multi-scale spatial features by means of dense convolution and residual connection; The second feature extraction branch captures the color information and texture information of the real-time panoramic image at different scales by combining a color histogram and a gray-level co-occurrence matrix as a multi-scale color-texture feature.
3. The road safety risk identification and early warning system based on image semantic segmentation according to claim 2 is characterized in that: The first feature extraction branch includes a residual network unit, a dense convolution block, a multi-scale spatial fusion unit and an attention guidance unit; The residual network unit is used to downsample the real-time panoramic image and extract intermediate feature maps at different scales; The dense convolution block is composed of a plurality of depth-separable convolutional layers and residual connections with the same dilation rate, and is used to extract and transform the real-time panoramic image to obtain low-level features of the real-time panoramic image; The multi-scale spatial fusion unit is used to perform matrix element addition on the features output by the residual network unit and the features output by the dense convolution block in the same dimension; The attention guiding unit is used to perform multi-angle screening and aggregation on the features output by the multi-scale spatial fusion unit from the channel level, pixel level and global level to obtain the final extracted multi-scale spatial features.
4. The road safety risk identification and early warning system based on image semantic segmentation according to claim 2 is characterized in that: The second feature extraction branch includes a scale division unit, a color extraction unit, a texture extraction unit and a multi-scale color and texture fusion unit; The scale division unit is used to divide the real-time panoramic image into images of different scales; The color extraction unit is used to convert images of different scales from RGB space to HSV space, calculate the color distribution of the image, and form the color features of the image; The texture extraction unit is used to convert images of different scales into grayscale images, calculate the local pixel relationship value of the grayscale image, and count the histogram of the grayscale relationship of the grayscale image to obtain the LBP texture feature of the image; The multi-scale color and texture fusion unit is used to fuse the color features of each scale image with the LBP texture features to obtain multi-scale color and texture features.
5. The road safety risk identification and early warning system based on image semantic segmentation according to claim 1 is characterized in that: The cross-modal relationship enhancement module includes a relationship matrix construction unit and a dynamic fusion unit; The relationship matrix construction unit is used to respectively splice the spatial features and color and texture features of the real-time panoramic image at multiple scales to obtain a spatial feature splicing result and a color and texture feature splicing result, and then construct a cross-modal feature relationship matrix according to the cosine similarity between the spatial feature splicing result and the color and texture feature splicing result; The dynamic fusion unit is used to identify high similarity regions, medium similarity regions and low similarity regions according to the cross-modal feature relationship matrix; If the cosine similarity between the spatial feature stitching result and the color and texture feature stitching result of a certain area in the real-time panoramic image is between 0.8 and 1.0, the current area is regarded as a high similarity area. For the high similarity area, it is suitable for the recognition of color-sensitive targets, and the color and texture feature stitching result is taken as the leading factor; If the cosine similarity between the spatial feature stitching result and the color and texture feature stitching result of a certain area in the real-time panoramic image is between 0.3 and 0.8, the current area is regarded as a medium similarity area, and for the medium similarity area, the stitching results of the two modal features are stitched or weighted fused; If the cosine similarity between the spatial feature stitching result and the color and texture feature stitching result of a certain area in the real-time panoramic image is less than 0.3, the current area is regarded as a low-similarity area. For the low-similarity area, it is suitable for identifying geometric feature targets, with the spatial feature stitching result as the dominant one.
6. The road safety risk identification and early warning system based on image semantic segmentation according to claim 5 is characterized in that: The expression of the cross-modal feature relationship matrix is: ; ; in, , represents the cross-modal feature relationship matrix, whose dimension is ; ; represents the inner product, represents the L2 norm; Represents the i-th row vector of the cross-modal feature relationship matrix; Represents the j-th row vector of the cross-modal feature relationship matrix; Indicates the result of spatial feature splicing; Indicates the result of color and texture feature splicing; Indicates the dimension of the feature concatenation result, represents the length of the tensor, Represents the tensor width, and the tensor length is equivalent to the number of sampling points, that is, Each point has a length of The vector describes the features.
7. The road safety risk identification and early warning system based on image semantic segmentation according to claim 1 is characterized in that: The semantic segmentation module divides the multi-level detection problem of road safety into three subtasks, each of which corresponds to a prediction network; The prediction network corresponding to the first subtask identifies and segments the object categories contained in the road in the panoramic image around the vehicle; The prediction network corresponding to the second subtask identifies whether the identified object category will cause safety hazards to road safety; The prediction network corresponding to the third subtask detects the level of road safety hazards.
8. The road safety risk identification and early warning system based on image semantic segmentation according to claim 7 is characterized in that: The loss functions of the three prediction networks are expressed as: ; ; ; in, Represents the total number of pixels in the panoramic image. , , Represent the labels of the three subtasks respectively. , , Represent the parameters of the three prediction networks respectively, Represents the i-th pixel in the input multimodal fusion feature; , , Represent the prediction results of the three prediction networks respectively; The losses of the three prediction networks are weighted summed as the final loss of the semantic segmentation module: ; in, , , Represent the weights of the three losses respectively.
9. The road safety risk identification and early warning system based on image semantic segmentation according to claim 7, characterized in that: The object categories identified by the prediction network corresponding to the first subtask include at least potholes, road water, road reflections, overpass reflections, low-altitude floating objects, traffic signs, road cracks, tiny sharp objects on the road, dynamic obstacles, static obstacles, accident areas and construction areas; the prediction network corresponding to the second subtask analyzes the location of the objects identified by the first prediction network to predict whether there are safety hazards; the prediction network corresponding to the third subtask comprehensively analyzes the safety hazard judgment results of each object category output by the second prediction network, and outputs the safety hazard level, including normal passage, mild safety hazard, moderate safety hazard and severe safety hazard.
10. A road safety risk identification and early warning method based on image semantic segmentation, characterized in that: A system according to any one of claims 1 to 9, comprising the following steps: Obtain a real-time panoramic image of a preset range around the driving vehicle; Extracting multi-scale color texture features and multi-scale spatial features in the real-time panoramic image; Calculating the cosine similarity between color texture features and spatial features at multiple scales, constructing a cross-modal feature relationship matrix, and dynamically fusing multi-scale features according to the cross-modal feature relationship matrix to obtain multi-modal fusion features; From the perspective of task decomposition, a hierarchical analysis is performed on the multi-level detection problem of road safety, and the multi-modal fusion features are detected to obtain road safety detection results; Based on the results of road safety inspections, graded warnings are issued for road safety.
Citation Information
Patent Citations
Point cloud weak supervision semantic segmentation method based on multi-modal and multi-scale affinity relationship
CN116091775A
Vehicle potential safety hazard early warning method, device and equipment based on multi-sensor fusion
CN119099640A
Furniture material detection and risk assessment system
CN119919398A
Cited By
Method, device and equipment for identifying landslide-caused road burying interaction target and medium
CN122336614A