Road image detection method and device and storage medium

By injecting salt and pepper noise into the road model training image and using a reparameterized vision transformer and weighted feature fusion enhancement feature pyramid network, combining knowledge graphs and key point detection algorithms, the problem of inaccurate target detection in complex traffic environments is solved, and higher detection accuracy and robustness are achieved.

CN120388334APending Publication Date: 2025-07-29WUHAN POLYTECHNIC UNIVERSITY
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510304761.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing road image detection methods cause inaccurate target detection due to environmental changes in complex urban traffic environments, especially in noise, target occlusion and multi-scale object detection problems.

Method used

By injecting salt and pepper noise into the road model training image, the features are extracted using a reparameterized vision transformer, and combined with the weighted feature fusion enhanced feature pyramid network and the road object detection algorithm that integrates knowledge graphs and key point detection, object detection and model training are performed.

Benefits of technology

It improves the accuracy and robustness of target detection in noise interference and complex traffic environments, can better identify and locate targets such as vehicles, pedestrians, traffic signs, etc., and improves detection accuracy and position prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388334A_ABST
    Figure CN120388334A_ABST
Patent Text Reader

Abstract

The invention discloses a road image detection method and device and a storage medium, relates to the technical field of image recognition, and discloses a road image detection method, which comprises the following steps: acquiring a road model training image, and injecting salt and pepper noise into the road model training image to obtain an enhanced image; performing feature extraction on the enhanced image according to a re-parameterized visual converter to obtain a basic feature map; performing feature fusion on the basic feature map according to a weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map; performing target detection on the weighted fusion multi-scale feature map according to a road target detection algorithm fusing a knowledge map and key point detection to obtain a target road detection result; and training a road detection model according to the target road detection result, and detecting a real-time road image according to the road detection model. The accuracy and robustness of road image detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a road image detection method, device, and storage medium. Background Art

[0002] With the acceleration of economic development and urbanization, the rapid increase in urban population and the popularization of new energy vehicles have led to a substantial increase in the number of vehicles, further exacerbating urban traffic pressure, and problems such as traffic congestion, safety issues, and environmental pollution have become increasingly serious. Traditional manual inspection methods not only have high costs, low efficiency, but also poor reliability, and cannot effectively meet the growing traffic monitoring needs. Although infrared or radar detection technology can accurately measure the position, speed, and type of vehicles, and has the ability to work all-weather and strong anti-interference, at the same time, the cost is also very high, and there are certain limitations in complex traffic environments.

[0003] In recent years, with the progress of deep learning technology, object detection algorithms have made significant developments, especially in terms of accuracy and efficiency. Most traditional object detection methods rely on the anchor box mechanism, which can achieve good object localization, but in complex scenarios, the anchor box method may face a trade-off between accuracy and computational efficiency. Especially in the case of object occlusion, illumination changes, and complex backgrounds, the detection accuracy may be affected.

[0004] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention

[0005] The main purpose of this application is to provide a road image detection method, device, and storage medium, aiming to solve the technical problem of inaccurate road target detection caused by environmental changes in complex urban traffic environments.

[0006] To achieve the above purpose, this application proposes a road image detection method, and the method includes:

[0007] Obtain a road model training image, and inject salt-and-pepper noise into the road model training image to obtain an enhanced image;

[0008] Extract features from the enhanced image according to the reparameterized vision transformer to obtain a basic feature map;

[0009] Perform feature fusion on the basic feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map;

[0010] Perform object detection on the weighted fusion multi-scale feature map according to the road target detection algorithm that combines a fusion knowledge graph and key point detection to obtain a target road detection result;

[0011] Train a road detection model based on the target road detection result, and detect a real-time road image based on the road detection model.

[0012] In one embodiment, the step of performing object detection on the weighted fusion multi-scale feature map according to the road object detection algorithm that fuses a knowledge graph and key point detection to obtain a target road detection result includes:

[0013] Calculate the Gaussian heat map value and target offset of the weighted fusion multi-scale feature map according to the road object detection algorithm that fuses a knowledge graph and key point detection;

[0014] Perform linearly deformable convolution on the weighted fusion multi-scale feature map according to a linearly deformable convolution detection head to obtain a detection feature map;

[0015] Perform strip pooling on the detection feature map to obtain a pooled feature map;

[0016] Identify the object category in the pooled feature map according to the relationship prior knowledge graph to obtain the target object category;

[0017] Obtain a target road detection result according to the Gaussian heat map value, the target offset, and the target object category.

[0018] In one embodiment, the step of performing strip pooling on the detection feature map to obtain a pooled feature map includes:

[0019] Obtain the feature map height, feature map width, first preset channel feature value, and second preset channel feature value according to the detection feature map;

[0020] Calculate a preset channel row feature value according to the first preset channel feature value and the feature map height;

[0021] Calculate a preset channel column feature value according to the second preset channel feature value and the feature map width;

[0022] Perform group normalization and one-dimensional convolution on the preset channel row feature value and the preset channel column feature value respectively to obtain a row prediction feature and a column prediction feature;

[0023] Obtain a pooled feature map according to the row prediction feature, the column prediction feature, and channel information.

[0024] In one embodiment, the step of identifying the object category in the pooled feature map according to the relationship prior knowledge graph to obtain the target object category includes:

[0025] Establish a traffic knowledge graph according to the relationship prior knowledge graph and a common sense knowledge base;

[0026] Obtain the wandering nodes according to the pooled feature map, and obtain the transition probability matrix according to the wandering nodes;

[0027] Obtain the restart probability, the preset number of wandering times, the wandering probability, and the initial probability value according to the restart random walk algorithm;

[0028] Obtain the transition matrix according to the restart probability, the transition probability matrix, the preset number of wandering times, the wandering probability, and the initial probability value, and obtain the semantic consistency matrix according to the transition matrix;

[0029] Obtain the target object category according to the initial label, the target label, the hot spot ratio, the traffic knowledge graph, and the semantic consistency matrix.

[0030] In one embodiment, the step of calculating the Gaussian heatmap value and the target offset of the weighted fusion multi-scale feature map according to the road target detection algorithm that combines the knowledge graph and key point detection includes:

[0031] According to the road target detection algorithm that combines the knowledge graph and key point detection and the weighted fusion multi-scale feature map, obtain the horizontal standard deviation, the vertical standard deviation, the abscissa of the center point, the ordinate of the center point, the prediction probability, the focal loss adjustment parameter, the polynomial loss coefficient, the bounding box size loss, and the bounding box offset loss;

[0032] Calculate the focal loss according to the prediction probability and the focal loss adjustment parameter;

[0033] Calculate the polynomial loss according to the focal loss, the prediction probability, the focal loss adjustment parameter, and the polynomial loss coefficient;

[0034] Calculate the target offset according to the bounding box size loss, the bounding box offset loss, and the polynomial loss;

[0035] Calculate the Gaussian heatmap value according to the horizontal standard deviation, the vertical standard deviation, the abscissa of the center point, and the ordinate of the center point.

[0036] In one embodiment, the step of performing feature fusion on the base feature map according to the weighted feature fusion enhanced feature pyramid network to obtain the weighted fusion multi-scale feature map includes:

[0037] Preprocess the base feature map according to the first convolutional layer and the second convolutional layer of the weighted feature fusion enhanced feature pyramid network to obtain a preprocessed feature map;

[0038] Perform downsampling on the preprocessed feature map according to the reparameterized vision transformer block and the reparameterized vision transformer compression and excitation block to obtain a downsampled feature map;

[0039] Perform feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer;

[0040] Perform multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map.

[0041] In one embodiment, the step of performing feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer includes:

[0042] Extract the feature maps of the downsampled feature map to obtain a first feature map and a second feature map;

[0043] Perform convolution and upsampling on the first feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer;

[0044] Perform convolution and upsampling on the first fusion layer and the second feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a second fusion layer.

[0045] In one embodiment, the step of performing multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map includes:

[0046] Generate weights according to the weighted feature fusion enhanced feature pyramid network to obtain a first weight and a second weight;

[0047] Obtain a first fusion layer feature vector and a second fusion layer feature vector according to the first fusion layer and the second fusion layer;

[0048] Obtain a weighted fusion multi-scale feature map according to the first weight, the first fusion layer feature vector, the second weight, and the second fusion layer feature vector.

[0049] In addition, to achieve the above object, the present application also proposes a road image detection device, and the road image detection device includes: a data acquisition module, configured to acquire a road model training image and a real-time road image, and inject salt-and-pepper noise into the road model training image to obtain an enhanced image;

[0050] A feature extraction module, configured to perform feature extraction on the enhanced image according to the reparameterized vision transformer to obtain a basic feature map;

[0051] A feature fusion module, configured to perform feature fusion on the base feature map according to weighted feature fusion to enhance the feature pyramid network, so as to obtain a weighted fusion multi-scale feature map;

[0052] An object detection module, configured to perform object detection on the weighted fusion multi-scale feature map according to a road object detection algorithm that fuses a knowledge graph and key point detection, so as to obtain an object road detection result;

[0053] An image detection module, configured to train a road detection model according to the object road detection result, and perform detection on the real-time road image according to the trained road detection model.

[0054] In addition, to achieve the above object, the present application also proposes a road image detection device, where the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the road image detection method as described above.

[0055] In addition, to achieve the above object, the present application also proposes a storage medium, where the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the road image detection method as described above are implemented.

[0056] In addition, to achieve the above object, the present application also provides a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the road image detection method as described above are implemented.

[0057] One or more technical solutions proposed by the present application have at least the following technical effects:

[0058] By injecting salt-and-pepper noise into the road model training images to simulate the noise interference during the image transmission process, the robustness of the model in a noisy environment is enhanced, overfitting is avoided, and the generalization ability of the model is increased. The reparameterized vision transformer is used to extract features from the images, improving the feature extraction process, enabling the model to effectively extract useful feature information from complex road images, and enhancing the object recognition and localization capabilities. The weighted feature fusion enhanced feature pyramid network is used to fuse the feature maps obtained from different scales, enabling the model to better handle objects of different sizes and scales, and improving the detection capabilities for small objects, large objects, distant objects, and objects in different directions. During the object detection process, the road object detection algorithm that fuses the knowledge graph and key point detection, by introducing the relational prior knowledge graph, enables the model to not only identify the objects in the image but also understand the spatial and semantic relationships between the objects, further improving the classification accuracy of the objects and the accuracy of position prediction. It solves the problems of poor detection accuracy and robustness of existing road image detection methods in complex environments, such as roads, pedestrians, vehicles, etc. in traffic scenes, due to noise, object scale differences, and unclear relationships between objects. Compared with the existing technologies, the road image detection method adopting this integrated technology has significant advantages in complex traffic environments, especially in terms of noise interference, object occlusion, and multi-scale object detection, improving the overall effect of road object detection and being able to more accurately identify and locate objects such as vehicles, pedestrians, and traffic signs on the road. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing the embodiments consistent with this application, and are used together with the specification to explain the principles of this application.

[0060] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0061] Figure 1 It is a schematic flowchart provided for Embodiment 1 of the road image detection method of this application;

[0062] FIG. 2(a) is a visualization schematic diagram of 0% noise addition provided for Embodiment 1 of the road image detection method of this application;

[0063] FIG. 2(b) is a visualization schematic diagram of 15% noise addition provided for Embodiment 1 of the road image detection method of this application;

[0064] Figure 3Schematic diagram of the noise addition detection effect provided by the first embodiment of the road image detection method of the present application;

[0065] Fig. 4(a) is a schematic diagram of a common convolutional detection head provided by the first embodiment of the road image detection method of the present application;

[0066] Fig. 4(b) is a schematic diagram of a linear deformable convolutional detection head provided by the first embodiment of the road image detection method of the present application;

[0067] Figure 5 Schematic diagram of the relational prior knowledge graph structure provided by the first embodiment of the road image detection method of the present application;

[0068] Figure 6 Flow chart provided by the second embodiment of the road image detection method of the present application;

[0069] Figure 7 Schematic diagram of the module structure of the road image detection device according to the embodiment of the present application;

[0070] Figure 8 Schematic diagram of the device structure of the hardware operating environment involved in the road image detection method according to the embodiment of the present application.

[0071] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners

[0072] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application, and are not used to limit the present application.

[0073] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the accompanying drawings of the specification and specific implementation manners.

[0074] The main solution of the embodiment of the present application is: obtaining a road model training image, injecting salt and pepper noise into the road model training image to obtain an enhanced image; extracting features from the enhanced image according to a reparameterized vision transformer to obtain a basic feature map; performing feature fusion on the basic feature map according to a weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map; performing object detection on the weighted fusion multi-scale feature map according to a road object detection algorithm that combines a fusion knowledge graph and key point detection to obtain a target road detection result; training a road detection model according to the target road detection result, and detecting a real-time road image according to the road detection model.

[0075] In this embodiment, for the convenience of description, the following will be described with an identification road image detection device as the execution subject.

[0076] Due to the inaccurate road target detection caused by environmental changes in the complex urban traffic environment of the prior art, the present application provides a solution. By injecting salt-and-pepper noise into the road model training images to simulate the noise interference during the image transmission process, the robustness of the model in a noisy environment is improved, overfitting is avoided, and the generalization ability of the model is increased. The reparameterized vision transformer is used to extract features from the images, improving the feature extraction process, enabling the model to effectively extract useful feature information from complex road images, and enhancing the target recognition and localization capabilities. The weighted feature fusion enhanced feature pyramid network is used to fuse the feature maps obtained from different scales, enabling the model to better handle targets of different sizes and scales, and enhancing the detection capabilities for small targets, large targets, distant targets, and targets in different directions. During the target detection process, the road target detection algorithm that fuses the knowledge graph and key point detection makes the model not only able to identify the targets in the image, but also understand the spatial and semantic relationships between the targets by introducing the relational prior knowledge graph, further improving the classification accuracy of the targets and the accuracy of position prediction. This solves the problem of poor detection accuracy and robustness of the existing road image detection methods in complex environments, such as the targets of roads, pedestrians, vehicles, etc. in traffic scenes due to noise, target scale differences, and unclear relationships between targets. Compared with the prior art, the road image detection method adopting this integrated technology has significant advantages in complex traffic environments, especially in terms of noise interference, target occlusion, and multi-scale target detection, improving the overall effect of road target detection and enabling more accurate identification and localization of targets such as vehicles, pedestrians, and traffic signs on the road.

[0077] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a road image detection device, etc. that can implement the above functions. Hereinafter, taking the road image detection device as an example, this embodiment and the following embodiments will be described.

[0078] Based on this, an embodiment of the present application provides a road image detection method, referring to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the road image detection method of the present application.

[0079] In this embodiment, the road image detection method includes steps S10 to S50:

[0080] Step S10, obtaining road model training images, and injecting salt-and-pepper noise into the road model training images to obtain enhanced images;

[0081] It should be noted that the road model training images are image datasets used to train road detection models, including different road scenarios, including but not limited to targets such as lanes, traffic signs, vehicles, pedestrians, etc. These road model training images are used to train the model to help the model learn how to identify and detect road targets in different environments.

[0082] Additionally, salt-and-pepper noise is a type of image noise that randomly sets some pixels in the image to black or white. Black pixels represent pepper, and white pixels represent salt. Salt-and-pepper noise can simulate interference during image transmission. By injecting salt-and-pepper noise into the training images, the model can learn how to handle noise and interference in the image during training. Additionally, the enhanced image is the training image after being processed with salt-and-pepper noise. By injecting noise, the quality of the image is damaged to a certain extent, and the model can be trained in a more challenging environment, thereby improving the robustness of the model in practical applications.

[0083] It can be understood that for each training image, there is a 0.5 probability of deciding whether to apply salt-and-pepper noise processing to the image; during salt-and-pepper noise processing, each pixel has a 5% probability of being replaced with a black or white pixel, which will generate "salt" and "pepper" noise in the image. Referring to FIGS. 2(a) and 2(b), where FIGS. 2(a) and 2(b) are respectively the 0% noise addition visualization schematic diagram and the 15% noise addition visualization schematic diagram provided by Embodiment 1 of the road image detection method of the present application. FIG. 2(a) shows the original road image without adding any salt-and-pepper noise. FIG. 2(b) shows the road image with 15% salt-and-pepper noise added. The noise in the image has significantly increased, causing some pixels to be randomly set to black or white, thus affecting the quality of the image and the clarity of the target. The image with salt-and-pepper noise added as a training image can help the model better adapt to these imperfect inputs. The target detection boxes in the image can still accurately mark the positions of the targets, demonstrating the robustness and detection ability of the model in a noisy environment.

[0084] Refer to Figure 3 , Figure 3 which is the noise addition detection effect schematic diagram provided by Embodiment 1 of the road image detection method of the present application.

[0085] As Figure 3As shown, there are significant differences in the detection accuracy between the case without adding noise and the case with added noise under different noise levels. The abscissa in the figure represents the noise content of the image, ranging from 0 to 0.30. The larger the value, the stronger the noise and the worse the image quality. The ordinate represents the object detection accuracy of the model under different noise conditions. mAP50 refers to the average accuracy rate of the model when the IoU threshold is 50%. The higher the value, the higher the detection accuracy of the model. The blue curve represents the change in the detection accuracy of the model without noise enhancement with the increase in noise level. It can be seen that as the noise level increases, the detection accuracy drops rapidly, indicating that the model has poor robustness to noise. The red curve represents the detection accuracy of the model with salt-and-pepper noise enhancement. Although the noise level increases, the red curve is more stable compared to the blue curve, and the decline in detection accuracy is smaller, indicating that the robustness of the model is significantly improved after adding noise enhancement.

[0086] Step S20: Extract features from the enhanced image according to the reparameterized vision transformer to obtain a basic feature map.

[0087] It should be noted that the reparameterized vision transformer (RepViT) is an improved architecture based on the vision transformer (ViT), aiming to improve the efficiency and performance of the model through reparameterization technology. Compared with the traditional vision transformer, RepViT adopts a lightweight structure, combines the advantages of the convolutional neural network (CNN) and the transformer, reduces the computational amount and model parameters, and at the same time maintains good performance in visual tasks. Through the reparameterization of the network structure, RepViT can execute visual tasks with lower latency and higher efficiency, and is suitable for resource-constrained environments. Additionally, the basic feature map is the preliminary image features extracted by the deep neural network, usually the intermediate representation after the image undergoes convolutional operations, pooling, etc. The basic feature map retains the main feature information in the image, such as edges, textures, shapes, etc., and provides basic data for further object detection and classification tasks.

[0088] It can be understood that by using the RepViT network, the key information in the enhanced image is effectively extracted to form a basic feature map. RepViT combines the advantages of the convolutional network and the transformer, improves the efficiency by reducing the computational amount and the number of parameters, and ensures high-quality feature extraction. The obtained basic feature map is the basic data for subsequent object detection, contains the high-level semantic information of the image, and can help the model better identify the objects in the image.

[0089] Step S30: Perform feature fusion on the basic feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map.

[0090] It can be understood that although RepVit has few parameters and fast training, it still has deficiencies when directly used for feature extraction. The size of the feature map output by the backbone network has decreased from 512×512 of the input image to 16×16, and the downsampling factor is 32. This results in the loss of a large amount of spatial information during the downsampling process, the decline of the model's ability to extract target features, and thus the reduction of the model's detection accuracy. To make up for the spatial information lost by the RepVit network, the output parts of the last three layers of the RepVit network are extracted as effective feature layers, the feature of each feature layer is enhanced, and then each feature map is upsampled to a size of 128×128 for weighted fusion.

[0091] It should be noted that the Weighted Feature Fusion Enhanced Feature Pyramid Network (WF-AFPN) is a multi-scale feature fusion method. By fusing feature maps from different scales in a weighted manner, it enhances the model's detection ability for targets of different sizes, especially for the recognition of small and large targets, and improves the detection accuracy. Additionally, the weighted fusion multi-scale feature map is obtained by fusing feature maps from different scales in a weighted manner. Weighted fusion helps to enhance the multi-scale information in the image, enabling the model to establish connections between targets of different sizes and improving the accuracy of target detection.

[0092] In addition, it can be understood that the Weighted Feature Fusion Enhanced Feature Pyramid Network (WF-AFPN) is used to perform feature fusion on the basic feature map. WF-AFPN combines feature maps of multiple scales and fuses different scale information in the image on the basis of weighting. In this way, the model can establish connections between targets of different sizes and improve the detection accuracy for targets of different sizes. The finally output weighted fusion multi-scale feature map contains feature information from different scales and enhances the perception ability for various targets.

[0093] Step S40, perform target detection on the weighted fusion multi-scale feature map according to the road target detection algorithm that fuses the knowledge graph and key point detection, and obtain the target road detection result;

[0094] It should be noted that the road target detection algorithm that fuses the knowledge graph and key point detection (KGKPD) is an algorithm that combines prior knowledge (such as the relationships between targets such as roads, vehicles, and pedestrians) with target detection technology. Through the semantic relationship information in the knowledge graph and the technology of key point detection, the algorithm can better understand the spatial and semantic relationships between targets, thereby improving the accuracy of target detection. Additionally, the target road detection result is the result that contains the target position, category, and other relevant information such as the size and direction of the detection box after being processed by the detection algorithm.

[0095] It is understandable that KGKPD uses the feature information from the weighted fusion multi-scale feature map and combines the prior knowledge in the knowledge graph to identify the targets in the image. By using key point detection, the algorithm can locate the precise positions of the targets, and finally obtain the target road detection results, that is, the positions, categories, and other detection information of each target in the image.

[0096] In a feasible implementation manner, step S40 may include steps S41 to S45:

[0097] Step S41, calculate the Gaussian heat map value and the target offset of the weighted fusion multi-scale feature map according to the road target detection algorithm that fuses the knowledge graph and key point detection;

[0098] It should be noted that the Gaussian heat map value is the heat map data generated by Gaussian distribution in target detection, which is used to represent the central position of the target. Each pixel value in this heat map represents the possibility or confidence of the target at that position. By simulating the distribution of the center points of the target bounding boxes through Gaussian distribution, it helps to locate the target center. The higher the value of this heat map, the closer the central position of the target is to this point, and vice versa, it means the target is farther away or does not exist.

[0099] In addition, the target offset refers to the value used to adjust the position of the predicted bounding box in target detection. It represents the spatial deviation between the detected bounding box and the actual target bounding box. The target offset is usually predicted through a regression model to ensure that the detection box can accurately enclose the target object. This helps to improve the detection accuracy, especially for adjustment when the prediction of the bounding box is not completely accurate.

[0100] It is understandable that by using the road target detection algorithm that fuses the knowledge graph and key point detection, richer image features are obtained through weighted fusion of the multi-scale feature map. Calculate the Gaussian heat map value in the generated feature map, and the Gaussian heat map value is used to indicate the central position of the target. At the same time, calculate the target offset to accurately adjust the detection box of the target and ensure accurate positioning of the target.

[0101] In a feasible implementation manner, step S41 may include steps S411 to S415:

[0102] Step S411, according to the road target detection algorithm that fuses the knowledge graph and key point detection and the weighted fusion multi-scale feature map, obtain the horizontal standard deviation, vertical standard deviation, abscissa of the center point, ordinate of the center point, prediction probability, focal loss adjustment parameter, polynomial loss coefficient, bounding box size loss, and bounding box offset loss;

[0103] It should be noted that the horizontal standard deviation is the horizontal scale in the heat map generated by the Gaussian distribution. It represents the degree of expansion of the target bounding box in the horizontal direction. A larger horizontal standard deviation indicates that the target is wider, while a smaller horizontal standard deviation indicates that the target is narrower. In object detection, the standard deviation is usually associated with the size of the target, helping to determine the actual position and range of the target.

[0104] In addition, the vertical standard deviation is the vertical scale in the heat map generated by the Gaussian distribution. It represents the degree of expansion of the target bounding box in the vertical direction. Similar to the horizontal standard deviation, a larger vertical standard deviation indicates that the target is taller, while a smaller vertical standard deviation indicates that the target is shorter. Through the vertical standard deviation, the height of the target can be more accurately located.

[0105] In addition, the abscissa and ordinate of the center point are the horizontal and vertical coordinates of the center of the target bounding box in the image, respectively.

[0106] In addition, the prediction probability refers to the confidence level of the model that the target belongs to a certain category. In object detection, the model calculates the features of each region in the image and outputs the probability values that the region belongs to different categories. A higher prediction probability indicates that the model has stronger confidence in the existence of the target.

[0107] In addition, the focal loss adjustment parameter is a parameter used to adjust the attention degree of the focal loss function to difficult samples and easy-to-classify samples. When facing the problem of class imbalance, the focal loss can focus on difficult-to-classify samples by adjusting this parameter, thereby improving the recognition ability of the model for minority class samples.

[0108] In addition, the polynomial loss coefficient is a parameter in the polynomial loss function, which is used to adjust the weighted influence of the loss. In object detection, the polynomial loss function is used to further optimize the classification and localization of the target, especially when dealing with class imbalance.

[0109] In addition, the bounding box size loss is a loss function used to measure the gap between the predicted bounding box size and the actual target bounding box. This loss usually includes the differences in width and height. The goal is to minimize the difference between the predicted box and the true box as much as possible, thereby improving the accuracy of object detection.

[0110] In addition, the bounding box offset loss is used to calculate the deviation between the center point of the bounding box predicted by the object detection model and the true bounding box. By regressing the offset of the predicted bounding box, the model can more accurately locate the target.

[0111] Step S412, calculate the focal loss according to the prediction probability and the focal loss adjustment parameter;

[0112] It should be noted that focal loss is a loss function used to handle the problem of class imbalance, especially in object detection tasks. Focal loss assigns higher weights to difficult samples when calculating the loss, reducing the influence of easy-to-classify samples, so that the model pays more attention to difficult-to-classify objects and improves its detection accuracy.

[0113] It can be understood that the focal loss is calculated by adjusting the parameters with the predicted probability and the focal loss adjustment parameter. The predicted probability determines the confidence of each object, while the focal loss adjustment parameter enhances the model's learning ability for minority classes or difficult-to-recognize objects by changing the model's attention to difficult samples. In this way, focal loss helps to improve the model's performance in the class imbalance scenario. The calculation formula of focal loss is as follows:

[0114] L FL = -(1 - P t ) γ log(P t )

[0115] In the formula, P t represents the predicted probability; γ represents the focal loss adjustment parameter. By setting γ, the sensitivity of the model to easy-to-classify samples and difficult samples (usually minority class samples) can be adjusted. When γ increases, the model will pay more attention to those difficult-to-classify samples. (1 - P t ) γ will reduce the weights of those high-confidence samples and increase the weights of low-confidence samples. That is, let the model pay more attention to the misclassified samples that are easy to classify during training, especially in the case of class imbalance.

[0116] Step S413, calculate the polynomial loss according to the focal loss, the predicted probability, the focal loss adjustment parameter, and the polynomial loss coefficient;

[0117] It should be noted that polynomial loss (Poly Loss) is a loss function, usually used to solve the problem of class imbalance. Compared with the traditional cross-entropy loss, polynomial loss will introduce a coefficient to adjust the weights of samples, especially increasing the emphasis on difficult-to-classify samples. By adjusting the coefficient of the loss function, the model is made more sensitive to the class imbalance situation.

[0118] It can be understood that the polynomial loss is calculated by combining the focal loss, the predicted probability, the focal loss adjustment parameter, and the polynomial loss coefficient. The polynomial loss coefficient controls the degree of attention of the loss function to difficult samples. By increasing the loss weights of difficult-to-classify samples, the model can pay more attention to the learning of minority class objects, thus improving the class imbalance problem. The calculation formula of polynomial loss is as follows:

[0119] L poly-1 = LFL +β(1 - P t ) γ+1

[0120] Wherein, L FL represents the focal loss; γ represents the predicted probability; β represents the focal loss adjustment parameter; represents the polynomial loss coefficient. β is usually set to a negative value to adjust the contribution of the focal loss function. In this embodiment, β = -1.

[0121] Step S414, calculate the target offset according to the bounding box size loss, the bounding box offset loss, and the polynomial loss;

[0122] It can be understood that the target offset is calculated by combining the bounding box size loss, the bounding box offset loss, and the polynomial loss. These loss functions work together to adjust the bounding box prediction of the model, making the positioning of the target more accurate. By this method, the accuracy of target detection can be significantly improved, especially in the case where the shape of the target changes or the distance between targets is small. The calculation formula of the target offset is as follows:

[0123] L = L poly-1 + 0.1L size + L offsets

[0124] Wherein, L poly-1 represents the polynomial loss; L size represents the bounding box size loss; L offsets represents the bounding box offset loss.

[0125] Step S415, calculate the Gaussian heatmap value according to the horizontal standard deviation, the vertical standard deviation, the abscissa of the center point, and the ordinate of the center point.

[0126] It can be understood that based on parameters such as the horizontal and vertical standard deviations and the center point coordinates, the Gaussian heatmap value is calculated. These values can indicate the center position of the target and its possible distribution in the image. Through the Gaussian heatmap, the model can more accurately locate the center of the target and make accurate predictions about the size and shape of the target, further improving the accuracy of target detection. The calculation formula of the Gaussian heatmap value is as follows:

[0127]

[0128] Wherein, σ x represents the horizontal standard deviation; σ y represents the vertical standard deviation; x0 represents the abscissa of the center point; y0 represents the ordinate of the center point.

[0129] In addition, it can be understood that the obtained heatmap is

[0130] Step S42: Perform linear deformable convolution on the weighted fusion multi-scale feature map according to the linear deformable convolution detection head to obtain a detection feature map;

[0131] It should be noted that linear deformable convolution (LDConv) is an improved convolution operation that can flexibly adjust the positions of sampling points under different convolution kernel sizes. A significant feature of this convolution operation is that the number of sampling points is linearly related to the size of the convolution kernel, rather than the square relationship in traditional convolution. Therefore, linear deformable convolution performs more efficiently when dealing with irregular shapes and background noise, and is especially suitable for scenarios where the target morphology changes greatly.

[0132] It can be understood that the weighted fusion feature map in step S41 is processed using linear deformable convolution. Through this convolution operation, the model can flexibly adjust the position and size of the convolution kernel to more precisely capture the features and shape of the target. The output of this step is a detection feature map that contains the specific position and shape information of the target in the image.

[0133] Refer to FIGS. 4(a) and 4(b), where FIGS. 4(a) and 4(b) are schematic diagrams of a conventional convolution detection head and a linear deformable convolution detection head provided in the first embodiment of the road image detection method of the present application, respectively.

[0134] As shown in FIG. 4(a), the figure shows that conventional convolution extracts features by sliding a convolution kernel on the input image. The convolution kernel operates at each position of the image to gradually generate a feature map. Through the operation of conventional convolution, low-level features in the image, such as edges and textures, can be obtained. As shown in FIG. 4(b), different from conventional convolution, when processing an image, linear deformable convolution can flexibly adjust the position and shape of the convolution kernel according to the shape of the target. The convolution kernel in the figure presents a deformed shape and can better adapt to the changes of the target, such as different target sizes and shapes, thereby improving the detection ability of the model in complex scenarios. This method can capture the changes in the target shape better than conventional convolution, improving the flexibility and accuracy of target detection.

[0135] Step S43: Perform strip pooling on the detection feature map to obtain a pooled feature map;

[0136] It should be noted that strip pooling is a feature pooling method used to enhance the spatial feature capture ability of convolutional neural networks. In strip pooling, each channel in the feature map performs pooling operations in the row and column directions respectively, thereby capturing long-range spatial dependencies. Strip pooling helps to enhance target features, especially for image information with long-range dependencies, and can better capture the spatial structure of the target.

[0137] It can be understood that strip pooling operations are performed on the detected feature map obtained in step S42. By pooling the rows and columns of each channel, the model can capture long-range spatial dependency information, thereby better understanding the target and background features in the image. The pooled feature map has higher spatial dependency and can provide more robust object recognition.

[0138] In a feasible implementation, step S43 may include steps S431 to S435:

[0139] Step S431, according to the detected feature map, obtain the feature map height, feature map width, first preset channel feature value, and second preset channel feature value;

[0140] It should be noted that the feature map height and feature map width respectively refer to the spatial dimensions of the detected feature map, usually represented as the number of pixels in the vertical and horizontal directions of the image. The height and width determine the resolution of the feature map and affect the detail level of subsequent feature processing.

[0141] In addition, the first preset channel feature value and the second preset channel feature value respectively represent the feature values of specific channels calculated through certain layers in the network, usually extracted by the model at specific layers or nodes and processed according to preset criteria. The first and second preset channel feature values respectively represent the specific features extracted in a certain channel dimension and are used for subsequent calculations.

[0142] Step S432, calculate the preset channel row feature value according to the first preset channel feature value and the feature map height;

[0143] It should be noted that the preset channel row feature value is the result obtained by calculating the first preset channel feature value and the feature map height. It usually represents the important features in the row dimension of the feature map and is a further extraction and representation of the features in the row direction.

[0144] It can be understood that the first preset channel feature value is combined with the height of the feature map to calculate the preset channel row feature value. This feature value is particularly important in subsequent processing. It helps to determine the feature processing and feature extraction in the row direction. In this way, the model can further refine and enhance the vertical features in the image. The calculation formula for the preset channel row feature value is as follows:

[0145]

[0146] Wherein, x c (h, i) represents the first preset channel eigenvalue; H represents the height of the feature map. In this embodiment, the first preset channel eigenvalue represents the eigenvalue of the preset channel c at the h-th row and the i-th column.

[0147] Step S433, calculate the preset channel column eigenvalue according to the second preset channel eigenvalue and the width of the feature map;

[0148] It should be noted that the preset channel column eigenvalue is the result obtained by calculating the second preset channel eigenvalue and the width of the feature map. It represents the features in the column dimension of the feature map and is used to represent the information in the horizontal direction.

[0149] It can be understood that the second preset channel eigenvalue is combined with the width of the feature map to calculate the preset channel column eigenvalue. This eigenvalue helps to extract the horizontal features in the image and is the basis for the model to perform more refined feature learning. By calculating the column features, the model can accurately identify the local features of the target object. The calculation formula of the preset channel column eigenvalue is as follows:

[0150]

[0151] Wherein, x c (j, w) represents the second preset channel eigenvalue; W represents the width of the feature map. In this embodiment, the second preset channel eigenvalue represents the eigenvalue of the preset channel c at the j-th row and the w-th column.

[0152] Step S434, perform group normalization and one-dimensional convolution on the preset channel row eigenvalue and the preset channel column eigenvalue respectively to obtain the row prediction feature and the column prediction feature;

[0153] It should be noted that group normalization is a normalization method that divides the feature map into multiple groups and performs normalization operations on each group. This method is particularly effective when dealing with large-scale inputs and can maintain the stability of network training.

[0154] In addition, one-dimensional convolution is a convolution operation that is usually used to process data in a single direction. Different from two-dimensional convolution, one-dimensional convolution only performs a sliding operation in one direction and can more efficiently extract the features in a certain dimension.

[0155] It can be understood that the row prediction feature and the column prediction feature are respectively the results obtained by further processing the preset channel row and column feature values, representing the predicted values of the features in the image. These features will be used for subsequent pooling operations to assist in object detection and localization. The calculation formulas for the row prediction feature and the column prediction feature are as follows:

[0156] y h = σ(G n (F h (z h )))

[0157] In the formula, z h represents the preset channel row feature value; σ represents the activation function; F h represents the one-dimensional convolution operation applied along the height direction; G n represents group normalization.

[0158] y w = σ(G n (F w (z w )))

[0159] In the formula, represents the preset channel column feature value; σ represents the activation function; F w represents the one-dimensional convolution operation applied along the width direction; G n represents group normalization.

[0160] Step S435: Obtain a pooled feature map according to the row prediction feature, the column prediction feature, and the channel information.

[0161] It can be understood that by combining the prediction feature, the column prediction feature, and the channel information, a pooled feature map is obtained. The pooled feature map contains the key information for object detection and reduces the spatial resolution through the pooling operation, enhancing the computational efficiency of the model and its adaptability to complex backgrounds. This feature map provides the necessary information support for subsequent object detection. The feature calculation formula for the pooled feature map is as follows:

[0162] Y = x c × y h × y w

[0163] In the formula, y h represents the row prediction feature; y w represents the column prediction feature; X c represents the channel information; Y represents the final output feature, and based on this feature, a pooled feature map can be obtained.

[0164] Step S44: Identify the object category in the pooled feature map according to the relationship prior knowledge graph to obtain the target object category;

[0165] It should be noted that the Relation Prior Knowledge Graph (RPKG) encodes knowledge and semantic relationships into a graph structure to assist object detection during model inference. The nodes in the graph represent entities, and the edges represent the relationships between entities, constituting a knowledge structure for enhancing the understanding ability of object detection algorithms. By combining with the detection results, the knowledge graph can intervene in the output of the model, helping the model more accurately identify the relationships between objects, thereby improving detection accuracy.

[0166] In addition, the target object category refers to the type of target object recognized by the model during object detection. Each target object is assigned a category label according to its visual features, such as "car", "pedestrian", etc. Category recognition is crucial for object detection as it helps the system understand the specific meaning of each target in the image, providing clear information for subsequent processing.

[0167] It can be understood that the relation prior knowledge graph is used to combine with the pooled feature map to perform category recognition on the objects in the image. Through the prior knowledge in the graph, the model can infer the category of the object, for example, determining whether an object in the image is a "pedestrian" or a "vehicle". This method helps improve the accuracy of object detection in complex scenarios.

[0168] In a feasible implementation manner, step S44 may include steps S441 to S445:

[0169] Step S441, establish a traffic knowledge graph according to the relation prior knowledge graph and the common sense knowledge base;

[0170] It should be noted that the common sense knowledge base is a database containing extensive background knowledge, covering the general cognition and common sense in human daily life. This knowledge usually exists in the form of natural language and has a low degree of structuring. The information in the common sense knowledge base can be used for reasoning, judgment, and supplementing missing information to improve the understanding and decision-making ability of the model. In this embodiment, the common sense knowledge base can be ConceptNet. ConceptNet connects different concepts (such as "car", "pedestrian", etc.) and their relationships (such as "is", "has", "belongs to", etc.) in the form of a graph structure to help the computer understand common sense information in daily life. An example table of the content of ConceptNet is shown in Table 1 below.

[0171] Table 1

[0172]

[0173] It can be understood that the data in the knowledge base will be converted into the form of triples, and each triple contains three parts: subject, predicate, and object. These triples are used to represent the relationships between concepts. For example, the relationship "apple - belongs_to - fruit" can be represented as a triple: "apple - is_a - fruit". During the conversion process, the corresponding relationships between the symbols involved and their specific meanings will be stored. The table of the corresponding relationships between symbols and specific meanings is shown in Table 2 below.

[0174] Table 2

[0175] Symbol Meaning / r Relationship, relation, edge / c Concept, concept, node / en English / zh Chinese uri Concatenate the (relationship, start point, end point) triple into a string, which can be used as an index json Additional information

[0176] During the conversion process, all data representing negative relationships (such as "does_not_belong_to", "is_not", etc.) will be removed. The purpose of doing this is to only retain the data of positive or neutral relationships, avoid unnecessary interference, and ensure that the subsequent constructed semantic consistency matrix can more accurately reflect the positive relationships. In order to adapt the ConceptNet knowledge base to the construction process of the semantic consistency matrix, it is necessary to perform format conversion on it, that is, convert the data into English triples and exclude all entries representing negative relationships for subsequent storage.

[0177] In addition, the traffic knowledge graph is a knowledge graph specifically for the traffic field constructed based on the relationship prior knowledge graph and the common sense knowledge base. It contains various entities involved in the traffic system, such as vehicles, roads, traffic signs, etc., and the relationships between them, such as the relationship between roads and traffic flow, roads and traffic signals, etc. The traffic knowledge graph helps the model better understand the semantic relationships in the traffic environment, enhancing the model's reasoning ability and accuracy.

[0178] It can be understood that the traffic knowledge graph is constructed by combining the relationship prior knowledge graph and the common sense knowledge base. This graph can represent the entities in the traffic field and the relationships between them in the form of a graph structure, helping the model understand the complex relationships in the traffic environment. This provides basic support for subsequent object detection and semantic reasoning, enabling the model to perform more accurate reasoning and decision-making.

[0179] Refer to Figure 5 , Figure 5 which is the schematic diagram of the relationship prior knowledge graph structure provided for the first embodiment of the road image detection method of this application.

[0180] Such as Figure 5As shown, on the left side of the figure, an S matrix is presented. This matrix is the semantic consistency matrix, which contains multiple data blocks. These data blocks are distinguished by green and blue, representing different data sets or feature dimensions. The middle part of the picture shows the generation process of the classification matrix. The data blocks are connected to the node graph, forming a network structure, where each node represents an entity or feature, and the edges between the nodes indicate the relationships between them. Through this graphical network structure, the model can understand the semantic relationships between entities and provide support for the subsequent construction of the knowledge graph. The lower part of the figure describes the construction process of the knowledge graph. The green and blue areas still exist, indicating the entities and relationships extracted from the original data by the knowledge graph. Through graph construction, the data blocks transform from the initial simple form into a more complex knowledge structure, and the graph begins to possess more semantic information. The right side of the figure shows the process of feature enhancement. Here, the data blocks in the figure interact with other elements, and through feature enhancement operations, certain important features in the graph are further strengthened. The main purpose of this stage is to enhance the data features through weighting or other means, enabling the model to more accurately utilize these features during the target classification and reasoning processes. Finally, the right side of the figure shows the final "RPKG structure" after processing. The color of the data blocks turns red, representing the final output after the aforementioned processing and enhancement. Through this structure, the model can more accurately identify and classify the target object, and finally output the category of the target object.

[0181] Step S442: Obtain the wandering nodes according to the pooled feature map, and obtain the transition probability matrix according to the wandering nodes;

[0182] It should be noted that the wandering nodes refer to the nodes that serve as the starting point or traversing point in the random walk process in the graph structure. In graph theory, the random walk usually starts from an initial node and transfers to other nodes with a certain probability. The wandering nodes play an important role in the knowledge graph and are used to obtain the semantic relationships between the targets.

[0183] In addition, the transition probability matrix is a matrix that describes the transition probabilities between the nodes in the graph. In the random walk process, each element in the matrix represents the transition probability from one node to another. It reflects the relative importance between different nodes in the graph and helps calculate the node probabilities in the wandering path.

[0184] It can be understood that the wandering nodes are obtained through the pooled feature map, and the transition probability matrix is calculated based on these nodes. The pooled feature map provides important abstract features in the image. The wandering nodes are selected from it, and the relationships between the nodes are described by the transition probability matrix, which provides a data basis for the subsequent random walk process.

[0185] Step S443: Obtain the restart probability, the preset number of walks, the walking probability, and the initial probability value according to the restart random walk algorithm;

[0186] It should be noted that the restart random walk algorithm (Random Walk with Restart, RWR) is a graph algorithm that simulates the random walk process starting from a certain node. And at each step of the walk, there is a certain probability of "restarting" to the starting node. Compared with the ordinary random walk algorithm, RWR introduces a restart mechanism, enabling the walk path to return to the starting node randomly, which helps to strengthen the importance of certain nodes in the graph.

[0187] In addition, the restart probability refers to the probability that the walk "restarts" from the current node to the starting node during the restart random walk process. A higher restart probability means that the walk process is more likely to return to the starting node, thereby enhancing the semantic consistency between nodes in the graph.

[0188] In addition, the preset number of walks refers to the set number of walks when performing the restart random walk. Generally, the more walks there are, the deeper the algorithm understands the graph structure, and thus it can better capture the relationships between nodes in the graph.

[0189] In addition, the walking probability refers to the probability of moving from one node to another in each random walk. It is usually determined by the transition probability matrix of the graph and is used to describe the strength of the relationship between nodes.

[0190] In addition, the initial probability value refers to the starting probability of each node when performing the random walk. It is usually set to the same initial value, indicating the possibility of starting the walk from each node.

[0191] Step S444: Obtain the transition matrix according to the restart probability, the transition probability matrix, the preset number of walks, the walking probability, and the initial probability value, and obtain the semantic consistency matrix according to the transition matrix;

[0192] It should be noted that the transition matrix is a matrix obtained by calculating factors such as the transition probability matrix, the walking probability, and the restart probability. It describes the relative transition probability between nodes during the walk process and helps to capture the semantic relationships between nodes.

[0193] It can be understood that through the restart probability, the transition probability matrix, the number of walks, the walking probability, and the initial probability, the RWR algorithm starts the random walk. Each walk will be selected according to the probability distribution of the current node and the transition rules, and the "restart" mechanism ensures that there is a certain probability of returning to the starting point. As the number of walks increases, the model's understanding of the nodes in the graph gradually deepens, and finally the transition matrix is formed. The calculation formula of the transition matrix is as follows:

[0194]

[0195] Where k represents the preset number of walks; α represents the restart probability, α∈(0,1). The higher the restart probability, the more likely the random walk is to restart from the starting node; W i,j represents the transition probability matrix, specifically the transition probability matrix from node I to node j; and Both represent the probability of wandering, among which, represents the probability of walking at node i for the kth time, represents the probability of node i in the k-1th walk; e represents the initial probability value and is the starting vector. The transfer matrix can be obtained by calculating the transfer matrix formula. The transfer matrix expression is as follows:

[0196]

[0197] In the formula, the transfer matrix R i,j It represents the probability of transitioning from one state category to another when the operator is in a certain state category; represents the starting probability; α represents the restart probability; represents the probability of node i in k walks. As the number of random walks increases, the similarity R between nodes i,j will converge to a stable value.

[0198] It should be noted that the semantic consistency matrix is derived from the transfer matrix and is used to represent the semantic consistency between nodes in the graph. It reflects the degree of association between nodes at the semantic level and can help the model better understand the knowledge and relationships in the graph.

[0199] It can be understood that since the required semantic consistency matrix is a symmetric matrix, the semantic consistency matrix can be obtained according to node i and node j. The semantic consistency matrix calculation formula is as follows:

[0200]

[0201] In the formula, S i,j and S j,i They represent the semantic consistency matrix between node i and node j, and between node j and node i. Since the semantic consistency matrix is symmetric, that is, S i,j =S j,i ; R i,j and R j,i They represent the transfer matrices from node i to node j and from node j to node i respectively.

[0202] Step S445 , obtaining the target object category according to the initial label, the target label, the hotspot proportion, the traffic knowledge graph, and the semantic consistency matrix.

[0203] It should be noted that the initial labels refer to the class labels assigned to each target object for the input data when the model starts training. The initial labels usually come from manual annotation or the output of a previous classification model. They provide a starting point for the classification task of the model and help the model learn how to identify target objects from the input data.

[0204] In addition, the target labels refer to the predicted labels obtained after the inference phase when the model training is completed. These labels represent the classification results of the model for the target objects. The target labels are usually the class identifiers output by the model, which are compared with the initial labels to evaluate the accuracy and performance of the model.

[0205] In addition, the hotspot ratio refers to a parameter in the object detection process where the model assigns higher weights to certain regions (i.e., "hotspot" regions) in the image. These "hotspot" regions usually refer to the regions where the target objects are located, and the model adjusts its attention according to the importance of these regions. The hotspot ratio is usually determined through the heatmap or attention map of the image, with the goal of enhancing the model's attention to key regions and improving the accuracy of object detection.

[0206] It can be understood that by integrating the initial labels, target labels, hotspot ratio, traffic knowledge graph, and semantic consistency matrix, a classification matrix is calculated, and based on the classification matrix, the precise classification results of the target objects can be obtained. This process enables the model, when performing object detection, not only to identify the presence of objects but also to accurately determine the object categories in a complex environment, improving the robustness and accuracy of the object detection system. The calculation formula of the classification matrix is as follows:

[0207]

[0208] In the formula, b represents the initial label; b′ represents the target label; e represents the hotspot ratio, H represents all the hotspots, e ∈ (0,1) is used to weigh the ratio of the two items, and in this embodiment, e = 0.75; represents the traffic knowledge graph, which can specifically be all the class labels; S b,b′ represents the semantic consistency matrix.

[0209] Step S45, obtaining the target road detection result according to the Gaussian heatmap value, the target offset, and the target object category.

[0210] It can be understood that by comprehensively analyzing the Gaussian heatmap value, the target offset, and the target object category, the target road detection result is finally obtained. These results usually include the position, category of the target, and other features such as direction and size, and are used for subsequent object tracking, analysis, or action decision-making.

[0211] Step S50, train a road detection model based on the target road detection result, and detect a real-time road image according to the road detection model.

[0212] It should be noted that the road detection model is a deep learning model trained to detect road-related targets such as vehicles, pedestrians, and traffic signs in images. Through the training dataset, the model can identify and locate the targets in the image.

[0213] In addition, the real-time road image refers to the image data obtained in real time from road monitoring devices, cameras, etc. Since it is captured in real time, these images have high timeliness and can reflect information such as the current traffic conditions and road environment.

[0214] It can be understood that by training on a large number of labeled road image data, the model can learn how to identify different types of road targets and generate accurate detection results. After training, the model can be applied to the detection task of real-time road images. The real-time road images are input into the trained model, and the model analyzes these images to quickly obtain the detection results of road targets. This provides real-time data support for traffic monitoring, autonomous driving, etc., and helps to make decisions quickly.

[0215] This embodiment provides a road image detection method. By adopting the technical means of combining the reparameterized vision transformer (RepViT) and the weighted feature fusion enhanced feature pyramid network (WF-AFPN), it solves the robustness problem of traditional object detection models under complex environments and noise interference. By injecting salt-and-pepper noise to enhance the training images, it improves the detection accuracy of the model in multi-noise scenarios, and finally realizes high-precision road target detection under multi-scale and multi-object conditions. This method effectively improves the adaptability and accuracy of the object detection model in practical applications. Especially when processing real-time road images, it significantly enhances the detection ability and real-time responsiveness of the model.

[0216] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 6 , steps S31 to S34 are further included in step S30 of the road image detection method:

[0217] Step S31, preprocess the basic feature map according to the first convolutional layer and the second convolutional layer of the weighted feature fusion enhanced feature pyramid network to obtain a preprocessed feature map;

[0218] It should be noted that the first convolutional layer and the second convolutional layer are two hierarchical structures in a convolutional neural network (CNN), which are used to extract different levels of features from the input image. In the first convolutional layer, usually lower-level features (such as edges, textures, etc.) are extracted, while the second convolutional layer further processes and refines these low-level features to capture more complex patterns. After preprocessing, a preprocessed feature map can be obtained.

[0219] In this embodiment, the first convolutional layer and the second convolutional layer have a stride of 2 and a size of 3×3. The number of filters in the first convolutional layer is set to 24, and the stride of the second convolutional layer is set to 48. This design helps the model capture local features and gradually increases the receptive field through multi-layer stacking, thereby learning a richer feature representation.

[0220] It can be understood that weighted feature fusion is used to enhance the first convolutional layer and the second convolutional layer of the feature pyramid network to preprocess the base feature map. Through convolutional operations, the model can extract more and more hierarchical feature information from the base feature map. The purpose of preprocessing is to make the features of the image richer so that subsequent processing steps can better perform object recognition and localization.

[0221] Step S32: Downsample the preprocessed feature map according to the reparameterized vision transformer block and the reparameterized vision transformer squeeze-and-excitation block to obtain a downsampled feature map;

[0222] It should be noted that the reparameterized vision transformer block (RepViT Block) is an improved form of the vision transformer (ViT), aiming to improve the performance of the model by reducing the computational amount and improving the efficiency. The RepViT block optimizes the traditional transformer structure through reparameterization technology, enabling it to reduce the computational overhead while ensuring the accuracy when processing images.

[0223] In addition, the reparameterized vision transformer squeeze-and-excitation block (RepViT SE Block) is a further optimization of the RepViT block, introducing the squeeze-and-excitation (SE) mechanism. This mechanism improves the model's attention to important features by introducing adaptive adjustment in the feature channel dimension. In this embodiment, the 1st, 3rd, 5th... blocks in each stage use SE layers, and this interleaved placement method aims to maximize the improvement in accuracy while controlling the increase in latency.

[0224] In addition, the downsampled feature map is a feature map processed by reducing the spatial resolution of the image. Usually, downsampling is achieved through convolutional operations, pooling operations, or transposed convolutional operations, aiming to reduce the computational amount and compress the image data while retaining key information.

[0225] Step S33: Perform feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer;

[0226] It can be understood that since the size of the input image is 512×512 and it undergoes five downsamplings by the RepVit network, the size of the final feature map is 16×16. If the 16×16 feature map in CenterNet is directly upsampled to 128×128 using transposed convolution, it may cause loss or blurring of feature information, making the upsampled feature map unable to retain the important information in the original feature map, thus affecting the detection result. At the same time, it may lead to insufficient detection accuracy for small targets. The Asymptotic Feature Pyramid Network (AFPN) solves the problem of information loss during feature transmission by allowing direct interaction between non-adjacent layers and introduces an adaptive spatial fusion operation to handle information conflicts between features at different levels. To further improve the effect of feature fusion, the WF-AFPN module is proposed. It introduces the multi-scale fusion strategy of AFPN, combines it with the SE module to obtain a multi-scale feature map with channel weights, and finally uses transposed convolution to upsample feature maps of different scales to the same scale and introduces learnable weight parameters in the adaptive spatial fusion to form the Weight Adaptive spatial fusion operation (WASF), further improving the effect of adaptive adjustment and fusion. And depth convolution is used to replace the original conventional convolution, reducing the computational amount and model parameters while maintaining the performance of the network.

[0227] It should be noted that the first fusion layer and the second fusion layer refer to the result layers obtained through feature fusion, which represent the combination of feature information at different levels. The first fusion layer usually retains more low-level features, while the second fusion layer fuses more high-level features for further processing and analysis.

[0228] It can be understood that the weighted feature fusion enhanced feature pyramid network obtains the first fusion layer and the second fusion layer through feature fusion of the downsampled feature map. This process combines feature maps of different scales and different levels to form a more comprehensive and in-depth feature representation. Through feature fusion, the model can simultaneously capture the detailed information and global semantic information of the image, providing strong support for object detection.

[0229] In a feasible implementation manner, step S33 may include steps S331 to S333:

[0230] Step S331: Extract the feature maps of the downsampled feature map to obtain a first feature map and a second feature map;

[0231] It should be noted that the first feature map and the second feature map are two feature maps with different scales extracted from the downsampled feature map. The first feature map contains lower-level local features, while the second feature map contains higher-level semantic information. By extracting the downsampled feature map, more diverse feature representations can be obtained, providing multi-dimensional information for subsequent fusion and processing.

[0232] It can be understood that feature maps at different levels are extracted from the main network and are respectively labeled as C3, C4, and C5. These feature maps represent semantic information at different levels. Among them, C3 and C4 belong to lower-level features with finer-grained local information, while C5 belongs to higher-level features with stronger global semantic information.

[0233] Step S332: According to the weighted feature fusion enhanced feature pyramid network, perform convolution and upsampling on the first feature map to obtain a first fusion layer;

[0234] It should be noted that convolution is a commonly used operation in deep learning for extracting features of images. In the convolution operation, the convolution kernel (or filter) slides through different parts of the input image, performs weighted summation of local regions, and generates a new feature map. Convolution helps to capture local patterns in the image, such as edges, textures, etc.

[0235] In addition, upsampling is an operation to restore the spatial dimensions of an image by increasing its resolution. Common upsampling methods include transposed convolution (or deconvolution) and interpolation methods. Upsampling is usually used in the decoding stage of a convolutional neural network to restore the spatial resolution of the feature map to a larger scale, so as to capture more spatial information.

[0236] In addition, the first fusion layer is a feature layer obtained by performing convolution and upsampling on the first feature map. This layer is usually used to further process the detailed information in the image and fuse this information into a higher-level feature representation. The first fusion layer provides the low-level features of the image and the spatial information enhanced through convolution and upsampling operations.

[0237] It can be understood that the first fusion layer is successively subjected to 1×1 and 3×3 depthwise convolutions, then the SE module is used to strengthen the feature representation. Immediately afterwards, one output is upsampled to become the P3 layer, and this P3 layer is the first fusion layer.

[0238] Step S333: According to the weighted feature fusion enhanced feature pyramid network, perform convolution and upsampling on the first fusion layer and the second feature map to obtain a second fusion layer.

[0239] It should be noted that the second fusion layer is a feature layer obtained by performing convolution and upsampling on the first fusion layer and the second feature map. It represents the result of combining the first fusion layer and the second feature map, and usually contains more hierarchical information. The second fusion layer is crucial for the multi-scale feature fusion of the model, and it helps the model better understand complex target shapes.

[0240] It can be understood that the first fusion layer and the second feature map generate the second fusion layer through convolution and upsampling operations. Through the convolution operation, the features of the first fusion layer and the second feature map are fused to further extract deeper information. The upsampling restores the spatial resolution of the feature map and enhances the ability to accurately locate and recognize targets. The second fusion layer provides a more comprehensive feature representation for subsequent object detection, enabling the model to process image information at multiple scales. The upsampled result becomes the P3 layer, and another part is fused with the C4 layer through the WASF module to become the L3 layer, which is the second fusion layer and participates in the next round of input. This process is repeated until all feature layers are fused.

[0241] Step S34: Perform multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map.

[0242] It can be understood that the multi-scale weighted fusion of the first fusion layer and the second fusion layer is performed by the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map. This process combines information at multiple scales and weights it according to the contribution of each scale, so that the final feature map can provide sufficient target information at each scale, thereby improving the model's detection ability for targets of different sizes.

[0243] In a feasible implementation manner, step S34 may include steps S341 to S343:

[0244] Step S341: Generate weights according to the weighted feature fusion enhanced feature pyramid network to obtain a first weight and a second weight;

[0245] It should be noted that the first weight and the second weight are the weights generated by the weighted feature fusion enhanced feature pyramid network, corresponding to the features of the first fusion layer and the second fusion layer respectively. These weights reflect the importance of the two fusion layers in the feature fusion process, and the feature map with a higher weight will receive more attention in the fusion process.

[0246] It is understandable that the model uses weighted feature fusion to enhance the Feature Pyramid Network to generate the first weight and the second weight. These weights are allocated according to the contributions of the respective feature maps, indicating the importance of each feature map in the final fusion. The first weight and the second weight are respectively used to measure the importance of the first fusion layer and the second fusion layer in the weighted fusion process, thereby helping the model determine how to combine features at different levels.

[0247] Step S342: Obtain a first fusion layer feature vector and a second fusion layer feature vector according to the first fusion layer and the second fusion layer;

[0248] It should be noted that the first fusion layer feature vector and the second fusion layer feature vector are respectively feature representations obtained from the first fusion layer and the second fusion layer. By converting the feature maps of each fusion layer into feature vectors, the model can more conveniently perform further processing and calculations. These feature vectors contain multi-scale information of the image, providing a necessary basis for subsequent weighted fusion.

[0249] It is understandable that according to the feature maps of the first fusion layer and the second fusion layer, through feature extraction operations, they are converted into feature vectors. In this way, the multi-dimensional features of the first fusion layer and the second fusion layer are compressed into one-dimensional vectors, facilitating subsequent weighted fusion. Each feature vector represents important pattern information in the image, enabling the model to perform more effective learning and decision-making in subsequent steps.

[0250] Step S343: Obtain a weighted fusion multi-scale feature map according to the first weight, the first fusion layer feature vector, the second weight, and the second fusion layer feature vector.

[0251] It is understandable that by performing weighted fusion on the first weight, the first fusion layer feature vector, the second weight, and the second fusion layer feature vector, a weighted fusion multi-scale feature map is obtained. The model weights each feature map according to its weight, and the feature map with a higher weight plays a greater role in the final fusion. By fusing the weighted feature vectors, a comprehensive feature map containing multi-scale information is obtained, which can better represent the target features in the image, especially for the recognition and localization of multi-scale targets with higher precision. The calculation formula for the feature vector of the weighted fusion multi-scale feature map is as follows:

[0252]

[0253] In the formula, and respectively represent the first weight and the second weight, represents the weight of the higher layer, and respectively represent the feature space weights of three different layers at the m-th layer, and satisfy and respectively represent the feature vector of the first fusion layer and the feature vector of the second fusion layer. represents the feature vector of a higher layer; W i , i ∈ 1, 2, 3 represent learnable weight parameters.

[0254] This embodiment provides a road image detection method. By combining a weighted feature fusion enhanced feature pyramid network (WF-AFPN), it solves the problems of information loss caused by downsampling and insufficient feature fusion in multi-scale object detection, and achieves the beneficial effects of improving object detection accuracy and reducing computational complexity through multi-scale weighted fusion. By introducing an adaptive spatial fusion and channel weighting mechanism, the fusion process of the feature map is further optimized, enabling the model to effectively extract object features at different scales. Especially when dealing with objects of different sizes and complex backgrounds, the robustness and accuracy of the model are significantly improved.

[0255] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the road image detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.

[0256] This application also provides a road image detection device. Please refer to Figure 7 , the road image detection device includes:

[0257] A data acquisition module 10, configured to acquire a road model training image and a real-time road image, and inject salt-and-pepper noise into the road model training image to obtain an enhanced image;

[0258] A feature extraction module 20, configured to perform feature extraction on the enhanced image according to a reparameterized vision transformer to obtain a basic feature map;

[0259] A feature fusion module 30, configured to perform feature fusion on the basic feature map according to a weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map;

[0260] An object detection module 40, configured to perform object detection on the weighted fusion multi-scale feature map according to a road object detection algorithm that combines a fusion knowledge graph and key point detection to obtain a target road detection result;

[0261] An image detection module 50, configured to train a road detection model according to the target road detection result and detect the real-time road image according to the trained road detection model.

[0262] The road image detection device provided by this application adopts the road image detection method in the above-mentioned embodiment, and can solve the technical problem that the road target detection is inaccurate due to environmental changes in a complex urban traffic environment. Compared with the prior art, the beneficial effects of the road image detection device provided by this application are the same as those of the road image detection method provided by the above-mentioned embodiment, and other technical features in the road image detection device are the same as the features disclosed in the method of the above-mentioned embodiment, which will not be elaborated here.

[0263] In one embodiment, the target detection module 40 is further configured to calculate the Gaussian heatmap value and the target offset of the weighted fusion multi-scale feature map according to the road target detection algorithm that fuses the knowledge graph and key point detection; perform linear deformable convolution on the weighted fusion multi-scale feature map by using the linear deformable convolution detection head to obtain a detection feature map; perform strip pooling on the detection feature map to obtain a pooled feature map; identify the object category in the pooled feature map according to the relationship prior knowledge graph to obtain the target object category; and obtain the target road detection result according to the Gaussian heatmap value, the target offset, and the target object category.

[0264] In one embodiment, the target detection module 40 is further configured to obtain the feature map height, the feature map width, the first preset channel feature value, and the second preset channel feature value according to the detection feature map; calculate the preset channel row feature value according to the first preset channel feature value and the feature map height; calculate the preset channel column feature value according to the second preset channel feature value and the feature map width; perform group normalization and one-dimensional convolution on the preset channel row feature value and the preset channel column feature value respectively to obtain a row prediction feature and a column prediction feature; and obtain the pooled feature map according to the row prediction feature, the column prediction feature, and the channel information.

[0265] In one embodiment, the target detection module 40 is further configured to establish a traffic knowledge graph according to the relationship prior knowledge graph and the common sense knowledge base; obtain a wandering node according to the pooled feature map, and obtain a transition probability matrix according to the wandering node; obtain a restart probability, a preset number of wandering times, a wandering probability, and an initial probability value according to the restart random walk algorithm; obtain a transition matrix according to the restart probability, the transition probability matrix, the preset number of wandering times, the wandering probability, and the initial probability value, and obtain a semantic consistency matrix according to the transition matrix; and obtain the target object category according to the initial label, the target label, the hot spot ratio, the traffic knowledge graph, and the semantic consistency matrix.

[0266] In one embodiment, the target detection module 40 is further configured to obtain a horizontal standard deviation, a vertical standard deviation, an abscissa of the center point, an ordinate of the center point, a prediction probability, a focal loss adjustment parameter, a polynomial loss coefficient, a bounding box size loss, and a bounding box offset loss according to a road target detection algorithm that fuses a knowledge graph and key point detection and the weighted fusion multi-scale feature map; calculate a focal loss according to the prediction probability and the focal loss adjustment parameter; calculate a polynomial loss according to the focal loss, the prediction probability, the focal loss adjustment parameter, and the polynomial loss coefficient; calculate a target offset amount according to the bounding box size loss, the bounding box offset loss, and the polynomial loss; and calculate a Gaussian heatmap value according to the horizontal standard deviation, the vertical standard deviation, the abscissa of the center point, and the ordinate of the center point.

[0267] In one embodiment, the feature fusion module 30 is further configured to preprocess the basic feature map according to the first convolutional layer and the second convolutional layer of the weighted feature fusion enhanced feature pyramid network to obtain a preprocessed feature map; perform downsampling on the preprocessed feature map according to the reparameterized vision transformer block and the reparameterized vision transformer squeeze-and-excitation block to obtain a downsampled feature map; perform feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer; and perform multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map.

[0268] In one embodiment, the feature fusion module 30 is further configured to extract the feature maps of the downsampled feature map to obtain a first feature map and a second feature map; perform convolution and upsampling on the first feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer; and perform convolution and upsampling on the first fusion layer and the second feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a second fusion layer.

[0269] In one embodiment, the feature fusion module 30 is further configured to generate weights according to the weighted feature fusion enhanced feature pyramid network to obtain a first weight and a second weight; obtain a first fusion layer feature vector and a second fusion layer feature vector according to the first fusion layer and the second fusion layer; and obtain a weighted fusion multi-scale feature map according to the first weight, the first fusion layer feature vector, the second weight, and the second fusion layer feature vector.

[0270] The present application provides a road image detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the road image detection method in the first embodiment above.

[0271] Reference is made below Figure 8 , which shows a schematic structural diagram of a road image detection device suitable for implementing the embodiments of the present application. The road image detection device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 8 The road image detection device shown is only an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.

[0272] As Figure 8As shown, the road image detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the ROM (Read Only Memory) 1002 or the program loaded from the storage device 1003 into the RAM (Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the road image detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the road image detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a road image detection device with various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems can be alternatively implemented or had.

[0273] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0274] The road image detection device provided by the present application adopts the road image detection method in the above embodiments, and can solve the technical problem that the road target detection is inaccurate due to environmental changes in a complex urban traffic environment. Compared with the prior art, the beneficial effects of the road image detection device provided by the present application are the same as those of the road image detection method provided by the above embodiments, and other technical features in the road image detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0275] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0276] As described above, the above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0277] This application provides a computer-readable storage medium with computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the road image detection method in the above embodiments.

[0278] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0279] The above computer-readable storage medium can be included in the road image detection device; it can also exist separately without being assembled into the road image detection device.

[0280] The above computer-readable storage medium carries one or more programs, which, when executed by a road image detection device, cause the road image detection device to: obtain road model training images, inject salt-and-pepper noise into the road model training images to obtain enhanced images; extract features from the enhanced images according to a reparameterized vision transformer to obtain basic feature maps; perform feature fusion on the basic feature maps according to a weighted feature fusion enhanced feature pyramid network to obtain weighted fusion multi-scale feature maps; perform object detection on the weighted fusion multi-scale feature maps according to a road object detection algorithm that combines a fusion knowledge graph and key point detection to obtain target road detection results; train a road detection model according to the target road detection results, and detect real-time road images according to the road detection model.

[0281] Computer program code for performing the operations of this application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).

[0282] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0283] The modules involved in the embodiments of the present application can be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.

[0284] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned road image detection method, which can solve the technical problem of inaccurate road target detection caused by environmental changes in a complex urban traffic environment. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the road image detection method provided by the above embodiments, and will not be elaborated here.

[0285] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the road image detection method as described above are implemented.

[0286] The computer program product provided by the present application can solve the technical problem of inaccurate road target detection caused by environmental changes in a complex urban traffic environment. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the road image detection method provided by the above embodiments, and will not be elaborated here.

[0287] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application under the technical concept of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A road image detection method, characterized in that, The method includes: Obtaining a road model training image and injecting salt-and-pepper noise into the road model training image to obtain an enhanced image; Performing feature extraction on the enhanced image according to a reparameterized vision transformer to obtain a basic feature map; Performing feature fusion on the basic feature map according to a weighted feature fusion enhanced feature pyramid network to obtain a weighted fusion multi-scale feature map; Performing object detection on the weighted fusion multi-scale feature map according to a road object detection algorithm that fuses a knowledge graph and key point detection to obtain a target road detection result; Training a road detection model according to the target road detection result and detecting a real-time road image according to the road detection model.

2. The method according to claim 1, characterized in that The step of performing object detection on the weighted fusion multi-scale feature map according to a road object detection algorithm that fuses a knowledge graph and key point detection to obtain a target road detection result includes: Calculating the Gaussian heatmap value and target offset of the weighted fusion multi-scale feature map according to a road object detection algorithm that fuses a knowledge graph and key point detection; Performing linearly deformable convolution on the weighted fusion multi-scale feature map according to a linearly deformable convolution detection head to obtain a detection feature map; Performing strip pooling on the detection feature map to obtain a pooled feature map; Identifying the object category in the pooled feature map according to a relational prior knowledge graph to obtain a target object category; Obtaining a target road detection result according to the Gaussian heatmap value, the target offset, and the target object category.

3. The method according to claim 2, wherein The step of performing strip pooling on the detection feature map to obtain a pooled feature map includes: Obtaining the feature map height, feature map width, first preset channel feature value, and second preset channel feature value according to the detection feature map; Calculating a preset channel row feature value according to the first preset channel feature value and the feature map height; Calculating a preset channel column feature value according to the second preset channel feature value and the feature map width; Performing group normalization and one-dimensional convolution on the preset channel row feature value and the preset channel column feature value respectively to obtain a row prediction feature and a column prediction feature; Obtaining a pooled feature map according to the row prediction feature, the column prediction feature, and channel information.

4. The method according to claim 2, wherein The step of identifying the object category in the pooled feature map according to a relational prior knowledge graph to obtain a target object category includes: Establishing a traffic knowledge graph according to a relational prior knowledge graph and a common sense knowledge base; Obtaining wandering nodes according to the pooled feature map and obtaining a transition probability matrix according to the wandering nodes; Obtaining a restart probability, a preset number of wandering times, a wandering probability, and an initial probability value according to the restart random walk algorithm; Obtaining a transition matrix according to the restart probability, the transition probability matrix, the preset number of wandering times, the wandering probability, and the initial probability value, and obtaining a semantic consistency matrix according to the transition matrix; Obtaining a target object category according to an initial label, a target label, a hot spot ratio, the traffic knowledge graph, and the semantic consistency matrix.

5. The method according to claim 2, characterized in that, The steps of calculating the Gaussian heatmap value and the target offset of the weighted fusion multi-scale feature map according to the road target detection algorithm integrating the knowledge graph and key point detection include: Based on the road target detection algorithm integrating the knowledge graph and key point detection and the weighted fusion multi-scale feature map, obtain the horizontal standard deviation, vertical standard deviation, abscissa of the center point, ordinate of the center point, prediction probability, focal loss adjustment parameter, polynomial loss coefficient, bounding box size loss, and bounding box offset loss; Calculate the focal loss according to the prediction probability and the focal loss adjustment parameter; Calculate the polynomial loss according to the focal loss, the prediction probability, the focal loss adjustment parameter, and the polynomial loss coefficient; Calculate the target offset according to the bounding box size loss, the bounding box offset loss, and the polynomial loss; Calculate the Gaussian heatmap value according to the horizontal standard deviation, the vertical standard deviation, the abscissa of the center point, and the ordinate of the center point.

6. The method according to claim 1, wherein The steps of performing feature fusion on the basic feature map according to the weighted feature fusion enhanced feature pyramid network to obtain the weighted fusion multi-scale feature map include: Preprocess the basic feature map according to the first convolutional layer and the second convolutional layer of the weighted feature fusion enhanced feature pyramid network to obtain a preprocessed feature map; Perform downsampling on the preprocessed feature map according to the reparameterized vision transformer block and the reparameterized vision transformer squeeze-and-excitation block to obtain a downsampled feature map; Perform feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer; Perform multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain the weighted fusion multi-scale feature map.

7. The method according to claim 6, wherein The steps of performing feature fusion on the downsampled feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer and a second fusion layer include: Extract the feature maps of the downsampled feature map to obtain a first feature map and a second feature map; Perform convolution and upsampling on the first feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a first fusion layer; Perform convolution and upsampling on the first fusion layer and the second feature map according to the weighted feature fusion enhanced feature pyramid network to obtain a second fusion layer.

8. The method according to claim 6, wherein The steps of performing multi-scale weighted fusion on the first fusion layer and the second fusion layer according to the weighted feature fusion enhanced feature pyramid network to obtain the weighted fusion multi-scale feature map include: Generate weights according to the weighted feature fusion enhanced feature pyramid network to obtain a first weight and a second weight; Obtain a first fusion layer feature vector and a second fusion layer feature vector according to the first fusion layer and the second fusion layer; Obtain the weighted fusion multi-scale feature map according to the first weight, the first fusion layer feature vector, the second weight, and the second fusion layer feature vector.

9. A road image detection device, characterized in that, The device includes: A data acquisition module, configured to acquire road model training images and real-time road images, and inject salt-and-pepper noise into the road model training images to obtain enhanced images; A feature extraction module, configured to extract features from the enhanced images according to a reparameterized vision transformer to obtain basic feature maps; A feature fusion module, configured to perform feature fusion on the basic feature maps according to a weighted feature fusion enhanced feature pyramid network to obtain weighted fusion multi-scale feature maps; An object detection module, configured to perform object detection on the weighted fusion multi-scale feature maps according to a road object detection algorithm that fuses a knowledge graph and key point detection to obtain target road detection results; An image detection module, configured to train a road detection model according to the target road detection results, and detect the real-time road images according to the trained road detection model.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the road image detection method according to any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Multi-scale road vehicle detection method based on re-parameterized visual converter

    CN120564164A