Road Median Detection Method and Device Based on Global Image Receptive Field

By introducing multi-scale enhancement modules of image global receptive field and ResNet residual network structure, the problem of insufficient receptive field in traditional methods is solved, and high-precision automatic detection of road isolation zones is realized, which is suitable for intelligent traffic management.

CN118155161BActive Publication Date: 2025-07-11BEIJING TRANSPORTATION RES CENT +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410105601.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-25
Publication Date
2025-07-11
Estimated Expiration
2044-01-25

AI Technical Summary

Technical Problem

The existing road isolation zone detection methods rely on manual annotation, which is difficult to meet the needs of smart transportation. The lack of receptive fields of traditional convolutional neural networks leads to incomplete information capture and loss of key information, making it difficult to accurately identify complex and diverse road isolation zones.

Method used

The road isolation zone detection model based on the image global receptive field is adopted, and the ResNet residual network structure is used to introduce multi-scale enhancement modules and residual connections. The receptive field of the network layer is increased through the multi-scale feature enhancement module, and end-to-end automatic detection is carried out in combination with deep learning technology.

Benefits of technology

It improves the accuracy and robustness of road isolation belt identification, and can accurately identify the non-isolation belt form in complex traffic scenarios, providing a more effective solution and is suitable for smart traffic management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118155161B_ABST
    Figure CN118155161B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for detecting road isolation belts based on the global receptive field of images, belonging to the technical field of road detection based on machine vision. An image of a road isolation belt to be detected is acquired; the acquired image is processed by a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type; wherein, the road isolation belt detection model based on the global receptive field of the image is trained based on the training data; the predicted road isolation belt form is obtained through the forward propagation process of successive convolutions, the loss between the predicted result and the true label is calculated, and backpropagation is performed through the loss to update the model weights until the set number of iteration rounds is reached. The present invention increases the receptive field of the network layer, accurately judges the motor-vehicle and non-motor-vehicle isolation form of the scene, and improves the comprehensiveness and accuracy of the prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of road detection based on machine vision, and particularly relates to a road median detection method and device based on the global receptive field of an image. Background Art

[0002] The purpose of setting a road median between motor vehicle lanes and non-motor vehicle lanes is to enhance traffic safety and regularize traffic flow. As a physical separator, the median can effectively divide motor vehicle lanes and non-motor vehicle lanes, reduce the cross-interference between motor vehicles and non-motor vehicles, and decrease the probability of traffic accidents. In addition, the road median helps to maintain traffic order, promote smooth traffic, realize the diversion of motor vehicles, non-motor vehicles and pedestrians, and form a highly ordered and safe traffic system. It lays a solid foundation for the efficient operation of the urban slow traffic system.

[0003] Detecting the road median between motor vehicle lanes and non-motor vehicle lanes is a key task in the field of slow traffic. Nowadays, the forms of road medians are diverse, including fences, solid lines, green belts and cement, etc. Accurately identifying the type of road median is crucial for traffic management and urban planning. Different types of road medians may have different impacts on traffic flow, and automatic identification can provide valuable information for traffic planning and safety. With the continuous development of intelligent transportation technology, it becomes more feasible to use image processing and artificial intelligence technology for automatic identification of road medians.

[0004] In the field of computer vision, image classification has always been a very crucial task, and its main goal is to classify the input image into different predefined categories. To achieve high-precision image classification, researchers have proposed a variety of different methods, including traditional classical methods and the latest deep learning technologies. The core idea of these methods is usually to extract features from the image and then use these features for classification to determine the category to which the image belongs.

[0005] With the rapid development of deep learning technology, deep neural networks are increasingly widely used in image classification tasks. In particular, the convolutional neural network (CNN) has become the preferred method for image classification tasks. However, for some specific image classification tasks, such as the recognition of road median forms, traditional convolutional neural networks may face some challenges. Because the types of road medians are diverse, and they may have different shapes, colors and textures in different scenarios. Traditional CNNs may be limited when dealing with this diversity, so more advanced methods are needed to solve this problem.

[0006] In this context, the residual neural network based on global receptive field has emerged. This is a deep learning method that introduces residual connections to enable the network to learn residual information in the image, so as to better handle various complex and diverse image features. The global receptive field means that each neuron in the network has a comprehensive perception and understanding of the overall information of the input image. This expansion of the receptive field enables the residual neural network to better capture the contextual information and details in the image, thereby improving the accuracy and robustness of the image classification task.

[0007] At present, there is no practical application scheme similar to the present invention for identifying the form of road isolation strips. However, it is worth pointing out that this task can essentially be regarded as converting the detection problem into an image classification problem. In terms of the network structure of the algorithm model, ResNet (Residual Network) proposes a residual block structure, which solves the problem that the network is difficult to train due to gradient disappearance or gradient explosion in deep neural networks, and the problem that deeper layers lead to a decrease in network performance, making it perform well in image classification tasks.

[0008] The network structure of ResNet is based on a deep convolutional neural network, but its core innovation lies in the introduction of the residual block structure. In this structure, each residual block contains two branches, one is the identity mapping, and the other is the learned residual. This design enables the network to learn residual information, that is, the difference of the image relative to the identity mapping, rather than directly learning the complete feature mapping. This solves the gradient vanishing and gradient exploding problems in traditional deep networks, allowing deeper network structures to be trained and optimized.

[0009] In image classification tasks, ResNet has been widely used due to its deep scalability and high performance. By stacking multiple residual blocks, the network can achieve a very deep structure and effectively capture the complex features in the image. In addition, due to the design of the residual block, increasing the network depth no longer leads to performance degradation, but improves classification accuracy.

[0010] In summary, the existing technology lacks automatic detection methods and devices for road median strips using deep learning mechanisms. Most current methods still rely on manual annotation for classification, which is difficult to meet the needs in the context of smart transportation. When convolutional neural networks are directly used to classify traffic images, the actual receptive field is often small, and the theoretical receptive field is not enough to cover the entire road median strip, which may lead to incomplete capture of information. In addition, the downsampling process (such as pooling) may lead to the loss of a large amount of key information, and in the task of machine non-median strip recognition, the median strip usually spans the entire image, so a larger receptive field is required to achieve more accurate prediction output. Summary of the Invention

[0011] The object of the present invention is to provide a method and device for detecting road isolation belts based on the global receptive field of images, so as to solve at least one of the technical problems existing in the above-mentioned background art.

[0012] In order to achieve the above object, the present invention adopts the following technical solutions:

[0013] On the one hand, the present invention provides a method for training a road isolation belt detection model based on the global receptive field of images, including:

[0014] Obtain training data; wherein, the training data includes multiple road isolation belt images and labels for annotating the types of isolation belts in the images;

[0015] Train the road isolation belt detection model based on the global receptive field of images using the training data; wherein, the predicted form of the road isolation belt is obtained through the forward propagation process of layer-by-layer convolution, the loss between the prediction result and the true label is calculated, and the loss is used for backpropagation to update the model weights until the set number of iteration rounds is reached; wherein, the road isolation belt detection model based on the global receptive field of images includes a multi-scale enhancement module, and the multi-scale enhancement module consists of a 1x1 convolutional layer and batch normalization for reducing the number of input channels, a series of 3x3 convolutional layers and batch normalization layers for feature extraction and multi-scale information, followed by a 1x1 convolutional layer and batch normalization layer for increasing the number of output channels, and finally a ReLU activation function, and an optional downsampling part for reducing the size of the input feature map.

[0016] Optionally, the network structure of the road isolation belt detection model based on the global receptive field of images adopts a ResNet residual network structure, and a set of smaller filter banks are introduced to replace a single 3×3 convolutional kernel with n channels; wherein, each filter bank contains w channels, and n = s×w, where s is the number of filters in the filter bank; these filter banks are interconnected in a hierarchical manner similar to residual connections to increase the scale diversity of the output features.

[0017] Optionally, for the input features, a 1×1 convolutional layer is used to adjust the number of channels; the input features are divided into three levels; each group of filters first extracts features from a group of input feature maps, and the output features of the previous group are passed to another group of input features Figure 1 and passed to the next group of filters until all input feature maps are processed; the feature maps of all groups are connected together and input into a 1×1 convolutional layer to achieve complete information fusion.

[0018] Optionally, by traversing the training dataset, calculate the loss of each training sample and perform backpropagation to update the model parameters; every certain number of iterations, the model will perform forward propagation on the test dataset to calculate the test loss and accuracy to evaluate the performance of the model on unseen data until the specified number of iterations is reached.

[0019] Optionally, obtain training data, including: classifying traffic images according to different forms of road dividers: ". / facility strip" indicates that the road divider in this type of image is cement or green belt, ". / fence" indicates that the road divider in this type of image is a fence, ". / line" indicates that the road divider in this type of image is a solid line, and ". / else" indicates that the form of the road divider in this type of image is other than the above three cases; use code to read and process the dataset images to obtain the training data and test dataset.

[0020] Optionally, use code to read and process the dataset images to obtain the training data and test data, including:

[0021] Use the glob library and the glob.glob method to obtain the file paths of all images in the folder containing image files, and these image files are under the path specified in the code;

[0022] Define four image categories: 'else', 'facility strip', 'fence', and 'line', and the mapping relationship between them and integer labels;

[0023] Create a custom dataset class MyDataset, which inherits from torch.utils.data.Dataset and is used to process image data. In the initialization, pass in the image path, label, and data transformation function;

[0024] The data transformation function transform defines a series of transformation operations, including resizing the image size to 256x256 pixels and converting the image to a tensor;

[0025] Create a data loader dataloader to divide the dataset into batches, and each batch contains 16 images and corresponding labels;

[0026] Divide the dataset: randomly permute all image paths and labels, and divide the dataset into a training set and a test set.

[0027] In a second aspect, the present invention provides a method for detecting road dividers based on the global receptive field of an image, including:

[0028] Obtain the road divider image to be detected;

[0029] Process the acquired image using a road median detection model based on the global receptive field of the image to obtain the detection result of the road median type; wherein, the road median detection model based on the global receptive field of the image is obtained according to the model training method described in the first aspect.

[0030] In a third aspect, the present invention provides a road median detection device based on the global receptive field of the image, including:

[0031] An acquisition module, configured to acquire a road median image to be detected;

[0032] A processing module, configured to process the acquired image using a road median detection model based on the global receptive field of the image to obtain the detection result of the road median type; wherein, the road median detection model based on the global receptive field of the image is obtained according to the model training method described in the first aspect.

[0033] In a fourth aspect, the present invention provides a non-transitory computer-readable storage medium, which is used to store computer instructions. When the computer instructions are executed by a processor, the method for detecting a road median based on the global receptive field of the image as described in the second aspect is implemented.

[0034] In a fifth aspect, the present invention provides a computer device, including a memory and a processor, where the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the method for detecting a road median based on the global receptive field of the image as described in the second aspect.

[0035] In a sixth aspect, the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the method for detecting a road median based on the global receptive field of the image as described in the second aspect.

[0036] Term Explanation:

[0037] Receptive Field: The receptive field refers to the range of input data that a neuron in a neural network perceives. In a convolutional neural network (CNN), the receptive field represents the size of the input image area that a neuron is concerned about. Smaller receptive fields are usually used to capture local features, while larger receptive fields are used to understand global structures and context information.

[0038] Residual Neural Network (ResNet): A residual neural network is a deep learning architecture designed to address the problems of vanishing gradients and exploding gradients in deep neural networks. By introducing the structure of residual blocks, it allows for skip connections, enabling the network to be deeper and easier to train. The key innovation of ResNet is to use the identity mapping as a reference model, and further layers can learn residual information through residual connections, thus improving network performance.

[0039] Road median strip: A road median strip is a traffic facility on the road, usually an area used to separate different driving directions or lanes, which can take different forms, such as solid lines, dashed lines, fences, cement, or green belts. The road median strip helps to improve traffic safety, guide the flow of vehicles, and separate different types of traffic.

[0040] Slow traffic: Slow traffic refers to traffic modes with relatively low speeds in urban or road traffic, usually including non-motor vehicles such as walking, cycling, skateboarding, etc., as well as pedestrians. Slow traffic plays an important role in urban sustainable development and reducing traffic congestion.

[0041] Advantages of the present invention: The residual neural network technology based on the global receptive field is different from previous network models. Through the proposed multi-scale feature enhancement module, features can be represented at a finer-grained level, increasing the receptive field of the network layer, and enabling the development of a more accurate end-to-end recognition technology and device for non-median strip forms. By inputting a traffic image, a reliable prediction output can be obtained, accurately judging the non-median strip form of the scene, including but not limited to fences, solid lines, green belts, or other forms, improving the comprehensiveness and accuracy of the recognition results, and providing an effective and innovative solution to the problem of non-median strip recognition in complex road scenarios.

[0042] The advantages of additional aspects of the present invention will be more clearly given in the following description part, or can be understood through the practice of the present invention. Brief Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is a flowchart of the training method for the road median strip detection model based on the global receptive field of images according to the embodiments of the present invention.

[0045] Figure 2 Flowchart of the road isolation belt detection method based on the global receptive field of an image according to an embodiment of the present invention.

[0046] Figure 3 Schematic diagram of the ResNet residual network structure in the prior art.

[0047] Figure 4 Schematic diagram of the improved ResNet residual network structure according to an embodiment of the present invention.

[0048] Figure 5 Image with the inspection result marked output by the model according to an embodiment of the present invention. Detailed implementation manners

[0049] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The implementation manners described below through the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be construed as a limitation to the present invention.

[0050] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art in the field to which the present invention belongs.

[0051] It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless defined as such herein.

[0052] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the phrase "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements and / or groups thereof.

[0053] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0054] To facilitate the understanding of the present invention, the following takes specific embodiments in conjunction with the accompanying drawings to further explain and illustrate the present invention, and the specific embodiments do not constitute a limitation to the embodiments of the present invention.

[0055] Those skilled in the art should understand that the drawings are only schematic diagrams of the embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.

[0056] Embodiment 1

[0057] In this Embodiment 1, first, a road isolation belt detection device based on the global receptive field of an image is provided, including: an acquisition module for acquiring an image of a road isolation belt to be detected; a processing module for processing the acquired image by using a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type.

[0058] In this embodiment, based on the above device, a road isolation belt detection method based on the global receptive field of an image is implemented, including: using the acquisition module to acquire an image of a road isolation belt to be detected; using the processing module to process the acquired image by using a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type; wherein, the road isolation belt detection model based on the global receptive field of the image is obtained according to the training method of the road isolation belt detection model based on the global receptive field of the image.

[0059] In this embodiment, a method for training a road isolation belt detection model based on the global receptive field of an image includes: obtaining training data; wherein, the training data includes multiple road isolation belt images and labels for annotating the types of isolation belts in the images; training the road isolation belt detection model based on the global receptive field of the image using the training data; wherein, the predicted form of the road isolation belt is obtained through the forward propagation process of successive convolutions, the loss between the prediction result and the true label is calculated, and backpropagation is performed through the loss to update the model weights until the set number of iteration rounds is reached; wherein, the road isolation belt detection model based on the global receptive field of the image includes a multi-scale enhancement module, and the multi-scale enhancement module consists of a 1x1 convolutional layer and batch normalization for reducing the number of input channels, a series of 3x3 convolutional layers and batch normalization layers for feature extraction and multi-scale information, followed by a 1x1 convolutional layer and batch normalization layer for increasing the number of output channels, and finally a ReLU activation function, and an optional downsampling part for reducing the size of the input feature map.

[0060] Wherein, the network structure of the road isolation belt detection model based on the global receptive field of the image adopts a ResNet residual network structure, and a set of smaller filter banks are introduced to replace a single 3×3 convolutional kernel with n channels; wherein, each filter bank contains w channels, and n = s×w, where s is the number of filters in the filter bank; these filter banks are interrelated in a hierarchical manner similar to residual connections to increase the scale diversity of the output features.

[0061] For the input features, a 1×1 convolutional layer is used to adjust the number of channels; the input features are divided into three levels; each group of filters first extracts features from a group of input feature maps, and the output features of the previous group are passed together with another group of input features Figure 1 to the next group of filters until all input feature maps are processed; the feature maps of all groups are concatenated and input into a 1×1 convolutional layer to achieve complete information fusion.

[0062] By traversing the training data set, the loss of each training sample is calculated and backpropagation is performed to update the model parameters; every certain number of iterations, the model performs forward propagation on the test data set, calculates the test loss and accuracy to evaluate the performance of the model on unseen data until the specified number of iterations is reached.

[0063] Obtain training data, including: classifying traffic images according to different forms of road dividers: ". / facility strip" indicates that the road divider in this type of image is cement or green belt, ". / fence" indicates that the road divider in this type of image is a fence, ". / line" indicates that the road divider in this type of image is a solid line, and ". / else" indicates that the form of the road divider in this type of image is other situations except the above three; use code to read and process the dataset images to obtain training data and test datasets.

[0064] Using code to read and process the dataset images, obtaining training data and test data includes:

[0065] Use the glob library and the glob.glob method to obtain the file paths of all images in the folder containing image files, and these image files are in the path specified by the code;

[0066] Define four image categories: 'else', 'facility strip', 'fence', and 'line', and the mapping relationship between them and integer labels;

[0067] Create a custom dataset class MyDataset, which inherits from torch.utils.data.Dataset and is used to process image data. In the initialization, pass in the image path, label, and data transformation function;

[0068] The data transformation function transform defines a series of transformation operations, including resizing the image to 256x256 pixels and converting the image to a tensor;

[0069] Create a data loader dataloader to divide the dataset into batches, and each batch contains 16 images and corresponding labels;

[0070] Divide the dataset: randomly permute all image paths and labels, and divide the dataset into a training set and a test set.

[0071] Example 2

[0072] In this Example 2, a method for detecting road dividers based on the global receptive field of an image is provided, and this method uses a road divider detection model based on the global receptive field of an image.

[0073] First, in this embodiment, to train the road median detection model based on the global receptive field of the image, a series of configuration tasks need to be carried out. These configurations include installing the Linux operating system, creating a development environment for Python 3.7 or higher, and configuring the deep learning framework PyTorch 1.4 or higher. Since the algorithm adopted is based on a deep learning model, it is recommended to perform model training in a GPU environment, and it is necessary to install the GPU version of PyTorch 1.4 or higher, as well as the corresponding version of the CUDA parallel computing architecture.

[0074] The basic process of training the road median detection model in the slow traffic scenario is as Figure 1 shown. In the model training stage, the training set constructed from the traffic images to be detected is input, and the predicted road median form is obtained through the forward propagation process with successive convolutions. Then, the loss between the predicted result and the true label is calculated, and backpropagation is performed through the loss to update the model weights. This process is executed iteratively until the set number of epochs is reached.

[0075] In the model testing stage, the test set data is loaded, and the trained model is used to generate the predicted results. Subsequently, the calculation of evaluation metrics is carried out, and the model performance is judged based on these metrics. If the model fails to meet the expected requirements, it will return to the training stage for further adjustment and training. Once the expected performance level is achieved, the model weights are saved, completing the entire process of the technological invention and obtaining the final solution.

[0076] The model obtained through the above model training method is used to detect the road median form in traffic images, as Figure 2 shown. It adopts the residual neural network technology of deep learning, realizes the end-to-end automatic detection of the road median, and significantly improves the efficiency of judging the separation form of motor vehicles and non-motor vehicles on the road in the slow traffic scenario. This innovation enables us to more effectively cope with complex traffic scenarios and provides important support for urban traffic management. Compared with the original ResNet residual network structure as Figure 3 shown, a set of smaller filter banks are introduced to replace a single 3×3 convolutional kernel with n channels. Each filter bank contains w channels, where n = s×w, s is the number of filters in the filter bank, and s = 3 in this embodiment. These filter banks are interconnected in a hierarchical manner similar to the residual connection (as Figure 4 shown) to increase the scale diversity of the output features, and this module is named "Bottleneck-Multi".

[0077] Specifically, for the input features, first, the 1×1 convolutional layer is used to adjust the number of channels. Then, these features are divided into three levels, as Figure 4As shown. Each group of filters first extracts features from a set of input feature maps. Then, the output features of the previous group are passed to the next group of filters together with another set of input features Figure 1 This process is repeated until all input feature maps are processed. Finally, the feature maps of all groups are concatenated and input into a 1×1 convolutional layer to achieve complete information fusion. In this embodiment, the letter I is used to represent the input features, and the letter O is used to represent the output features.

[0078] Using this module, the features can either be directly output through a single convolution or convolved again after being added to the input feature maps of the next group. Therefore, the output can be obtained from a single set of input features through a single convolution or from two sets of input features through convolutions at different levels. This enables the resulting output features to obtain richer combinations at different scales, thereby increasing the receptive field of the network layer. In addition, compared with the original residual structure module, the newly proposed multi-scale feature enhancement module reduces the computational complexity and the number of parameters. In the original network structure, the 3×3 convolutional layer brings relatively high computational complexity. For example, when the number of input channels is 512, the computational complexity of the original network is 512×3×3×512. In contrast, in the new module we proposed, the features are divided into three groups, and each 3×3 convolutional layer only requires 1 / 8 of the original computational complexity, thus effectively reducing the computational burden.

[0079] Such an improvement enhances the receptive field of each network layer and exhibits multi-scale features at a finer-grained level, providing more accurate performance for the identification of road isolation belts. This is particularly important for the task of separating motor vehicles from non-motor vehicles, because according to prior knowledge, the isolation belt usually runs through the entire image, and only based on the global receptive field of the image can a more accurate identification of the form of the isolation belt be achieved. This innovative design makes the network model more suitable for complex road scenarios, improves the accuracy and robustness of separating motor vehicles from non-motor vehicles, and brings an important breakthrough to the field of intelligent transportation.

[0080] Specifically, during model training, the inputs of the algorithm are: 1. Slow traffic scene image data: including the training set and the test set, stored in different folders according to different road isolation belts, and the folder name where the image is located is used as the true label to calculate the loss value with the predicted value; 2. Model algorithm hyperparameters: including the cropping size of the image, the batch size, the number of iterations, and the learning rate during training, etc. The output of the algorithm is: obtaining the parameter weights of the trained model algorithm that meets the performance evaluation criteria and the images with the inspection results marked (such as Figure 5 shown).

[0081] Execution steps:

[0082] One: Input image preprocessing stage

[0083] Step 1-1: Divide traffic images into four parts according to different forms of road dividers: 1. ". / facility strip" indicates that the road divider in this type of image is cement or a green belt; 2. ". / fence" indicates that the road divider in this type of image is a fence; 3. ". / line" indicates that the road divider in this type of image is a solid line; 4. ". / else" indicates that the form of the road divider in this type of image is other situations except the above three cases.

[0084] Step 1-2: Use code to read and process dataset images for use by the deep learning model, mainly including the following steps:

[0085] (1) Use the glob library and the glob.glob method to obtain the file paths of all images in the folder containing image files, and these image files are in the path specified in the code.

[0086] (2) Define four image classes: 'else', 'facility strip', 'fence', and 'line', and the mapping relationship between them and integer labels.

[0087] (3) Create a custom dataset class MyDataset, which inherits from torch.utils.data.Dataset and is used to process image data. In the initialization, pass in the image path, label, and data transformation (conversion) function.

[0088] (4) The data transformation function transform defines a series of transformation operations, including resizing the image size to 256x256 pixels and converting the image to a tensor (Tensor).

[0089] (5) Create a data loader dataloader to divide the dataset into batches, with each batch containing 16 images and corresponding labels. This is to effectively load and process the data.

[0090] (6) Divide the dataset. First, randomly permute (shuffle) all image paths and labels. Then, divide the dataset into a training set and a test set, with 80% of the data for training and 20% for testing.

[0091] (7) Finally, create instances of the training dataset (train_ds) and the test dataset (test_ds), as well as the corresponding data loaders (train_dl and test_dl).

[0092] II. Model Training Stage

[0093] Step 2-1: Define the network structure. First, we need to define our newly proposed multi-scale enhancement module "Bottleneck-Multi". This is the basic component of the network model and is used to form the entire network structure. This module consists of a 1x1 convolutional layer and batch normalization to reduce the number of input channels, a series of 3x3 convolutional layers and batch normalization layers for feature extraction and multi-scale information, followed by a 1x1 convolutional layer and batch normalization layer to increase the number of output channels, and finally a ReLU activation function. An optional downsampling part is used to reduce the size of the input feature map. Then, based on this module, a new "ResNet" class and model function are constructed.

[0094] Step 2-2: Set the optimizer, loss function, and various hyperparameters for training the network model. The cross-entropy loss function is selected to define the loss function. The Adam optimizer is used to optimize the weights of the model, and the learning rate is 0.003. The model is trained by iterating a specified number of epochs (here it is 120). The model is transferred to the GPU (if available) for accelerated calculation.

[0095] Step 2-3: The model first traverses the training dataset, calculates the loss of each training sample, and performs backpropagation to update the model parameters. This process is to train the model to better fit the training data. Then, every certain number of iterations, the model will perform forward propagation on the test dataset, calculate the test loss and accuracy, to evaluate the performance of the model on unseen data. These steps are alternated until the specified number of iterations (epochs) is reached.

[0096] Step 2-4: Save the trained model to the file "belt_model.pth".

[0097] III. Model prediction stage

[0098] Step 3-1: Traverse all image files in the specified folder and obtain the file names in the folder using os.listdir.

[0099] Step 3-2: Load the image and open the image using the PIL library. Use transforms to preprocess the image, resize it to a tensor of size (256, 256), and normalize it. Define the network model, which has the same structure as the model used during training, but remove the last classification layer and modify the classifier on this basis to match the task. Load the trained model, which was obtained through training before and the file name is "belt_model.pth".

[0100] Step 3-3: Pass the processed image to the model for prediction. This process uses model.eval() to ensure that the model is in evaluation mode. Make a prediction through the model to obtain the output probability distribution. output.argmax(1) gets the index of the prediction result, and then looks up the corresponding class label in data_class according to the index.

[0101] Step 3-4: Print out the prediction result, indicating which class the image is classified into, and add the class label to the image. Finally, save the image with the prediction result to the folder ". / result" with the same file name as the original image.

[0102] Embodiment 3

[0103] Embodiment 3 provides a non-transitory computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the method for detecting a road isolation belt based on the global receptive field of an image as described above is implemented. The method includes:

[0104] Use an acquisition module to acquire an image of a road isolation belt to be detected; use a processing module to process the acquired image using a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type; wherein, the road isolation belt detection model based on the global receptive field of the image is obtained according to the training method of the road isolation belt detection model based on the global receptive field of the image.

[0105] Embodiment 4

[0106] Embodiment 4 provides a computer device including a memory and a processor. The processor and the memory communicate with each other. The memory stores program instructions executable by the processor. The processor calls the program instructions to execute the method for detecting a road isolation belt based on the global receptive field of the image as described above. The method includes:

[0107] Use an acquisition module to acquire an image of a road isolation belt to be detected; use a processing module to process the acquired image using a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type; wherein, the road isolation belt detection model based on the global receptive field of the image is obtained according to the training method of the road isolation belt detection model based on the global receptive field of the image.

[0108] Embodiment 5

[0109] Embodiment 5 of the present invention provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory, so that the electronic device executes instructions for implementing the road isolation belt detection method based on the global receptive field of the image as described above. The method includes:

[0110] Use an acquisition module to acquire an image of the road isolation belt to be detected; use a processing module to process the acquired image with a road isolation belt detection model based on the global receptive field of the image to obtain a detection result of the road isolation belt type; wherein, the road isolation belt detection model based on the global receptive field of the image is obtained according to the training method of the road isolation belt detection model based on the global receptive field of the image.

[0111] In summary, the training method of the road isolation belt detection model based on the global receptive field of the image and the isolation belt detection method provided by the present invention apply deep learning technology to the detection task of road isolation belts for the first time, which makes the detection process more automated and no longer rely on a large number of manual identifications, thereby improving efficiency and adaptability. This innovation not only makes the detection of road isolation belts more convenient, but also is widely applicable in the context of intelligent transportation, providing a powerful tool for traffic management and intelligent transportation systems, and is expected to significantly improve traffic safety and efficiency. The introduction of the multi-scale feature enhancement module helps to improve the model's ability to represent multi-scale information and globalize the receptive field. Compared with traditional methods, this module can better capture details in the image, thereby improving the detection accuracy of road isolation belts. By improving the residual network structure, especially in terms of the receptive field of the model, the requirements of the road isolation belt recognition task are better met. It can handle the situation where the isolation belt usually runs through the entire image, thereby more accurately predicting the output.

[0112] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0113] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in one or more flows Figure 1 one flow or more flows and / or blocks Figure 1 or one block or more blocks.

[0114] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in one or more flows Figure 1 one flow or more flows and / or blocks Figure 1 or one block or more blocks.

[0115] These computer program instructions can also be loaded onto a computer or other programmable data processing device to perform a series of operation steps on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one flow or more flows and / or blocks Figure 1 or one block or more blocks.

[0116] Although the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that based on the technical solutions disclosed in the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts should be covered within the protection scope of the present invention.

Claims

1. A training method for a road isolation belt detection model based on the global receptive field of an image, characterized in that, Including: Obtain training data; wherein, the training data includes multiple road median images and labels for annotating the median type in the images; Train the road median detection model based on the image global receptive field using the training data; wherein, the predicted road median form is obtained through the forward propagation process of successive convolutions, the loss between the prediction result and the true label is calculated, and backpropagation is performed using the loss to update the model weights until the set number of iteration rounds is reached; wherein, the road median detection model based on the image global receptive field includes a multi-scale enhancement module, and the multi-scale enhancement module consists of a 1x1 convolutional layer and a batch normalization layer for reducing the number of input channels, a series of 3x3 convolutional layers and batch normalization layers for feature extraction and multi-scale information, followed by a 1x1 convolutional layer and a batch normalization layer for increasing the number of output channels, and finally a ReLU activation function, and the downsampling part is used to reduce the size of the input feature map; Among them, the network structure of the road median detection model based on the image global receptive field adopts the ResNet residual network structure, and a set of filter banks are introduced to replace a single 3×3 convolutional kernel with n channels; wherein, each filter bank contains w channels, and n = s×w, where s is the number of filters in the filter bank; these filter banks are interconnected in a hierarchical manner with residual connections to increase the scale diversity of the output features; for the input features, a 1×1 convolutional layer is used to adjust the number of channels; the input features are divided into three levels; each group of filters first extracts features from a group of input feature maps, and the output features of the previous group are passed to the next group of filters together with another group of input feature maps until all input feature maps are processed; the feature maps of all groups are concatenated and input into a 1×1 convolutional layer to achieve complete information fusion.

2. The method for training a road isolation belt detection model based on the global receptive field of an image according to claim 1, wherein By traversing the training dataset, calculate the loss of each training sample and perform backpropagation to update the model parameters; every certain number of iterations, the model performs forward propagation on the test dataset, calculates the test loss and accuracy to evaluate the performance of the model on unseen data until the specified number of iterations is reached.

3. The method for training a road isolation belt detection model based on the global receptive field of an image according to claim 1, wherein Obtain training data, including: classifying traffic images according to different forms of road medians into: ". / facilitystrip" indicating that the road median in the image is a cement or green belt, ". / fence" indicating that the road median in the image is a fence, ". / line" indicating that the road median in the image is a solid line, and ". / else" indicating that the road median form in the image is other situations except the above three; use code to read and process the dataset images to obtain the training data and the test dataset.

4. The method for training a road isolation belt detection model based on the global receptive field of an image according to claim 3, wherein Using code to read and process the dataset images to obtain the training data and the test data includes: Use the glob library and the glob.glob method to obtain the file paths of all images in the folder containing the image files, and these image files are under the path specified in the code; Define four image categories: 'else', 'facility strip', 'fence', and 'line', as well as the mapping relationship between them and integer labels; Create a custom dataset class MyDataset, which inherits from torch.utils.data.Dataset and is used to process image data. In the initialization, pass in the image path, labels, and data transformation functions; The data transformation function transform defines a series of transformation operations, including resizing the image to 256x256 pixels and converting the image to a tensor; Create a data loader dataloader to divide the dataset into batches, with each batch containing 16 images and corresponding labels; Divide the dataset: randomly permute all image paths and labels, and divide the dataset into a training set and a test set.

5. A road isolation belt detection method based on the global receptive field of an image, characterized in that, Include: Obtain the image of the road isolation belt to be detected; Process the obtained image using the road isolation belt detection model based on the global receptive field of the image to obtain the detection result of the road isolation belt type; among them, the road isolation belt detection model based on the global receptive field of the image is obtained according to the model training method described in any one of claims 1-4.

6. A road isolation belt detection device based on the global receptive field of an image, characterized in that, Include: An obtaining module for obtaining the image of the road isolation belt to be detected; A processing module for processing the obtained image using the road isolation belt detection model based on the global receptive field of the image to obtain the detection result of the road isolation belt type; among them, the road isolation belt detection model based on the global receptive field of the image is obtained according to the model training method described in any one of claims 1-4.

7. A computer device, characterized in that, Includes a memory and a processor, the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor calls the program instructions to execute the road isolation belt detection method based on the global receptive field of the image as described in claim 5.

8. An electronic device, characterized in that, Include: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes the instructions for implementing the road isolation belt detection method based on the global receptive field of the image as described in claim 5.

Citation Information

Patent Citations

  • Road guardrail anomaly detection method based on artificial intelligence and image processing

    CN111797803A

  • Real-time gesture detection method and device

    CN114612832A