Road area recognition method and related equipment
By building a lightweight TC-PC backbone network and decoder network, the problem of high computational cost of road area identification in the prior art is solved, and efficient and accurate road area identification is achieved.
Patent Information
- Application Number
- CN202111472168.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-03
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2041-12-03
AI Technical Summary
The existing road area identification method requires a large number of parameters to achieve a high recognition accuracy rate, resulting in excessive calculation costs.
A road area recognition method is proposed, using the constructed TC-PC backbone network to extract the road image to be detected, and processing the feature map with the decoder network, reducing the amount of parameters and improving the recognition efficiency.
While ensuring recognition accuracy, the time cost during the calculation process is reduced and the recognition efficiency is improved.
Smart Images

Figure CN114419577B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and in particular, to a road area recognition method and related devices. Background Art
[0002] Road area recognition is mainly applied to the field of autonomous driving. Its main goal is to automatically identify the road area in an image given a road image. Road area recognition will help the vehicle control system determine the driving direction and speed.
[0003] Existing road area recognition is mainly based on the principle of semantic segmentation. Specifically, common road area recognition methods can be divided into three categories according to the data type: methods based on RGB images, methods combining RGB images with Lidar images or depth maps, and methods based on point clouds. Existing road area recognition methods require a large number of parameters to achieve a high recognition accuracy, resulting in too high a computational cost. Summary of the Invention
[0004] In view of this, the purpose of this application is to propose a road area recognition method and related devices to solve the above technical problems.
[0005] Based on the above purpose, this application provides a road area recognition method, including:
[0006] Obtain the road image to be detected;
[0007] Construct a road area recognition model, where the road area recognition model includes a TC-PC backbone network and a decoder network;
[0008] Use the TC-PC backbone network to extract features from the road image to be detected to obtain a road image feature map;
[0009] Use the decoder network to process the road image feature map to obtain a road area recognition result.
[0010] Further, the TC-PC backbone network includes a TC convolution module, and the TC convolution module includes a layer normalization layer, an attention mechanism layer, a first batch normalization layer, and a first Conv2D convolution layer connected in sequence.
[0011] Further, using the TC convolution module to process the road image to be detected includes:
[0012] Divide the road image to be detected into multiple windows of a preset size;
[0013] Use the attention mechanism to process the road image data in each window to obtain local features of the road image to be detected;
[0014] Fuse the local features to obtain a first road image feature map.
[0015] Further, the TC-PC backbone network includes a PC adjustment module, and the PC adjustment module includes a patch Merging layer, a second batch normalization layer, and a second Conv2D convolutional layer connected in sequence. Among them, the patch Merging layer is used to reduce the width and height of the input road image data by half, and the second Conv2D convolutional layer is used to adjust the number of channels of the road image data;
[0016] A residual branch is also connected after the patch Merging layer, and the residual branch is used to weightedly superimpose the output of the patch Merging layer on the output of the second Conv2D convolutional layer.
[0017] Further, using the PC adjustment module to process the first road image feature map includes:
[0018] Reduce the width and height of the first road image feature map to half of the current value to obtain a second road image feature map;
[0019] Output the second road image feature map through the output channels after quantity expansion to obtain the road image feature map.
[0020] Further, the attention mechanism of the attention mechanism layer is represented by the following formula:
[0021]
[0022] Among them, Q represents the query, K represents the key, V represents the value, and d k represents the feature dimension of Q, K, and V.
[0023] Further, in the TC convolutional module:
[0024] The size of the window is w*w, and the convolutional kernel of the first Conv2D convolutional layer is k*k, and the stride is 1, where k≠1.
[0025] Based on the same inventive concept, the second aspect of the present application provides a road area recognition device, including:
[0026] An acquisition module: configured to acquire a road image to be detected;
[0027] A construction module: configured to construct a road area recognition model, and the road area recognition model includes a TC-PC backbone network and a decoder network;
[0028] Extraction module: Configured to extract features of the road image to be detected by using the TC-PC backbone network, and obtain a road image feature map;
[0029] Processing module: Configured to process the road image feature map by using the decoder network, and obtain a road area recognition result.
[0030] Based on the same inventive concept, a third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method described in the first aspect is implemented.
[0031] Based on the same inventive concept, a fourth aspect of the present application provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method described in the first aspect.
[0032] As can be seen from the above, the road area recognition method and related devices provided by the present application extract features of the road image to be detected based on the constructed TC-PC backbone network, which can reduce the number of parameters, have a fast convergence speed, effectively improve the recognition efficiency while ensuring the recognition accuracy, and reduce the time cost during the operation process. Description of the Drawings
[0033] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0034] Figure 1 It is a flowchart of the road area recognition method according to an embodiment of the present application;
[0035] Figure 2 It is a schematic structural diagram of the TC convolution module according to an embodiment of the present application;
[0036] Figure 3 It is a flowchart of processing the road image to be detected by using the TC convolution module according to an embodiment of the present application;
[0037] Figure 4 It is a schematic structural diagram of the PC adjustment module according to an embodiment of the present application;
[0038] Figure 5 It is a flowchart of processing the first road image feature map by using the PC adjustment module according to an embodiment of the present application;
[0039] Figure 6Schematic structural diagram of the road area recognition device according to an embodiment of the present application;
[0040] Figure 7 Schematic structural diagram of the electronic device according to an embodiment of the present application. Detailed implementation manners
[0041] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0042] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the ordinary meanings understood by those of ordinary skill in the art to which the present application belongs. The terms "first", "second" and similar words used in the embodiments of the present application do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects.
[0043] As described in the background art section, the technical solutions for road area recognition in the related art are still difficult to meet the requirements. The existing road area recognition methods can be divided into three categories according to the data type: methods based on RGB images, methods combining RGB images with Lidar images or depth images, and methods based on point clouds. The applicant found during the implementation of the present application that the main problems existing in the existing road area recognition methods are as follows: in the method based on RGB images, when predicting the road area, the input data only includes RGB images, and the information provided by RGB images is not sufficient for the existing network models to make comprehensive and accurate judgments, resulting in difficulty in breaking through the accuracy; the input data of the method combining RGB images with Lidar images or depth images is to add Lidar images or depth images on the basis of RGB images. Through the data fusion method, the network model can obtain key information from the depth images or Lidar images when analyzing RBG images, so as to make more accurate judgments. Although the prediction accuracy is relatively high, it will lead to a high computational cost; the input data of the method based on point clouds is no longer images, but point cloud data collected by lidar. This type of algorithm uses the mainstream models in the field of point clouds for encoding and decoding, so as to classify the road area. The method based on point clouds is restricted by the capabilities of the mainstream models. In addition, the proportion of road data in the point cloud data is smaller than that of images, and the training and prediction processes are more difficult.
[0044] In view of this, the embodiments of the present application provide a road area recognition method, which constructs a lightweight TC-PC backbone network to extract the features of the road image to be detected, can reduce the number of parameters, and further reduce the amount of calculation, while meeting the accuracy of road area recognition, reducing the time cost of calculation.
[0045] The technical solution of the present application will be described in detail below through specific embodiments.
[0046] Referring to Figure 1 , a road area recognition method provided by an embodiment of the present application includes the following steps:
[0047] Step S101, obtaining a road image to be detected.
[0048] In this step, the road image to be detected includes an RGB image and a depth map. The RGB image refers to an image obtained by the changes of three color channels of red (R), green (G), and blue (B) and their superposition with each other to obtain various colors; the depth image refers to an image of the distance (depth) information from the image collector to each point in the scene, and the information directly reflects the geometric shape of the visible surface in the scene. The depth image can be obtained by other means such as lidar depth imaging and depth cameras.
[0049] It should be noted that the surface normal information in the depth image can be obtained by processing the depth image using a Surface Normal Estimator (SNE).
[0050] Step S102, constructing a road area recognition model, where the road area recognition model includes a TC (Transformer_Conv)-PC (PatchMerging_Conv) backbone network and a decoder network.
[0051] In this step, the TC-PC backbone network is the encoder network. The encoding process of the encoder network is the process in which the resolution of the feature map decreases from large to small during the forward propagation of the convolutional layer, which can map data from low dimension to high dimension and encode the data; the decoding process of the decoder network is the process in which the resolution of the feature map increases from small to large during the forward propagation of the convolutional layer, which can map data from high dimension to low dimension and decode the data to achieve information conversion.
[0052] Step S103, using the TC-PC backbone network to extract features from the road image to be detected to obtain a road image feature map.
[0053] In this step, the TC-PC backbone network extracts more image features of the original road image data with fewer parameters, so that the subsequent decoder network can reconstruct the original road image and ensure the recognition accuracy.
[0054] Step S104, using the decoder network to process the road image feature map to obtain a road area recognition result.
[0055] In this step, the feature map of the road image to be predicted can be binarized to obtain the binary image of the road area and the binary image of the background in the road image to be predicted. Then, the binary image of the road area and the binary image of the background are fused to obtain an optimized binary image of the road area, and further the recognition result of the road area is obtained.
[0056] It can be seen that a road area recognition method provided by this embodiment extracts the features of the road image to be detected based on the constructed TC-PC backbone network, which can reduce the number of parameters, has a fast convergence speed, effectively improves the recognition efficiency while ensuring the recognition accuracy, and reduces the time cost in the operation process.
[0057] In some embodiments, combined with Figure 2 , the TC-PC backbone network includes up to TC convolution modules. The TC convolution module includes a layer normalization layer (Layer Normalization, LN), an attention mechanism layer (W-MSA), a first batch normalization layer (Batch Normalization, BN), and a first Conv2D convolution layer connected in sequence. The first batch normalization layer and the first Conv2D convolution layer can achieve parameter sharing between windows.
[0058] In this embodiment, the attention mechanism layer selects W-MSA (window based Multi-head Self-Attention) as the attention mechanism. Hierarchical attention network, recurrent attention network, global attention model, local attention model, self-attention model, or multi-head attention model, etc. can be selected as the attention mechanism according to the actual situation, and no specific limitation is made here.
[0059] In some embodiments, combined with Figure 3 , using the TC convolution module to process the road image to be detected includes the following steps:
[0060] Step S301, divide the road image to be detected into multiple windows of a preset size.
[0061] In this step, the size of the window can be set to 6*6, 12*12, 18*18, 24*24, etc., but the window size cannot be 1*1. If it is 1*1, the image data inside each window cannot be effectively processed subsequently. In addition, the size of the input image data is an integer multiple of the window size. For example, if the size of the input image data is 224*224, the corresponding window size can be 8*8, but not 13*13.
[0062] Step S302, use the attention mechanism to process the road image data in each window to obtain the local features of the road image to be detected.
[0063] Step S303: Fuse the local features to obtain the first road image feature map.
[0064] In some embodiments, referring to Figure 4 , the TC-PC backbone network includes a PC adjustment module. The PC adjustment module includes a patch Merging layer, a second batch normalization layer, and a second Conv2D convolutional layer connected in sequence. The second Conv2D convolutional layer is used to adjust the number of channels of the road image data, and the patch Merging layer is used to reduce both the width and height of the input road image data by half. Specifically, elements can be selected at intervals of 1 unit in the row direction and the column direction.
[0065] A residual branch is also connected after the patch Merging layer. The residual branch is used to weighted-sum and stack the output of the patch Merging layer onto the output of the second Conv2D convolutional layer. Specifically, when the number of input channels and output channels of the second Conv2D convolutional layer is the same, the residual branch plays a role; otherwise, the residual branch does not work.
[0066] In some embodiments, in combination with Figure 5 , using the PC adjustment module to process the first road image feature map includes the following steps:
[0067] Step S501: Reduce the width and height of the first road image feature map to half of the current size to obtain a second road image feature map.
[0068] In this step, reducing the width and height of the first road image feature map to half of the current size, that is, reducing the resolution of the first road image feature map, can save the amount of computation.
[0069] Step S502: Output the second road image feature map through the output channels after quantity expansion to obtain the road image feature map.
[0070] In some embodiments, the attention mechanism of the attention mechanism layer is represented by the following formula:
[0071]
[0072] where Q represents the query, K represents the key, V represents the value, and d k represents the feature dimension of Q, K, and V.
[0073] In some embodiments, in the TC convolutional module:
[0074] The size of the window is w*w, and the convolutional kernel of the first Conv2D convolutional layer is k*k with a stride of 1, where k≠1.
[0075] In this embodiment, when the convolution operation area of the Conv2D convolution layer is the same as the size of the window, the Conv2D convolution layer cannot share information between windows. For example, when the window size is set to 6*6, the convolution kernel size, stride, and padding of the Conv2D convolution layer cannot be set to (6,6,0), while setting the convolution kernel size, stride, and padding parameters of Conv2D to (5,1,1), (3,1,1), etc. can achieve information sharing between windows. In addition, when the convolution kernel size is 1*1, the convolution operation cannot process data from different windows simultaneously, so the convolution kernel size cannot be 1*1.
[0076] It should be noted that the method of the embodiment of the present application can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiment of the present application, and these multiple devices will interact with each other to complete the described method.
[0077] It should be noted that some embodiments of the present application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0078] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application also provides a road area recognition device.
[0079] Referring to Figure 6 , the road area recognition device includes:
[0080] An acquisition module 601: configured to acquire a road image to be detected;
[0081] A construction module 602: configured to construct a road area recognition model, where the road area recognition model includes a TC-PC backbone network and a decoder network;
[0082] An extraction module 603: configured to extract features from the road image to be detected by using the TC-PC backbone network to obtain a road image feature map;
[0083] A processing module 604: configured to process the road image feature map by using the decoder network to obtain a road area recognition result.
[0084] As an alternative embodiment, the TC-PC backbone network includes a TC convolution module, and the TC convolution module includes a layer normalization layer, an attention mechanism layer, a first batch normalization layer, and a first Conv2D convolution layer connected in sequence.
[0085] As an alternative embodiment, the extraction module 603 is specifically configured to divide the road image to be detected into multiple windows of a preset size; use an attention mechanism to process the road image data in each window to obtain local features of the road image to be detected; and fuse the local features to obtain a first road image feature map.
[0086] As an alternative embodiment, the TC-PC backbone network includes a PC adjustment module, and the PC adjustment module includes a patch Merging layer, a second batch normalization layer, and a second Conv2D convolution layer connected in sequence. Among them, the patch Merging layer is used to reduce both the width and height of the input road image data by half, and the second Conv2D convolution layer is used to adjust the number of channels of the road image data; a residual branch is further connected after the patch Merging layer, and the residual branch is used to weight and superimpose the output of the patch Merging layer on the output of the second Conv2D convolution layer.
[0087] As an alternative embodiment, the extraction module 603 is further specifically configured to reduce the width and height of the first road image feature map to half of the current size to obtain a second road image feature map; output the second road image feature map through the output channels after quantity expansion to obtain the road image feature map.
[0088] As an alternative embodiment, the attention mechanism of the attention mechanism layer is represented by the following formula:
[0089]
[0090] where Q represents a query, K represents a key, V represents a value, and d k represents the feature dimension of Q, K, and V.
[0091] As an alternative embodiment, in the TC convolution module: the size of the window is w*w, the convolution kernel of the first Conv2D convolution layer is k*k, and the stride is 1, where k≠1.
[0092] For the convenience of description, when describing the above device, various modules are described separately according to their functions. Of course, when implementing the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0093] The device of the above embodiment is used to implement the corresponding road area recognition method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0094] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the road area recognition method described in any of the above embodiments.
[0095] Figure 7 FIG. shows a more specific schematic diagram of the hardware structure of the electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. Among them, the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other inside the device through the bus 1050.
[0096] The processor 1010 may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0097] The memory 1020 may be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 1020 may store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 1020 and are called and executed by the processor 1010.
[0098] The input / output interface 1030 is used to connect to an input / output module to implement information input and output. The input / output module may be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0099] The communication interface 1040 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0100] The bus 1050 includes a path for transmitting information between various components of the device (such as the processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040).
[0101] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0102] The electronic device in the above embodiment is used to implement the corresponding road area recognition method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0103] Based on the same inventive concept, corresponding to the method in any of the above embodiments, the present application also provides a non-transitory computer-readable storage medium. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the road area recognition method described in any of the foregoing embodiments.
[0104] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be achieved by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.
[0105] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the road area recognition method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0106] Those of ordinary skill in the art should understand that: The discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present application (including the claims) is limited to these examples; Under the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present application as described above, and they are not provided in detail for the sake of brevity.
[0107] In addition, for the sake of simplicity of description and discussion, and in order not to make the embodiments of the present application difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. In addition, the device may be shown in block diagram form in order to avoid making the embodiments of the present application difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present application are to be implemented (that is, these details should be completely within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present application, it will be apparent to those skilled in the art that the embodiments of the present application can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0108] Although the present application has been described in connection with specific embodiments of the present application, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0109] The embodiments of the present application are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included within the protection scope of the present application.
Claims
1. A road area recognition method, characterized in that, comprising: Obtaining a road image to be detected; Constructing a road area recognition model, the road area recognition model including a TC-PC backbone network and a decoder network; Using the TC-PC backbone network to perform feature extraction on the road image to be detected to obtain a road image feature map; Using the decoder network to process the road image feature map to obtain a road area recognition result; The TC-PC backbone network includes a TC convolution module, and the TC convolution module includes a layer normalization layer, an attention mechanism layer, a first batch normalization layer, and a first Conv2D convolution layer connected in sequence; The TC-PC backbone network includes a PC adjustment module, and the PC adjustment module includes a patch Merging layer, a second batch normalization layer, and a second Conv2D convolution layer connected in sequence, wherein the patch Merging layer is used to reduce both the width and height of the input road image data by half, and the second Conv2D convolution layer is used to adjust the number of channels of the road image data; A residual branch is further connected after the patch Merging layer, and the residual branch is used to weightedly superimpose the output of the patch Merging layer onto the output of the second Conv2D convolution layer; Wherein the attention mechanism layer is a W-MSA attention mechanism.
2. The method according to claim 1, characterized in that, Using the TC convolution module to process the road image to be detected includes: Dividing the road image to be detected into multiple windows of a preset size; Using the attention mechanism to process the road image data in each window to obtain local features of the road image to be detected; Fusing the local features to obtain a first road image feature map.
3. The method according to claim 1, characterized in that, Using the PC adjustment module to process the first road image feature map includes: Reducing the width and height of the first road image feature map to half of the current size to obtain a second road image feature map; Outputting the second road image feature map through the output channels after quantity expansion to obtain the road image feature map.
4. The method according to claim 1, characterized in that, The attention mechanism of the attention mechanism layer is represented by the following formula: Among them, Q represents a query, K represents a key, V represents a value, d k represents Q, K, V the characteristic dimension of.
5. The method according to claim 2, characterized in that, In the TC convolution module: The size of the window is w*w, and the convolution kernel of the first Conv2D convolutional layer is k*k, with a stride of 1, where k 1.
6. A road area recognition device, characterized in that, comprising: An acquisition module: configured to acquire a road image to be detected; A construction module: configured to construct a road area recognition model, the road area recognition model including a TC-PC backbone network and a decoder network; An extraction module: configured to use the TC-PC backbone network to perform feature extraction on the road image to be detected to obtain a road image feature map; A processing module: configured to use the decoder network to process the road image feature map to obtain a road area recognition result; The TC-PC backbone network includes a TC convolution module, and the TC convolution module includes a layer normalization layer, an attention mechanism layer, a first batch normalization layer, and a first Conv2D convolution layer connected in sequence; The TC-PC backbone network includes a PC adjustment module, and the PC adjustment module includes a patch Merging layer, a second batch normalization layer, and a second Conv2D convolution layer connected in sequence. Among them, the patch Merging layer is used to reduce both the width and height of the input road image data by half, and the second Conv2D convolution layer is used to adjust the number of channels of the road image data; A residual branch is further connected after the patch Merging layer, and the residual branch is used to weightedly superimpose the output of the patch Merging layer onto the output of the second Conv2D convolution layer; Among them, the attention mechanism layer is a W-MSA attention mechanism.
7. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, when the processor executes the program, the method described in any one of claims 1 to 5 is implemented.
8. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium stores computer instructions, characterized in that, the computer instructions are used to cause the computer to execute the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Open-pit mine road model construction method based on P-LinkNet network
CN111242231A
Road crack image recognition method, device and system based on convolutional neural network
CN111597932A