Point prompt road illegal parking area segmentation method and system based on Transform mechanism

Through the point prompt technology based on the Transformer mechanism, the point information and image embedding vectors in road video are integrated, and the problem of low segmentation accuracy of road illegal parking areas in the existing technology is solved, and higher accuracy and efficiency are achieved, which is suitable for complex road environments.

CN119942485APending Publication Date: 2025-05-06SHANGHAI INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510022973.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has the problem of low accuracy in the segmentation of illegal road parking areas, especially when facing light changes, diversified stop markings, occlusions and complex road geometry, the segmentation accuracy and robustness are insufficient.

Method used

The point-prompt road illegal stop area segmentation method based on the Transformer mechanism is adopted. Through pre-processing, illegal stop area detection, segmentation block embedding and mask decoding, point information and image embedding vectors are integrated, and the segmentation mask is generated using the instant self-attention and cross-attention mechanism.

Benefits of technology

It improves the accuracy and efficiency of illegal parking area segmentation, enhances adaptability to complex environments, and realizes efficient coordination with the autonomous driving system, which is suitable for most road scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942485A_ABST
    Figure CN119942485A_ABST
Patent Text Reader

Abstract

The invention provides a point prompt road illegal parking area segmentation method and system based on a Transform mechanism, and is applied to the technical field of road illegal parking area segmentation, detected point information and embedded vectors are integrated, all feature maps become more powerful and accurate semantically, the accuracy of subsequent illegal parking area segmentation is improved, and the accuracy of road illegal parking area segmentation is improved. According to the method, a complicated illegal parking area segmentation task is processed by using a transformation operation based on a Transform mechanism, and a segmentation result can be flexibly adjusted according to a prompt provided by a user, so that the efficiency is improved while the detection accuracy of an illegal parking area segmentation model is improved. The method achieves the automatic segmentation of the illegal parking region, gives consideration to the segmentation precision and detection speed of the illegal parking region, can be suitable for most road scenes, and is higher in robustness and practicality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of road illegal parking area segmentation, and specifically to a point prompt road illegal parking area segmentation method and system based on a Transformer mechanism. Background Art

[0002] As a key link in autonomous driving technology, the road no-parking area segmentation method has its technical foundation derived from the development and needs of multiple fields. The accurate identification and understanding of no-parking areas are crucial for autonomous driving systems to navigate safely and efficiently in complex and changing environments. This technology not only helps vehicles accurately locate drivable areas, but also effectively detects and analyzes no-parking area markers, providing accurate road environment information for vehicle decision-making systems. In recent years, the rapid development of deep learning, especially the major breakthroughs of convolutional neural networks (CNN) in the field of image recognition and segmentation, has greatly promoted the advancement of automatic segmentation technology for road no-parking areas. However, these methods have low accuracy in lane parking area segmentation.

[0003] Based on this, a new technical solution is needed. Summary of the invention

[0004] In view of this, the present application provides a method and system for segmenting illegal parking areas on point-prompted roads based on a Transformer mechanism.

[0005] This application provides the following technical solutions:

[0006] According to a Transformer mechanism-based point-prompted road illegal parking area segmentation method provided in this application, the following steps are included:

[0007] Preprocessing step: preprocess the input road video to obtain a video image;

[0008] Illegal parking area detection step: perform lane illegal parking area detection on the video image to obtain point information of the illegal parking area;

[0009] Illegal parking area segmentation step: segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedded vector; convert the point information into a second embedded vector, integrate the first embedded vector and the second embedded vector to obtain integrated embedded information;

[0010] Mask decoding step: perform feature fusion on the integrated embedded information, generate a segmentation mask through transformation, the segmentation mask covers the illegal parking area, and outputs the segmentation video of the illegal parking area.

[0011] Preferably, in the preprocessing step, the video data set is subjected to frame extraction processing before the illegal parking area detection step, and the image size is unified and the RGB values ​​are normalized; frame-by-frame processing is performed before the illegal parking area segmentation step, and the image size is unified and the RGB values ​​are normalized.

[0012] Preferably, in the illegal parking area detection step, for the video image, image parameters of the feature map are input, video image features are extracted, the feature map is adaptively aggregated, the feature map is divided into grids through a pooling operation, and features are extracted from the grids, and multi-scale feature fusion is performed on the extracted features to obtain an illegal parking area detection picture with a detection frame, and the point information of the illegal parking area is obtained according to the illegal parking area detection picture with the detection frame.

[0013] Preferably, extracting features from video images includes:

[0014] Perform a convolution operation:

[0015] Conv(x) = W*x+b;

[0016] Among them, Conv represents convolution; W represents the convolution kernel weight; * represents the convolution operation; x represents the input tensor; b represents the bias;

[0017] Perform batch normalization:

[0018]

[0019] Among them, BN represents batch normalization; y represents the output of the convolution operation; μ represents the mean; σ 2 represents variance; ε represents a constant to prevent division by zero; γ and β represent learnable parameters;

[0020] Use activation function:

[0021]

[0022] Among them, SiLU represents the activation function; z represents the output after batch normalization, σ(z) represents the sigmoid function; e represents a natural constant.

[0023] Preferably, performing multi-scale feature fusion on the extracted features includes:

[0024] Perform upsampling operation:

[0025] x up =Upsample(x T , sacle_factor = s);

[0026] Among them, x T represents the input feature map; s represents the upsampling factor; x upIt is the feature map obtained after upsampling; Upsample means upsampling; sacle_factor means scaling factor;

[0027] Adjust the number of channels and fusion features of the feature map:

[0028] x conv =Conv(x up ,out_channels=C out , kernel_size=k, stride=1, padding=p);

[0029] Among them, x conv Represents the feature map after convolution; out_channels and C out Indicates the number of output channels, kernel_size and k indicate the convolution kernel size; stride indicates the step size; padding and p indicate padding;

[0030] Concatenate the convolutional feature map with the low-level feature map:

[0031] x fuse =cat(x low ,x conv , dim=1);

[0032] Among them, x fuse Represents the fused feature map; x low Represents a low-level feature map; cat represents the operation of connecting tensors on a preset dimension; dim represents the dimension parameter of the concatenation operation.

[0033] Preferably, obtaining point information of the illegal parking area according to the illegal parking area detection picture with the detection frame includes:

[0034] Get the center point of the detection box, select the points near the center point of the box as output points, and obtain the point information of the illegal parking area:

[0035]

[0036] Where rand_x is the calculated horizontal coordinate of the random point, and rand_y is the calculated vertical coordinate of the random point; random.uniform is a random floating point number used to generate the horizontal and vertical coordinates of the center point of the box within the preset range; x c Indicates the horizontal coordinate of the center point of the box; y c Indicates the vertical coordinate of the center point of the box.

[0037] Preferably, in the illegal parking area segmentation step, the video image is segmented to obtain a plurality of segments, each segment being an independent visual token; each visual token is converted into a high-dimensional vector through a linear embedding process, mapped to a vector space for processing, and a first embedding vector is obtained; the point information is converted into a second embedding vector, which is combined with the first embedding vector to form a feature sequence of the Transformer.

[0038] Preferably, in the mask decoding step, the integrated embedded information is fused through a transformation operation to obtain the fused illegal parking area features and processed to generate a segmentation mask, which represents the segmentation prediction of the illegal parking area in the image; an immediate self-attention and cross-attention mechanism is adopted, and the global dependency is captured when processing the features of the illegal parking area position through immediate self-attention, and the connection is established between the features of different illegal parking areas through cross-attention to interact between the features; bidirectional mapping is performed to adjust the illegal parking area segmentation result.

[0039] Preferably, for each illegal parking area input element, a similarity score is calculated, and the information is integrated by weighted average to calculate the Q, K and V matrices:

[0040] Q=XW Q ,K=XW K ,V=XW V

[0041] Among them, Q is the query vector, which indicates the role of the input at the current position in the self-attention mechanism; K is the key vector, which is used to calculate the correlation between the input at the current position and the input at other positions; V is the value vector, which contains the information of the input at the current position and is used to calculate the self-attention weight; X represents the given input sequence; W Q , W K , W V is the learned weight matrix;

[0042] Calculate the attention score:

[0043]

[0044] Among them, d k represents the dimension of the vector used to scale the inner product; T represents transpose; softmax represents the normalized exponential function; Attention represents the attention score;

[0045] Perform a nonlinear transformation:

[0046] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 ;

[0047] Where FFN(x) represents a nonlinear transformation; W 1 , W 2 、b 1 and b 2 All represent the learned parameters;

[0048] Normalize:

[0049] Output=LayerNorm(x+SubLayer(x));

[0050] Among them, Output represents output; LaverNorm represents layer normalization; Sublayer(x) represents the output of the sublayer.

[0051] According to the present application, a point-prompt road illegal parking area segmentation system based on the Transformer mechanism is provided, which includes the following modules:

[0052] Preprocessing module: preprocess the input road video to obtain video images;

[0053] Illegal parking area detection module: detects illegal parking areas in lanes on video images and obtains point information of illegal parking areas;

[0054] Illegal parking area segmentation module: segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedded vector; convert the point information into a second embedded vector, integrate the first embedded vector and the second embedded vector to obtain integrated embedded information;

[0055] Mask decoding module: The integrated embedded information is feature-fused, and a segmentation mask is generated through transformation. The segmentation mask covers the illegal parking area and outputs the segmentation video of the illegal parking area.

[0056] Compared with the prior art, the beneficial effects that can be achieved by at least one of the above technical solutions adopted in the present application include at least:

[0057] This application integrates the detected point information with the embedded vector, making all feature maps more powerful and accurate in semantics, improving the accuracy of subsequent illegal parking area segmentation, using transformation operations based on the Transformer mechanism to handle complex illegal parking area segmentation tasks, and being able to flexibly adjust the segmentation results according to the prompts provided by the user, so that the illegal parking area segmentation model improves detection accuracy while also improving efficiency. This application realizes the automated segmentation of illegal parking areas while taking into account the segmentation accuracy and detection speed of illegal parking areas, can adapt to most road scenes, and has high robustness and practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0059] Figure 1 It is the flow chart of the illegal parking area detection and point information extraction module in this application;

[0060] Figure 2 It is a diagram of the illegal parking area segmentation structure of the fusion point information extraction module in this application;

[0061] Figure 3 This is a flowchart of the point prompt road illegal parking area division of this application. DETAILED DESCRIPTION

[0062] The embodiments of the present application are described in detail below with reference to the accompanying drawings.

[0063] The following describes the implementation methods of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. The present application can also be implemented or applied through other different specific implementation methods, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, in the absence of conflict, the following embodiments and the features in the embodiments can be combined with each other. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work belong to the scope of protection of the present application.

[0064] It should be noted that various aspects of the embodiments within the scope of the appended claims are described below. It should be apparent that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is merely illustrative. Based on the present application, it should be understood by those skilled in the art that an aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number and aspect described herein can be used to implement the device and / or practice the method. In addition, other structures and / or functionalities other than one or more of the aspects described herein can be used to implement this device and / or practice this method.

[0065] It should also be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present application. The drawings only show components related to the present application rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0066] Additionally, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, it will be understood by those skilled in the art that the examples can be practiced without these specific details.

[0067] Through in-depth research and improvement exploration on the segmentation of illegal parking areas on roads, the applicant found that the methods in the existing technology still face severe challenges in segmentation accuracy and robustness when facing fluctuations in lighting conditions, the diversity of materials and colors of no-parking signs, occlusion phenomena, wear and aging of signs, and complex road geometries. The evolution of no-parking area segmentation technology, in addition to relying on the continuous innovation of deep learning algorithms, also needs to cope with multiple challenges in practical applications. Real-time performance, environmental adaptability, and seamless integration with autonomous driving systems are the keys to promoting the maturity of this technology. While ensuring high-precision segmentation, how to improve the system response speed, enhance adaptability to environmental changes, and achieve efficient collaboration with the autonomous driving decision-making system is the focus of current research and development.

[0068] Based on this, the technical solutions provided by the embodiments of the present application are described below in conjunction with the accompanying drawings.

[0069] The embodiment of this specification proposes a method for segmenting a point-prompted road no-parking area based on a Transformer mechanism, such as Figure 1 , Figure 2 as well as Figure 3 As shown, the following steps are included:

[0070] Preprocessing step (step S1): preprocess the input road video to obtain a video image. Specifically, perform data preprocessing on the input road video, extract frames from the video, uniformly scale the image size, and normalize it. Input the video to be detected into the image preprocessing module of the illegal parking area detection model, segment the video image, and perform corresponding pixel conversion to facilitate subsequent image processing; that is, input the video to be segmented into the illegal parking area detection module and the illegal parking area segmentation model, respectively, and perform preprocessing to obtain the illegal parking area image data for subsequent point information extraction and image embedding vector conversion.

[0071] The dataset used is UAVid, which is a dataset designed specifically for the semantic segmentation task of drone videos in urban scenes. It contains 2,898 infrared thermal images extracted from 43,470 frames, which are captured by drones from different scenes (such as schools, parking lots, roads, playgrounds, etc.), covering a wide range and including multiple object categories such as people and vehicles. Accuracy is the main evaluation indicator of this dataset.

[0072] Illegal parking area detection step (step S2): lane illegal parking area detection is performed on the video image to obtain point information of the illegal parking area.

[0073] Specifically, the image is subjected to lane illegal parking area detection and random point information of the illegal parking area is obtained. The processed video image is input into the illegal parking area detection module, and after the illegal parking area is detected, the illegal parking area location point information in the image is calculated and saved and transmitted to the illegal parking area segmentation module. Figure 1 The middle part shows that after the network structure detects the illegal parking area, it calculates the regional location point information in the image and saves it and transmits it to the illegal parking area segmentation module.

[0074] Illegal parking area segmentation step (step S3): segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedded vector; convert the point information into a second embedded vector, integrate the first embedded vector and the second embedded vector to obtain integrated embedded information.

[0075] Specifically, the obtained illegal parking area point information is fed into the illegal parking area segmentation model for preprocessing transformation; the illegal parking area segmentation module first converts the input image into a series of embedded vectors, and also converts the obtained point information into embedded vectors, that is, in the embedded vector conversion module, the illegal parking area segmentation module first divides the input image into multiple small blocks, each of which is regarded as an independent visual token. These tokens are then linearly embedded into a vector of a fixed size, and the obtained point information is also converted into an embedded vector through the prompt encoder. The two parts of the embedded vector are integrated so that the model can only segment the illegal parking area in the image.

[0076] Mask decoding step (step S4): feature fusion is performed on the integrated embedded information, and a segmentation mask is generated through transformation. The segmentation mask covers the illegal parking area, and a segmentation video of the illegal parking area is output.

[0077] Specifically, based on the transformation operation of the Transformer mechanism, the input lane video is segmented frame by frame, and the illegal parking area segmentation result with the highest confidence is selected. Finally, these illegal parking area segmentation results are reintegrated and the illegal parking area segmentation video information is output.

[0078] The mask decoding module transforms the embedded information integrated in the previous stage to generate a segmentation mask, which accurately covers the specified illegal parking area object. Finally, the segmented illegal parking area video is output; that is, the mask decoding module performs feature fusion on the embedded information integrated in the previous stage, and finally generates a segmentation mask through a series of transformations. At the same time, it uses the immediate self-attention and cross-attention mechanism to update all embeddings, and works in two directions (from prompt to image embedding and back) to accurately map to the final mask, and finally outputs the complete illegal parking area segmentation video.

[0079] In one embodiment, if Figure 1 , Figure 2 and Figure 3 As shown, in the preprocessing step, the video data set is subjected to frame extraction processing before the illegal parking area detection step, the image size is unified and the RGB values ​​are normalized; and frame-by-frame processing is performed before the illegal parking area segmentation step, the image size is unified and the RGB values ​​are normalized.

[0080] Specifically, in the image preprocessing module, the input video data is first processed by frame extraction, and the illegal parking area detection module will adjust the image size to a uniform resolution of 640×640. In the illegal parking area segmentation module, the image size is uniformly adjusted to a resolution of 1024×1024. That is, the high-altitude lane video dataset used is mainly urban and highway scenes shot at different altitudes in good weather. The illegal parking area detection module will extract the video dataset and adjust the image size to a uniform resolution of 640×640. The illegal parking area segmentation module processes frame by frame and adjusts the image size to a uniform resolution of 1024×1024. In the preprocessing of these two modules, the image RGB values ​​will be normalized to 0~1.

[0081] In one embodiment, if Figure 1 and Figure 3 As shown, in the illegal parking area detection step, for the video image, the image parameters of the feature map are input, the video image features are extracted, the feature map is adaptively aggregated, the feature map is divided into grids through pooling operations, and features are extracted from the grids. Multi-scale feature fusion is performed on the extracted features to obtain an illegal parking area detection picture with a detection frame, and the point information of the illegal parking area is obtained according to the illegal parking area detection picture with a detection frame. Upsampling+Feature Fusion means upsampling+feature fusion.

[0082] In one embodiment, if Figure 1 and Figure 3 As shown, the video image feature extraction includes:

[0083] Perform a convolution operation:

[0084] Conv(x) = W*x+b;

[0085] Among them, Conv represents convolution; W represents the convolution kernel weight; * represents the convolution operation; x represents the input tensor; b represents the bias.

[0086] Perform batch normalization:

[0087]

[0088] Among them, BN represents batch normalization; y represents the output of the convolution operation; μ represents the mean; σ 2 represents the variance; ε represents a constant to prevent division by zero; γ and β represent learnable parameters.

[0089] Use activation function:

[0090]

[0091] Among them, SiLU represents the activation function; z represents the output after batch normalization, σ(z) represents the sigmoid function; e represents a natural constant.

[0092] In one embodiment, if Figure 1 and Figure 3 As shown, the multi-scale feature fusion of the extracted features includes:

[0093] Perform upsampling operation:

[0094] x up =Upsample(x T , sacle_factor = s);

[0095] Among them, x T represents the input feature map; s represents the upsampling factor; x up It is the feature map obtained after upsampling; Upsample means upsampling; sacle_factor means the scaling factor.

[0096] Adjust the number of channels and fusion features of the feature map:

[0097] x conv =Conv(x up ,out_channels=C out , kernel_size=k, stride=1, padding=p);

[0098] Among them, x conv Represents the feature map after convolution; out_channels and C outIt represents the number of output channels, kernel_size and k represent the convolution kernel size; stride represents the step size; padding and p represent padding.

[0099] Concatenate the convolutional feature map with the low-level feature map:

[0100] x fuse =cat(x low ,x conv , dim=1);

[0101] Among them, x fuse Represents the fused feature map; x low Represents a low-level feature map; cat represents the operation of connecting tensors on a preset dimension; the dim parameter refers to the dimension of the concatenation operation.

[0102] In one embodiment, if Figure 1 and Figure 3 As shown, the point information of the illegal parking area is obtained according to the illegal parking area detection picture with the detection frame, including:

[0103] Get the center point of the detection box, select the points near the center point of the box as output points, and obtain the point information of the illegal parking area:

[0104]

[0105] Among them, rand_x and rand_y are the calculated coordinates of the random point, and random.uniform is used to generate a random floating point number within the specified range of the horizontal and vertical coordinates of the center point of the box. c Indicates the horizontal coordinate of the center point of the box; y c Indicates the vertical coordinate of the center point of the box.

[0106] In one embodiment, if Figure 2 and Figure 3 As shown, in the illegal parking area segmentation step, the video image is segmented to obtain multiple segments, each segmented block is used as an independent visual token; each visual token is converted into a high-dimensional vector through a linear embedding process, mapped to the vector space for processing, and a first embedding vector is obtained; the point information is converted into a second embedding vector, which is combined with the first embedding vector to form a feature sequence of the Transformer. Conv Laye represents a convolutional layer. Feature Map represents a feature map. Image embedding represents an image embedding. Transformer Block represents a transformation module, which is responsible for processing an input sequence and generating an output sequence.

[0107] Specifically, step S3 includes the following steps:

[0108] Step S301: The illegal parking area segmentation module will first perform a preliminary segmentation on the input image and then divide it into multiple small blocks. That is, the illegal parking area segmentation module will first perform a preliminary segmentation on the input image and then further divide the processed image into multiple small blocks. Each small block is regarded as an independent visual token. These tokens represent local features in the image and are the basis for subsequent illegal parking area segmentation.

[0109] Step S302: Each visual token is converted into a high-dimensional vector through a linear embedding process, retaining the unique characteristics of the token and mapping it to a vector space suitable for further processing. The embedded vectors are standardized to ensure that the vectors of different tokens are comparable, which facilitates subsequent integration and analysis.

[0110] Step S303: At the same time, the point information obtained in the illegal parking area detection module (step S2) is converted into an embedded vector, and combined with the above output feature vector to form a Transformer input sequence, that is, a new feature sequence. This sequence contains both image information and user prompt information.

[0111] In one embodiment, if Figure 2 and Figure 3 As shown, in the mask decoding step, the integrated embedded information is fused through a transformation operation to obtain the fused illegal parking area features and processed to generate a segmentation mask, which represents the segmentation prediction of the illegal parking area in the image; the immediate self-attention and cross-attention mechanism are adopted, and the global dependency is captured when processing the features of the illegal parking area position through the immediate self-attention, and the connection is established between the features of different illegal parking areas through the cross-attention to interact between the features; bidirectional mapping is performed to adjust the illegal parking area segmentation result.

[0112] In one embodiment, if Figure 2 and Figure 3 As shown, for each illegal parking area input element, the similarity score is calculated, and the information is integrated by weighted average to calculate the Query (Q), Key (K) and Value (V) matrices:

[0113] Q=XW Q ,K=XW K ,V=XW V

[0114] Among them, Query (Q) is the query vector representing the role of the input at the current position in the self-attention mechanism, Key (K) is the key vector used to calculate the correlation between the input at the current position and the input at other positions, and Value (V) is the value vector containing the information of the input at the current position, which is used to calculate the self-attention weight. X represents a given input sequence; W Q , WK , W V is the learned weight matrix.

[0115] Calculate the attention score:

[0116]

[0117] Among them, d k represents the dimension of the vector used to scale the inner product; T represents transpose; softmax represents the normalized exponential function; Attention represents the attention score.

[0118] Perform a nonlinear transformation:

[0119] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 ;

[0120] Here, FFN(x) refers to the nonlinear transformation of the input x by the Position-wise Feed-Forward Network. 1 , W 2 、b 1 and b 2 They all represent the learned parameters.

[0121] Normalize:

[0122] Output=LayerNorm(x+SubLayer(x));

[0123] Among them, Output represents output; LaverNorm represents layer normalization, where Sublayer(x) is the output of the sublayer and x is the input of the sublayer.

[0124] Specifically, step S4 includes the following steps:

[0125] Step S401: The integrated Transformer illegal parking area embedding sequence will be sent to the mask decoding module, where the features of the illegal parking area will be fused through a series of transformation operations (convolution, pooling, non-linear activation function, etc.), that is, the integrated illegal parking area embedding information will be sent to the mask decoding module, where different features will be fused through a series of transformation operations (convolution, pooling, non-linear activation function, etc.); the fused illegal parking area features are further processed to finally generate a segmentation mask. This mask represents the model's segmentation prediction of the illegal parking area in the image.

[0126] Step S402: In order to improve the accuracy and detail of illegal parking area segmentation, the illegal parking area segmentation model adopts the immediate self-attention and cross-attention mechanisms. Self-attention allows the model to consider other positions in the entire feature map when processing the features of the illegal parking area position, thereby capturing global dependencies. Cross-attention is used to establish connections between features of different levels or different types of illegal parking areas to enhance the interaction between features.

[0127] The self-attention mechanism captures global information by calculating the relationship between each illegal parking area and other elements in the input sequence. For each illegal parking area input element, calculate its similarity score with all other elements, and then integrate this information through the weighted average method to generate a new representation. Specifically, given an input sequence X, calculate the Query (Q), Key (K) and Value (V) matrices, where W Q , W K , W V is the learned weight matrix. It is calculated by the following steps:

[0128] Q=XW Q ,K=XW K ,V=XW V

[0129] Assume d k is the dimension of the vector, used to scale the inner product. Then calculate the attention score:

[0130]

[0131] After each attention layer, the Transformer includes a feed-forward neural network layer that applies a nonlinear transformation to each position independently, W 1 , W 2 , b 1 , b 2 are the learned parameters:

[0132] FFN(x)=max(0,xW 1 +b 1 )W 2 +b 2 .

[0133] The Transformer applies residual connections and layer normalization after each sub-layer (self-attention and feed-forward network):

[0134] Output=LayerNorm(x+SubLayer(x)).

[0135] Step S403: The model works in two directions, from the prompt (the point entered by the user) to the image embedding and then back to the mask. This bidirectional mapping helps the model better understand the user's intention and adjust the parking area segmentation result accordingly to achieve a more accurate segmentation effect.

[0136] The embodiment of the present invention also provides a point-prompt road illegal parking area segmentation system based on the Transformer mechanism, such as Figure 1 , Figure 2 as well as Figure 3 As shown, it includes the following modules:

[0137] Preprocessing module: preprocess the input road video to obtain video images.

[0138] Illegal parking area detection module: detects illegal parking areas in lanes of video images and obtains point information of illegal parking areas.

[0139] Illegal parking area segmentation module: segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedded vector; convert the point information into a second embedded vector, integrate the first embedded vector and the second embedded vector to obtain integrated embedded information.

[0140] Mask decoding module: The integrated embedded information is feature-fused, and a segmentation mask is generated through transformation. The segmentation mask covers the illegal parking area and outputs the segmentation video of the illegal parking area.

[0141] In one embodiment, if Figure 1 and Figure 3 As shown in the figure, in the illegal parking area detection module, for a given image to be detected, the input image parameters are 640x640x3, and these three numbers represent the height, width, and number of channels of the feature map. First, the image is feature extracted. The feature extraction modules mainly include the Conv module, the C3 module (residual module), and the SPPF module (spatial pyramid pooling module). The Conv module is mainly composed of a convolutional layer, a BN layer, and a SiLU activation function. The C3 module adaptively aggregates the previous feature map, and the SPPF module obtains a multi-scale feature representation.

[0142] Let W be the convolution kernel weight, * represents the convolution operation, x is the input tensor, and b is the bias. The formula for the convolution operation in the Conv module is as follows:

[0143] Conv(x)=W*x+b(1.1).

[0144] Let y be the output of the convolution operation, μ and σ 2 are the mean and variance respectively, ε is a small constant to prevent division by zero, γ and β are learnable parameters, then the formula for batch normalization is as follows:

[0145]

[0146] The activation function used is SiLU (Sigmoid-Weighted Linear Unit).

[0147] Let z be the output after batch normalization, σ(z) be the sigmoid function, then the activation function SiLU (Sigmoid-Weighted Linear Unit) is as follows:

[0148]

[0149] The C3 module aggregates the features of the feature map obtained in the previous step (a part of the feature map that has been processed) and the unprocessed part of the feature map, that is, splices them together. This ensures that more detail information is retained through jump connections while ensuring the depth of the network. The SPPF module divides the input feature map into grids of different sizes through pooling operations at different scales, and extracts features from each grid. In this way, global and local feature information can be captured at the same time, that is, the model can capture global and local feature information at the same time, improving its adaptability to targets of different scales. There are three parallel maximum pooling operations MaxPool2d, which obtain more comprehensive spatial feature information through weighted fusion of global features and local features.

[0150] After that, it is necessary to perform multi-scale feature fusion on the extracted features and pass them to the prediction layer. In this process, multiple upsampling, concatenation, point and dot product are required to aggregate the features. First, upsampling is performed. Assuming that the input feature map is x, the upsampling factor is s, x up is the feature map obtained after upsampling, then the upsampling operation can be expressed as:

[0151] x up =Upsample(x,sacle_factor=s)(1.4).

[0152] After upsampling, a convolutional layer is added to adjust the number of channels of the feature map and the fusion features. The number of input channels is C in , the number of output channels is C out , the convolution kernel size is kxk, which is a set of 1x1 and 3x3 convolution kernels, with a step size of 1 and a padding of p (default is 0), x conv is the feature map after convolution, then the convolution operation can be expressed as:

[0153] x conv =Conv(x up,out_channels=C out , kernel_size=k, stride=1, padding=p) (1.5).

[0154] After the previous step, the convolution feature map needs to be spliced ​​with the low-level feature map. The fusion operation is usually achieved by element-by-element addition or splicing. Let x fuse is the fused feature map, x low is the low-level feature map, x high is a high-level feature map, where cat is the operation of connecting tensors in a specified dimension. The above operation can be expressed as follows, that is, finally concatenating the convolutional feature map with the low-level feature map:

[0155] x fuse =cat(x low ,x conv ,dim=1) (1.6).

[0156] The last part is responsible for the final regression prediction, that is, the illegal parking area in the image needs to be regressed and predicted, and the location and category of the target are detected using the feature maps extracted from the above two parts. First, the standard convolution layer is used to further extract features, and the 1x1 convolution layer and 3x3 convolution layer are used to reduce the channel dimension and increase the receptive field and feature extraction capabilities, and the decoding and non-maximum suppression are used to remove redundant bounding boxes and improve the detection accuracy to obtain the final illegal parking area detection result.

[0157] At this point, we can get a detection image of the illegal parking area with a detection frame. In order to obtain the point information of the illegal parking area in the image, we get the center point of the frame (x c ,y c ), in order to ensure that the point information obtained is more in line with the illegal parking area segmentation standard, the point near the center of the box is selected as the output point, that is, the illegal parking area detection picture with a detection box is finally obtained, and the center point of the box is set to (x c ,y c ), then the formula for obtaining the illegal parking area point information is as follows:

[0158]

[0159] The above formula can finally obtain the point information in the illegal parking area detection.

[0160] This application adds a parking violation detection module, which integrates the detected point information with the embedded vector obtained by the parking violation segmentation module, making all feature maps more powerful and accurate in semantics, and improving the accuracy of subsequent parking violation segmentation. In the parking violation segmentation module, the transformation operation based on the Transformer mechanism is used to handle complex parking violation segmentation tasks, and the segmentation results can be flexibly adjusted according to the prompts provided by the user. This makes the parking violation segmentation model improve the detection accuracy while also improving the efficiency.

[0161] The present application relates to the field of computer vision, and is a point-prompted road no-parking area segmentation method based on the Transformer mechanism, which belongs to the field of semantic segmentation. The Transformer mechanism is a neural network model based on the self-attention mechanism, which is used to process sequence data. The present application improves the accuracy of lane illegal parking area segmentation, and improves the performance of multi-scale problem detection. The present application realizes the automated segmentation of illegal parking areas while taking into account the segmentation accuracy and detection speed of illegal parking areas, can adapt to most road scenes, and has high robustness and practicality.

[0162] In this specification, the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the embodiments described later, the description is relatively simple, and the relevant parts can be referred to the partial description of the previous embodiments.

[0163] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A point-prompt road illegal parking area segmentation method based on Transformer mechanism, characterized in that: The steps include: Preprocessing step: preprocess the input road video to obtain a video image; Illegal parking area detection step: perform lane illegal parking area detection on the video image to obtain point information of the illegal parking area; Illegal parking area segmentation step: segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedding vector; convert the point information into a second embedding vector, integrate the first embedding vector and the second embedding vector to obtain integrated embedding information; Mask decoding step: perform feature fusion on the integrated embedded information, generate a segmentation mask through transformation, the segmentation mask covers the illegal parking area, and outputs the segmentation video of the illegal parking area.

2. According to the Transformer mechanism-based point-prompted road illegal parking area segmentation method of claim 1, it is characterized in that: In the preprocessing step, the video data set is subjected to frame extraction processing before the illegal parking area detection step, and the image size is unified and the RGB values ​​are normalized; and frame-by-frame processing is performed before the illegal parking area segmentation step, and the image size is unified and the RGB values ​​are normalized.

3. The point-prompted road illegal parking area segmentation method based on Transformer mechanism according to claim 1 is characterized in that: In the illegal parking area detection step, for the video image, the image parameters of the feature map are input, the video image features are extracted, the feature map is adaptively aggregated, the feature map is divided into grids through pooling operations, and features are extracted from the grids. Multi-scale feature fusion is performed on the extracted features to obtain an illegal parking area detection picture with a detection frame, and the point information of the illegal parking area is obtained according to the illegal parking area detection picture with the detection frame.

4. According to claim 3, the method for segmenting illegal parking areas on point-prompted roads based on the Transformer mechanism is characterized in that: Video image feature extraction includes: Perform a convolution operation: Conv(x) = W*x+b; Among them, Conv represents convolution; W represents the convolution kernel weight; * represents the convolution operation; x represents the input tensor; b represents the bias; Perform batch normalization: Among them, BN represents batch normalization; y represents the output of the convolution operation; μ represents the mean; σ 2 represents variance; ε represents a constant to prevent division by zero; γ and β represent learnable parameters; Use activation function: Among them, SiLU represents the activation function; z represents the output after batch normalization, σ(z) represents the sigmoid function; e represents a natural constant.

5. According to claim 3, the method for segmenting illegal parking areas on point-prompted roads based on the Transformer mechanism is characterized in that: Multi-scale feature fusion of the extracted features includes: Perform upsampling operation: x up =Upsample(x T ,sacle_factor=s); Among them, x T represents the input feature map; s represents the upsampling factor; x up It is the feature map obtained after upsampling; Upsample means upsampling; sacle_factor means scaling factor; Adjust the number of channels and fusion features of the feature map: x conv =Conv(x up ,out_channels=C out ,kernel_size=k,stride=1,padding=p); Among them, x conv Represents the feature map after convolution; out_channels and C out Indicates the number of output channels, kernel_size and k indicate the convolution kernel size; stride indicates the step size; padding and p indicate padding; Concatenate the convolutional feature map with the low-level feature map: x fuse =cat(x low ,x conv ,dim=1); Among them, x fuse Represents the fused feature map; x low Represents a low-level feature map; cat represents the operation of connecting tensors on a preset dimension; dim represents the dimension parameter of the concatenation operation.

6. According to the Transformer mechanism-based point-prompted road illegal parking area segmentation method of claim 3, it is characterized in that: According to the illegal parking area detection picture with the detection frame, the point information of the illegal parking area is obtained, including: Get the center point of the detection box, select the points near the center point of the box as output points, and obtain the point information of the illegal parking area: Where rand_x is the calculated horizontal coordinate of the random point, and rand_y is the calculated vertical coordinate of the random point; random.uniform is a random floating point number used to generate the horizontal and vertical coordinates of the center point of the box within the preset range; x c Indicates the horizontal coordinate of the center point of the box; y c Indicates the vertical coordinate of the center point of the box.

7. The point-prompted road illegal parking area segmentation method based on Transformer mechanism according to claim 1 is characterized in that: In the illegal parking area segmentation step, the video image is segmented to obtain a plurality of segments, each segment being an independent visual token; each visual token is converted into a high-dimensional vector through a linear embedding process, mapped to a vector space for processing, and a first embedding vector is obtained; the point information is converted into a second embedding vector, which is combined with the first embedding vector to form a feature sequence of the Transformer.

8. The method for segmenting illegal parking areas on point-prompted roads based on the Transformer mechanism according to claim 1 is characterized in that: In the mask decoding step, the integrated embedded information is fused through a transformation operation to obtain the fused illegal parking area features and, after processing, generate a segmentation mask, which represents the segmentation prediction of the illegal parking area in the image; an immediate self-attention and cross-attention mechanism is adopted, and the global dependency is captured when processing the features of the illegal parking area position through immediate self-attention, and the connection is established between the features of different illegal parking areas through cross-attention to interact between the features; bidirectional mapping is performed to adjust the illegal parking area segmentation result.

9. The method for segmenting illegal parking areas on point-prompted roads based on the Transformer mechanism according to claim 8 is characterized in that: For each illegal parking area input element, calculate the similarity score, integrate the information through weighted average, and calculate the Q, K and V matrices: Q=XW Q ,K=XW K ,V=XW V Among them, Q is the query vector, which indicates the role of the input at the current position in the self-attention mechanism; K is the key vector, which is used to calculate the correlation between the input at the current position and the input at other positions; V is the value vector, which contains the information of the input at the current position and is used to calculate the self-attention weight; X represents the given input sequence; W Q , W K , W V is the learned weight matrix; Calculate the attention score: Among them, d k represents the dimension of the vector used to scale the inner product; T represents transpose; softmax represents the normalized exponential function; Attention represents the attention score; Perform a nonlinear transformation: FFN(x)=max(0,xW1+b1)W2+b2; Where FFN(x) represents a nonlinear transformation; W1, W2, b1, and b2 represent learned parameters; Normalize: Output=LayerNorm(x+SubLayer(x)); Among them, Output represents output; LaverNorm represents layer normalization; Sublayer(x) represents the output of the sublayer.

10. A point-prompt road illegal parking area segmentation system based on Transformer mechanism, characterized in that: Includes the following modules: Preprocessing module: preprocess the input road video to obtain video images; Illegal parking area detection module: detects illegal parking areas in lanes on video images and obtains point information of illegal parking areas; Illegal parking area segmentation module: segment the video image to obtain multiple segments, embed each segment into a vector to obtain a first embedded vector; convert the point information into a second embedded vector, integrate the first embedded vector and the second embedded vector to obtain integrated embedded information; Mask decoding module: The integrated embedded information is feature-fused, and a segmentation mask is generated through transformation. The segmentation mask covers the illegal parking area and outputs the segmentation video of the illegal parking area.