Traffic scene image data analysis and target segmentation method and device
By introducing a dual-frequency domain visual state space module and a spiral visual state space decoder into the Mamba model, the existing Mamba model lacks perception of image details and global structure and lacks understanding of traffic scenes, and achieves more efficient traffic emergency object segmentation performance.
Patent Information
- Application Number
- CN202510017381.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-06
AI Technical Summary
The scanning mechanism in the existing Mamba model needs to be strengthened to perceive image details and global structure, and the lack of in-depth understanding of traffic scenes, resulting in poor performance in traffic emergency object segmentation tasks.
The dual-frequency domain visual state space module is used to process high/low frequency information through shift window scanning and expansion scanning mechanisms, and combined with the spiral visual state space decoder, helical selective scanning strategy combined with driving attention priors for decoding step by step, enhancing the understanding and segmentation ability of traffic scenes.
It effectively improves the model's perception of image details and global structure, enhances its in-depth understanding of traffic scenes, and significantly improves the performance of traffic emergency object segmentation.
Smart Images

Figure CN120107282A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic scene image processing, and in particular to a method and device for traffic scene image data analysis and target segmentation. Background Art
[0002] As the number of vehicles on the road increases, the traffic environment becomes more and more complex and congested. Accurately analyzing traffic scene images captured by on-board cameras and identifying emergency objects (such as pedestrians, vehicles, and obstacles) with potential traffic accident risks plays a vital role in enabling vehicles to quickly take countermeasures such as avoidance, deceleration, or braking. Although the human visual system can automatically focus on important objects in complex and dynamic scenes, it often has difficulty responding to sudden unexpected events in a timely manner. Therefore, using computer vision technology to parse traffic images and outline the areas of emergency objects can effectively support Advanced Driver-Assistance Systems (ADAS) to make adaptive risk avoidance responses.
[0003] As we all know, the salient object detection (SOD) task aims to imitate the human visual perception system to analyze and segment the most conspicuous objects in the image. From 2015 to 2022, models based on convolutional neural networks (CNNs) have significantly improved the performance of SOD based on the powerful multi-level feature learning ability of CNNs. In recent years, Transformer-based models have gradually become the mainstream framework for SOD and have achieved excellent performance. Therefore, it is a natural idea to use the SOD algorithm to implement traffic emergency object segmentation (TEOS). From this perspective, some studies have attempted to perform salient object segmentation in traffic scenes, but due to the lack of a dedicated traffic SOD dataset, they only extracted images related to traffic scenes from the existing general SOD dataset for experimental verification. However, in some scenarios, salient objects and emergency objects are essentially different. SOD focuses on visual saliency, while TEOS requires deep traffic context awareness (such as road conditions and traffic regulations). It is essentially driven by semantics and needs to filter out visually misleading elements. In other words, in a traffic driving environment, the most visually prominent object may not be the most critical to traffic safety, and may even mislead the driver's attention. Therefore, directly applying the SOD algorithm to TEOS is not the best choice. In addition, another task related to emergency target segmentation in traffic scenes is the Driving Attention Prediction (DAP) task, which aims to design a human-centric ADAS, analyze and predict the driver's attention perception behavior and avoid unsafe operations. For example, some studies have introduced semantic context features of driving scenes to help find key objects / areas that attract the driver's attention. However, (1) due to differences in driving habits and safety awareness, the driver's gaze point is subjective; (2) the DAP task usually outputs a heat map of visual attention points rather than a pixel-level prediction of the complete object. In complex scenes, detecting the complete object is more critical, so the applicability of the DAP task in complex traffic scenes is limited. In contrast, TEOS reduces the reliance on the driver to notice and respond to emergency situations, thereby significantly reducing the probability of accidents caused by human errors (such as inattention or misjudgment of danger).
[0004] Due to the lack of open source TEOS datasets, there is little investment in related research, which is the primary problem that needs to be solved at present. In addition, in the field of image data processing, the performance of Mamba-based models has gradually surpassed CNN and Transformer models. As a key component of the Mamba visual model, an effective scanning mechanism can improve model performance and facilitate the training process. For example, Cross Scan is the most widely used scanning method, which unfolds along four different paths and flattens the image blocks, which can be regarded as a fusion of two bidirectional scans; Continuous Scan processes adjacent markers between columns (or rows) instead of jumping to the opposite marker like Cross Scan; Hilbert Scan scans along a tortuous path based on the Hilbert matrix; Hierarchical Scan uses different convolution kernel sizes to capture semantic knowledge from a global to local or from a macro to micro perspective; Spatiotemporal scan includes two 3D scanning methods, namely space-first scanning and time-first scanning. However, the current scanning strategy ignores the special domain information of traffic scenes, and the modeling of local and global information needs to be further strengthened. Based on the above background, achieving high-performance Mamba-based traffic emergency object segmentation still faces the following challenges: the scanning mechanism in the existing Mamba model needs to strengthen its perception of image details and global structure; the existing scanning mechanism lacks a deep understanding of traffic scenes. Summary of the invention
[0005] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a method and device for traffic scene image data analysis and target segmentation are provided. The present invention aims to solve the problem that the scanning mechanism in the existing Mamba model needs to strengthen the perception of image details and global structure and lacks an in-depth understanding of traffic scenes.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for traffic scene image data analysis and target segmentation comprises using a pre-trained traffic scene object segmentation model to obtain a target segmentation result for an input image of a traffic scene, wherein the traffic scene object segmentation model comprises an S-stage visual state space encoder, S-1 dual-frequency domain visual state space modules, an S-1-stage spiral visual state space decoder, an upsampling module and a segmentation head, wherein the S-stage visual state space encoder is used to extract S levels of RGB features, the S-1 dual-frequency domain visual state space module is used to perform high-frequency and low-frequency frequency domain enhancement on the first S-1 levels of RGB features through a shift window scanning and an expansion scanning mechanism, and then jump-connect to the input end of a corresponding level of spiral visual state space decoder, and the spiral visual state space decoder of the S-1 stages adopts a spiral selective scanning strategy combined with a priori of driving attention to perform step-by-step decoding, and finally the first-stage spiral visual state space decoder outputs the output through the upsampling module and the segmentation head to obtain a final prediction result segmentation map.
[0007] Optionally, when the high-frequency and low-frequency frequency domain enhancement are performed on the RGB features of the first S-1 levels, respectively, the dual-frequency domain visual state space module performs high-frequency and low-frequency frequency domain enhancement on the RGB features of any s-th level, respectively, including: S101, through offline discrete cosine transform DCT, the input RGB features are mapped to the frequency domain and decomposed into high-frequency components and low-frequency components in the frequency domain and sent to the parallel high-frequency branch and low-frequency branch respectively; S102, in the high frequency branch and the low frequency branch, respectively, performing layer normalization, linearization and deep convolution extraction on the high frequency component and the low frequency component to obtain a high frequency feature map and a low frequency feature map; S103, dividing the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy, and then scanning the window translation inside the window along the horizontal or vertical direction to generate an expanded window sequence, generating a cross-window global sequence along the same direction, and then sequentially passing through the selective state space model, layer normalization, and linearization to obtain the output features of the high-frequency branch; dividing the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy, and then scanning the expansion window inside the expansion window along the horizontal or vertical direction to generate an expanded expansion window sequence, generating a cross-window global sequence along the same direction, and then sequentially passing through the selective state space model, layer normalization, and linearization to obtain the output features of the low-frequency branch; S104, concatenate the output features of the high-frequency branch and the output features of the low-frequency branch, linearize them, and multiply them with the features of the input dual-frequency domain visual state space module, and then pass them through the feedforward network FFN to obtain the high- and low-frequency frequency domain enhanced features.
[0008] Optionally, in step S103, the function expression for dividing the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy is: , , , , in, For forward horizontal window scanning, For the merge operation, Indicates the window size, and are the height and width of the input image, for Round down, for Round down, are the coordinates of the window, The two-dimensional feature map of the input No. Row, No. The elements of the column, For forward vertical window scanning, For the general The scan results are indexed in reverse order. For reverse horizontal scanning, For the general The scan results are indexed in reverse order. For reverse longitudinal scanning, Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
[0009] Optionally, in step S103, the function expression for dividing the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy is: , , , , in, For the forward lateral expansion scan, For the merge operation, is the expansion rate, From 0 to Change, indicating that this operation performs expansion sampling on every possible starting position, For After expanding the rows into a one-dimensional sequence, the scan elements are expanded with a step size R. and are the height and width of the input image, for Round up, For the forward longitudinal expansion scan, For reverse lateral expansion scan, For the general The scan results are indexed in reverse order. For reverse longitudinal expansion scanning, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
[0010] Optionally, the spiral visual state space decoder includes a layer normalization module, a linearization module, a deep convolution module, a SiLU activation function module, a spiral visual state space module, a layer normalization module, a GELU activation function module, a linearization module, a jump connection module and a domain visual dual-frequency state space module connected in sequence, one input of the jump connection module comes from the GELU activation function module, and the other input comes from the original input of the spiral visual state space decoder.
[0011] Optionally, the function expression of the spiral visual state space module is: , , , , in, Based on and The spiral scanning sequence generated by two point sets, For the merge operation, is the height of the input image, for Round down, Based on and Two point sets select specific sets of pixels from an image, For arrive The point set of For arrive These point sets are generated by the Bresenham algorithm and are used to determine the spiral sampling path; and The start and end points should be relative to and Offset to get For arrive The point set of For arrive The point set of Based on and The spiral scanning sequence generated by two point sets, for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
[0012] Optionally, the S-stage visual state space encoder includes a block embedding layer and S stacked visual state space modules, the block embedding layer is used to divide the input image into non-overlapping image blocks that do not use position encoding, and the input end of the last S-1 level visual state space module in the S stacked visual state space modules is connected in series with a downsampling module, and the S stacked visual state space modules respectively output S levels of RGB features with different resolution ratios to the input image ~ , and the RGB features of the first S-1 levels ~ Each of them is output to a corresponding dual-frequency domain visual state space module.
[0013] In addition, the present invention also provides a system for traffic scene image data analysis and target segmentation, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the method for traffic scene image data analysis and target segmentation.
[0014] In addition, the present invention also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the method of traffic scene image data analysis and target segmentation through a processor.
[0015] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the method of traffic scene image data analysis and target segmentation through a processor.
[0016] Compared with the prior art, the present invention mainly has the following advantages: 1. The present invention includes a dual-frequency domain visual state space module that utilizes a shifted window scanning and an expansion scanning mechanism to process high / low frequency information to enhance the perception of details and global structures, and can solve the problem that the scanning mechanism in the existing Mamba model needs to enhance the perception of image details and global structures.
[0017] 2. The present invention includes a spiral visual state space decoder that adopts a spiral selective scanning strategy combined with driving attention prior to effectively capture global multi-directional information to emphasize key areas in traffic scenes, thereby effectively solving the problem that the scanning mechanism in the existing Mamba model lacks an in-depth understanding of traffic scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of the network structure of the traffic scene object segmentation model (Tramba model) in an embodiment of the present invention.
[0019] Figure 2 Schematic diagram of the network structure of the dual-frequency domain visual state space module in an embodiment of the present invention.
[0020] Figure 3 Schematic diagram of the network structure of the spiral visual state space decoder in an embodiment of the present invention.
[0021] Figure 4 The quantitative comparison results of the Tramba model in the embodiment of the present invention and 23 competing models are shown in FIG.
[0022] Figure 5This is the first part of the comparison results between the Tramba model and other SOD models in the embodiment of the present invention.
[0023] Figure 6 This is the second part of the comparison results between the Tramba model and other SOD models in the embodiment of the present invention.
[0024] Figure 7 Result of ablation experiment of Tramba model in the embodiment of the present invention. DETAILED DESCRIPTION
[0025] In order to enable those skilled in the art to better understand the technical solution of the present invention, the technical solution of the present invention will be further described in detail below in conjunction with the accompanying drawings in the embodiments of the present invention.
[0026] This embodiment provides a method for traffic scene image data analysis and target segmentation, including using a pre-trained traffic scene object segmentation model Tramba (Tramba model) to obtain a target segmentation result for an input image of a traffic scene. The traffic scene object segmentation model Tramba in this embodiment is improved based on the Mamba model, based on the U-shaped Mamba structure and the VMamba backbone network of the Mamba model, and on this basis, the decoder and the dual-frequency domain visual state space module between the VMamba backbone network and the decoder are improved. Specifically, Figure 1As shown, the traffic scene object segmentation model Tramba in this embodiment includes S stages of visual state space (Visual State Space, VSS) encoders, S-1 dual frequency domain visual state space modules (DFVSS), S-1 stages of spiral visual state space (Helix Visual State Space, HVSS) decoders, upsampling modules and segmentation heads, wherein the S stage visual state space encoders are used to extract S levels of RGB features, and the S-1 dual frequency domain visual state space modules are used to perform high and low frequency frequency domain enhancement on the first S-1 levels of RGB features through shift window scanning and expansion scanning mechanisms, and then jump to the input end of the corresponding level of spiral visual state space decoder, and the spiral visual state space decoder of the S-1 stage adopts a spiral selective scanning strategy combined with a priori of driving attention (that is, the driver focuses on the middle center area in front of the road to ensure driving safety) to perform decoding step by step, and finally the first level spiral visual state space decoder outputs the output through the upsampling module and the segmentation head to obtain the final prediction result segmentation map. Among them, the stage number S of the visual state space encoder can be determined according to actual needs. For example, as an optional implementation, in this embodiment, S=4, that is, the Tramba model includes 4 stages of visual state space encoders, 3 dual-frequency domain visual state space modules and 3 stages of spiral visual state space decoders.
[0027] like Figure 1 As shown, in this embodiment, the S-stage visual state space encoder includes a block embedding layer and S stacked visual state space modules (VSS blocks). The block embedding layer is used to transform the input image The image is divided into non-overlapping image blocks that do not use position encoding, and a downsampling module is connected in series to the input end of the last S-1 level visual state space module in the S stacked visual state space modules, and the S stacked visual state space modules respectively output S levels of RGB features with different resolution ratios to the input image. ~ , and the RGB features of the first S-1 levels ~ They are output to a corresponding dual-frequency domain visual state space module respectively, where and are the height and width of the input image respectively. In this embodiment, the downsampling module has a magnification of 1 / 2, so the feature map output by the S stacked visual state space modules is equal to the input image The resolution ratios are 1 / 4, 1 / 8, 1 / 16 and 1 / 32 respectively. After the dual frequency domain visual state space module is cascaded in the bottom three encoding stages, the intermediate features with enhanced frequency domain information are obtained. , whose resolution is the same as the original resolution. On the other hand, multiple HVSS blocks are stacked at each decoding stage, whose input is obtained by connecting the high-level features in the skip connection And the intermediate layer features after upsampling The fusion features obtained by splicing Each segmentation header layer consists of a Convolution is used to Mapping to prediction results , to achieve deep supervision, and The final prediction result.
[0028] Considering the distinguishability of foreground and background features in the frequency domain, a new dual-frequency domain visual state space (DFVSS) module is proposed in this embodiment to combine the complementary advantages of the RGB domain and the frequency domain. Specifically, DFVSS first introduces the offline discrete cosine transform (DCT) to map the RGB features to the frequency domain and decompose them into high-frequency and low-frequency components. In order to further enhance the frequency representation and model the interaction of frequency information, DFVSS contains two parallel branches: high-frequency VSS (High-Frequency VSS, HFVSS branch / high-frequency branch) based on the sliding window scanning strategy (Shifted Window 2D-Selective-Scan, Window-SS2D) and low-frequency VSS (Low-Frequency VSS, LFVSS branch / low-frequency branch) based on the dilation scanning strategy (Dilation 2D-Selective-Scan, Dilation-SS2D). As shown in Figure 2 As shown in FIG. 1 , when the RGB features of the first S-1 levels are enhanced in the high and low frequency domains respectively, the dual-frequency domain visual state space module performs high and low frequency frequency domain enhancement on the RGB features of any s-th level respectively, including: S101, through offline discrete cosine transform DCT, the input RGB features are mapped to the frequency domain and decomposed into high-frequency components and low-frequency components in the frequency domain and sent to the parallel high-frequency branch and low-frequency branch respectively; S102, in the high frequency branch and the low frequency branch, respectively, performing layer normalization, linearization and deep convolution extraction on the high frequency component and the low frequency component to obtain a high frequency feature map and a low frequency feature map; S103, divide the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy, and then scan the window translation inside the window in the horizontal or vertical direction to generate an expanded window sequence, generate a cross-window global sequence along the same direction, and then pass through the selective state space model (Selective State-Space Model, S6, a well-known framework for efficiently modeling sequence data by selectively scanning and updating the state space), layer normalization, and linearization to obtain the output features of the high-frequency branch; divide the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy, and then scan the expansion window inside the expansion window in the horizontal or vertical direction to generate an expanded expansion window sequence, generate a cross-window global sequence along the same direction, and then pass through the selective state space model, layer normalization, and linearization to obtain the output features of the low-frequency branch; S104, concatenate the output features of the high-frequency branch and the output features of the low-frequency branch, linearize them, and multiply them with the features of the input dual-frequency domain visual state space module, and then pass them through the feedforward network FFN to obtain the high- and low-frequency frequency domain enhanced features.
[0029] In this embodiment, the HFVSS branch introduces a novel sliding window scanning mechanism, which first divides the feature map into non-overlapping translation windows (e.g. or ), and then scan inside the window along the horizontal or vertical direction respectively to generate a local expansion sequence of image blocks, and finally generate a cross-window global sequence along the same direction. In this way, the structural relationship in the local sequence of image blocks is prioritized, thereby enhancing the ability to capture rapidly changing local edge and texture features. Different from local window self-attention, the HFVSS branch promotes information interaction between local windows without sliding, effectively alleviating the problems of structural information loss and receptive field degradation. For two-dimensional input , the function expression of dividing the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy in step S103 is: , , , , in, For forward horizontal window scanning, For the merge operation, Indicates the window size, and are the height and width of the input image, for Round down, for Round down, are the coordinates of the window, The two-dimensional feature map of the input No. Row, No. The elements of the column, For forward vertical window scanning, For the general The scan results are indexed in reverse order. For reverse horizontal scanning, For the general The scan results are indexed in reverse order. For reverse longitudinal scanning, Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
[0030] To further enhance the understanding of global context, this embodiment proposes a dilation scanning mechanism in the LFVSS branch, which preferentially models long-distance dependencies between image patches through sparse connections. Dilation scanning introduces gaps (or dilations) between sampling points, which enables exponential expansion of the receptive field at a lower cost. The degree of dilation is controlled by the dilation rate. For a two-dimensional input In step S103 of this embodiment, the function expression for dividing the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy is: , , , , in, For the forward lateral expansion scan, For the merge operation, is the expansion rate, From 0 to changes, indicating that this operation performs an expansion sampling on every possible starting position (ensuring comprehensive coverage), For After expanding the rows into a one-dimensional sequence, the scan elements are expanded with a step size R. and are the height and width of the input image, for Round up, For the forward longitudinal expansion scan, For reverse lateral expansion scan, For the general The scan results are indexed in reverse order. For reverse longitudinal expansion scanning, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , Recognizing the importance of different frequency bands in accurate object localization and segmentation, the method of this embodiment concatenates and fuses the outputs from the dual-branch structure to allow full interaction between different spectra. This comprehensive frequency domain representation is then used to enhance RGB features.
[0031] Some recent studies have focused on enhancing Mamba's 2D processing capabilities by incorporating additional scanning strategies into 2D-Selective-Scan (SS2D). However, these methods are still insufficient to fully capture rich directional information and may encounter difficulties in traffic scenes with complex backgrounds, dynamic changes, and diverse target features. In order to effectively address the lack of directionality of traditional scanning mechanisms and enhance attention to key areas of traffic scenes, this embodiment proposes a novel Helix Visual State Space (HVSS) module, which is inspired by the tendency of the human visual system to preferentially focus on the center of the visual field to enhance balance. Existing studies have confirmed that experienced drivers mainly focus on the road ahead to ensure safe driving, that is, urgent objects are concentrated in the central area of the field of view. Specifically, the Helix Visual State Space HVSS uses a designed spiral scan to scan short sequences along the central axis, which can be regarded as "image slices" along one direction, with each slice covering the central area. After each scan, the axis rotates by the distance of two image blocks, and the scan continues until a circle is completed, forming two sequences of forward and backward slices. The remaining gap provides another pair of sequences. Then, the context information of each sequence is captured in parallel and the output sequence is merged, ensuring that the context information from all directions at each position is fully integrated. Figure 3As shown, the spiral visual state space decoder in this embodiment includes a layer normalization module, a linearization module, a deep convolution module, a SiLU activation function module, a spiral visual state space module, a layer normalization module, a GELU activation function module, a linearization module, a jump connection module and a domain visual dual-frequency state space module (existing module) connected in sequence, one input of the jump connection module comes from the GELU activation function module, and the other input comes from the original input of the spiral visual state space decoder. In this embodiment, the function expression of the spiral visual state space module is: , , , , in, Based on and The spiral scanning sequence generated by two point sets, For the merge operation, is the height of the input image, for Round down, Based on and Two point sets select specific sets of pixels from an image, For arrive The point set of For arrive These point sets are generated by the Bresenham algorithm and are used to determine the spiral sampling path; and The start and end points should be relative to and Offset to get For arrive The point set of For arrive The point set of Based on and The spiral scanning sequence generated by two point sets, for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
[0032] In order to verify the method of traffic scene image data analysis and target segmentation in this embodiment, the traffic scene object segmentation model Tramba (Tramba for short) in this embodiment is implemented based on PyTorch on an NVIDIA RTX 4090 GPU (24 GB). During the Tramba model training process, the encoder weights are initialized using the pre-trained VMamba-B backbone network. The size of each image is scaled to 384×384, and data enhancement methods such as cropping and flipping are used. In this embodiment, binary cross-entropy loss (Binary Cross-Entropy, BCE) and intersection over Union loss (Intersection over Union, IoU) are used to train the model, and a deep supervision strategy is used throughout the process. The Adam optimizer is used, and the initial learning rate is set to 1e-4, which is kept unchanged in the first 60 iterations and then reduced by 5 times in the remaining 20 iterations, for a total of 80 training iterations. For model evaluation, eight indicators are used in this embodiment to compare all models, including: maximum / average / adaptive F-value ( , , )、Maximum / average / adaptive E value( , , )、S value and Mean Absolute Error, ). Moreover, in this embodiment, the first large-scale traffic emergency object segmentation dataset is first established, named TEOS10K-TE dataset, which contains 13,753 images taken by vehicle cameras with pixel-level annotations. In order to maintain the high diversity of the TEOS10K-TE dataset, images from various real-life traffic scenes are collected in this embodiment, including urban intersections, highways, and rural roads, covering different weather conditions (such as rainy days, snowy days, and haze) and lighting conditions (such as sunny days and low light). In order to ensure the quality of annotation, three people with rich driving experience and strong safety awareness were selected in this embodiment to perform pixel-level annotation. Since there is currently no direct model for the TEOS task, this embodiment compares the proposed Tramba model with 23 state-of-the-art salient object detection models (SOD), including 12 models based on convolutional neural networks CNN, 10 models based on Transformer and one model based on Mamba, such as Figure 4 To ensure fairness, all SOD models are fully trained on the TEOS dataset until the loss is stable, based on their public source code and default network parameter settings. Figure 4 The quantitative comparison results of the Tramba model with 23 existing models are shown, where the number of parameters and FLOPs values compare the complexity of each model. Figure 4 It can be seen that the Tramba model proposed in this embodiment outperforms all CNN or Transformer-based SOD models in all indicators. Compared with the Mamba-based model VMamba, the Tramba model still shows a significant performance advantage. These outstanding results show that the Tramba model not only effectively mimics human visual attention to salient areas, but also effectively distinguishes the semantic priority between salient objects and urgent objects in traffic scenes. In addition, this embodiment visualizes the comparison results of the Tramba model and other SOD models in several challenging traffic scenes. Figure 5 and Figure 6 See Figure 5 and Figure 6 It can be seen that the Tramba model of this embodiment can accurately locate the position of emergency objects and maintain the integrity of the objects with fine edge details, thereby generating high-quality result maps in various traffic scenarios.
[0033] In addition, this embodiment also uses ablation experiments to evaluate the role of the main modules in the Tramba model in this embodiment. This embodiment first starts with a baseline model without a dual-frequency visual state space module (dual-frequency VSS) and a spiral visual state space module (spiral VSS). Its performance is as follows Figure 7 Next, the experiment added the high-frequency branch (high-frequency VSS) and low-frequency branch (low-frequency VSS) in the proposed dual-frequency domain visual state space module, and the spiral scanning strategy in the spiral visual state space module (spiral VSS), and the results are shown in Figure 7 From line 2 to line 4. Figure 7 It can be seen from the performance comparison results shown that: 1) The Tramba model in this embodiment uses frequency features as an auxiliary to bring obvious performance improvement; 2) The high-frequency branch (based on the sliding window scanning strategy) and low-frequency branch (based on the expansion scanning strategy) proposed by the Tramba model in this embodiment effectively learn rich frequency perception information containing high-frequency and low-frequency clues; 3) The spiral visual state space module (spiral VSS) of the Tramba model in this embodiment enriches the direction encoding and central area focus, effectively improving the performance of the model in complex traffic scenarios.
[0034] In addition, this embodiment also provides a system for traffic scene image data analysis and target segmentation, including a microprocessor and a memory connected to each other, and the microprocessor is programmed or configured to execute the traffic scene image data analysis and target segmentation method.
[0035] In addition, this embodiment also provides a computer-readable storage medium, in which a computer program or instruction is stored. The computer program or instruction is programmed or configured to execute the method of traffic scene image data analysis and target segmentation through a processor.
[0036] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the method of traffic scene image data analysis and target segmentation through a processor.
[0037] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application may be in the form of methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0038] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A method for traffic scene image data analysis and target segmentation, characterized in that: The method comprises using a pre-trained traffic scene object segmentation model to obtain a target segmentation result from an input image of a traffic scene, wherein the traffic scene object segmentation model comprises an S-stage visual state space encoder, S-1 dual-frequency domain visual state space modules, an S-1-stage spiral visual state space decoder, an upsampling module and a segmentation head, wherein the S-stage visual state space encoder is used to extract S levels of RGB features, the S-1 dual-frequency domain visual state space module is used to perform high-frequency and low-frequency frequency domain enhancement on the first S-1 levels of RGB features through a shift window scanning and an expansion scanning mechanism, and then jump-connect the RGB features to the input end of a corresponding level of spiral visual state space decoder, and the spiral visual state space decoder of the S-1 stages adopts a spiral selective scanning strategy combined with a prior of driving attention to perform decoding step by step, and finally the first-stage spiral visual state space decoder outputs the output through an upsampling module and a segmentation head to obtain a final prediction result segmentation map.
2. The method for traffic scene image data analysis and target segmentation according to claim 1, characterized in that: When the RGB features of the first S-1 levels are enhanced in the high and low frequency domains respectively, the dual frequency domain visual state space module performs high and low frequency frequency domain enhancement on the RGB features of any s-th level respectively, including: S101, through offline discrete cosine transform DCT, the input RGB features are mapped to the frequency domain and decomposed into high-frequency components and low-frequency components in the frequency domain and sent to the parallel high-frequency branch and low-frequency branch respectively; S102, in the high frequency branch and the low frequency branch, respectively, performing layer normalization, linearization and deep convolution extraction on the high frequency component and the low frequency component to obtain a high frequency feature map and a low frequency feature map; S103, dividing the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy, and then scanning the window translation inside the window along the horizontal or vertical direction to generate an expanded window sequence, generating a cross-window global sequence along the same direction, and then sequentially passing through the selective state space model, layer normalization, and linearization to obtain the output features of the high-frequency branch; dividing the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy, and then scanning the expansion window inside the expansion window along the horizontal or vertical direction to generate an expanded expansion window sequence, generating a cross-window global sequence along the same direction, and then sequentially passing through the selective state space model, layer normalization, and linearization to obtain the output features of the low-frequency branch; S104, concatenate the output features of the high-frequency branch and the output features of the low-frequency branch, linearize them, and multiply them with the features of the input dual-frequency domain visual state space module, and then pass them through the feedforward network FFN to obtain the high- and low-frequency frequency domain enhanced features.
3. The method for traffic scene image data analysis and target segmentation according to claim 2, characterized in that: In step S103, the function expression for dividing the high-frequency feature map into non-overlapping translation windows based on the sliding window scanning strategy is: , , , , in, For forward horizontal window scanning, For the merge operation, Indicates the window size, and are the height and width of the input image, for Round down, for Round down, are the coordinates of the window, The two-dimensional feature map of the input No. Row, No. The elements of the column, For forward vertical window scanning, For the general The scan results are indexed in reverse order. For reverse horizontal scanning, For the general The scan results are indexed in reverse order. For reverse longitudinal scanning, Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
4. The method for traffic scene image data analysis and target segmentation according to claim 2, characterized in that: The function expression for dividing the low-frequency feature map into non-overlapping expansion windows based on the expansion scanning strategy in step S103 is: , , , , in, For the forward lateral expansion scan, For the merge operation, is the expansion rate, From 0 to Change, indicating that this operation performs expansion sampling on every possible starting position, For After expanding the rows into a one-dimensional sequence, the scan elements are expanded with a step size R. and are the height and width of the input image, for Round up, For the forward longitudinal expansion scan, For reverse lateral expansion scan, For the general The scan results are indexed in reverse order. For reverse longitudinal expansion scanning, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
5. The method for traffic scene image data analysis and target segmentation according to claim 1, characterized in that: The spiral visual state space decoder includes a layer normalization module, a linearization module, a deep convolution module, a SiLU activation function module, a spiral visual state space module, a layer normalization module, a GELU activation function module, a linearization module, a jump connection module and a domain visual dual-frequency state space module which are connected in sequence. One input of the jump connection module comes from the GELU activation function module, and the other input comes from the original input of the spiral visual state space decoder.
6. The method for traffic scene image data analysis and target segmentation according to claim 5, characterized in that: The function expression of the spiral visual state space module is: , , , , in, Based on and The spiral scanning sequence generated by two point sets, For the merge operation, is the height of the input image, for Round down, Based on and Two point sets select specific sets of pixels from an image, For arrive The point set of For arrive These point sets are generated by the Bresenham algorithm and are used to determine the spiral sampling path; and The start and end points should be relative to and Offset to get For arrive The point set of For arrive The point set, Based on and The spiral scanning sequence generated by two point sets, for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. for The reverse spiral scanning sequence, For the general The scan results are indexed in reverse order. Equivalent to and The product of "and" " respectively represent forward scanning and reverse scanning, the input two-dimensional feature map Any element in row i and column j ,in is the number of feature channels, , .
7. The method for traffic scene image data analysis and target segmentation according to claim 1, characterized in that: The S-stage visual state space encoder includes a block embedding layer and S stacked visual state space modules, wherein the block embedding layer is used to divide the input image into image blocks that do not overlap and do not use position encoding, and the input end of the last S-1 level visual state space module in the S stacked visual state space modules is connected in series with a downsampling module, and the S stacked visual state space modules respectively output S levels of RGB features with different resolution ratios to the input image. ~ , and the RGB features of the first S-1 levels ~ Each of them is output to a corresponding dual-frequency domain visual state space module.
8. A system for analyzing traffic scene image data and segmenting targets, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the method for traffic scene image data analysis and target segmentation as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the method for analyzing traffic scene image data and segmenting targets as described in any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the method for analyzing traffic scene image data and segmenting targets as described in any one of claims 1 to 7 through a processor.
Citation Information
Cited By
Remote sensing image landslide identification method based on visual state space and frequency domain enhancement
CN121170486A
Non-local information compensation Mama image deblurring method and system
CN121437325A
A method and system for deblurring Mamba images with nonlocal information compensation
CN121437325B
Wafer image super-resolution method and related equipment
CN121481846A