A method, apparatus, storage medium, and electronic device for generating a topological scene graph

By enhancing and de-raining low-light and raining images, combining natural language supervision and Fourier learning strategies, a robust traffic scene map is generated, which solves the problem of image acquisition differences in autonomous driving and improves global path planning capabilities and safety.

CN120031770BActive Publication Date: 2025-07-08HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510519727.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-08
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

The existing autonomous driving technology has great differences in image acquisition in different driving environments, which leads to decision-making problems and lacks global path planning capabilities.

Method used

The topological scene graph generation method is adopted to generate robust traffic scene graphs by enhancing and de-raining low-light and raining images, combining natural language supervision and Fourier learning strategies, and providing global topological understanding.

Benefits of technology

It improves the robustness of image acquisition, enhances the safety of autonomous driving and global path planning capabilities, and adapts to different driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120031770B_ABST
    Figure CN120031770B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, storage medium and electronic device for generating a topological scene graph, which classifies the images around the vehicle at the target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than the brightness threshold and without rain features, and the rainy images are images with rain features; performs light enhancement processing on the low-light images to obtain low-light corrected images; performs rain removal processing on the rainy images to obtain rain-corrected images; performs fusion processing on the low-light corrected images, rain-corrected images and normal images to obtain the topological scene graph at the target moment, where the normal images are images other than the low-light images and rainy images in the images around the vehicle. It not only improves the robustness of the images collected in different driving environments, but also proposes a traffic topological scene graph, providing a new idea for global path planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more particularly, to a method, apparatus, storage medium, and electronic device for generating a topological scene graph. Background Art

[0002] In recent years, with the continuous development of technologies in the field of artificial intelligence, autonomous driving technology has become an important driving force for the development of the automotive industry. Currently, it is in the transition from assisted driving to conditional autonomous driving and also in the stage of starting highly autonomous driving, with broad development prospects.

[0003] The safety issue of autonomous driving is an important aspect that cannot be ignored during the application process. However, the existing methods only perform the same processing on the image information collected at the vehicle end, without considering that the images collected in different driving environments may vary greatly and do not meet the judgment criteria, which brings difficulties to the decision-making of autonomous driving. Summary of the Invention

[0004] The purpose of the present invention is to provide a method, apparatus, storage medium, and electronic device for generating a topological scene graph to improve the above problems.

[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present invention are as follows:

[0006] In a first aspect, an embodiment of the present invention provides a method for generating a topological scene graph, the method including:

[0007] Classify the images around the vehicle at the target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than the brightness threshold and without rainy features, and the rainy images are images with rainy features;

[0008] Perform light enhancement processing on the low-light images to obtain low-light corrected images;

[0009] Perform rain removal processing on the rainy images to obtain rain-corrected images;

[0010] Perform fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment, where the normal images are the images in the images around the vehicle other than the low-light images and the rainy images.

[0011] In a second aspect, an embodiment of the present invention provides a topological scene graph generating apparatus, the apparatus including:

[0012] A first processing unit, configured to classify the images around the vehicle at a target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than a brightness threshold and not including rainy features, and the rainy images are images including rainy features;

[0013] The first processing unit is further configured to perform light enhancement processing on the low-light images to obtain low-light corrected images;

[0014] The first processing unit is further configured to perform rain removal processing on the rainy images to obtain rain-corrected images;

[0015] A second processing unit, configured to perform fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment, where the normal images are images in the images around the vehicle other than the low-light images and the rainy images.

[0016] In a third aspect, an embodiment of the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.

[0017] In a fourth aspect, an embodiment of the present invention provides an electronic device, where the electronic device includes: a processor and a memory, and the memory is used to store one or more programs; when the one or more programs are executed by the processor, the above method is implemented.

[0018] Compared with the prior art, a topological scene graph generation method, device, storage medium, and electronic device provided by an embodiment of the present invention classify the images around the vehicle at a target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than a brightness threshold and not including rainy features, and the rainy images are images including rainy features; perform light enhancement processing on the low-light images to obtain low-light corrected images; perform rain removal processing on the rainy images to obtain rain-corrected images; perform fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment, where the normal images are images in the images around the vehicle other than the low-light images and the rainy images. It not only increases the robustness of the images collected in different driving environments, but also proposes a traffic topological scene graph, providing a new idea for global path planning.

[0019] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically enumerates preferred embodiments and cooperates with the attached drawings for detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for the embodiments. It should be understood that the following drawings only show certain embodiments of the present invention, and therefore should not be regarded as a limitation of the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention;

[0022] Figure 2 It is a schematic flowchart of a method for generating a topology scenario graph provided by an embodiment of the present invention;

[0023] Figure 3 It is a schematic diagram of units of a topology scenario graph generation device provided by an embodiment of the present invention.

[0024] In the figure: 10 - processor; 11 - memory; 12 - bus; 13 - communication interface; 501 - first processing unit; 502 - second processing unit. Detailed implementation manners

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0027] It should be noted that: similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present invention, terms such as "first" and "second" are only used for differential description and cannot be understood as indicating or implying relative importance.

[0028] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0029] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the inventive product is usually placed during use. It is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0030] In the description of the present invention, it should also be noted that unless otherwise clearly specified and limited, the terms "set" and "connect" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0031] The following will describe in detail some embodiments of the present invention with reference to the drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0032] In some scenarios, it is not considered that the images collected in different driving environments may be unclear, such as in low light or rainy conditions. This will inevitably have a great impact on the perception ability of autonomous driving and pose a great potential safety hazard. In addition, in terms of the decision-making of autonomous driving, the previous methods focused on point-to-point driving decisions and ignored the global path planning. This will not only affect the driving efficiency, but also cannot cope with various emergencies.

[0033] To solve the above problems, embodiments of the present invention provide correction mechanisms for images in different driving environments respectively: images in low-light environments are sent to an image enhancement module based on natural language supervision, and images in rainy environments are sent to an image enhancement module based on Fourier learning strategies, ensuring the robustness of the corrected images. In addition, embodiments of the present invention also propose a traffic topology scene graph, which provides a global topological understanding of the entire traffic scene, including the connection relationships between lanes and the impact of lane signals on lanes. Such a global view is crucial for global path planning.

[0034] Specifically, embodiments of the present invention provide a method for generating a topology scene graph, which can also be understood as a method for robust traffic scene graph understanding based on natural language supervision and Fourier learning strategies. It not only enhances the robustness of images collected in different driving environments but also proposes a traffic topology scene graph, providing new ideas for global path planning.

[0035] Embodiments of the present invention provide an electronic device, which can be a vehicle computer device, a service device, a mobile phone device, etc. Please refer to Figure 1 , the structural schematic diagram of the electronic device. The electronic device includes a processor 10, a memory 11, and a bus 12. The processor 10 and the memory 11 are connected through the bus 12, and the processor 10 is used to execute an executable module stored in the memory 11, such as a computer program.

[0036] The processor 10 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the topology scene graph generation method can be completed by the integrated logic circuit in the hardware of the processor 10 or instructions in software form. The above-mentioned processor 10 can be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it can also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0037] The memory 11 may include a high-speed random access memory (RAM: Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory.

[0038] The bus 12 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Figure 1 Although only one bidirectional arrow is used in the figure, it does not mean that there is only one bus 12 or only one type of bus 12 .

[0039] The memory 11 is used to store programs, such as programs corresponding to the topological scene graph generation device. The topological scene graph generation device includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or solidified in the operating system (OS) of the electronic device. After receiving the execution instruction, the processor 10 executes the program to implement the topological scene graph generation method.

[0040] Possibly, the electronic device provided by the embodiment of the present invention further includes a communication interface 13. The communication interface 13 is connected to the processor 10 via a bus.

[0041] It should be understood that Figure 1 The structure shown is only a schematic diagram of a portion of the electronic device. The electronic device may also include Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown. Figure 1 Each component shown in the figure can be implemented by hardware, software or a combination thereof.

[0042] A topological scene graph generation method provided by an embodiment of the present invention can be applied to, but not limited to, Figure 1 For detailed procedures, please refer to the electronic equipment shown in Figure 2 , a topological scene graph generation method includes: S10, S20, S30 and S40, which are specifically described as follows.

[0043] S10, classifying the images around the car at the target moment to determine low-light images and rainy images.

[0044] The low-light image is an image whose brightness is lower than a brightness threshold and does not include rain features, and the rain image is an image that includes rain features.

[0045] In an optional implementation, cameras are arranged at the front, rear, left front, left rear, right front and right rear of the car to collect images around the car in real time.

[0046] Optionally, the images around the vehicle at the target moment are input into a classifier for classification. The classifier consists of a convolutional neural network. After multiple layers of convolution and pooling, the features of the image will be gradually abstracted. Finally, through the fully connected layer, these abstracted features are mapped to three category labels: "normal image", "low-light image", and "rainy image" for subsequent targeted operations.

[0047] S20, perform low-light enhancement on the low-light image to obtain a low-light corrected image.

[0048] In an alternative implementation, the low-light image collected in a low-light environment is fed into a low-light enhancement module based on natural language supervision for low-light enhancement to obtain a low-light corrected image.

[0049] Optionally, S20, performing low-light enhancement on the low-light image to obtain a low-light corrected image includes: S21, S22, S23, and S24, which are specifically described as follows.

[0050] S21, input the low-light image into a 3×3 convolutional layer to extract shallow features .

[0051] Among them, , , H represents the length of the image, W represents the width of the image, and 3 and C are the number of channels corresponding to the image or feature.

[0052] S22, use the feature enhancement unit to process the shallow features .

[0053] Among them, the feature enhancement unit includes multiple feature enhancement modules. The input feature of the first feature enhancement module is the shallow feature , and the input feature of the i-th feature enhancement module is the output feature of the (i - 1)-th feature enhancement module, where 2 ≤ i ≤ the total number of feature enhancement modules in the feature enhancement unit.

[0054] In an alternative implementation, the feature enhancement unit includes 3 feature enhancement modules. The output of the first feature enhancement module is , will be input into the second feature enhancement module to obtain , will be input into the third feature enhancement module to obtain .

[0055] Optionally, the feature enhancement module includes an information fusion attention part and a natural language supervision part. The information fusion attention part includes a channel attention block, a pixel attention block, and a cross-layer attention fusion block. On this basis, the embodiments of the present invention also provide an optional implementation manner on how the feature enhancement module works. Please refer to the following. S22, use the feature enhancement unit to process the shallow features as follows, including S221, S222, S223, S224, S225, and S226, which are specifically described as follows.

[0056] S221, process the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module .

[0057] Optionally, S221, process the input feature of the s-th feature enhancement module through the channel attention block to obtain the output of the channel attention block of the s-th feature enhancement module , including S221-1, S221-2, and S221-3, which are specifically as follows.

[0058] S221-1, perform global average pooling on the input feature of the s-th feature enhancement module to obtain the first global average pooling result .

[0059]

[0060] Wherein, represents the first global average pooling result (including the global features of each channel), H represents the length of the input feature (also the length of the image), W represents the width of the input feature (also the width of the image), represents the input feature the pixel value of the pixel point with coordinates (m, n) in.

[0061] S221-2, according to the first global average pooling result , determine the feature channel weight corresponding to the s-th feature enhancement module.

[0062] Optionally, in order to determine the weights of each channel, the first global average pooling result will pass through two convolutional layers, and then be activated using the Sigmoid and ReLU functions to obtain the feature channel weight .

[0063]

[0064] Among them, represents the feature channel weight corresponding to the s-th feature enhancement module.

[0065] S221-3. According to the input feature and the feature channel weight corresponding to the s-th feature enhancement module to determine the output of the channel attention block of the s-th feature enhancement module .

[0066] Optionally, the output of the channel attention block of the s-th feature enhancement module , and then obtain the output after enhancing important channel features and suppressing unimportant channel features.

[0067] S222. Process the output of the channel attention block of the s-th feature enhancement module through a pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module .

[0068] Optionally, S222. Process the output of the channel attention block of the s-th feature enhancement module through a pixel attention block to obtain the output of the pixel attention block of the s-th feature enhancement module , including S222-1, S222-2, and S222-3, which are specifically described as follows.

[0069] S222-1. According to the output of the channel attention block , obtain the feature pixel weight corresponding to the s-th feature enhancement module .

[0070] Optionally, the pixel attention layer first inputs the output of the channel attention block into two convolutional layers, and then activates them using the Sigmoid and ReLU functions to obtain the feature pixel weight corresponding to the s-th feature enhancement module .

[0071]

[0072] Among them, represents the feature pixel weight corresponding to the s-th feature enhancement module.

[0073] S222-2. According to the feature pixel weight corresponding to the s-th feature enhancement module and the output of the channel attention block to determine the output of the pixel attention block of the s-th feature enhancement module .

[0074] The output after obtaining the feature information of enhancing the low-light region and the high-frequency region can be obtained. , .

[0075] S223. Use the cross-layer attention fusion block to process the input features , the output of the channel attention block and the residual connection result of the input features , as well as the output of the pixel attention block and the residual connection result of the input features to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module .

[0076] Optionally, S223. Use the cross-layer attention fusion block to process the input features , the output of the channel attention block and the residual connection result of the input features , as well as the output of the pixel attention block and the residual connection result of the input features to generate the output of the cross-layer attention fusion block of the s-th feature enhancement module , including: S223-1, S223-2, S223-3, S223-4, and S223-5, which are specifically described as follows.

[0077] S223-1. Stack the input features , the output of the channel attention block and the residual connection result of the input features , as well as the output of the pixel attention block and the residual connection result of the input features front and back to obtain the reshaping result .

[0078] Among them, .

[0079] S223-2. Use a 1×1 convolutional layer to integrate the reshaping result to obtain the integrated information.

[0080] S223-3. Based on the integrated information, use a 3×3 depth convolution to generate the first query matrix , the first key matrix and the first value matrix .

[0081] S223-4. According to the first query matrix , the first key matrix and the first value matrix , obtain cross-layer attention weighted features.

[0082]

[0083]

[0084] Among them, represents the cross-layer attention weighted feature, is a 3×3 layer-related attention matrix, is the scale factor, and performs two-dimensional reshaping on the first query matrix , the first key matrix and the first value matrix respectively to obtain two-dimensional , and , two-dimensional , , .

[0085] S223-5, according to the cross-layer attention weighted feature and the reshaping result , determine the output of the cross-layer attention fusion block of the s-th feature enhancement module.

[0086]

[0087] Among them, represents the output of the cross-layer attention fusion block of the s-th feature enhancement module.

[0088] S224, divide the output of the cross-layer attention fusion block into n levels to obtain the image feature .

[0089] Among them, represents the i-th level feature, which can be used as the image feature input of the natural language supervision part.

[0090] S225, input the low-light image into the text feature extraction model to obtain the text feature of the low-light image .

[0091] Among them, represents the i-th text feature, which is used to describe the text of the image. The text feature extraction model can be but is not limited to a pre-trained CLIP model.

[0092] S226, use the cross-attention mechanism as the fusion layer for the text feature and the image feature Fuse them to obtain the output features of the feature enhancement module.

[0093] Optionally, in S226, use the cross-attention mechanism as the fusion layer to fuse the text features and the image features to obtain the output features of the feature enhancement module, including: S226-1, S226-2, and S226-3, which are specifically described as follows.

[0094] S226-1, convert the text features into the vector form through the encoder .

[0095] It should be understood that the vector form is more convenient to process.

[0096] S226-2, based on the vector form and the image features , generate the second query matrix , the second key matrix , and the second value matrix .

[0097] Among them, , , , among which, represents the projection matrix corresponding to the query matrix, represents the projection matrix corresponding to the key matrix, represents the projection matrix corresponding to the value matrix, represents the intermediate representation form of the image features .

[0098] S226-3, obtain the attention-weighted features according to the second query matrix , the second key matrix and the second value matrix .

[0099]

[0100] Among them, B is the pre-configured relative position encoding matrix, is the scaling factor.

[0101] It should be noted that the attention-weighted features are the output of the natural language supervision part and also the final output of the feature enhancement module.

[0102] S23, fuse the output features of each feature enhancement module to obtain the first fused feature .

[0103] Taking the number of feature enhancement modules as 3 as an example, finally, the outputs from the three feature enhancement modules 、 、 are input into the feature fusion module. Arrange them in sequence to form .

[0104] S24, process the first fused feature with two layers of 3×3 convolutional layers and input the processing result into a 1×1 convolutional layer to output the low-light corrected image.

[0105] The final output channel number is 3, and the corrected low-light corrected image .

[0106] S30, perform rain removal on the rainy image to obtain the rain-corrected image.

[0107] Optionally, send the rainy image collected in the rainy environment into the light enhancement module based on the Fourier learning strategy to obtain the rain-corrected image.

[0108] Optionally, S30, perform rain removal on the rainy image to obtain the rain-corrected image, including: S31, S32, and S33, which are specifically described as follows.

[0109] S31, input the rainy image into a 3×3 convolutional layer to extract shallow features .

[0110] Among them, .

[0111] S32, input the shallow features into the multi-scale U-Net architecture to obtain deeper features.

[0112] Among them, the multi-scale U-Net architecture includes multiple groups of Fourier residual state space blocks. The input feature of the first group of Fourier residual state space blocks is the shallow feature , and the input feature of the i-th group of Fourier residual state space blocks is the output feature of the (i - 1)-th group of Fourier residual state space blocks, where 2 ≤ i ≤ the total number of Fourier residual state space blocks.

[0113] In an alternative embodiment, each group of Fourier residual state-space blocks includes a Fourier-space interaction state-space model and a Fourier-channel evolution state-space model. Based on this, in relation to the content in S32, the embodiments of the present invention also provide an alternative embodiment. Please refer to the following text. S32, the shallow features are input into a multi-scale U-Net architecture to obtain deeper features, including: S32A to S32O, which are specifically described as follows.

[0114] S32A, in the Fourier-space interaction state-space model, the input features are normalized using a normalization technique (LayerNorm) to generate normalized features .

[0115] By performing the normalization process, the problem of gradient vanishing or explosion can be reduced.

[0116] S32B, in the Fourier branch, the normalized features are subjected to a fast Fourier transform to transform the normalized features from the current spatial domain to the Fourier space to obtain the amplitude spectrum and the phase spectrum P .

[0117] Among them, the amplitude spectrum and the phase spectrum P are the representations of the normalized features in the frequency domain. The amplitude spectrum represents the intensity information of the frequency, and the phase spectrum P represents the phase information of the frequency.

[0118] It should be noted that in order to promote the interaction between spatial information and frequency information, the Fourier branch and the spatial branch are processed collaboratively.

[0119] S32C, the amplitude spectrum and the phase spectrum P are respectively scanned and rearranged using zigzag encoding to obtain a one-dimensional amplitude sequence and a one-dimensional phase sequence .

[0120] In order to orderly correlate the relationships between different frequencies, scanning and rearrangement are required. Among them, the frequencies in the one-dimensional amplitude sequence and the one-dimensional phase sequence are arranged in ascending order from low to high.

[0121] S32D, the one-dimensional amplitude sequence and the one-dimensional phase sequence are fed into the frequency Mamba block to obtain the output of the Fourier branch 。

[0122] Optionally, S32D takes the one-dimensional amplitude sequence and the one-dimensional phase sequence as inputs to the frequency Mamba block to obtain the output of the Fourier branch , including: S32D1, S32D2, S32D3, and S32D4, which are described in detail below.

[0123] S32D1 passes the one-dimensional amplitude sequence and the one-dimensional phase sequence sequentially through a depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the first amplitude intermediate product and the first phase intermediate product .

[0124] Thus, the dynamic relationship between different frequencies is fully captured, with low-frequency components enhanced and high-frequency components suppressed.

[0125] S32D2 performs an inverse fast Fourier transform on the first amplitude intermediate product and the first phase intermediate product to obtain the second fused feature .

[0126] By performing an inverse fast Fourier transform, the features are transformed back to the current spatial domain.

[0127] S32D3 inputs the normalized feature into the activation function SiLU to obtain .

[0128] S32D4 performs an element-wise multiplication of the second fused feature with to obtain the output of the Fourier branch .

[0129] By performing an element-wise multiplication of the second fused feature with , the importance of the frequency components is dynamically adjusted, better removing raindrops and restoring background details, and finally obtaining the output of the Fourier branch .

[0130] S32E performs a 1×1 convolution operation on the normalized feature in the spatial branch to extract features between channels.

[0131] S32F inputs the features between channels into the spatial Mamba block to obtain the output of the spatial branch .

[0132] ;

[0133] Among them, represents the output of the spatial branch, and

[0134] S32G, based on the output of the Fourier branch and the output of the spatial branch , obtains the output of the Fourier spatial interaction state space model .

[0135] Optionally, stack the output of the Fourier branch and the output of the spatial branch front and back, perform a 1×1 convolution operation on the stacked result front and back, and the number of output channels is , obtains the output of the Fourier spatial interaction state space model , and at the same time serves as the input of the Fourier channel evolution state space model .

[0136] S32H, performs global average pooling on the input of the Fourier channel evolution state space model to obtain the second global average pooling result .

[0137]

[0138] Among them, represents the second global average pooling result (which effectively encapsulates the global information of the features), represents the length of represents the width of represents the feature at the pixel value of the pixel point with coordinates (m, n) in

[0139] S32I, performs channel Fourier transform on the second global average pooling result to obtain the channel Fourier transform result .

[0140]

[0141] Among them, is the number of channels of z is the index in the frequency domain, j is the imaginary unit, is the normalized value of the frequency, represents

[0142] S32J, obtain the amplitude component and the phase component according to the real part and the imaginary part in the channel Fourier transform result. Among them,

[0143]

[0144]

[0145] where represents the amplitude component, represents the phase component, represents the real part in the channel Fourier transform result , represents the imaginary part in the channel Fourier transform result .

[0146] S32K, sequentially pass the amplitude component and the phase component through the depthwise separable convolutional layer → SiLU activation function → SSM framework → LayerNorm normalization layer to obtain the second amplitude intermediate product and the second phase intermediate product .

[0147] Enhance the low-frequency components and suppress the high-frequency components to obtain a feature sequence 、 with stronger representation ability.

[0148] S32L, use the inverse channel Fourier transform to convert the second amplitude intermediate product and the second phase intermediate product back to the current spatial domain to obtain the third fusion feature.

[0149] S32M, input the input of the Fourier channel evolution state space model into the activation function SiLU to obtain .

[0150] S32N, perform an element-wise multiplication of the third fusion feature and to obtain the output of the Fourier channel evolution state space model.

[0151] By performing an element-wise multiplication of the third fusion feature and , dynamically adjust the importance of the frequency components, better remove raindrops and restore background details, and obtain .

[0152] S32O, output of the Fourier space interaction state space model and output Multiply to obtain the output of the Fourier residual state space block .

[0153]

[0154] Among them, represents the output of the Fourier residual state space block.

[0155] S33, input the output of the last Fourier residual state space block into a 1×1 convolutional layer to output a rain-corrected image with 3 output channels.

[0156] After several similar Fourier residual state space blocks, finally passing through a 1×1 convolutional layer with 3 output channels, the corrected image is finally obtained .

[0157] S40, fuse the low-light corrected image, the rain-corrected image, and the normal image to obtain the topological scene graph at the target time.

[0158] Among them, the normal image is the image in the images around the car except for the low-light image and the rain image.

[0159] Based on Figure 2 , regarding the content in S40, the embodiment of the present invention also provides an optional implementation manner. Please refer to the following. S40, fuse the low-light corrected image, the rain-corrected image, and the normal image to obtain the topological scene graph at the target time, including: S41, S42, S43, and S44, which are specifically described as follows.

[0160] S41, construct a query matrix corresponding to the target image .

[0161] Among them, the target image includes the low-light corrected image, the rain-corrected image, and the normal image.

[0162] Optionally, S41, construct a query matrix corresponding to the target image , including: S411, S412, and S413, which are specifically described as follows.

[0163] S411, extract the two-dimensional features of the target image .

[0164] Optionally, use the ResNet-50 network to extract the two-dimensional features of the target image .

[0165] S412, map the two-dimensional features to the bird's-eye view space to generate a grid-like bird's-eye view feature .

[0166] Optionally, use BEVFormer to map two-dimensional features to the bird's-eye view space (BEV), generating grid-based bird's-eye view features .

[0167] S413, use the bird's-eye view features as the input of the deformable attention mechanism to obtain an output query matrix .

[0168] Among them, , is the pre-trained initialized query matrix.

[0169] S42, input the query matrix into the lane detection head to obtain the predicted nodes corresponding to the topological scene graph.

[0170] Among them, the predicted nodes corresponding to the topological scene graph include an ordered point sequence of the lane centerline and a lane category sequence , represents the predicted nodes corresponding to the topological scene graph.

[0171] , , represents the ordered point sequence of the lane centerline, represents the i-th ordered point of the lane centerline, , represents the i-th lane category.

[0172] S43, combine the query matrix with the ordered point sequence of the lane centerline to determine the connection probability between the starting lane and the ending lane.

[0173] Optionally, S43, combine the query matrix with the ordered point sequence of the lane centerline to determine the connection probability between the starting lane and the ending lane, including: S431, S432, S433, and S434, which are specifically described as follows.

[0174] S431, input the query matrix and the ordered point sequence of the lane centerline into the lane aggregation layer, and use the geometric distance between the lane centerlines to guide the aggregation of global structural information to obtain global aggregation information .

[0175] Optionally, S431, input the query matrix and the ordered point sequence of the lane centerline Input into the lane aggregation layer, and use the geometric distance between lane centerlines to guide the aggregation of global structural information to obtain global aggregation information , including: S431-1, S431-2, and S431-3, which are specifically described as follows.

[0176] S431-1, the lane aggregation layer calculates the spatial adjacency matrix according to the geometric distance between lane centerlines .

[0177] In the geometric-guided self-attention mechanism of the lane aggregation layer, the spatial adjacency matrix is calculated according to the geometric distance between lanes, and the geometric distance between lanes is converted into a weight matrix (spatial adjacency matrix).

[0178] The formula for the spatial adjacency matrix is:

[0179]

[0180] where represents the spatial adjacency matrix, represents the end point of the ordered point sequence of the centerline of the i-th lane (the end point is ), represents the start point of the ordered point sequence of the centerline of the j-th lane (the end point is 0), represents and the geometric distance between is a constant, is the normalization operation.

[0181] S431-2, perform a linear transformation on the query matrix to obtain the input feature 1 of the geometric-guided self-attention unit of the -th layer.

[0182] S431-3, the geometric-guided self-attention unit of the l -th layer uses the spatial adjacency matrix to transform its input feature to obtain its corresponding output feature .

[0183] When l > 1, the input feature l of the geometric-guided self-attention unit of the -th layer is the output feature of the geometric-guided self-attention unit of the l-1 -th layer.

[0184] Among them, the formula for the output feature of the geometric-guided self-attention unit of the l -th layer is:

[0185]

[0186]

[0187] Among them, is the weight of the l -th pre-trained linear transformation layer, which is used to generate the query matrix, key matrix, and value matrix respectively. is the weight of the linear transformation layer in the l -th pre-trained layer for generating the query matrix. is the weight of the linear transformation layer in the l -th pre-trained layer for generating the key matrix. is the weight of the linear transformation layer in the l -th pre-trained layer for generating the value matrix. respectively represent the query matrix, key matrix, and value matrix. is the scaling factor. is the activation function. represents the weight of the geometric-guided self-attention unit in the l -th layer. represents the output feature of the geometric-guided self-attention unit in the l -th layer. represents the input feature of the geometric-guided self-attention unit in the l -th layer. The output feature of the geometric-guided self-attention unit in the L-th layer is the global aggregation information. , where L represents the total number of layers of the geometric-guided self-attention unit.

[0188] S432. The global aggregation information is processed through a feed-forward network, and the processing result of the feed-forward network is connected to the input feature of the geometric-guided self-attention unit in the L-th layer with a residual connection, and the result of the residual connection is processed through a normalization layer to obtain the input M of the edge prediction head.

[0189] Among them, the input of the edge prediction head .

[0190] S433. Based on the input of the edge prediction head, the starting lane feature and the ending lane feature are obtained. Among them, , , represents the multi-layer perceptron of the starting lane feature. represents the multi-layer perceptron of the ending lane feature.

[0191] In the edge prediction head, the work to be done is to predict the connection relationships between the nodes in the topological graph. Since edge prediction involves whether there is a connection relationship between each pair of lanes, it is necessary to separate the lane features to distinguish the starting lane and the ending lane, and two independent multi-layer perceptrons are used to process the starting lane features respectively and the ending lane features .

[0192] S434. According to the starting lane features and the ending lane features , determine the connection probability between the starting lane and the ending lane.

[0193] Optionally, the calculation formula for the connection probability between the starting lane s and the ending lane e is:[[]]

[0194] , where represents the connection probability between the starting lane s and the ending lane e, are respectively and constituent elements of.

[0195] S44. Generate a topological scene graph according to the predicted nodes corresponding to the topological scene graph and the connection probability between the starting lane and the ending lane.

[0196] The embodiment of the present invention provides a topological scene graph generation, which is a robust traffic scene graph understanding method based on natural language supervision and Fourier learning strategy, not only increasing the robustness of the images collected in different driving environments, but also proposing a traffic topological scene graph, providing a new idea for global path planning.

[0197] Please refer to Figure 3 , Figure 3 which is a topological scene graph generation device provided by the embodiment of the present invention. Optionally, this topological scene graph generation device is applied to the electronic device described above.

[0198] The topological scene graph generation device includes: a first processing unit 501 and a second processing unit 502.

[0199] The first processing unit 501 is used to classify the images around the vehicle at the target moment to determine low-light images and rainy images. Among them, a low-light image is an image with a brightness lower than the brightness threshold and does not include rainy features, and a rainy image is an image including rainy features.

[0200] The first processing unit 501 is also used to perform light enhancement processing on the low-light image to obtain a low-light corrected image.

[0201] The first processing unit 501 is also used to perform rain removal processing on the rain image to obtain a rain-corrected image.

[0202] The second processing unit 502 is used to perform fusion processing on the low-light corrected image, the rain-corrected image, and the normal image to obtain a topological scene graph at the target moment, where the normal image is the image in the images around the vehicle excluding the low-light image and the rain image.

[0203] Optionally, the first processing unit 501 may execute the above S10, S20, and S30, and the second processing unit 502 may execute the above S40.

[0204] It should be noted that the topological scene graph generation device provided in this embodiment can execute the method flow shown in the above method flow embodiment to achieve the corresponding technical effects. For a brief description, for the parts not mentioned in this embodiment, reference may be made to the corresponding content in the above embodiment.

[0205] The embodiment of the present invention also provides a storage medium, which stores computer instructions and programs. When the computer instructions and programs are read and run, they execute the topological scene graph generation method of the above embodiment. The storage medium may include memory, flash memory, registers, or a combination thereof, etc.

[0206] The following provides an electronic device, which may be an on-board computer device, a service device, a mobile phone device, etc. As shown in Figure 1 it can implement the above topological scene graph generation method; specifically, the electronic device includes: a processor 10, a memory 11, and a bus 12. The processor 10 may be a CPU. The memory 11 is used to store one or more programs. When the one or more programs are executed by the processor 10, the topological scene graph generation method of the above embodiment is executed.

[0207] In summary, the topological scene graph generation method, device, storage medium, and electronic device provided by the embodiment of the present invention classify the images around the vehicle at the target moment to determine the low-light image and the rain image, where the low-light image is an image with a brightness lower than the brightness threshold and does not include rain features, and the rain image is an image including rain features; perform light enhancement processing on the low-light image to obtain a low-light corrected image; perform rain removal processing on the rain image to obtain a rain-corrected image; perform fusion processing on the low-light corrected image, the rain-corrected image, and the normal image to obtain a topological scene graph at the target moment, where the normal image is the image in the images around the vehicle excluding the low-light image and the rain image. It not only increases the robustness of the images collected in different driving environments, but also proposes a traffic topological scene graph, providing a new idea for global path planning.

[0208] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0209] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

Claims

1. A method for generating a topological scene graph, characterized in that, The method includes: Classify the images around the vehicle at the target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than the brightness threshold and without rainy features, and the rainy images are images with rainy features; Perform light enhancement processing on the low-light images to obtain low-light corrected images; Perform rain removal processing on the rainy images to obtain rain-corrected images; Perform fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment, where the normal images are images in the images around the vehicle other than the low-light images and the rainy images; The performing fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment includes: Construct a query matrix corresponding to the target image , where the target image includes the low-light corrected image, the rain-corrected image, and the normal image; Input the query matrix into the lane detection head to obtain the predicted nodes corresponding to the topological scene graph; wherein, the predicted nodes corresponding to the topological scene graph include an ordered point sequence of the lane centerline and a lane category sequence , denotes the predicted nodes corresponding to the topological scene graph; Combined query matrix With an ordered point sequence of the lane centerline , determine the connection probability between the starting lane and the ending lane; Based on the predicted nodes corresponding to the topological scenario graph and the connection probability between the starting lane and the ending lane, generate a topological scenario graph.

2. The topological scenario graph generation method according to claim 1, wherein Constructing a query matrix corresponding to the target image , including: Extract the two-dimensional features of the target image , Map two-dimensional features to the bird's-eye view space to generate grid bird's-eye view features ; Take the bird's-eye view feature as the input of the deformable attention mechanism to obtain an output query matrix , where , is the pre-trained initialized query matrix.

3. The topological scene graph generation method according to claim 1, characterized in that, The combined query matrix and the ordered point sequence of the lane centerline , to determine the connection probability between the starting lane and the ending lane, including: Input the query matrix and the ordered point sequence of the lane centerline into the lane aggregation layer, and use the geometric distance between lane centerlines to guide the aggregation of global structural information to obtain global aggregation information ; Process the globally aggregated information through a feed-forward network and perform a residual connection between the processing result of the feed-forward network and the input features of the geometric-guided self-attention unit at the L-th layer and process the result of the residual connection through a normalization layer to obtain the input M of the edge prediction head; Among them, the input of the edge prediction head ; Based on the input of the side prediction head obtain the starting lane feature and the ending lane feature , where and . denotes the multi-layer perceptron for the starting lane feature denotes the multi-layer perceptron for the ending lane feature; Based on the starting lane characteristics and the ending lane characteristics , determine the connection probability between the starting lane and the ending lane.

4. The topological scenario graph generation method according to claim 3, wherein, The formula for the connection probability between the starting lane s and the ending lane e is: , where represents the connection probability between the starting lane s and the ending lane e, are respectively and constituent elements of.

5. The topological scene graph generation method according to claim 3, characterized in that, The query matrix and the ordered point sequence of the lane centerline are input into the lane aggregation layer, and the geometric distance between the lane centerlines is used to guide the aggregation of global structural information to obtain global aggregation information , including: The lane aggregation layer calculates a spatial adjacency matrix based on the geometric distance between lane centerlines ; Perform a linear transformation on the query matrix to obtain the input features of the geometrically-guided self-attention unit of the 1 ith layer ; The l layer geometric-guided self-attention unit utilizes the spatial adjacency matrix to transform its input features and obtain its corresponding output features ; when l > 1, the input features l of the layer geometric-guided self-attention unit l-1 are the output features of the l-1 layer geometric-guided self-attention unit; Among them, the l formula for the output features of the layer geometric guided self-attention unit is: Among them, is the weight of the l -th layer of pre-trained linear transformation, which is used to generate the query matrix, key matrix and value matrix respectively. represent the query matrix, key matrix and value matrix respectively. is the scaling factor. is the activation function. represents the weight of the geometric-guided self-attention unit of the l -th layer. represents the output feature of the geometric-guided self-attention unit of the l -th layer. represents the input feature of the geometric-guided self-attention unit of the l -th layer. The output feature of the L-th layer of geometric-guided self-attention unit is the global aggregation information. , where L represents the total number of layers of the geometric-guided self-attention unit.

6. The topological scenario graph generation method according to claim 5, wherein The formula for the spatial adjacency matrix is: Among them, represents the spatial adjacency matrix, represents the end point of the ordered point sequence of the center line of the i-th lane, represents the starting point of the ordered point sequence of the center line of the j-th lane, represents and the geometric distance between, is a constant, is the normalization operation.

7. A topological scene graph generation device, characterized in that, The device includes: A first processing unit for classifying the images around the vehicle at the target moment to determine low-light images and rainy images, where the low-light images are images with a brightness lower than the brightness threshold and without rainy features, and the rainy images are images with rainy features; The first processing unit is further configured to perform light enhancement processing on the low-light images to obtain low-light corrected images; The first processing unit is further configured to perform rain removal processing on the rainy images to obtain rain-corrected images; A second processing unit for performing fusion processing on the low-light corrected images, the rain-corrected images, and normal images to obtain a topological scene graph at the target moment, where the normal images are images in the images around the vehicle other than the low-light images and the rainy images; Performing fusion processing on the low-light corrected image, the rain corrected image, and the normal image to obtain a topological scene graph at the target moment, including: constructing a query matrix corresponding to the target image , where the target image includes the low-light corrected image, the rain corrected image, and the normal image; inputting the query matrix into a lane detection head to obtain predicted nodes corresponding to the topological scene graph; where the predicted nodes corresponding to the topological scene graph include an ordered point sequence of the lane centerline and a lane category sequence , represents the predicted nodes corresponding to the topological scene graph; combining the query matrix with the ordered point sequence of the lane centerline to determine the connection probability between the starting lane and the ending lane; generating a topological scene graph according to the predicted nodes corresponding to the topological scene graph and the connection probability between the starting lane and the ending lane.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method according to any one of claims 1-6.

9. An electronic device, characterized in that, It includes: A processor and a memory, where the memory is used to store one or more programs; When the one or more programs are executed by the processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Post-processing method for multi-stage iteration cooperative representation of rain-removed image

    CN117218008A

  • Lane line detection method and system combining standard definition map and satellite map

    CN119785326A