Roadside parking scene analysis method and system based on multi-mode perception fusion

By generating environmental condition marking information and performing feature coding methods, and extracting feature information of different mode data in combination with the backbone network model, flexible fusion of roadside parking scene data is achieved, solving the problem of inflexible data fusion in traditional methods and improving the accuracy of analysis.

CN119964128APending Publication Date: 2025-05-09AIPARK TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510421144.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The existing fusion methods of multiple sensor data are processed uniformly under all conditions, which limits the flexibility of sensor data of different modalities, resulting in poor practical application performance and difficulty in large-scale application in roadside scenarios.

Method used

By using the RGB image data of the visible light camera to generate environmental condition marking information, feature encoding and backbone network model extract feature information of different modal data, and flexible fusion of modal data is carried out according to environmental conditions.

Benefits of technology

It realizes flexible application of data in different modes, improves the accuracy of road-side parking scenario analysis, and overcomes the problem of inflexible data fusion in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964128A_ABST
    Figure CN119964128A_ABST
Patent Text Reader

Abstract

The invention discloses a roadside parking scene analysis method and system based on multi-modal perception fusion, and relates to the field of intelligent parking management.The method comprises the steps that RGB image data of a visible light camera is utilized to classify external environment conditions, a condition mark is marked to guide fusion of modal data of various sensors, and a multi-modal sensing fusion model is obtained; when model prediction is carried out, sensor data adaptive to the current environment condition is selected according to the environment condition of RGB image prediction, fusion is carried out, and flexible application of different modal data is realized; and meanwhile, visible light modal data, event camera data, laser radar modal data and millimeter wave radar data are utilized to realize data fusion of different combination modes on the modal data in different scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent parking management, and in particular to a roadside parking scene analysis method and system based on multimodal perception fusion. Background Art

[0002] In recent years, with the development of intelligent technology, various intelligent parking management technologies are being tried and constructed. Among them, setting up parking spaces on the roadside, using intelligent sensing equipment to collect video images and other data, and processing and analyzing the data through computer vision algorithms to achieve intelligent parking management is currently a mainstream solution. At present, it is mainly achieved by installing high-position video cameras on the roadside to collect data, and then using visual algorithms to perform license plate detection, license plate recognition, target detection, semantic segmentation and other algorithms to achieve intelligent parking charging and management.

[0003] However, compared with other spectra, the visible light imaging range is significantly narrower, and it is only effective in environments with good lighting and high visibility. It may fail at night, in special weather such as rain and fog, or in environments with obstructions. As a result, the understanding and analysis of traffic scenes in special environments cannot be achieved. Therefore, many schemes currently propose to use multi-modal data from multiple sensor devices including visible light cameras, event cameras, lidar, millimeter-wave radar, etc. to analyze roadside parking scenes. Data of different modalities can play different advantages in different environments and weather conditions. The fusion of multiple modal data can provide more data information support for the analysis of complex scenes, which plays an important role in improving the accuracy of traffic scene analysis. However, current methods for fusing data from multiple sensors usually uniformly process data collected by all sensors under all conditions, which limits the flexibility of using data from sensors of different modalities as input, resulting in poor performance in actual applications of data collection from multiple sensors. In actual roadside scenarios, it is difficult to implement large-scale application of sensors of different modalities for roadside parking management. Summary of the invention

[0004] In order to solve the above technical problems, the present invention provides a roadside parking scene analysis method and system based on multimodal perception fusion, which can solve the problem that the actual application performance of the existing multiple sensor data collection is poor, and it is difficult to implement roadside parking management using large-scale applications of different modal sensors in actual roadside scenarios.

[0005] To achieve the above object, the present invention provides a roadside parking scene analysis method based on multimodal perception fusion, the method comprising: Generate environmental condition marking information corresponding to the current roadside parking scene based on the collected RGB image; Performing feature encoding on the environmental condition mark information; Extract feature information of different modal data from roadside parking scene data through a preset backbone network model; fusing the feature information of the different modal data according to the environmental condition label information after the feature decoding; The roadside parking scene is analyzed according to the fused feature information.

[0006] Furthermore, the step of generating environmental condition marking information corresponding to the current roadside parking scene according to the collected RGB image includes: Analyze the time and weather information of the current roadside parking scene based on the collected RGB images; Environmental condition marking information corresponding to the current roadside parking scene is generated according to the time information and weather information of the current roadside parking scene.

[0007] Furthermore, the step of characteristic encoding the environmental condition mark information includes: The environmental condition mark information text is converted into digital representation information through a preset word segmentation module; The converted digital representation information is feature extracted and encoded through the preset word segmentation module.

[0008] Furthermore, the step of extracting feature information of different modal data from the roadside parking scene data by using a preset backbone network model includes: Extracting shared feature information from roadside parking scene data through a shared backbone network model in a preset backbone network model; The feature information of different modal data is extracted from the roadside parking scene data through the lightweight feature adapter model corresponding to each modality in the preset backbone network model.

[0009] Furthermore, the step of fusing the feature information of the different modal data according to the environmental condition tag information after the feature decoding includes: According to the formula The fusion is performed, where the fused features are expressed as , They respectively represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera, and are the weight parameters corresponding to different environmental marking information.

[0010] Furthermore, the present invention provides a roadside parking scene analysis system based on multimodal perception fusion, the system comprising: A generation module, used to generate environmental condition marking information corresponding to the current roadside parking scene based on the collected RGB image; An encoding module, used for performing feature encoding on the environmental condition mark information; An extraction module, used to extract feature information of different modal data from roadside parking scene data through a preset backbone network model; A fusion module, used for fusing the feature information of the different modal data according to the environmental condition label information after the feature decoding; The parsing module parses the roadside parking scene according to the fused feature information.

[0011] Furthermore, the generation module is specifically used to parse the time information and weather information of the current roadside parking scene according to the collected RGB image; and generate environmental condition marking information corresponding to the current roadside parking scene according to the time information and weather information of the current roadside parking scene.

[0012] Furthermore, the encoding module is specifically used to convert the environmental condition mark information text into digital representation information through a preset word segmenter module; and perform feature extraction and encoding on the converted digital representation information through the preset word segmenter module.

[0013] Furthermore, the extraction module is specifically used to extract shared feature information from roadside parking scene data through a shared backbone network model in a preset backbone network model; and to extract feature information of different modal data from roadside parking scene data through a lightweight feature adapter model corresponding to each modality in the preset backbone network model.

[0014] Furthermore, the fusion module is specifically used according to the formula The fusion is performed, where the fused features are expressed as , They respectively represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera, and are the weight parameters corresponding to different environmental marking information.

[0015] The present invention provides a roadside parking scene analysis method and system based on multimodal perception fusion, which utilizes the RGB image data of a visible light camera to classify external environmental conditions, and marks a condition tag to guide the fusion of multiple sensor modal data. When performing model prediction, the sensor data adapted to the current environmental conditions is selected and fused according to the environmental conditions predicted by the RGB image, thereby realizing flexible application of different modal data. At the same time, visible light modal data, event camera data, lidar modal data, and millimeter wave radar data are utilized to realize data fusion of different combinations of the above modal data in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1It is a flow chart of a roadside parking scene analysis method based on multimodal perception fusion provided by the present invention; Figure 2 It is a schematic diagram of a roadside parking scene analysis system based on multimodal perception fusion provided by the present invention. DETAILED DESCRIPTION

[0017] The device structure and implementation of the present invention are further described in detail below through the accompanying drawings and embodiments.

[0018] The present invention provides a roadside parking scene analysis method based on multimodal perception fusion, such as Figure 1 As shown, the specific steps include: 101. Generate environmental condition marking information corresponding to the current roadside parking scene based on the collected RGB image.

[0019] Specifically, the time information and weather information of the current roadside parking scene are analyzed according to the collected RGB image; and the environmental condition marking information corresponding to the current roadside parking scene is generated according to the time information and weather information of the current roadside parking scene.

[0020] For example, the time of a day is divided into day and night; the weather conditions are divided into sunny, rainy, foggy, snowy, etc.; by combining two aspects of the above environmental conditions, the environmental condition annotations are divided into the following categories: sunny during the day, rainy during the day, foggy during the day, snowy during the day, sunny at night, rainy at night, foggy at night, snowy at night; text annotation is performed on each data according to the above eight environmental condition categories.

[0021] For the embodiment of the present invention, before step 101, it may also include: installing multimodal perception data acquisition equipment at different traffic locations, including visible light cameras, radars, laser radars, event cameras, etc., to obtain multimodal perception data; the different traffic locations include major urban roads, intersections, roadside parking spaces, key areas such as parking spaces at the entrances of schools and hospitals, etc., covering dynamic change data of vehicles, pedestrians and other targets in different time periods in the above-mentioned different scenarios; since different sensors may have different perception ranges, resolutions and sensing capabilities, etc., therefore, in order to subsequently analyze the multimodal data To perform fusion, it is necessary to first align the acquired multimodal perception data including RGB images, radar data, lidar data, and event camera data; the data alignment mainly includes two aspects, namely the time dimension and the space dimension: the data alignment in the space dimension is to unify the images of data acquisition of different spectral perceptions into the same coordinate system; the data alignment in the time dimension is that the data acquisition frequencies of sensors of different spectra may be different, so it is necessary to align the data in the time dimension of perception devices of different spectra to ensure that the acquisition data of multiple modal devices are the perception data recorded in the same scene at the same time; The multimodal data is semantically segmented and labeled using a polygon tool, and the labeled categories include but are not limited to target labels such as pedestrians, vehicles, non-motor vehicles, lane lines, traffic signs, and green plants. The vehicles specifically include cars, buses, taxis, and other categories.

[0022] 102. Perform feature encoding on the environmental condition mark information.

[0023] Specifically, the environmental condition mark information text is converted into digital representation information through a preset word segmenter module; and the converted digital representation information is feature extracted and encoded through the preset word segmenter module.

[0024] For example, the prompt text encoding module includes two parts: a word segmenter module and an encoder module. Specifically, the word segmenter module is used to convert the annotated text into a digital representation for the subsequent model to understand the text, and a commonly used word segmenter module of a large language model is adopted; the encoder module is used to extract and encode the converted digital representation, and a multi-layer transformer network structure is adopted to achieve this; specifically, the encoded text is represented as.

[0025] 103. The feature information of different modal data is extracted from the roadside parking scene data through the preset backbone network model.

[0026] Specifically, shared feature information is extracted from the roadside parking scene data through a shared backbone network model in a preset backbone network model; feature information of different modal data is extracted from the roadside parking scene data through a lightweight feature adapter model corresponding to each modality in the preset backbone network model.

[0027] It should be noted that most multimodal fusion methods require each sensor modality to have a separate backbone network for feature encoding, which leads to high computational complexity, cumbersome training process, and difficulty in application in actual real-time traffic scenarios. Therefore, the present invention proposes to use a unified backbone network to extract features for multiple modal data, and use feature adapters to retain the unique features of each modal data; The backbone network includes two parts, namely, a backbone network shared by all modal data and a separate lightweight feature adapter module for each modal data; for the shared backbone network, it includes but is not limited to using a network based on a convolutional neural network or a transformer structure to perform shared feature extraction; for each modality's separate lightweight feature adapter module, a multilayer perceptron network based on a residual network structure is used to implement it, which can fully extract and effectively prevent the overfitting of the model; specifically, the output of each modality feature adapter module is defined as , which represent the output of four different modal data: visible light camera, lidar, radar, and event camera.

[0028] For the embodiments of the present invention, by utilizing such a backbone network, in addition to improving the model operation efficiency, there are two additional advantages: 1) using a shared backbone network to change and promote non-RGB modality data, such as lidar, radar and event camera modalities, is to map the data into a feature space compatible with the RGB modality; 2) the feature adapter of each modality allows the extraction of mode-specific information, providing additional feature information for the RGB modality data; making different modality data complement each other, thereby promoting the subsequent fusion of different modality data in different scenarios.

[0029] 104. Fuse the feature information of the different modal data according to the environmental condition label information after the feature decoding.

[0030] Specifically, according to the formula The fusion is performed, where the fused features are expressed as , They respectively represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera, and are the weight parameters corresponding to different environmental marking information.

[0031] It should be noted that for the four weight parameters According to the representation, different calculation weight values ​​are used according to different environmental condition marking information, and the sum of the four weight parameters is 1; according to the definition of the above steps, the environmental condition marking is divided into the following categories: sunny day, rainy day, foggy day, snowy day, sunny night, rainy night, foggy night, snowy night; for example, when the environmental condition is marked as sunny day, it is used ; When the environmental condition is marked as rainy at night, use ; The definition of different weight parameters can be adjusted according to the actual application effect to achieve the best fusion process of different modal data.

[0032] 105. Analyze the roadside parking scene according to the fused feature information.

[0033] Specifically, the decoder input is the fused features obtained in the above steps. The decoder adopts a lightweight multi-layer perceptron network structure to decode the fused features of different modalities and generate the final semantic segmentation mask.

[0034] Furthermore, a loss function is constructed to perform model training; specifically, the loss function includes two aspects. One is the loss function for labeling learning under different environmental conditions, which uses a contrast loss function and is expressed as:

[0035] Where k represents the number of categories, Represents the model output result, + represents a positive sample, Represents the temperature coefficient, which is a hyperparameter used to control the distribution shape of the model output; for the multimodal semantic segmentation model loss function , using the commonly used segmentation loss function and the cross entropy loss function for calculation; , where C represents the number of target segmentation categories, It indicates the probability that the predicted sample belongs to class c. Represents the true label category distribution; An embodiment of the present invention provides a roadside parking scene analysis method based on multimodal perception fusion, which uses RGB image data of a visible light camera to classify external environmental conditions, and marks a condition tag to guide the fusion of multiple sensor modal data. When performing model prediction, sensor data that is adapted to the current environmental conditions is selected and fused according to the environmental conditions predicted by the RGB image, thereby realizing flexible application of different modal data. At the same time, visible light modal data, event camera data, lidar modal data, and millimeter wave radar data are used to realize data fusion of different combinations of the above modal data in different scenarios.

[0036] As Figure 1The specific implementation of the method shown in the figure, the embodiment of the present invention provides a roadside parking scene analysis system based on multi-modal perception fusion, such as Figure 2 As shown, the system includes: a generating module 21, which is used to generate environmental condition marking information corresponding to the current roadside parking scene according to the collected RGB image; An encoding module 22, used for performing feature encoding on the environmental condition mark information; An extraction module 23, used to extract feature information of different modal data from roadside parking scene data through a preset backbone network model; A fusion module 24, configured to fuse the feature information of the different modal data according to the environmental condition label information after the feature decoding; The parsing module 25 is configured to parse the roadside parking scene according to the fused feature information.

[0037] Furthermore, the generation module 21 is specifically used to analyze the time information and weather information of the current roadside parking scene according to the collected RGB image; and generate environmental condition marking information corresponding to the current roadside parking scene according to the time information and weather information of the current roadside parking scene.

[0038] Furthermore, the encoding module 22 is specifically used to convert the environmental condition mark information text into digital representation information through a preset word segmenter module; and perform feature extraction and encoding on the converted digital representation information through the preset word segmenter module.

[0039] Furthermore, the extraction module 23 is specifically used to extract shared feature information from roadside parking scene data through a shared backbone network model in a preset backbone network model; and to extract feature information of different modal data from roadside parking scene data through a lightweight feature adapter model corresponding to each modality in the preset backbone network model.

[0040] Furthermore, the fusion module 24 is specifically configured to: The fusion is performed, where the fused features are expressed as , They respectively represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera, and are the weight parameters corresponding to different environmental marking information.

[0041] An embodiment of the present invention provides a roadside parking scene analysis system based on multimodal perception fusion, which uses RGB image data of a visible light camera to classify external environmental conditions, and marks a condition tag to guide the fusion of multiple sensor modal data. When performing model prediction, the sensor data that is adapted to the current environmental conditions is selected and fused according to the environmental conditions predicted by the RGB image, thereby realizing flexible application of different modal data; at the same time, visible light modal data, event camera data, lidar modal data, and millimeter wave radar data are used to realize data fusion of different combinations of the above modal data in different scenarios.

[0042] It should be understood that the specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, it should be understood that the specific order or hierarchy of steps in the process can be rearranged without departing from the scope of protection of the present disclosure. The attached method claims present the elements of the various steps in an exemplary order and are not intended to be limited to the specific order or hierarchy described.

[0043] In the above detailed description, various features are grouped together in a single embodiment to simplify the disclosure. This method of disclosure should not be interpreted as reflecting an intention that the embodiments of the claimed subject matter require more features than are clearly stated in each claim. On the contrary, as reflected in the appended claims, the invention is in a state of having less than all the features of the disclosed individual embodiments. Therefore, the appended claims are hereby expressly incorporated into the detailed description, with each claim standing on its own as a separate preferred embodiment of the invention.

[0044] The disclosed embodiments are described above to enable any person skilled in the art to implement or use the present invention. Various modifications of these embodiments are obvious to those skilled in the art, and the general principles defined herein may also be applied to other embodiments without departing from the spirit and scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in this application.

[0045] The above description includes examples of one or more embodiments. Of course, it is impossible to describe all possible combinations of components or methods for the purpose of describing the above embodiments, but it should be recognized by those skilled in the art that the various embodiments may be further combined and arranged. Therefore, the embodiments described herein are intended to cover all such changes, modifications and variations that fall within the scope of protection of the appended claims. In addition, with respect to the term "comprising" used in the specification or claims, the word is covered in a manner similar to the term "including", just as "including," is explained as a transitional word in the claims. In addition, any term "or" used in the specification of the claims is intended to mean "non-exclusive or".

[0046] Those skilled in the art may also understand that the various illustrative logical blocks, units, and steps listed in the embodiments of the present invention may be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly demonstrate the interchangeability of hardware and software, the various illustrative components, units, and steps mentioned above have generally described their functions. Whether such functions are implemented by hardware or software depends on the specific application and the design requirements of the entire system. Those skilled in the art may use various methods to implement the described functions for each specific application, but such implementation should not be understood as exceeding the scope of protection of the embodiments of the present invention.

[0047] The various illustrative logic blocks or units described in the embodiments of the present invention can be implemented or operated by a general-purpose processor, a digital signal processor, an application-specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component, or any combination of the above. The general-purpose processor can be a microprocessor, and optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, a plurality of microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.

[0048] The steps of the method or algorithm described in the embodiments of the present invention can be directly embedded in hardware, a software module executed by a processor, or a combination of the two. The software module can be stored in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or other storage media of any form in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and can write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be arranged in an ASIC, and the ASIC can be arranged in a user terminal. Optionally, the processor and the storage medium can also be arranged in different components in the user terminal.

[0049] In one or more exemplary designs, the above functions described in the embodiments of the present invention can be implemented in hardware, software, firmware, or any combination of the three. If implemented in software, these functions can be stored on a computer-readable medium, or transmitted in the form of one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media that facilitate the transfer of computer programs from one place to another. The storage medium can be any available medium that can be accessed by any general or special computer. For example, such computer-readable media can include but are not limited to RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage devices, or any other medium that can be used to carry or store program codes in the form of instructions or data structures and other forms that can be read by general or special computers, or general or special processors. In addition, any connection can be appropriately defined as a computer-readable medium, for example, if the software is transmitted from a website site, server or other remote resource through a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL) or wirelessly, such as infrared, wireless and microwave, it is also included in the defined computer-readable medium. The disk and disc include compact disk, laser disk, optical disk, DVD, floppy disk and Blu-ray disk. Disks usually copy data magnetically, while discs usually copy data optically with lasers. Combinations of the above may also be included in computer-readable media.

[0050] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A roadside parking scene analysis method based on multimodal perception fusion, characterized in that: The method comprises: Generate environmental condition marking information corresponding to the current roadside parking scene based on the collected RGB image; Performing feature encoding on the environmental condition mark information; Extract feature information of different modal data from roadside parking scene data through a preset backbone network model; fusing the feature information of the different modal data according to the environmental condition label information after the feature decoding; The roadside parking scene is analyzed according to the fused feature information.

2. According to claim 1, a roadside parking scene analysis method based on multimodal perception fusion is characterized in that: The step of generating environmental condition marking information corresponding to the current roadside parking scene according to the collected RGB image includes: Analyze the time and weather information of the current roadside parking scene based on the collected RGB images; Environmental condition marking information corresponding to the current roadside parking scene is generated according to the time information and weather information of the current roadside parking scene.

3. The method for analyzing roadside parking scenes based on multimodal perception fusion according to claim 1 is characterized in that: The step of characteristic encoding the environmental condition mark information comprises: The environmental condition mark information text is converted into digital representation information through a preset word segmentation module; The converted digital representation information is feature extracted and encoded through the preset word segmentation module.

4. The method for analyzing roadside parking scenes based on multimodal perception fusion according to claim 1 is characterized in that: The step of extracting feature information of different modal data from roadside parking scene data by using a preset backbone network model includes: Extracting shared feature information from roadside parking scene data through a shared backbone network model in a preset backbone network model; The feature information of different modal data is extracted from the roadside parking scene data through the lightweight feature adapter model corresponding to each modality in the preset backbone network model.

5. A roadside parking scene analysis method based on multimodal perception fusion according to claim 1 or 4, characterized in that: The step of fusing the feature information of the different modal data according to the feature-decoded environmental condition tag information comprises: According to the formula The fusion is performed, where the fused features are expressed as , They represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera. is the weight parameter corresponding to the feature information of the visible light camera modal data, is the weight parameter corresponding to the characteristic information of the lidar modal data, is the weight parameter corresponding to the characteristic information of radar modal data, is a weight parameter corresponding to the feature information of the event camera modal data, and the weight parameters corresponding to the feature information of different modal data are configured according to the environmental condition labeling information.

6. A roadside parking scene analysis system based on multimodal perception fusion, characterized in that: The system comprises: A generation module, used to generate environmental condition marking information corresponding to the current roadside parking scene based on the collected RGB image; An encoding module, used for performing feature encoding on the environmental condition mark information; An extraction module, used to extract feature information of different modal data from roadside parking scene data through a preset backbone network model; A fusion module, used for fusing the feature information of the different modal data according to the environmental condition label information after the feature decoding; The parsing module parses the roadside parking scene according to the fused feature information.

7. The roadside parking scene analysis system based on multimodal perception fusion according to claim 6 is characterized in that: The generation module is specifically used to analyze the time information and weather information of the current roadside parking scene according to the collected RGB image; and generate environmental condition marking information corresponding to the current roadside parking scene according to the time information and weather information of the current roadside parking scene.

8. The roadside parking scene analysis system based on multimodal perception fusion according to claim 6 is characterized in that: The encoding module is specifically used to convert the environmental condition mark information text into digital representation information through a preset word segmenter module; and perform feature extraction and encoding on the converted digital representation information through the preset word segmenter module.

9. The roadside parking scene analysis system based on multimodal perception fusion according to claim 6 is characterized in that: The extraction module is specifically used to extract shared feature information from roadside parking scene data through a shared backbone network model in a preset backbone network model; and to extract feature information of different modal data from roadside parking scene data through a lightweight feature adapter model corresponding to each modality in the preset backbone network model.

10. A roadside parking scene analysis system based on multimodal perception fusion according to claim 6 or 9, characterized in that: The fusion module is specifically used to The fusion is performed, where the fused features are expressed as , They represent the feature information of four different modal data: visible light camera, lidar, radar, and event camera. is the weight parameter corresponding to the feature information of the visible light camera modal data, is the weight parameter corresponding to the characteristic information of the lidar modal data, is the weight parameter corresponding to the characteristic information of radar modal data, is a weight parameter corresponding to the feature information of the event camera modal data, and the weight parameters corresponding to the feature information of different modal data are configured according to the environmental condition labeling information.

Citation Information

Patent Citations

  • Roadside parking management method and system based on three-dimensional vehicle attitude

    CN115909228A

  • Temperature adjusting system and method based on multi-mode sensing fusion

    CN118752968A

  • Automatic driving multi-mode dynamic fusion method and device driven by environmental element information

    CN119150226A

  • Intelligent driving environment sensing method and system based on multi-modal data fusion

    CN119665998A