A roadside parking high-low fusion multi-modal recognition method and system
By acquiring radar and visual perception data in real time for collaborative judgment and dynamically switching sensor dominance, the problem of insufficient resource scheduling in the existing roadside parking management system is solved, realizing efficient utilization of sensor resources and accuracy and reliability of parking management, and reducing operating costs.
Patent Information
- Application Number
- CN202511433297.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-09
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2045-10-09
AI Technical Summary
The specific problems that existing technologies cannot solve efficiently are: in large-scale application scenarios, existing roadside parking management systems suffer from excessive system load and significantly increased operating costs due to insufficient sensor resource scheduling strategies.
By acquiring radar and visual perception data in real time to make collaborative judgments on vehicle perception, generating blind spot trigger signals, extracting vehicle feature vectors from visual perception data, calling the high-level monitoring system for vehicle re-identification, and dynamically switching sensor dominance, the system can achieve efficient utilization of sensor resources.
This approach enables efficient utilization of sensor resources, reduces system computational load and communication bandwidth consumption, improves the accuracy and reliability of parking management, reduces billing disputes and maintenance costs, and enhances the system's management efficiency and maintenance level.
Smart Images

Figure CN120913286B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of roadside parking management, in particular to a roadside parking high-low fusion multi-modal recognition method and system. BACKGROUND
[0002] As an important part of urban static traffic management, the intelligent level of roadside parking management directly affects the efficiency of road resource utilization and user experience. Currently, the usual means of roadside parking management is to use low curb machines for parking space state perception. Although some existing technologies have appeared high-low sensor fusion monitoring schemes to achieve more accurate roadside parking management, there are obvious deficiencies in sensor resource scheduling strategies in existing technologies. Most systems use fixed collaborative mode and cannot dynamically allocate resources according to the actual state of the vehicle and scene requirements. Specifically, when the vehicle enters the detection blind area of the curb machine, the system often needs to continuously call the high-position camera for full tracking. Even if the vehicle has stopped and is in a stable state, the system still maintains comprehensive monitoring of the high-position camera, resulting in continuous consumption of computing resources and communication bandwidth. In addition, the existing scheme lacks an intelligent sensor switching mechanism. When the curb machine restores the recognition ability, the system cannot automatically return the monitoring dominance to the curb machine, causing long-term occupation of high-position system resources. This static resource allocation method will cause the system to be overloaded and the operating cost to increase significantly in large-scale application scenarios.
[0003] Therefore, there is a need for a new roadside parking management method that can dynamically adjust the sensor collaboration strategy according to the state of the vehicle, while ensuring recognition accuracy and achieving efficient use of system resources. SUMMARY
[0004] The purpose of the present application is to provide a roadside parking high-low fusion multi-modal recognition method and system, which solves the following technical problems:
[0005] Most high-low fusion parking monitoring systems use fixed collaborative mode, which will cause the system to be overloaded and the operating cost to increase significantly in large-scale application scenarios.
[0006] The purpose of the present application can be achieved by the following technical solutions:
[0007] A roadside parking high-low fusion multi-modal recognition method, comprising the following steps:
[0008] S1, real-time acquisition of radar perception data and visual perception data of any preset parking area corresponding curb machine in the target area, and vehicle perception collaborative determination according to the radar perception data and visual perception data, when the determination result exists conflict, then generate blind area trigger signal and record signal generation time t0;
[0009] S2, based on the generation time t0, the continuous video frame sequence in the visual perception data in the preset time window is reversely extracted, vehicle feature extraction processing is performed on the video frame sequence, and a vehicle feature vector is obtained;
[0010] S3, a high-position monitoring system video stream corresponding to the target area is called, vehicle re-identification is performed in the target area based on the vehicle feature vector using a feature matching algorithm, the position of the target vehicle is determined, and positioning and continuous tracking monitoring of the target vehicle are established;
[0011] S4, when the target vehicle is monitored to stop moving, the spatial position information and the vehicle body posture data of the target vehicle are collected and parking compliance is judged, if the judgment result is compliant parking, the curbstone machine is re-enabled for vehicle state verification, and the verification result is fused with the high-position monitoring data to generate a standardized event record containing the compliant parking start time, and the standardized event record is uploaded to the cloud platform to generate a billing record;
[0012] S5, if the judgment result is abnormal parking, an abnormal event identifier is generated and the spatio-temporal information of the target vehicle during the entire parking period is continuously recorded, and the spatio-temporal information is uploaded to the cloud platform to generate a billing record.
[0013] As a further scheme of the application: in S1, the specific process of generating the blind area trigger signal is:
[0014] When the target distance parameter in the radar perception data continuously presents a monotonic decreasing characteristic and the echo signal strength remains stable, and the recognition confidence of the continuous multiple frames of images in the visual perception data is lower than the preset threshold; or when the recognition confidence of the continuous multiple frames of images in the visual perception data continuously exceeds the preset threshold, and no effective target signal is detected in the radar perception data, it is determined that there is a recognition conflict, and after the spatio-temporal alignment of the radar perception data and the visual perception data is completed, if the duration of the recognition conflict exceeds the preset time window threshold, the blind area trigger signal is generated.
[0015] As a further scheme of the application: in S2, the specific process of obtaining the vehicle feature vector is:
[0016] Taking the generation time t0 as the end point, the continuous N frames of image sequence in the visual perception data within the preset time length are reversely extracted, N is a preset quantity threshold, after pretreatment of each frame of image such as illumination compensation and noise filtering, the pre-trained deep convolutional neural network is used to extract color features, texture features and structure features in the pretreated image at the same time, and each frame feature weight is calculated through a time sequence attention module to fuse the multi-frame feature information; the principal component analysis method is used for dimension reduction processing on the fused features; the dimension-reduced feature vector is normalized to obtain the vehicle feature vector.
[0017] As a further scheme of the present application: in the S3, the specific process of vehicle re-identification is:
[0018] Extract all vehicle region proposals in the video frame by a target detection algorithm and generate a candidate feature vector set; calculate the similarity matrix between the vehicle feature vector and any candidate feature vector set, and perform multi-feature dimension matching by using a weighted cosine similarity measure method;
[0019] By screening the maximum value of the similarity matrix, the highest similarity value of each candidate feature vector and the vehicle feature vector is obtained, and when the highest similarity value is greater than or equal to a preset threshold, it is confirmed that the re-identification is successful and the position coordinates of the corresponding vehicle region in the image coordinate system are obtained.
[0020] According to the camera calibration parameters, the image coordinates are converted to the world coordinate system to obtain the actual spatial position of the target vehicle, and a vehicle motion trajectory tracking chain is established.
[0021] As a further scheme of the present application: when there are two or more candidate feature vectors with the highest similarity value greater than or equal to the preset threshold, the actual spatial positions of the vehicle regions corresponding to each candidate feature vector are obtained, the Euclidean distances between each actual spatial position and the curb machine generating a blind area trigger signal are calculated, the vehicle corresponding to the candidate feature vector with the smallest Euclidean distance is selected as the target vehicle, and the spatial position coordinates corresponding to the candidate feature vector are used to establish a vehicle motion trajectory tracking chain.
[0022] As a further scheme of the present application: in the S4, the specific process of parking compliance determination is:
[0023] Based on the spatial position information and vehicle body posture data of the target vehicle, a three-dimensional detection box information of the target vehicle is generated, and the three-dimensional detection box information is matched and analyzed with the electronic fence data of the preset parking area. If the three-dimensional detection box is completely located within the electronic fence range, it is determined as a compliant parking; if the three-dimensional detection box is partially located within the electronic fence range, it is determined as an abnormal parking.
[0024] As a further scheme of the present application: in the S4, the specific generation process of the standardized event record is:
[0025] The curb machine perception terminal is re-enabled for vehicle perception and cooperative determination, and if the determination results are consistent, the monitoring authority switching operation is performed synchronously, and the vehicle monitoring task is transferred from the high-level monitoring system to the curb machine perception terminal;
[0026] The target vehicle position information in the high-level monitoring data is used to compensate the spatial error of the curb machine perception data, and the curb machine clock is synchronously corrected according to the accurate time stamp recorded by the high-level monitoring system;
[0027] When the degree of coincidence of the curb machine perception data and the high-level monitoring data after error compensation reaches the preset standard, a data fusion result containing spatial calibration parameters and time synchronization calibration parameters is generated, and a standardized parking event record containing a compliant parking start time after calibration is generated based on the fusion result.
[0028] As a further aspect of the application: if the determination result conflicts, a curb machine warning event identifier is generated, and the space-time information of the target vehicle during the entire parking process is continuously recorded through the high-level monitoring system, and the curb machine warning event identifier and the space-time information are uploaded to the cloud platform to generate billing records and fault warning information.
[0029] A roadside parking high-low position fusion multi-modal recognition system for implementing the above-mentioned roadside parking high-low position fusion multi-modal recognition method, comprising:
[0030] An abnormality triggering module is configured to acquire radar perception data and visual perception data of any preset parking area corresponding curb machine in a target area in real time, and perform vehicle perception collaborative determination based on the radar perception data and the visual perception data, and when the determination result conflicts, a blind area triggering signal is generated and the signal generation time t0 is recorded.
[0031] A feature recognition module is configured to extract a continuous video frame sequence in the visual perception data within a preset time window based on the generation time t0, perform vehicle feature extraction processing on the video frame sequence, and obtain a vehicle feature vector.
[0032] A target matching module is configured to call a video stream of a high-level monitoring system corresponding to the target area, perform vehicle re-identification in the target area based on the vehicle feature vector using a feature matching algorithm, determine a target vehicle position, and establish positioning and continuous tracking monitoring of the target vehicle.
[0033] A dynamic monitoring module is configured to acquire spatial position information and vehicle body posture data of the target vehicle when the target vehicle is detected to stop moving, and perform parking compliance determination, if the determination result is compliant parking, re-enable the curb machine to perform vehicle state verification, and generate a standardized event record containing a compliant parking start time after fusing the verification result and the high-level monitoring data, and upload the standardized event record to the cloud platform to generate billing records.
[0034] A result generation module is configured to generate an abnormal event identifier if the determination result is abnormal parking, and continuously record the space-time information of the target vehicle during the entire parking process, and upload the space-time information to the cloud platform to generate billing records.
[0035] The beneficial effects of the application are:
[0036] 1) The present application realizes efficient utilization of sensor resources through innovative vehicle perception collaborative decision-making and dynamic task switching mechanism. When the system determines that a vehicle enters the blind area of the curb machine through the multi-modal data conflict of the curb machine, the high-level monitoring system is automatically triggered to take over the tracking task, ensuring monitoring continuity. After the vehicle is stable and confirmed as a compliant stop, the high-level system is released by re-enabling the curb machine for state verification and returning the dominant right. Through the dynamic scheduling strategy based on vehicle state, the waste of resources caused by the continuous full-function monitoring of the high-level camera in the prior art is effectively avoided, significantly reducing the system's computational load, communication bandwidth occupation and energy consumption, providing a feasible technical foundation for large-scale urban deployment.
[0037] 2) The present application greatly improves the accuracy and reliability of parking management by establishing a perfect multi-source data fusion and verification mechanism. By using high and low sensor data to verify each other at key nodes, high-level data is used to supplement the visual field loss of the curb machine when the blind area is triggered, and the curb machine verifies the high-level monitoring results after compliance determination. When generating billing records, multi-source data is fused to form standardized event records. This effectively solves the problem of single sensor blind area, improves the accuracy of vehicle identity recognition, parking timestamp recording and compliance determination, provides an unalterable data evidence chain for parking billing and management decision-making, and significantly reduces billing disputes and management loopholes.
[0038] 3) The present application establishes a hierarchical processing and intelligent operation and maintenance mechanism for abnormal parking scenarios, significantly improving the management efficiency and operation and maintenance level of the system. When abnormal parking behavior is identified, the system generates a dedicated abnormal event identifier and starts an enhanced tracking mode to continuously record the spatiotemporal information of the vehicle during parking, forming a complete abnormal behavior evidence chain to provide data support for subsequent management decisions. When the curb machine is reactivated and the perception data conflict is still detected, the system generates a device warning event identifier and automatically switches to the high-level system to continue monitoring while synchronously uploading the warning information and spatiotemporal data to the cloud platform. This ensures the reliable operation of the billing process in the case of device abnormality and timely triggers the device fault warning mechanism to provide precise device state information and fault positioning basis for operation and maintenance personnel, greatly improving the operation response speed and processing efficiency, and realizing the intelligent and fine operation and maintenance of the parking management system. BRIEF DESCRIPTION OF DRAWINGS
[0039] The present application will be further described below with reference to the accompanying drawings.
[0040] Figure 1 is a multi-modal recognition method flowchart of a roadside parking high-low fusion of the present application.
[0041] Figure 2It is a roadside parking high-low fusion multi-modal recognition system structure schematic diagram of the present application. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.
[0043] Please refer to Figure 1 The present application is a roadside parking high-low fusion multi-modal recognition method, which comprises the following steps:
[0044] S1, real-time acquisition of radar sensing data and visual sensing data of any pre-set parking area corresponding to the curb machine in the target area, and vehicle sensing cooperative determination according to the radar sensing data and visual sensing data, when the determination result conflicts, a blind area trigger signal is generated and the signal generation time t0 is recorded;
[0045] S2, based on the generation time t0, the continuous video frame sequence in the visual sensing data in the pre-set time window is extracted reversely, vehicle feature extraction processing is performed on the video frame sequence, and a vehicle feature vector is obtained;
[0046] S3, calling the video stream of the high-level monitoring system corresponding to the target area, using feature matching algorithm to perform vehicle re-identification in the target area based on the vehicle feature vector, determining the position of the target vehicle, and establishing the positioning and continuous tracking monitoring of the target vehicle;
[0047] S4, when the target vehicle is monitored to stop moving, the spatial position information and vehicle body posture data of the target vehicle are collected and parking compliance determination is performed, if the determination result is compliant parking, the curb machine is re-enabled for vehicle state verification, and the verification result is fused with the high-level monitoring data to generate a standardized event record containing the compliant parking start time, and the standardized event record is uploaded to the cloud platform to generate a billing record;
[0048] S5, if the determination result is abnormal parking, an abnormal event identifier is generated and the spatio-temporal information of the target vehicle parking throughout is continuously recorded, and the spatio-temporal information is uploaded to the cloud platform to generate a billing record.
[0049] 1) The application realizes efficient utilization of sensor resources through innovative vehicle perception collaborative decision-making and dynamic task switching mechanism. When the system determines that a vehicle enters the blind area of the curb machine through the multi-modal data conflict of the curb machine, the high-level monitoring system is automatically triggered to take over the tracking task, ensuring monitoring continuity. After the vehicle is parked and confirmed as a compliant stop, the high-level system is released by re-enabling the curb machine for state verification and returning the authority. Through dynamic scheduling strategy based on vehicle state, the waste of resources caused by continuous full-function monitoring of high-level camera in existing technology is effectively avoided, significantly reducing the system's computing load, communication bandwidth occupation and energy consumption, providing a feasible technical basis for large-scale urban deployment.
[0050] 2) The application greatly improves the accuracy and reliability of parking management by establishing a perfect multi-source data fusion and verification mechanism. By using high and low sensor data to verify each other at key nodes, high-level data is used to supplement the visual field loss of the curb machine when the blind area is triggered, and the curb machine verifies the high-level monitoring results after compliance determination. When generating billing records, multi-source data is fused to form standardized event records. This effectively solves the problem of single sensor blind area, improves the accuracy of vehicle identity recognition, parking timestamp recording and compliance determination, provides an unalterable data evidence chain for parking billing and management decision-making, and significantly reduces billing disputes and management loopholes.
[0051] 3) The application establishes a hierarchical processing and intelligent operation and maintenance mechanism for abnormal parking scenarios, significantly improving the management efficiency and operation and maintenance level of the system. When abnormal parking behavior is identified, the system generates a dedicated abnormal event identifier and starts enhanced tracking mode to continuously record the spatiotemporal information of the vehicle during parking, forming a complete abnormal behavior evidence chain to provide data support for subsequent management decisions. When the curb machine is reactivated and the perception data conflict is still detected, the system generates a device warning event identifier and automatically switches to the high-level system to continue monitoring while uploading the warning information and spatiotemporal data to the cloud platform. This ensures the reliable operation of the billing process in the case of device abnormalities, and timely triggers the device fault warning mechanism to provide precise device state information and fault positioning basis for maintenance personnel, greatly improving the response speed and processing efficiency of maintenance, and realizing intelligent and fine-grained operation and maintenance of the parking management system.
[0052] In a preferred embodiment of the application, the specific process of generating a blind area trigger signal in S1 is as follows:
[0053] When the target distance parameter in the radar perception data continuously presents a monotonic decreasing characteristic and the echo signal strength remains stable, while the recognition confidence of the continuous multiple frames of image in the visual perception data is lower than the preset threshold value; or when the recognition confidence of the continuous multiple frames of image in the visual perception data continuously exceeds the preset threshold value, while no effective target signal is detected in the radar perception data, it is determined that there is a recognition conflict, and under the condition that the radar perception data and the visual perception data are completed in space-time alignment, if the duration of the recognition conflict exceeds the preset time window threshold value, a blind area triggering signal is generated.
[0054] The system judges the working state of the sensor by comparing the space-time consistency of the radar perception data and the visual perception data in real time. When the radar sensor detects that the target distance parameter continuously decreases and the signal strength is stable, it indicates that an object is continuously approaching the curb machine. The recognition confidence of the continuous multiple frames of image in the synchronously collected visual perception data is lower than the preset threshold value, indicating that the camera cannot recognize the vehicle features. This contradiction between radar and visual data indicates that the vehicle may have entered the optical blind area of the camera.
[0055] Another case is that the visual perception data continuously displays a high confidence vehicle recognition result, but the radar sensor does not detect an effective target signal. This anomaly may be caused by the radar being blocked by a metal obstacle or hardware failure. The system ensures the comparability of the data of the two sensors in the space-time dimension through the establishment of a time stamp synchronization mechanism and coordinate system normalization processing. When any of the above conflict modes lasts for more than a preset time window, it is determined as a reliable blind area triggering condition, and then a blind area triggering signal is generated.
[0056] The radar is sensitive to the existence and motion state of objects based on the principle of electromagnetic wave reflection, but does not have recognition capability. The visual sensor can recognize vehicle features based on the principle of optical imaging, but is easily disturbed by the environment. When the output results of the two are continuously contradictory, it often reveals the limitations or abnormal state of a single sensor. By setting a time window threshold value, temporary shielding and other accidental factors can be effectively filtered out, ensuring that a blind area alarm will only be triggered in a continuous abnormal situation, intelligently distinguishing between sensor blind areas and temporary environmental disturbances, providing accurate triggering basis for the subsequent high-level monitoring system intervention, thereby ensuring the perception reliability and decision accuracy of the entire parking management system, avoiding missed detection or misjudgment problems caused by sensor limitations, and ultimately achieving seamless monitoring of the entire parking process.
[0057] In another preferred embodiment of the present application, the specific process of obtaining the vehicle feature vector in S2 is:
[0058] With the generation time t0 as the terminal point, the continuous N frame image sequence in the preset time length in the visual perception data is extracted in reverse, N is a preset quantity threshold, after pretreatment of illumination compensation and noise filtering for each frame image, the pre-trained deep convolutional neural network is used to extract color features, texture features and structure features in the pretreated image, and the feature weight of each frame is calculated through the time sequence attention module, and the multi-frame feature information is weighted and fused; the fused features are processed by principal component analysis method; the feature vector after dimension reduction is normalized to obtain the vehicle feature vector.
[0059] With the blind area trigger signal generation time t0 as the reference point, the continuous image sequence in the previous period of time is extracted in reverse, and the selection of this time window is based on the prior knowledge of the time required for the vehicle to normally enter the parking space, to ensure that the complete visible state before the vehicle enters the blind area can be covered. The first illumination compensation processing is performed on each frame of image in the sequence, the influence of shadow and uneven illumination is eliminated by adjusting the image brightness and contrast, and then a filtering algorithm is used to suppress image noise. These pretreatment operations are to improve the quality of subsequent feature extraction. Then a deep convolutional neural network trained in advance on a large amount of vehicle image data is used to process each frame of image, and the network can automatically learn the color features (such as vehicle body main tone and color distribution), texture features (such as license plate character texture and vehicle lamp details) and structure features (such as vehicle contour and size ratio) of the vehicle through multi-layer convolution operation. Then the time sequence attention module is introduced to analyze the importance difference of different frames in the sequence, for example, the image of the front view of the vehicle usually contains more recognition information than the image of the serious occlusion, the module will assign higher weight to the frame with more information, and finally the features of all frames are weighted and fused according to the weight to form a comprehensive feature representation. In order to avoid high feature dimension and redundant information, principal component analysis method is used to reduce the dimension of the fused features, and the most distinctive feature components are retained. Finally, the feature vector after dimension reduction is normalized to make the module length uniform, which is convenient for subsequent similarity calculation and matching, and finally a compact feature vector containing both vehicle appearance characteristics and time sequence correlation is output.
[0060] The final vehicle feature vector can provide accurate feature basis for subsequent calling of high-bit monitoring system for vehicle re-identification, to ensure that the target vehicle position can be quickly and accurately determined and tracking is established, thereby laying a foundation for subsequent processes such as parking compliance judgment and monitoring resource dynamic allocation.
[0061] In another preferred embodiment of the present application, in the S3, the specific process of vehicle re-identification is:
[0062] The region proposals of all vehicles in the video frame are extracted by a target detection algorithm and a candidate feature vector set is generated; a similarity matrix between the vehicle feature vector and any candidate feature vector set is calculated, and a weighted cosine similarity measurement method is used for multi-feature dimension matching;
[0063] The highest similarity value between each candidate feature vector and the vehicle feature vector is obtained by maximum value screening of the similarity matrix, and when the highest similarity value is greater than or equal to a preset threshold, it is confirmed that the re-identification is successful and the position coordinates of the corresponding vehicle region in the image coordinate system are obtained.
[0064] The image coordinates are converted to the world coordinate system according to the camera calibration parameters, the actual spatial position of the target vehicle is obtained, and a vehicle motion trajectory tracking chain is established.
[0065] In another preferred embodiment of the application, the specific process of vehicle re-identification is as follows: first, real-time video stream is obtained from a high monitoring system, a target detection algorithm based on deep learning is used to analyze the video frame, the algorithm identifies all rectangular regions that may contain vehicles in the image through a convolutional neural network, these regions are called region proposals, each region proposal is subjected to image cropping and size standardization, and then input into another deep feature extraction network to generate corresponding candidate feature vectors, which together form a candidate feature vector set. Next, the similarity between the vehicle feature vector extracted in front of the blind area of the road machine and each vector in the candidate feature vector set is calculated, and a weighted cosine similarity measurement method is used, which assigns different weight coefficients to different dimensions such as color features, texture features and structural features, because different features have different importance in vehicle re-identification, for example, license plate texture features are usually more distinctive than color features. By scanning the matrix composed of all similarity values, the highest similarity value between each candidate vehicle and the query vehicle is found, and when the value exceeds a preset threshold, it indicates that a highly matched vehicle is found. At this time, according to the region proposal coordinates obtained in the target detection stage, the specific position of the vehicle in the image can be obtained. In order to convert the two-dimensional coordinates in the image to three-dimensional coordinates in the real world, camera calibration parameters are needed, which include the intrinsic parameters (such as focal length, principal point coordinates) and extrinsic parameters (such as camera installation height and angle) of the camera, and through the principle of perspective projection transformation, the image coordinates are mapped to the world coordinate system, and finally the accurate geographical position of the vehicle in the actual parking lot is obtained. Based on this position information, the system initializes a trajectory tracking chain, and through the position association and motion prediction between consecutive frames, a motion trajectory model of the vehicle is established.
[0066] The advantages of high and low sensors are fully utilized to achieve accurate matching and precise positioning of vehicle identity. The target detection algorithm is used to find the area of all possible vehicles, which can ensure that no potential target is missed and provide a basis for subsequent accurate matching. Weighted cosine similarity is used instead of simple similarity calculation because different features of the vehicle have different reliabilities in identification. For example, color is easily affected by light, while license plate texture is more stable. By assigning different weights, the accuracy of matching can be improved. The similarity threshold is set to filter out obviously mismatched vehicles and ensure the reliability of the identification result. Coordinate conversion is performed because image coordinates cannot be directly used for parking management and must be converted to real-world coordinates to determine the specific parking space of the vehicle. The establishment of the motion trajectory tracking chain enables continuous tracking of the vehicle and provides data support for subsequent parking behavior analysis and compliance determination. This method can ensure accurate identification and positioning of target vehicles even in complex road environments, providing reliable data input for the entire parking management system and ultimately achieving precise parking management and charging.
[0067] In another preferred embodiment of the present application, when the highest similarity values of two or more candidate feature vectors with the vehicle feature vector are greater than or equal to the preset threshold, the actual spatial positions of the vehicle regions corresponding to each candidate feature vector are obtained, the Euclidean distances between each actual spatial position and the curb machine that generates the blind area trigger signal are calculated, the vehicle corresponding to the candidate feature vector with the smallest Euclidean distance is selected as the target vehicle, and the spatial position coordinates corresponding to the candidate feature vector are used to establish a vehicle motion trajectory tracking chain.
[0068] When the similarity of multiple candidate feature vectors with the query vehicle feature vector exceeds the preset threshold, the system will start the multi-target conflict resolution mechanism. First, the actual spatial positions of the vehicle regions corresponding to each candidate feature vector are obtained by converting the image coordinates to the world coordinate system using the previously established camera calibration model, and each vehicle position can be represented by three-dimensional coordinates. Then, the Euclidean distances between each candidate vehicle position and the curb machine that triggers the blind area signal are calculated. The Euclidean distance is obtained by calculating the square sum of the differences between two points on each coordinate axis and then taking the square root, which accurately reflects the actual straight-line distance in space. The system selects the candidate vehicle with the smallest Euclidean distance as the final target because the vehicle is closest to the curb before entering the blind area, and according to the principle of motion continuity, this vehicle is most likely to be the one that triggered the blind area. Finally, the spatial position coordinates corresponding to this candidate feature vector are used to initialize the vehicle motion trajectory tracking chain, ensuring the accuracy of subsequent tracking.
[0069] The prior knowledge of the spatial position is fully utilized to solve the problem that the visual feature similarity cannot distinguish multiple candidate vehicles. When multiple vehicles are very similar in visual features, it is difficult to make an accurate judgment simply by relying on appearance matching, and the spatial position information provides an independent and reliable basis for discrimination. The vehicle closest to the curb is selected based on the physical fact that the vehicle needs to pass near the curb to enter its blind area, so the closest rut is most likely to trigger the blind area. This method can effectively solve the matching conflict problem in the case of multiple similar vehicles, improve the accuracy and reliability of re-identification, ensure that the system can correctly identify the real target vehicle, provide accurate data basis for subsequent parking behavior monitoring and management, and ultimately ensure the normal operation and accurate billing of the entire parking management system.
[0070] In another preferred embodiment of the present application, the specific process of parking compliance determination in S4 is:
[0071] Based on the spatial position information and the vehicle body posture data of the target vehicle, three-dimensional detection box information of the target vehicle is generated, and the three-dimensional detection box information is analyzed by spatial matching with the electronic fence data of the preset parking area. If the three-dimensional detection box is completely located within the electronic fence range, it is determined as compliant parking; if the three-dimensional detection box is partially located within the electronic fence range, it is determined as abnormal parking.
[0072] By matching the three-dimensional detection frame with the electronic fence, it is verified whether the vehicle is in the detection range of the curb machine blind area, and whether the curb machine can continuously and stably obtain accurate sensing data. This is the key basis for whether the monitoring dominance will be transferred from the high-level system to the curb machine. It can be understood that when the vehicle is parked in compliance, the curb machine is most likely to have no detection blind area, and its radar and visual sensing can be free from interference such as shielding and angle deviation, and can continuously output reliable data to lay a foundation for taking over the monitoring task. The "compliant parking" and "curb machine blind area and stable sensing" are directly related to form a clear monitoring dominance switching judgment standard: when the vehicle is parked in compliance, the curb machine has stable sensing capability, and after the dominance is transferred, it can rely on the near-detection advantage to realize accurate monitoring, avoiding resource consumption caused by continuous operation of the high-level system; when the vehicle is parked abnormally, the curb machine still has a blind area and unreliable sensing, and the dominance is not transferred and the high-level system continues to cover the blind area, ensuring that the vehicle monitoring is not interrupted, which provides a scientific basis for dynamic allocation of monitoring resources: only in the compliant scenario where the curb machine has stable sensing, the curb machine is allowed to take over, avoiding repeated switching of the system and waste of resources caused by invalid transfer; and ensuring the reliability of the whole process of parking management: in the compliant scenario, the curb machine stably senses to support subsequent vehicle state verification, data fusion and standardized billing record generation, and in the abnormal scenario, the high-level system continuously records the space-time information to support subsequent management, finally realizing the closed loop of "accurate identification-reliable monitoring-efficient billing" in roadside parking management, which meets the core goal of optimizing resource utilization and improving management accuracy through high-low level fusion in the document.
[0073] In another preferred embodiment of the present application, the specific generation process of the standardized event record in S4 is:
[0074] The curb machine sensing terminal is re-enabled for vehicle sensing cooperative judgment, and if the judgment results are consistent, the monitoring dominance switching operation is performed synchronously, and the vehicle monitoring task is transferred from the high-level monitoring system to the curb machine sensing terminal;
[0075] The spatial error of the curb machine sensing data is compensated by using the target vehicle position information in the high-level monitoring data, and the curb machine clock is synchronously corrected according to the accurate time stamp recorded by the high-level monitoring system;
[0076] When the consistency of the curb machine sensing data and the high-level monitoring data after error compensation reaches the preset standard, a data fusion result containing spatial calibration parameters and time synchronization calibration parameters is generated, and a standardized parking event record containing the calibrated compliant parking start time is generated based on the fusion result.
[0077] When the system determines that the vehicle is in compliance with the parking, first re-enable the curb machine perception terminal to monitor the vehicle state, compare the real-time data collected with the high-level monitoring system data to ensure consistency in vehicle identification and state judgment. This double verification mechanism can avoid misjudgment caused by single sensor error. After confirming data consistency, the system performs monitoring authority switching operation, gradually reduces the sampling frequency and data processing priority of the high-level monitoring system, while increasing the data collection weight of the curb machine perception terminal, to realize smooth transition of monitoring task. Next, use the accurate vehicle positioning information obtained by the high-level monitoring system to compensate the spatial error of the curb machine perception data. This is because the high-level camera has a wider field of view and more accurate coordinate measurement capability, which can correct the measurement deviation caused by the installation position and viewing angle limitation of the curb machine. At the same time, according to the high-precision clock source of the high-level monitoring system, the local clock of the curb machine is synchronized and corrected to eliminate the timestamp difference between devices. When the calibrated curb machine data and high-level monitoring data remain consistent within the preset tolerance range, the system generates a data fusion result containing spatial coordinate correction parameters and time synchronization parameters. Finally, based on this fusion result, a standardized parking event record is generated, which contains the calibrated accurate parking start time, vehicle precise positioning information and data source identification, forming a complete parking event evidence chain.
[0078] By fully utilizing the advantages of different sensors and ensuring the accuracy and reliability of data. By reactivating the curb machine and comparing the data, it can verify whether the sensor state has returned to normal, providing a basis for monitoring authority switching. Implementing monitoring authority switching is to optimize system resource allocation, transferring computing load to low-power curb machine terminal while ensuring monitoring continuity. Calibrating and compensating low-level data with high-level data is because high-level cameras have better global observation capability and can correct local measurement errors of curb machines. Time synchronization is to ensure that all sensor data has a unified time reference, avoiding time record errors caused by different device clocks. Generating standardized event records is to establish a complete and reliable parking evidence chain, providing accurate basis for billing and management decisions. This method can ensure that the system still produces accurate and reliable parking records in complex environments, ensuring the fairness and accuracy of billing, while improving the efficiency of system resources, providing technical support for the stable operation of large-scale parking management systems.
[0079] In another preferred embodiment of the present application, if the judgment result conflicts, a curb machine warning event identifier is generated, and the high-level monitoring system continuously records the spatio-temporal information of the target vehicle during the entire parking period, and uploads the curb machine warning event identifier and spatio-temporal information to the cloud platform to generate billing records and fault warning information.
[0080] When the curb machine is re-enabled, the system will start the abnormal processing procedure when the perception data of the curb machine is still inconsistent with the high-level monitoring data. First, the data comparison algorithm is used to analyze the difference between the two types of data in terms of vehicle presence state, position information or identification result, etc. When the difference exceeds the allowed range, a specific curb machine warning event identifier is generated, which includes device number, conflict type and timestamp, etc. At the same time, the system keeps the high-level monitoring system tracking the target vehicle, recording the complete spatio-temporal information of the vehicle during parking in a higher sampling frequency than the regular one, including vehicle position coordinates, attitude angle and time series data, etc. These spatio-temporal information and warning event identifier are bound by data association algorithm to form an event data package with complete context. Finally, the data package is uploaded to the cloud platform through an encrypted transmission protocol. The cloud platform first generates accurate billing records using reliable spatio-temporal information to ensure business continuity, and then analyzes the device health status according to the type and frequency of the warning event identifier. When a persistent abnormal pattern is detected, a device fault warning is automatically generated to notify the operation and maintenance personnel for corresponding processing.
[0081] The warning event identifier can accurately record the abnormal state of the device, providing structured data for subsequent analysis. The continuous operation of the high-level monitoring system ensures that reliable vehicle monitoring data is obtained even in the case of curb machine abnormalities, avoiding service interruption. Binding the warning information with spatio-temporal data for uploading ensures the integrity and traceability of the data, facilitating problem positioning and analysis. Using the cloud platform to handle billing and fault warnings simultaneously ensures both business and operation, ensuring the normal operation of the parking billing business and timely detection of potential device faults, significantly improving the reliability and maintainability of the system, reducing service interruptions caused by device faults, reducing operation and maintenance costs, and providing users with continuous and reliable parking service experience. Ultimately, it realizes the efficient and stable operation of the parking management system.
[0082] Referring to Figure 2 The application also includes a roadside parking high-low position fusion multi-modal recognition system for implementing the above-mentioned roadside parking high-low position fusion multi-modal recognition method, comprising:
[0083] An abnormal trigger module is configured to acquire radar perception data and visual perception data of a curb machine corresponding to any preset parking area in a target area in real time, and perform vehicle perception collaborative determination based on the radar perception data and visual perception data. When the determination result conflicts, a blind area trigger signal is generated and the signal generation time t0 is recorded.
[0084] A feature recognition module is configured to extract a continuous video frame sequence in the visual perception data within a preset time window based on the generation time t0, and perform vehicle feature extraction processing on the video frame sequence to obtain a vehicle feature vector.
[0085] The target matching module is configured to call a high-position monitoring system video stream corresponding to a target area, perform vehicle re-identification in the target area based on the vehicle feature vector by using a feature matching algorithm, determine a target vehicle position, and establish positioning and continuous tracking monitoring of the target vehicle.
[0086] The dynamic monitoring module is configured to, when it is monitored that the target vehicle stops moving, collect spatial position information and vehicle body posture data of the target vehicle and perform parking compliance determination, if the determination result is compliant parking, re-enable the curbstone machine to perform vehicle state verification, and generate a standardized event record containing a compliant parking start time after fusing the verification result and the high-position monitoring data, and upload the standardized event record to the cloud platform to generate a billing record.
[0087] The result generation module is configured to, if the determination result is abnormal parking, generate an abnormal event identifier and continuously record space-time information of the target vehicle parking throughout, and upload the space-time information to the cloud platform to generate a billing record.
[0088] The above describes one embodiment of the present application in detail, but the content described is only a preferred embodiment of the present application and cannot be considered as limiting the scope of the present application. Any equivalent changes and improvements made in the scope of the present application should still belong to the patent coverage of the present application.
Claims
1. A multimodal recognition method for roadside parking with high and low parking positions, characterized in that, Includes the following steps: S1. Real-time acquisition of radar perception data and visual perception data of curb motor corresponding to any preset parking area within the target area, and vehicle perception coordination determination based on the radar perception data and visual perception data. When there is a conflict in the determination results, a blind spot trigger signal is generated and the signal generation time t0 is recorded. S2, based on the generation time t0, extract the continuous video frame sequence from the visual perception data within the preset time window, and perform vehicle feature extraction processing on the video frame sequence to obtain the vehicle feature vector. S3, call the video stream of the high-level monitoring system corresponding to the target area, and use the feature matching algorithm based on the vehicle feature vector to re-identify the vehicle in the target area, determine the position of the target vehicle, and establish the positioning and continuous tracking monitoring of the target vehicle; S4. When the target vehicle stops moving, the spatial position information and body posture data of the target vehicle are collected and the parking compliance is determined. If the determination result is compliant parking, the curb machine is reactivated to verify the vehicle status. The verification result is then merged with the high-level monitoring data to generate a standardized event record containing the start time of compliant parking. The standardized event record is then uploaded to the cloud platform to generate a billing record. The specific process for generating standardized event logs is as follows: Reactivate the curb sensor terminal to perform vehicle perception and collaborative judgment. If the judgment results are consistent, simultaneously execute the monitoring leadership switch operation to transfer the vehicle monitoring task from the high-level monitoring system to the curb sensor terminal. Spatial error compensation is performed on the curb machine's sensing data using the target vehicle location information in the high-level monitoring data, and the curb machine's clock is synchronized and corrected based on the precise timestamp recorded by the high-level monitoring system. When the consistency between the curb machine's sensing data and the high-level monitoring data reaches a preset standard after error compensation, a data fusion result containing spatial calibration parameters and time synchronization calibration parameters is generated. Based on the fusion result, a standardized parking event record containing the start time of compliant parking after calibration is generated. S5. If the determination result is abnormal parking, an abnormal event identifier is generated and the spatiotemporal information of the target vehicle throughout the parking process is continuously recorded. The spatiotemporal information is then uploaded to the cloud platform to generate a billing record.
2. The multimodal recognition method for roadside parking height and low position fusion according to claim 1, characterized in that, In S1, the specific process of generating the blind zone trigger signal is as follows: When the target distance parameter in the radar sensing data continuously exhibits a monotonically decreasing characteristic and the echo signal strength remains stable, while the recognition confidence of multiple consecutive frames of images in the visual sensing data is lower than a preset threshold; or when the recognition confidence of multiple consecutive frames of images in the visual sensing data is continuously higher than a preset threshold, while no valid target signal is detected in the radar sensing data, it is determined that there is a recognition conflict. If the duration of the recognition conflict exceeds a preset time window threshold after the radar sensing data and visual sensing data are spatiotemporally aligned, a blind zone trigger signal is generated.
3. The multimodal recognition method for roadside parking height and low position fusion according to claim 1, characterized in that, In S2, the specific process of obtaining the vehicle feature vector is as follows: Starting from the generation time t0, a sequence of N consecutive images within a preset time period is extracted from the visual perception data, where N is a preset threshold number. After preprocessing each image with illumination compensation and noise filtering, a pre-trained deep convolutional neural network is used to simultaneously extract color, texture, and structural features from the preprocessed image. The feature weights of each frame are calculated using a temporal attention module, and the feature information from multiple frames is weighted and fused. Principal component analysis is used to reduce the dimensionality of the fused features. The reduced feature vectors are then normalized to obtain the vehicle feature vector.
4. The multimodal recognition method for roadside parking height and low position fusion according to claim 1, characterized in that, In S3, the specific process of vehicle re-identification is as follows: The region proposals of all vehicles in the video frame are extracted by the object detection algorithm and a candidate feature vector set is generated; the similarity matrix between the vehicle feature vector and any candidate feature vector set is calculated, and the weighted cosine similarity measure method is used for multi-feature dimension matching. By filtering the maximum value of the similarity matrix, the highest similarity value between each candidate feature vector and the vehicle feature vector is obtained. When the highest similarity value is greater than or equal to a preset threshold, the re-identification is confirmed to be successful and the position coordinates of the corresponding vehicle area in the image coordinate system are obtained. Based on the camera calibration parameters, the image coordinates are transformed to the world coordinate system to obtain the actual spatial position of the target vehicle and establish a vehicle motion trajectory tracking chain.
5. The multimodal recognition method for roadside parking height and low position fusion according to claim 4, characterized in that, When there are two or more candidate feature vectors whose highest similarity value with the vehicle feature vector is greater than or equal to a preset threshold, the actual spatial position of the vehicle area corresponding to each candidate feature vector is obtained, the Euclidean distance between each actual spatial position and the curb machine that generates the blind spot trigger signal is calculated, the vehicle corresponding to the candidate feature vector with the smallest Euclidean distance is selected as the target vehicle, and the vehicle motion trajectory tracking chain is established using the spatial position coordinates corresponding to the candidate feature vector.
6. The multimodal recognition method for roadside parking height and low position fusion according to claim 1, characterized in that, In S4, the specific process for determining parking compliance is as follows: Based on the spatial location information and body posture data of the target vehicle, a three-dimensional detection frame information of the target vehicle is generated. The three-dimensional detection frame information is then spatially matched with the electronic fence data of the preset parking area. If the three-dimensional detection frame is completely within the electronic fence area, it is determined to be compliant parking. If the 3D detection frame is located within the electronic fence area, it is determined to be an abnormal stop.
7. The multimodal recognition method for roadside parking height and low position fusion according to claim 1, characterized in that, It also includes generating a curb machine warning event identifier when there is a conflict in the judgment results, continuously recording the spatiotemporal information of the target vehicle throughout its parking process through a high-level monitoring system, and uploading the curb machine warning event identifier and spatiotemporal information to the cloud platform to generate billing records and fault warning information.
8. A multimodal recognition system for roadside parking with high and low positions, used to implement the multimodal recognition method for roadside parking with high and low positions as described in any one of claims 1-7, characterized in that, include: The abnormal triggering module is used to acquire radar perception data and visual perception data of the curb machine corresponding to any preset parking area in the target area in real time, and to perform vehicle perception collaborative judgment based on the radar perception data and visual perception data. When there is a conflict in the judgment results, a blind spot trigger signal is generated and the signal generation time t0 is recorded. The feature recognition module is used to extract a continuous video frame sequence from the visual perception data within a preset time window based on the generation time t0, and to perform vehicle feature extraction processing on the video frame sequence to obtain a vehicle feature vector. The target matching module is used to call the video stream of the high-level monitoring system corresponding to the target area, and use the feature matching algorithm based on the vehicle feature vector to re-identify the vehicle in the target area, determine the position of the target vehicle, and establish the positioning and continuous tracking monitoring of the target vehicle. The dynamic monitoring module is used to collect the spatial location information and vehicle posture data of the target vehicle when the target vehicle stops moving, and to determine the parking compliance. If the determination result is compliant parking, the curb machine is reactivated to verify the vehicle status. The verification result is then merged with the high-level monitoring data to generate a standardized event record containing the start time of compliant parking. The standardized event record is then uploaded to the cloud platform to generate a billing record. The result generation module is used to generate an abnormal event identifier and continuously record the spatiotemporal information of the target vehicle throughout its parking process if the determination result is abnormal parking, and upload the spatiotemporal information to the cloud platform to generate a billing record.
Citation Information
Patent Citations
Roadside parking unmanned charging management system and method based on curbstone machine
CN116486497A
Roadside parking management method based on multi-mode BEV fusion
CN118097353A