A method and device for monitoring traffic conditions in highway tunnel environment
Through hardware and software synchronization technology combined with focal length camera and lidar data, the limitations of a single sensor in the tunnel environment are solved, and comprehensive perception and monitoring of tunnel traffic conditions are achieved.
Patent Information
- Application Number
- CN202510157722.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The existing cameras and lidars have their own limitations when monitoring in highway tunnel environments. The cameras are greatly affected by light and are difficult to obtain distance information. Lidars cannot reflect object color and other categories of information, making it difficult to adapt to complex traffic scenes in tunnels.
The focal length camera and lidar data are obtained through hardware and software synchronization technology, image and point cloud features are extracted, depth estimation, cone representation and pooling operations are performed, and local feature extraction and global average pooling are combined to realize the fusion perception of image and point cloud data.
It realizes the effective integration of lidar and camera data, fully senses the characteristics of vehicles passing in the tunnel, and improves the accuracy and robustness of tunnel traffic conditions monitoring.
Smart Images

Figure CN119600555B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic environment perception, and in particular to a method and device for monitoring traffic conditions in a highway tunnel environment. Background Art
[0002] In the field of highway tunnel traffic condition monitoring, cameras and radars are the most commonly used monitoring sensors, especially LiDAR, which relies on point cloud scanning and is not easily affected by obstacles. It is widely used in the field of highway tunnel operation condition monitoring. However, due to the limitations of technical principles, both types of sensors have their inherent defects. For example, the disadvantage of cameras is that they are greatly affected by external lighting conditions and are difficult to adapt to the complex traffic environment of tunnels. At the same time, since the image data collected by the camera is mainly 2D data, it is also difficult to obtain the distance information of scenes and objects. In comparison, the data generated by LiDAR is 3D data, which can reflect the distance and position information of the detected object, but because LiDAR relies on the generated 3D point cloud for detection, it is naturally unable to reflect the color and other category information of the detected object. Summary of the invention
[0003] In order to solve the technical problem that a single sensor is difficult to adapt to the complex monitoring scene of a highway tunnel due to its own technical limitations, an embodiment of the present invention provides a method and device for monitoring the traffic conditions in a highway tunnel environment.
[0004] The technical solution of the embodiment of the present invention is achieved as follows:
[0005] The embodiment of the present invention provides a method for monitoring traffic conditions in a highway tunnel environment, the method comprising: acquiring image data collected by a focal camera and point cloud data collected by a laser radar; the focal camera and the laser radar are both located on the wall of the highway tunnel, and time alignment of the collected data is achieved based on hardware synchronization and software synchronization technology; extracting image features from the image data, screening feature points of the point cloud data according to curvature, and matching feature points of previous and next frames; combining the image features with depth information in the matched point cloud data, and performing depth estimation, cone representation and pooling operations on the combined features to obtain image BEV features; performing local feature extraction and maximum pooling operations on the point cloud data in the point cloud data, and fusing the image features to obtain laser radar BEV feature predictions; constructing a world coordinate system, and in the world coordinate system, cascading the image BEV features and the laser radar BEV feature predictions, extracting features, performing global average pooling and convolution prediction to obtain dynamic fusion features; the dynamic fusion features reflect the comprehensive perception of the features of vehicles passing through the highway tunnel.
[0006] In one embodiment, extracting image features from the image data includes: extracting image features from the image data using a first model; the first model is a ResNet network model; the image features are used to characterize vehicle model, vehicle color and license plate number features in the image data.
[0007] In one embodiment, the feature points of the point cloud data are screened according to the curvature, and the feature points of the previous and next frames are matched, including:
[0008] Preprocessing the original point cloud in the point cloud data, classifying the original point cloud according to the laser line, and calculating the curvature of each point on the laser line;
[0009] According to the curvature of each point on the laser line, the line point and the surface point are determined;
[0010] The line points and surface points of the previous and next frames are transferred to the same coordinate system and matched according to the nearest neighbor search;
[0011] The curvature of each point on the laser line is calculated using the following formula:
[0012] ;
[0013] in, represents the curvature, represents a selected set of surrounding points, Indicates the current point, Indicates neighboring points.
[0014] In one embodiment, the image features are combined with the depth information in the matched point cloud data, and the combined features are subjected to depth estimation, frustum representation and pooling operations to obtain image BEV features, including:
[0015] The image features are combined with the depth information in the matched point cloud data using the following calculation formula:
[0016] ;
[0017] in, Represents the combined features, Indicates that image features and depth information are combined in the feature dimension. Represents depth information, Represents image features;
[0018] The combined features are used to estimate depth using the following calculation formula:
[0019] ;
[0020] in, is the depth estimation network operation, represents the predicted depth map;
[0021] The depth map and image features are mapped to three-dimensional space using the following calculation formula to obtain the frustum features in three-dimensional space:
[0022] ;
[0023] in, The cone feature means placing the image features in three-dimensional space according to the predicted depth;
[0024] The cone features are pooled using the following calculation formula to obtain the image BEV features:
[0025] ;
[0026] in, represents a pooling operation that converts the cone feature into a BEV representation. Represents the BEV feature of the image.
[0027] In one embodiment, local feature extraction and maximum pooling operation are performed on the point cloud data in the point cloud image data, and image features are fused to obtain laser radar BEV feature prediction, including:
[0028] The local features of the point cloud data in the point cloud data are extracted using the following calculation formula to obtain the local features:
[0029] ;
[0030] in, Yes The local eigenvector of , is the input point cloud, is the number of points, For point Contains 3D coordinate information, and MLP is the backbone network of the VoxelNet deep learning model;
[0031] The local features are subjected to maximum pooling operation using the following calculation formula to obtain global features:
[0032] ;
[0033] in, represents the pooling operation, Represents global features;
[0034] The global feature and the image feature are concatenated and fused using the following calculation formula to obtain the fused feature:
[0035] ;
[0036] in, represents the fused features, Represents image features; Represents a splicing and fusion operation;
[0037] The fused features are converted into a bird's-eye view using the following calculation formula, and feature prediction is performed to obtain the lidar BEV feature prediction:
[0038] ;
[0039] in, It is the lidar BEV feature prediction, including the predicted target location, size and category; It is the feature map converted from the fused features to the bird's-eye view; Represents the prediction network.
[0040] In one embodiment, constructing a world coordinate system includes:
[0041] Construct the correspondence between the points in the world coordinate system and the points in the focal camera coordinate system:
[0042] ;
[0043] in, are the coordinates of the point in the focal camera coordinate system; is the coordinate of the point in the world coordinate system; is the rotation matrix of the focal camera coordinate system relative to the world coordinate system, is the translation matrix of the focal length camera coordinate system relative to the world coordinate system;
[0044] Construct the correspondence between points on the image plane and points in the focal camera coordinate system:
[0045] ;
[0046] in, are the coordinates of the point on the image plane; is the focal length camera internal parameter matrix;
[0047] Determine the coordinates of points in the point cloud data in the world coordinate system:
[0048] ;
[0049] in, is the coordinate of the point in the point cloud data in the world coordinate system, is the rotation matrix The inverse matrix of is a point in the point cloud data;
[0050] The correspondence between the points in the focal camera image plane and the points in the LiDAR point cloud data is constructed, and based on the correspondence, the Bundle optimization algorithm is used to minimize the following errors to achieve spatial alignment of the focal camera image features and the LiDAR point cloud features:
[0051] ;
[0052] in, and are the number of image feature points and points in the lidar point cloud, respectively.
[0053] In one embodiment, the image BEV feature and the laser radar BEV feature prediction are cascaded, feature extracted, globally averaged pooled, and convolutionally predicted to obtain a dynamic fusion feature; the dynamic fusion feature reflects the comprehensive perception of the characteristics of the vehicles passing through the highway tunnel, including:
[0054] The image BEV feature is concatenated with the lidar BEV feature prediction using the following calculation formula:
[0055] ;
[0056] in, represents the features after cascading, is the point cloud BEV feature, is the image BEV feature, , is the size of the cascaded feature map, Indicates a cascade operation;
[0057] Use the following calculation formula to extract the features after cascading to obtain the extracted features;
[0058] ;
[0059] in, represents the extracted features, represents a three-dimensional convolutional neural network;
[0060] Use the following calculation formula to perform global average pooling on the extracted features to obtain the pooled features:
[0061] ;
[0062] in, represents the features after pooling, is the number of channels after pooling;
[0063] The pooled features are predicted through the convolution layer to obtain dynamic fusion features, thereby realizing the dynamic fusion of point cloud BEV features and image BEV features.
[0064] An embodiment of the present invention also provides a device for monitoring traffic conditions in a highway tunnel environment, the device comprising a laser radar, a focal length camera and a data processing module; the laser radar is used to collect point cloud data; the focal length camera is used to collect image data; the data processing module is used to execute the steps of the above-mentioned method.
[0065] The method and device for monitoring the traffic conditions in a highway tunnel environment provided by an embodiment of the present invention obtain image data collected by a focal camera and point cloud data collected by a laser radar; the focal camera and the laser radar are both located on the wall of the highway tunnel, and time alignment of the collected data is achieved based on hardware synchronization and software synchronization technology; image features are extracted from the image data, feature points are screened according to curvature for the point cloud data, and feature points of previous and next frames are matched; the image features are combined with depth information in the matched point cloud data, and the combined features are subjected to depth estimation, cone representation and pooling operations to obtain image BEV features; local feature extraction and maximum pooling operations are performed on the point cloud data in the point cloud data, and the laser radar BEV feature prediction is obtained after the image features are fused; a world coordinate system is constructed, and in the world coordinate system, the image BEV features and the laser radar BEV feature prediction are cascaded, feature extracted, globally averaged pooled and convolutionally predicted to obtain dynamic fusion features; the dynamic fusion features reflect the comprehensive perception of the features of vehicles passing through the highway tunnel. The solution provided by the present invention can realize the effective fusion of laser radar point cloud data and focal length camera image data. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 A schematic diagram of a flow chart of a method for monitoring traffic conditions in a highway tunnel environment according to an embodiment of the present invention;
[0067] Figure 2 A schematic diagram of the complete process of a method for monitoring traffic conditions in a highway tunnel environment according to an embodiment of the present invention;
[0068] Figure 3 This is a schematic diagram of the structure of a perception device based on laser radar and camera data fusion according to an embodiment of the present invention;
[0069] Figure 4 1 is a diagram showing the internal structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0070] The present invention will be described in further detail below with reference to the accompanying drawings and embodiments.
[0071] The embodiment of the present invention provides a method for monitoring traffic conditions in a highway tunnel environment. Figure 1 As shown, the method includes:
[0072] Step 101: acquiring image data collected by a focal length camera and point cloud data collected by a laser radar; the focal length camera and the laser radar are both located on the wall of a highway tunnel, and time alignment of the collected data is achieved based on hardware synchronization and software synchronization technology;
[0073] Step 102: extracting image features from the image data, screening feature points of the point cloud data according to curvature, and matching feature points of previous and next frames;
[0074] Step 103: combining the image features with the depth information in the matched point cloud data, and performing depth estimation, frustum representation and pooling operations on the combined features to obtain image BEV features;
[0075] Step 104: extract local features and perform maximum pooling operations on the point cloud data in the point cloud image data, and fuse image features to obtain laser radar BEV feature prediction;
[0076] Step 105: construct a world coordinate system, in which the image BEV feature and the lidar BEV feature prediction are cascaded, feature extracted, globally averaged pooled and convolution predicted to obtain a dynamic fusion feature; the dynamic fusion feature reflects the comprehensive perception of the characteristics of vehicles passing through the highway tunnel.
[0077] This embodiment provides a method for intelligent perception of highway tunnels based on the fusion of laser radar and camera data, which can realize multi-source data fusion. The focal length camera and laser radar and other perception devices in this embodiment are arranged in the highway tunnel to collect traffic flow and passing vehicle information in the highway tunnel and perceive the traffic condition data of the highway tunnel. The method of this embodiment can realize the fusion and superposition of basic data perceived by different sensors, and finally output the fused perception results for monitoring the traffic condition of the highway tunnel and supporting the safety control and decision-making of the highway tunnel.
[0078] Specifically, see Figure 2 The complete process of the method of this embodiment includes the following steps:
[0079] Step 1: Ensure time alignment between the LiDAR and focal length camera based on hardware synchronization and software synchronization.
[0080] Step 2: Extract image features and point cloud features based on the focal length camera image data and the lidar sparse point cloud image.
[0081] Step 3: Extract deep features from the original image of the focal length camera, add the depth information of the sparse point cloud map of the laser radar, and perform depth estimation through image data and point cloud features. Generate the cone feature through the depth feature and image feature, and generate the image BEV feature after pooling.
[0082] Step 4: Input point cloud data, extract features through the backbone network, and fuse image features to guide the lidar BEV feature prediction.
[0083] Step 5: Construct a world coordinate system to align the focal camera image features with the lidar point cloud features.
[0084] Step 6: Cascade the point cloud and image BEV features according to the channel dimension, use the convolutional network to extract the cascaded features, and then realize dynamic feature fusion through global average pooling and convolution prediction.
[0085] Below, each of the above steps will be introduced in detail.
[0086] Step 1: Ensure time alignment between the LiDAR and focal length camera based on hardware synchronization and software synchronization.
[0087] The time alignment method includes performing ptp time synchronization based on IEEE 1588v2.0 PTP network protocol synchronization and realizing hardware synchronization based on GPS clock source.
[0088] The software synchronization principle is to use the Delay request-response mechanism of IEEE 1588v2.0 PTP. By calculating the time length of the data packet interaction between the sensor device and the network clock, the transmission path delay and the clock offset between the two are calculated.
[0089] Hardware synchronization mainly relies on the timing device with GPS clock source. The device sends a hardware pulse signal once per second through the PPS port. The sensor device receives the rising edge of the signal and parses the correct time information from the GPRMC data. It then sets and maintains this time base for continuous accumulation to achieve time synchronization with the GPS device.
[0090] Step 2: Extract image features and point cloud features based on the focal length camera image data and the lidar sparse point cloud image.
[0091] The image feature extraction uses ResNet as the backbone to extract a set of attributes that characterize the characteristics or content of the image, expressed as , mainly including the captured vehicle model, vehicle color, and license plate number features.
[0092] The point cloud feature extraction is performed based on A-LOAM, first receiving the original point cloud data from the laser radar hardware, then preprocessing the original data, and then classifying the input laser points according to the laser lines; Calculate the curvature of each point on the laser line. represents a selected set of surrounding points, Indicates the current point, Indicates neighboring points.
[0093] If the c value is large, it means that the gap between the current point and the surrounding points is large, and the curvature is high, which represents a line point; if the c value is small, it means that the gap between the current point and the surrounding points is small, and the curvature is low, which represents a surface point; after calculating the curvature of each point on the laser line, the line points and surface points can be filtered according to the curvature; the back end transfers the line points and surface points of the previous and next frames to the same coordinate system and matches them according to the nearest neighbor search.
[0094] Step 3: Based on the original image data of the focal length camera, the depth information of the sparse point cloud image of the laser radar is added to perform depth estimation to obtain the depth features of the image data. The cone features are generated through the depth features and image features, and after pooling processing, the image BEV features are generated.
[0095] Step 3.1: Combine the depth information extracted from the LiDAR point cloud features with the image features: , where The operation represents merging the depth information in the image features and the point cloud features in the feature dimension, and D represents the depth information in the point cloud features.
[0096] Step 3.2: Pass a depth estimation network Realize the prediction of each pixel depth value, where is the depth estimation network operation, Represents the predicted depth map.
[0097] Step 3.3: Combine the depth map information to map the image features to three-dimensional space, and convert the image features of the focal camera perspective into a frustum representation in three-dimensional space: , where It is the cone feature, which means that the image features are placed in the three-dimensional space according to the predicted depth.
[0098] Step 3.4: After obtaining the frustum features, these features are aggregated into BEV representation through pooling operation: , where VoxelPool represents a pooling operation that converts the cone features into BEV representation.
[0099] Step 4: Input point cloud data, extract local features of point cloud through backbone network, obtain global features through maximum pooling operation, and guide lidar BEV feature prediction by image fusion.
[0100] Step 4.1: Input the point cloud data collected by the LiDAR device. Each point contains three-dimensional coordinate (x, y, z) information. Use the backbone network of the VoxelNet deep learning model to process the raw point cloud data, use the shared MLP to extract local features, and use the maximum pooling to obtain global features. Specifically, assume is the input point cloud, where is the number of points. Local feature extraction can be expressed as: , where Yes The global feature is obtained by pooling operation. get: , where Represents the global features of point cloud data.
[0101] Step 4.2: Concatenate and fuse the image features extracted in step 2 with the global features in step 4.1. It can be expressed as .
[0102] Step 4.3: After fusing the image features, convert the features into a bird’s-eye view (BEV) perspective for feature prediction. is the feature map converted to the BEV perspective, and the prediction network can be expressed as: , where are the predicted object location, size, and category.
[0103] Step 5: Construct a world coordinate system to align the focal camera image features with the lidar point cloud features.
[0104] Step 5.1: Camera imaging model description: Construct a world coordinate system to describe the position and posture of the focal length camera and lidar in three-dimensional space. For point P in the world coordinate system , the point in the camera coordinate system of the original focal length camera , then the corresponding relationship can be expressed as:
[0105] ;
[0106] In the formula, is the rotation matrix of the camera coordinate system relative to the world coordinate system, is the translation matrix of the camera coordinate system relative to the world coordinate system.
[0107] In contrast, points on the image plane Point in camera coordinate system with focal length The relationship is:
[0108] ;
[0109] In the formula, is the focal length camera intrinsic parameter matrix.
[0110] Step 5.2: LiDAR coordinate transformation: For a point in the LiDAR point cloud , let its coordinates in the world coordinate system be , then the coordinates in the focal length camera coordinate system can be expressed as:
[0111] ;
[0112] In the formula, is the rotation matrix The inverse matrix of .
[0113] Step 5.3: Feature matching and spatial alignment: Use SIFT algorithm for feature matching to find the points in the focal camera image plane Points in the LiDAR point cloud At the same time, the Bundle optimization algorithm is used to minimize the following errors to achieve spatial alignment of focal camera image features and lidar point cloud features:
[0114] ;
[0115] In the formula, and are the number of image feature points and points in the lidar point cloud, respectively.
[0116] Step 6: Cascade the point cloud and image BEV features according to the channel dimension, use the convolutional network to extract the cascaded features, and then realize dynamic feature fusion through global average pooling and convolution prediction.
[0117] Step 6.1: Concatenate the point cloud features and the bird's eye view (BEV) features of the focal length camera image in the channel dimension: The point cloud BEV features in step 4 are calculated as , the focal length camera image BEV feature is , the cascade operation can be expressed as:
[0118] ;
[0119] In the formula, , is the size of the cascaded feature map.
[0120] Step 6.2: Use a three-dimensional convolutional neural network (3DCNN) to further extract features from the cascaded features: , capturing both local features and spatial hierarchical structures.
[0121] Step 6.3: Perform global average pooling on the features extracted by the 3D convolutional network: , where is the number of channels after pooling, which reduces the number of parameters and prevents overfitting while retaining important feature information.
[0122] Step 6.4: Predict the pooled features through the convolution layer to achieve dynamic fusion of point cloud BEV features and image BEV features.
[0123] The method provided by the present invention can realize the effective fusion of lidar point cloud data and focal camera image data, and realize the intelligent perception and management of tunnels, which is very necessary for monitoring traffic operation conditions in highway tunnel environments.
[0124] An embodiment of the present invention also provides a device for monitoring traffic conditions in a highway tunnel environment, the device comprising a laser radar, a focal length camera and a data processing module; the laser radar is used to collect point cloud data; the focal length camera is used to collect image data; the data processing module is used to execute the steps of the above-mentioned method.
[0125] Specifically, see Figure 3 The device of this embodiment can be specifically composed of a laser radar 1, a focal length camera 2, a communication and data transmission module 3, a data fusion module 4, and a signal processing and analysis module 5.
[0126] The laser radar 1 is used to collect three-dimensional information of vehicles passing through the tunnel, including vehicle shape and vehicle position.
[0127] The focal length camera 2 is used to collect real-time image information of vehicles passing through the tunnel, including vehicle driving conditions, traffic conditions, vehicle color, and license plate number.
[0128] The signal communication and control module 3 is used to receive the perception data of the laser radar 1 and the focal length camera 2, and transmit the data to the data fusion module 4.
[0129] The data fusion module 4 is used to fuse the data collected by the laser radar 1 and the focal length camera 2 to achieve comprehensive perception of the characteristics of vehicles passing through the tunnel.
[0130] The signal processing and analysis module 5 processes and analyzes the fused data, extracts key information of vehicles passing through the tunnel, including vehicle lane position, vehicle speed, and analyzes the traffic conditions in the tunnel. Generates a virtual image of the identified vehicle, and binds the identified vehicle shape, vehicle color, and license plate number information. At the same time, the analysis results are shared to the tunnel operation management platform 6 to realize intelligent perception and management of the tunnel.
[0131] In actual application, the laser radar 1, the focal length camera 2, and the communication and data transmission module 3 can be fixed on a mounting bracket and fixed on the wall of the highway tunnel at a certain height from the ground to monitor the traffic conditions in the highway tunnel. The perception data is transmitted to the data fusion module 4 through the communication and data transmission module 3. The data fusion module 4 and the signal processing and analysis module 5 are the computing bodies of the device and are set in the data room to fuse the image data and the point cloud data of the laser radar to realize real-time analysis and calculation of the precise perception information of the vehicle, and complete the binding of the vehicle identity information and the perception results based on the unification of time and space, including the license plate number, vehicle model, license plate color, lane and other information.
[0132] The laser radar 1 and focal length camera 2 in this embodiment are perception units. The data fusion module 3 in this embodiment can be encapsulated in a housing and connected to the laser radar 1 and focal length camera 2 through a transmission optical cable to obtain perception data. The laser radar 1, focal length camera 2, housing and the data fusion module 3 encapsulated therein can be uniformly fixed on a mounting bracket, and the mounting bracket has a fixed part, so that the device can be fixedly installed on the wall of a highway tunnel. The data fusion module 4 and the signal processing and analysis module 5 are relatively independent of other parts of the device. The two together constitute the computing module of the device, which is connected to the front-end device of the device through a network, and is used to fuse the calculation and analysis of the returned perception data, and finally share the fused perception results to the tunnel operation management platform 6 to realize the intelligent perception and management of the tunnel.
[0133] In addition, it should be noted that the laser radar in this embodiment does not specify a specific type of laser radar, and can include all laser radar perception sensors that can generate point cloud data, including but not limited to mechanical laser radar, semi-solid laser radar and solid-state laser radar.
[0134] The focal length camera in this embodiment is only one type of image sensing sensor, and alternative solutions include but are not limited to using a high-definition network camera, a bayonet camera, etc.
[0135] The installation method in this embodiment is to fix it on the road tunnel wall through a mounting bracket. Alternative installation methods include but are not limited to hanging installation and guide rail installation.
[0136] The perception application scenario in this embodiment is within a highway tunnel. Alternative application scenario solutions include, but are not limited to, perception systems for tunnel entrances and exits, perception systems for connection areas between tunnels, etc.
[0137] This embodiment breaks through the limitation of using a single perception sensor for traffic condition perception in the traditional highway tunnel perception field, and uses laser radar and focal length camera for highway tunnel traffic condition perception, effectively combining the advantages of different sensors. At the same time, this embodiment can first transmit the perception data of the perception sensor, and then perform data fusion analysis, so even if a single perception sensor is damaged due to some factors, it will not fail and completely lose its perception ability, and can still rely on the undamaged perception sensor to perceive the traffic condition based on a single data source.
[0138] The above-mentioned device provided in this embodiment belongs to the same concept as the above-mentioned method embodiment. The specific implementation process thereof is detailed in the method embodiment and will not be repeated here.
[0139] In order to implement the method of the embodiment of the present invention, the embodiment of the present invention also provides a computer program product, which includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the steps of the above method.
[0140] Based on the hardware implementation of the above program modules, and in order to implement the method of the embodiment of the present invention, the embodiment of the present invention further provides an electronic device (computer device). Specifically, in one embodiment, the computer device may be a terminal, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor A01, a network interface A02, a display screen A04, an input device A05 and a memory (not shown in the figure) connected through a system bus. Among them, the processor A01 of the computer device is used to provide computing and control capabilities. The memory of the computer device includes an internal memory A03 and a non-volatile storage medium A06. The non-volatile storage medium A06 stores an operating system B01 and a computer program B02. The internal memory A03 provides an environment for the operation of the operating system B01 and the computer program B02 in the non-volatile storage medium A06. The network interface A02 of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor A01, the method of any one of the above embodiments is implemented. The display screen A04 of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device A05 of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.
[0141] Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0142] The device provided by the embodiment of the present invention includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, the method of any one of the above embodiments is implemented.
[0143] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0144] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0145] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0147] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0148] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0149] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0150] It can be understood that the memory of the embodiment of the present invention can be a volatile memory or a non-volatile memory, and can also include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM, SyncLink Dynamic Random Access Memory), and direct RAMbus random access memory (DRRAM, Direct Rambus Random Access Memory).The memories described in the embodiments of the present invention are intended to include, but are not limited to, these and any other suitable types of memories.
[0151] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0152] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.
Claims
1. A method for monitoring traffic conditions in a highway tunnel environment, characterized in that: The method comprises: Acquire image data collected by a focal length camera and point cloud data collected by a laser radar; the focal length camera and the laser radar are both located on the wall of a highway tunnel, and time alignment of the collected data is achieved based on hardware synchronization and software synchronization technology; Extracting image features from the image data, filtering feature points of the point cloud data according to curvature, and matching feature points of previous and next frames; Combining the image features with the depth information in the matched point cloud data, and performing depth estimation, frustum representation and pooling operations on the combined features to obtain image BEV features; Performing local feature extraction and maximum pooling operations on the point cloud data in the point cloud image data, and fusing image features to obtain laser radar BEV feature prediction; Constructing a world coordinate system, in which the image BEV feature and the laser radar BEV feature prediction are cascaded, feature extracted, globally averaged pooled and convolutionally predicted to obtain a dynamic fusion feature; the dynamic fusion feature reflects the comprehensive perception of the characteristics of the vehicles passing through the highway tunnel; Among them, obtaining dynamic fusion features includes: The image BEV feature is concatenated with the lidar BEV feature prediction using the following calculation formula: in, represents the features after cascading, is the point cloud BEV feature, is the image BEV feature, , is the size of the cascaded feature map, Indicates a cascade operation; Use the following calculation formula to extract the features after cascading to obtain the extracted features; in, represents the extracted features, represents a three-dimensional convolutional neural network; Use the following calculation formula to perform global average pooling on the extracted features to obtain the pooled features: in, represents the features after pooling, is the number of channels after pooling; The pooled features are predicted through the convolution layer to obtain dynamic fusion features, thereby realizing the dynamic fusion of point cloud BEV features and image BEV features.
2. The method for monitoring traffic conditions in a highway tunnel environment according to claim 1, characterized in that: Extracting image features from the image data includes: Image features are extracted from the image data using a first model; the first model is a ResNet network model; the image features are used to characterize vehicle model, vehicle color and license plate number features in the image data.
3. The method for monitoring traffic conditions in a highway tunnel environment according to claim 1, characterized in that: The point cloud data is filtered for feature points according to the curvature, and feature points of previous and next frames are matched, including: Preprocessing the original point cloud in the point cloud data, classifying the original point cloud according to the laser line, and calculating the curvature of each point on the laser line; According to the curvature of each point on the laser line, the line point and the surface point are determined; The line points and surface points of the previous and next frames are transferred to the same coordinate system and matched according to the nearest neighbor search; The curvature of each point on the laser line is calculated using the following formula: in, represents the curvature, represents a selected set of surrounding points, Indicates the current point, Indicates neighboring points.
4. The method for monitoring traffic conditions in a highway tunnel environment according to claim 1, characterized in that: The image features are combined with the depth information in the matched point cloud data, and the combined features are subjected to depth estimation, cone representation and pooling operations to obtain image BEV features, including: The image features are combined with the depth information in the matched point cloud data using the following calculation formula: in, Represents the combined features, Indicates that image features and depth information are combined in the feature dimension. Represents depth information, Represents image features; The combined features are used to estimate depth using the following calculation formula: in, is the depth estimation network operation, represents the predicted depth map; The depth map and image features are mapped to three-dimensional space using the following calculation formula to obtain the frustum features in three-dimensional space: in, The cone feature means placing the image features in three-dimensional space according to the predicted depth; The cone features are pooled using the following calculation formula to obtain the image BEV features: in, represents a pooling operation that converts the cone feature into a BEV representation. Represents the BEV feature of the image.
5. The method for monitoring traffic conditions in a highway tunnel environment according to claim 1, characterized in that: Performing local feature extraction and maximum pooling operations on the point cloud data in the point cloud map data and fusing image features to obtain laser radar BEV feature prediction, including: The local features of the point cloud data in the point cloud data are extracted using the following calculation formula to obtain the local features: in, Yes The local eigenvector of , is the input point cloud, is the number of points, For point Contains 3D coordinate information, and MLP is the backbone network of the VoxelNet deep learning model; The local features are subjected to maximum pooling operation using the following calculation formula to obtain global features: in, represents the pooling operation, Represents global features; The global feature and the image feature are concatenated and fused using the following calculation formula to obtain the fused feature: in, represents the fused features, Represents image features; Represents a splicing and fusion operation; The fused features are converted into a bird's-eye view using the following calculation formula, and feature prediction is performed to obtain the lidar BEV feature prediction: in, It is the lidar BEV feature prediction, including the predicted target location, size and category; It is the feature map converted from the fused features to the bird's-eye view; Represents the prediction network.
6. The method for monitoring traffic conditions in a highway tunnel environment according to claim 1, characterized in that: Construct the world coordinate system, including: Construct the correspondence between the points in the world coordinate system and the points in the focal camera coordinate system: in, are the coordinates of the point in the focal camera coordinate system; is the coordinate of the point in the world coordinate system; is the rotation matrix of the focal camera coordinate system relative to the world coordinate system, is the translation matrix of the focal length camera coordinate system relative to the world coordinate system; Construct the correspondence between points on the image plane and points in the focal camera coordinate system: in, are the coordinates of the point on the image plane; is the focal length camera internal parameter matrix; Determine the coordinates of points in the point cloud data in the world coordinate system: in, is the coordinate of the point in the point cloud data in the world coordinate system, is the rotation matrix The inverse matrix of is a point in the point cloud data; The correspondence between the points in the focal camera image plane and the points in the LiDAR point cloud data is constructed, and based on the correspondence, the Bundle optimization algorithm is used to minimize the following errors to achieve spatial alignment of the focal camera image features and the LiDAR point cloud features: in, and are the number of image feature points and points in the lidar point cloud, respectively.
7. A road tunnel environment traffic condition monitoring device, characterized in that: The device includes a laser radar, a focal length camera and a data processing module; The laser radar is used to collect point cloud data; The focal length camera is used to collect image data; The data processing module is used to execute the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-mode BEV look-around perception method, device and equipment and storage medium
CN117392625A
Fusion-based object tracker using lidar point cloud and surrounding cameras for autonomous vehicles
US20230334673A1