Monocular vehicle positioning method and device combined with vehicle type identification technology

By combining vehicle model recognition and attitude estimation technologies, and utilizing a method that integrates deep learning and geometric reasoning, the problem of vehicle model differences and attitude angle influences in monocular vehicle localization methods is solved, achieving high-precision and low-cost vehicle localization, which is suitable for intelligent transportation systems.

CN121236716APending Publication Date: 2025-12-30KUN SHAN YUAN LE FU KE JI YOU XIAN GONG SI +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511453246.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2025-12-30

AI Technical Summary

Technical Problem

Existing monocular vehicle localization methods suffer from significant ranging errors when dealing with different vehicle models, ignore the impact of vehicle attitude angles on positioning accuracy, and lack effective utilization of vehicle type information, resulting in insufficient positioning accuracy.

Method used

By introducing vehicle model recognition technology and using a multi-branch directional regression network to obtain vehicle heading angle information, combined with vehicle model recognition technology and vehicle attitude angle, deep learning and a multi-branch directional regression network are used to obtain accurate vehicle positioning through vehicle model recognition. By using deep learning and a multi-branch directional regression network, and through vehicle positioning methods and technologies, combined with vehicle model recognition, accurate prior information on vehicle size is obtained, and combined with vehicle attitude angle information, the accuracy of monocular positioning is improved at low cost.

Benefits of technology

It achieves high-precision monocular vehicle positioning, reduces system costs, enhances attitude adaptability, and expands the application scope, making it suitable for scenarios such as intelligent traffic monitoring, autonomous driving assistance, and traffic flow analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236716A_ABST
    Figure CN121236716A_ABST
Patent Text Reader

Abstract

The invention relates to a low-cost monocular vehicle positioning method and device combined with a vehicle type recognition technology, and the method is constructed based on a pure vision algorithm, and specifically comprises the steps: obtaining a road traffic image collected by a monocular camera; performing feature extraction on the road traffic image to obtain visual features of the vehicle; the vehicle type is recognized according to the visual features, meanwhile, a vehicle attitude angle is estimated based on the visual features, and the vehicle attitude angle is a vehicle course angle and is obtained through a multi-branch direction regression network; according to the identified vehicle type, extracting size prior information of the vehicle type; and calculating position coordinate information of the vehicle in a ground coordinate system in combination with the vehicle attitude angle and size prior information. Compared with the prior art, the method has the advantages of high monocular positioning precision, low cost and the like by introducing the vehicle type recognition and attitude estimation technology and utilizing size prior information and vehicle attitude angle information of different vehicle types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent traffic control system technology, and in particular to a vehicle positioning method based on a monocular camera. Specifically, it is a monocular vehicle positioning method and device that combines vehicle model recognition technology, applicable to application scenarios such as unmanned intelligent agents, intelligent traffic monitoring, autonomous driving assistance systems, and traffic flow analysis. Background Technology

[0002] With the rapid development of intelligent transportation systems, accurate vehicle positioning technology has become a key technological foundation for applications such as traffic monitoring, autonomous driving, and intelligent parking. Traditional vehicle positioning methods mainly include GPS-based positioning, LiDAR-based positioning, and vision-based positioning. Among them, vision-based positioning methods have received widespread attention due to their advantages such as low cost and convenient deployment. Existing vision-based vehicle positioning technologies are mainly divided into two categories: binocular vision positioning and monocular vision positioning. Although binocular vision systems can obtain relatively accurate depth information through stereo matching, they require precise calibration and synchronization of two cameras, resulting in high system complexity, high cost, and strict requirements for the installation environment. In contrast, monocular vision positioning systems have advantages such as simple structure, low cost, and ease of deployment, making them more advantageous in practical applications.

[0003] However, existing monocular vehicle localization methods suffer from the following technical problems: First, traditional monocular localization methods typically assume that all vehicles have the same size parameters. This simplistic assumption leads to significant ranging errors when dealing with different vehicle types (such as sedans, SUVs, and trucks). Since different vehicle types vary significantly in length, width, and height, using uniform vehicle size prior information cannot meet the accuracy requirements of practical applications. Second, existing methods often neglect the impact of vehicle attitude angles on localization accuracy. In real-world traffic scenarios, the heading angle of the vehicle relative to the camera changes, especially when the vehicle is turning or changing lanes. This alters the actual length projection of the vehicle, directly affecting the accuracy of distance calculations based on geometric models. Third, most existing monocular localization algorithms rely on simple geometric relationships for distance estimation, lacking effective utilization of vehicle type information and failing to adaptively adjust to the characteristics of different vehicle types, thus limiting further improvements in localization accuracy.

[0004] A search revealed that the existing patent application CN110082779 A discloses a vehicle pose localization method and system based on 3D LiDAR. Although it determines the vehicle model, it still suffers from high costs due to the need to combine data obtained from 3D LiDAR and a monocular camera.

[0005] Therefore, there is an urgent need for a monocular vehicle localization method that can combine vehicle model recognition information, take into account the influence of vehicle posture, and has acceptable accuracy and low cost, in order to meet the requirements of intelligent transportation systems for vehicle localization accuracy and reliability. Summary of the Invention

[0006] The purpose of this invention is to provide a monocular vehicle localization method and apparatus that combines vehicle model recognition technology. By introducing vehicle model recognition and attitude estimation technology, and utilizing prior information on the size of different vehicle models and vehicle attitude angle information, the monocular localization accuracy can be improved at low cost.

[0007] The objective of this invention can be achieved through the following technical solutions: A monocular vehicle localization method incorporating vehicle model recognition technology includes the following steps: Acquire road traffic images captured by a monocular camera; Feature extraction is performed on the road traffic images to obtain the visual features of the vehicles; The vehicle model is identified based on the visual features, and the vehicle attitude angle is estimated based on the visual features. The vehicle attitude angle is the vehicle heading angle, which is obtained using a multi-branch directional regression network. Based on the identified vehicle model, extract the prior information on the size of that model; Based on the prior information of the vehicle's attitude angle and size, the vehicle's position coordinates in the ground coordinate system are calculated.

[0008] Furthermore, the visual features are extracted using a multi-scale convolutional neural network.

[0009] Furthermore, the vehicle model is identified using an SSD network.

[0010] Furthermore, the training datasets used when training the SSD network include the BIT-Vehicle dataset and the Peking University Vehicle Dataset.

[0011] Furthermore, based on the identified vehicle model, the prior size information of that model is extracted from the model database, which stores structured size information.

[0012] Furthermore, based on the similar triangle relationship, combined with the intrinsic and extrinsic parameters of the monocular camera and the prior information on the vehicle's attitude angle and size, the position coordinate information is calculated.

[0013] Furthermore, the process of calculating and obtaining the location coordinate information specifically includes: Let the camera position be The point, whose projection on the ground is Point, the projection of the lower edge of the vehicle onto the imaging plane is Point, calculate the optical axis of the camera based on the camera's focal length. The angle between the lines : Based on the position of the lower edge of the vehicle and the included angle In addition to the camera's installation height and tilt angle, the longitudinal distance between the vehicle and the camera's projection point is calculated. By combining the vehicle posture angle and vehicle length, the effective projection of the vehicle length in the longitudinal direction is calculated, and the longitudinal coordinates of the vehicle center relative to the camera projection point are obtained. By using the lateral projection position of the vehicle image center on the imaging plane, combined with the camera focal length and pitch angle, the lateral coordinates of the vehicle center relative to the camera projection point can be calculated. Complete the vehicle's positioning in the ground coordinate system.

[0014] Furthermore, the calculation of the longitudinal distance between the vehicle and the camera projection point is specifically as follows: When the lower edge of the vehicle is located in the lower half of the image, the image is determined by the camera height, pitch angle, and included angle. The tangent of the difference is obtained by calculation; When the lower edge of the vehicle is in the upper half of the image, the camera height, pitch angle, and included angle are used to determine this. The tangent of the sum is obtained by calculation.

[0015] The present invention also provides a computer-readable storage medium including one or more programs executable by one or more processors of an electronic device, said one or more programs including instructions for performing the monocular vehicle localization method combined with vehicle model recognition technology as described above.

[0016] The present invention also provides a monocular vehicle positioning device incorporating vehicle model recognition technology, comprising: The image acquisition module is used to acquire road traffic images captured by a monocular camera; The feature extraction module is used to extract features from the road traffic image to obtain the visual features of the vehicles; A vehicle model recognition module is used to identify the vehicle model based on the visual features. The attitude estimation module is used to estimate the vehicle attitude angle based on the visual features, wherein the vehicle attitude angle is the vehicle heading angle, which is obtained using a multi-branch directional regression network. The size query module is used to extract prior size information of the identified vehicle model. The position calculation module is used to calculate the vehicle's position coordinates in the ground coordinate system by combining the vehicle's attitude angle and size prior information.

[0017] Compared with the prior art, the present invention has the following beneficial effects: (1) Provide sufficient positioning accuracy: By combining vehicle model recognition technology, the true size prior information of different vehicle models is used to replace the traditional assumption of uniform vehicle size, thus ensuring the accuracy of monocular vision positioning.

[0018] (2) Enhance attitude adaptability: Vehicle attitude angle estimation is introduced, and vehicle heading angle information is obtained through a multi-branch directional regression network. This effectively corrects the impact of vehicle attitude changes on distance measurement and improves the positioning reliability in scenarios such as vehicle turning and lane changing.

[0019] (3) Optimize computation efficiency: The SSD network structure is used for vehicle model recognition, combined with an efficient vehicle model database retrieval mechanism, to achieve accurate positioning under real-time requirements and meet the needs of practical applications.

[0020] (4) Reduced system cost: Compared with binocular vision system, the present invention only requires a single camera to achieve high-precision positioning. At the same time, the method of the present invention is based on pure vision algorithm, which significantly reduces hardware cost and system complexity, and facilitates large-scale deployment and application.

[0021] (5) Expanding the scope of application: The positioning method provided by this invention is applicable to various application scenarios such as intelligent traffic monitoring, autonomous driving assistance, and traffic flow analysis, and has good application prospects and promotional value. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the method architecture of the present invention; Figure 2 This is a schematic diagram of a ranging algorithm that combines vehicle size information and orientation information in an embodiment of the present invention. Detailed Implementation

[0023] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0024] Terminology Definition Monocular Vehicle Localization: A technique that uses image information acquired by a single camera to determine the vehicle's position coordinates in three-dimensional space through computer vision and geometric reasoning methods.

[0025] Vehicle Model Recognition: The process of identifying a vehicle's specific brand, model, and other classification information by analyzing its visual characteristics.

[0026] Pose Angle: The directional angle of the vehicle relative to the camera coordinate system. In this invention, it specifically refers to the heading angle of the vehicle, which is the angle between the long axis of the vehicle and the projection of the camera's optical axis onto the ground.

[0027] Size Prior Information: Standard three-dimensional dimensional parameters of various vehicle models stored in a pre-established vehicle model database, including structured information such as length, width, and height.

[0028] Ground Coordinate System: A two-dimensional coordinate system established with the camera's projection point on the ground as the origin and the projection direction of the camera's optical axis on the ground as the positive X-axis.

[0029] This embodiment provides a monocular vehicle localization method that combines vehicle model recognition technology. It adopts a technical approach that combines deep learning and geometric reasoning. It obtains accurate prior information on vehicle size through vehicle model recognition and corrects geometric calculation errors by combining attitude estimation, thereby achieving high-precision monocular vehicle localization.

[0030] like Figure 1 As shown, the monocular vehicle localization method in this embodiment includes the following steps: First, image acquisition. Road traffic images are acquired using a monocular camera positioned above the road. During system initialization, a calibration program is needed to obtain the camera's intrinsic and extrinsic parameters (including focal length f, principal point coordinates, etc.), which will be used for subsequent geometric calculations.

[0031] Second, preliminary feature extraction. A convolutional neural network is used to extract features from the image to obtain the vehicle's visual feature information.

[0032] Specifically, the input image is first preprocessed, including size normalization. Then, a convolutional neural network is used as the backbone network for feature extraction. Through multiple layers of convolution, pooling, and activation operations, high-dimensional feature vectors containing information such as vehicle shape, texture, and color are extracted from the original image. The output of the feature extraction network is a multi-scale feature map, which is used for subsequent vehicle type recognition and pose estimation tasks.

[0033] Third, vehicle model recognition and attitude estimation. The visual features are input into two sub-networks, one for identifying vehicle model and the other for estimating vehicle attitude angles.

[0034] For the vehicle model recognition sub-network, the SSD (Single Shot MultiBox Detector) architecture is adopted. The SSD network can simultaneously complete vehicle detection and classification tasks in a single forward propagation, exhibiting good real-time performance. The network structure includes: a basic feature extraction layer, a multi-scale feature fusion layer, a classification prediction layer, and a bounding box regression layer. The classification prediction layer outputs the probability distribution of a vehicle belonging to each vehicle model category, and the final vehicle model recognition result is obtained through a softmax activation function. The training dataset includes the BIT-Vehicle dataset and the Peking University Vehicle Dataset, containing over 50,000 labeled images of different vehicle models, covering major vehicle categories such as sedans, SUVs, and trucks, to enhance recognition capabilities in multi-view scenarios.

[0035] For the attitude estimation sub-network, the vehicle attitude angle is the vehicle heading angle, obtained through a multi-bin network. The multi-bin network transforms the continuous angle regression problem into a discrete classification problem combined with fine-grained regression. Specifically, the 360-degree angle range is first divided into several bins (e.g., 12, each covering 30 degrees), with each bin corresponding to a classification branch. Then, fine-grained angle regression is performed within the defined bins to obtain accurate angle values. This method improves the accuracy and stability of angle estimation and can correct the angle between the vehicle's length direction and the camera's imaging axis, thereby improving the accuracy of longitudinal distance estimation.

[0036] Fourth, prior dimension information extraction. Based on the identified vehicle model, the system extracts the prior dimension information of that model from the vehicle model database. The vehicle model database includes vehicle brand, model, and structured dimension information such as length, width, and height, supporting the retrieval of corresponding prior dimension information by vehicle model. The system retrieves the corresponding dimension parameters, including vehicle length, based on the vehicle model category identified in the vehicle model recognition subnetwork, through querying. L ,width W This is the key information.

[0037] Fifth, position coordinate calculation. Combining the vehicle's attitude information and prior dimensional information, the vehicle's position coordinates in the ground coordinate system are calculated.

[0038] First, a distance calculation model based on the relationship of similar triangles is established, that is, using... Figure 2 (b) in The similarity relationship between two triangles connected by points guarantees the proportional relationship of their side lengths. Implement subsequent distance calculations. For example... Figure 2 As shown, let the camera position be... The point, whose projection on the ground is Point, the projection of the lower edge of the vehicle onto the imaging plane is Point, camera optical axis and The angle between the lines is Based on the camera's internal parameters and the vehicle's position in the image, the angle... The calculation formula is: in Projection point of the lower edge of the vehicle To the image center pixel distance, This refers to the camera's focal length.

[0039] The similar triangle relationship is constructed using the vehicle's three-dimensional dimensions and the camera's geometric model, combined with the camera's installation height. ,focal length Pitch angle The parameters are used for calculation. When the lower edge of the vehicle is located in the lower half of the image, the longitudinal distance between the vehicle and the camera projection point is calculated by the tangent of the difference between the camera height and the angle between them. When the vehicle edge is located in the upper half of the image, the sum of the included angles is used for calculation: Combined with the vehicle attitude angle obtained in step three And the vehicle length obtained in step four Calculate the effective projection of the vehicle length in the longitudinal direction: Therefore, the longitudinal coordinate of the vehicle center relative to the camera projection point is: The lateral coordinate is calculated by the lateral projection position of the vehicle image center onto the imaging plane. Combined with camera focal length With pitch angle The vehicle's lateral coordinates are then calculated. Finally, the relative two-dimensional plane coordinates of the vehicle in the ground coordinate system are obtained. This completes the vehicle's positioning in the ground coordinate system, and can also be achieved using the vehicle length obtained in step four. L and width W Further determine the actual two-dimensional collision range of the vehicle.

[0040] The above method can be applied to unmanned intelligent agents to more accurately locate vehicle intelligent agents and improve the reliability of independently completing specific tasks.

[0041] If the above methods are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0042] The experimental comparison and analysis based on the above methods are as follows: First, vehicle model recognition performance evaluation. The method of this invention was tested on a test dataset containing 13 common vehicle models, and the results are shown in Table 1. The overall average accuracy reached 67.92%, with most models achieving an accuracy exceeding 70%. Specifically, the Volkswagen Tiguan, Volkswagen Gol, and BMW 3 Series achieved 100% accuracy. A few models, such as the Audi A6, had lower recognition accuracy, mainly due to insufficient training samples and high visual similarity to other models (such as the Audi A4). The experiments demonstrate that the vehicle model recognition method of this invention can provide reliable vehicle classification results in most cases.

[0043] Table 1. Vehicle model recognition experiment results Second, a comparative analysis of distance measurement accuracy. This invention was compared with existing IPP (Inverse Perspective Projection) and SST (Similar Triangles) methods in a test range of 0-100 meters. The results are shown in Table 2. In the close-range range of 0-30 meters, the maximum error of this invention's method was 1.80 meters, with an average relative error of 2.93%. In the long-range range of 30-100 meters, the maximum error was 2.74 meters, with an average relative error of 2.40%. Compared with the SST method, this invention significantly reduces the distance measurement error, verifying the effectiveness of using prior information about vehicle dimensions for positioning. Although slightly lower in accuracy than the IPP method, the IPP method requires precise edge detection and a data fitting process specific to the environment, resulting in limited generalization ability. In contrast, this invention's method has better applicability and robustness. Table 2 Comparison of results from different distance measurement methods within the range of 0~100m In another embodiment, a monocular vehicle positioning device incorporating vehicle model recognition technology is also provided, comprising an image acquisition module, a feature extraction module, a vehicle model recognition module, a pose estimation module, a size query module, and a position calculation module, wherein: the image acquisition module is used to acquire road traffic images captured by a monocular camera; the feature extraction module is used to extract features from the road traffic images to obtain the visual features of the vehicle; the vehicle model recognition module is used to identify the vehicle model based on the visual features; the pose estimation module is used to estimate the vehicle pose angle based on the visual features; the size query module is used to extract prior size information of the identified vehicle model; and the position calculation module is used to calculate the vehicle's position coordinates in a ground coordinate system by combining the vehicle pose angle and the prior size information.

[0044] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0045] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A monocular vehicle positioning method combined with vehicle model recognition technology, characterized in that, The method comprises the following steps: acquiring a road traffic image collected by a monocular camera; extracting features of the road traffic image to obtain visual features of a vehicle; identifying a vehicle model according to the visual features, and estimating a vehicle attitude angle based on the visual features, wherein the vehicle attitude angle is a vehicle heading angle, and the vehicle attitude angle is obtained by using a multi-branch direction regression network; extracting size prior information of the vehicle model according to the identified vehicle model; combining the vehicle attitude angle and the size prior information to calculate position coordinate information of the vehicle in a ground coordinate system.

2. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 1, characterized in that, The visual features are extracted by using a multi-scale convolutional neural network.

3. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 1, characterized in that, The vehicle model is identified by using an SSD network.

4. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 3, characterized in that, A training data set used when the SSD network is trained comprises a BIT-Vehicle data set and a Peking University Vehicle Dataset.

5. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 1, wherein, The size prior information of the vehicle model is extracted from a vehicle model database according to the identified vehicle model, and the vehicle model database stores structured size information.

6. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 1, wherein, The position coordinate information is calculated according to a similar triangle relationship, in combination with intrinsic and extrinsic parameters of the monocular camera and the vehicle attitude angle and the size prior information.

7. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 6, wherein, The process of calculating the position coordinate information specifically comprises: Let the camera position be The point, whose projection on the ground is Point, the projection of the lower edge of the vehicle onto the imaging plane is Point, calculate the optical axis of the camera based on the camera's focal length. The angle between the lines : According to the vehicle lower edge position, the included angle And the installation height and the pitch angle of the camera, the longitudinal distance of the vehicle from the camera projection point is calculated; combining the vehicle attitude angle and the length of the vehicle model to calculate an effective projection of the length of the vehicle in a longitudinal direction, and obtaining a longitudinal coordinate of a center of the vehicle relative to a projection point of the camera; obtaining a transverse coordinate of the center of the vehicle relative to the projection point of the camera by converting the transverse projection position of the center of the vehicle in an imaging plane in combination with a focal length and a pitch angle of the camera; completing positioning of the vehicle in the ground coordinate system.

8. The monocular vehicle positioning method incorporating vehicle model recognition technology according to claim 7, wherein, The calculation of the longitudinal distance of the vehicle from the projection point of the camera specifically comprises: When the vehicle lower edge is in the lower half of the image, the difference between the camera height and the pitch angle and the included angle is calculated by the tangent value When the vehicle lower edge is in the upper half of the image, it is calculated from the tangent of the sum of the camera height and the pitch angle and the included angle .

9. A computer-readable storage medium, characterized in that, one or more programs for execution by one or more processors of an electronic device, the one or more programs comprising instructions for performing a monocular vehicle positioning method in combination with a vehicle model identification technique according to any one of claims 1-8.

10. A monocular vehicle positioning device incorporating a vehicle model recognition technique, characterized by, comprise: an image acquisition module configured to acquire a road traffic image collected by a monocular camera; a feature extraction module configured to extract features of the road traffic image to obtain visual features of a vehicle; a vehicle model identification module configured to identify a vehicle model according to the visual features; an attitude estimation module configured to estimate a vehicle attitude angle based on the visual features, wherein the vehicle attitude angle is a vehicle heading angle, and the vehicle attitude angle is obtained by using a multi-branch direction regression network; a size query module configured to extract size prior information of the vehicle model according to the identified vehicle model; a position calculation module configured to combine the vehicle attitude angle and the size prior information to calculate position coordinate information of the vehicle in a ground coordinate system.

Citation Information

Patent Citations

  • Vehicle attitude positioning method and system based on 3D laser radar

    CN110082779A