A method, device and equipment for generating road signs

By optimizing the timestamp, vehicle geographical location and camera heading angle, matching feature points to calculate the world geographical coordinates of street sign corner points, the problem of inaccurate street sign positions in high-precision maps is solved, and the accuracy and user experience of street sign positioning are improved.

CN114820783BActive Publication Date: 2025-06-20ZHIDAO NETWORK TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210359554.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-07
Publication Date
2025-06-20
Estimated Expiration
2042-04-07

AI Technical Summary

Technical Problem

In high-precision maps, traditional street sign collection methods cannot accurately locate the location of street signs, which affects the identification of roadside signs by high-speed vehicles.

Method used

By obtaining multi-frame target images and marking timestamps, the vehicle geographic location and camera heading angle are optimized, feature points are matched to obtain the optimized three-dimensional coordinates of street sign corner points, and the world geographic coordinates of street sign corner points are calculated based on camera parameters.

Benefits of technology

It improves the accuracy of the display position of street signs in high-precision maps, reduces measurement errors, and improves users' experience of using high-precision maps.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114820783B_ABST
    Figure CN114820783B_ABST
Patent Text Reader

Abstract

The present application relates to a road sign generation method, apparatus, and device. The method includes: obtaining multiple frames of target images and respectively marking timestamps for the multiple frames of target images, where each target image includes a target road sign, and the timestamp is a preset duration ahead of the corresponding image capture time; determining the camera heading angle corresponding to each target image according to the optimized geographical location of the vehicle corresponding to each timestamp; matching the feature points of two adjacent frames of target images to obtain the optimized three-dimensional coordinates of the corner points of the target road sign in the target image in the camera coordinate system; and determining the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle. The solution provided by the present application can improve the display position accuracy of road signs in a high-precision map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular, to a method, apparatus, and device for generating road signs. Background Art

[0002] Compared with traditional electronic maps, high-precision maps can simulate and display real road feature information such as lanes, guardrails, traffic signs, etc., and provide lane-level navigation.

[0003] In related technologies, various map surveying and mapping companies collect road feature information by using special surveying vehicles. These surveying vehicles usually use lidar technology to collect various road feature information, so as to create high-precision maps. However, taking road signs in traffic signs as an example, when collecting road signs, the traditional collection method can only accurately locate the position information of road signs to within a few meters, and cannot accurately locate the position of road signs in high-precision maps. Although such maps can also be used for navigation, since the position of the road signs displayed in the map is not accurate enough, this is not conducive to vehicles traveling at high speeds to identify the roadside road signs in a timely and accurate manner. Summary of the Invention

[0004] To solve or partially solve the problems existing in the related technologies, this application provides a method, apparatus, and device for generating road signs, which can improve the display position accuracy of road signs in high-precision maps.

[0005] The first aspect of this application provides a method for generating road signs, including:

[0006] Obtain multiple frames of target images and mark time stamps for each of the multiple frames of target images. Among them, each target image contains a target road sign, and the time stamp is a preset duration earlier than the corresponding image shooting time;

[0007] Determine the camera heading angle corresponding to each target image according to the optimized geographical position of the vehicle corresponding to each time stamp;

[0008] Match the feature points of two adjacent frames of the target images to obtain the optimized three-dimensional coordinates of the corner points of the target road sign in the camera coordinate system;

[0009] Determine the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical position of the vehicle, and the camera heading angle.

[0010] The second aspect of this application provides a road sign generating apparatus, which includes:

[0011] An obtaining module, configured to obtain multiple frames of target images and mark time stamps for each of the multiple frames of target images. Among them, each target image contains a target road sign, and the time stamp is a preset duration earlier than the corresponding image shooting time;

[0012] An angle determination module, configured to determine a camera heading angle corresponding to each of the target images according to the optimized geographical location of the vehicle corresponding to each of the timestamps.

[0013] A coordinate optimization module, configured to match feature points of two adjacent frames of the target images, and obtain optimized three-dimensional coordinates of corner points of a target road sign in the target images in a camera coordinate system.

[0014] A processing module, configured to determine world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

[0015] A third aspect of the present application provides an electronic device, including:

[0016] A processor; and

[0017] A memory, storing executable code thereon, which when executed by the processor, causes the processor to execute the method as described above.

[0018] A fourth aspect of the present application provides a computer-readable storage medium, storing executable code thereon, which when executed by a processor of an electronic device, causes the processor to execute the method as described above.

[0019] The technical solution provided by the present application may include the following beneficial effects:

[0020] The technical solution provided by the present application optimizes the timestamp, the optimized geographical location of the vehicle, the camera heading angle, and the optimized three-dimensional coordinates of the corner points associated with the target images respectively in different links, thereby gradually reducing the measurement error, making the total cumulative error smaller, and then improving the accuracy of the world geographical coordinates finally obtained by the corner points, so that the road sign can be positioned and displayed in the high-precision map according to the more accurate world geographical coordinates, improving the user experience of using the high-precision map.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] By describing the exemplary embodiments of the present application in more detail in conjunction with the drawings, the above and other objects, features, and advantages of the present application will become more obvious, wherein, in the exemplary embodiments of the present application, the same reference numerals generally represent the same components.

[0023] Figure 1 is a schematic flowchart of a road sign generation method shown in an embodiment of the present application;

[0024] Figure 2is a side view of a vehicle shown in an embodiment of the present application;

[0025] Figure 3 is Figure 2 a top view of the relative positions of the GPS positioning device and the camera in the vehicle body in

[0026] Figure 4 is another schematic flow chart of a road sign generation method shown in an embodiment of the present application;

[0027] Figure 5 is a schematic diagram for calculating the camera heading angle in an embodiment of the present application;

[0028] Figure 6 is a schematic diagram showing the correspondence between the absolute pose, feature points, and reprojection error shown in an embodiment of the present application;

[0029] Figure 7 is a schematic diagram for calculating the world geographical coordinates of corner points shown in an embodiment of the present application;

[0030] Figure 8 is a schematic structural diagram of a road sign generation device shown in an embodiment of the present application;

[0031] Figure 9 is a schematic structural diagram of an electronic device shown in an embodiment of the present application. Detailed Embodiments

[0032] The embodiments of the present application will be described in more detail below with reference to the accompanying drawings. Although the embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided to make the present application more thorough and complete, and to fully convey the scope of the present application to those skilled in the art.

[0033] The terms used in the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0034] It should be understood that although the terms "first", "second", "third", etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this application, "a plurality" means two or more unless otherwise specifically defined.

[0035] In the related art, the position of the road sign displayed in the high-precision map is not accurate enough, which affects the accuracy of the content displayed in the high-precision map.

[0036] In view of the above problems, an embodiment of this application provides a road sign generation method, which can improve the display position accuracy of the road sign in the high-precision map.

[0037] The technical solutions of the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0038] Figure 1 It is a schematic flowchart of the road sign generation method shown in the embodiment of this application.

[0039] See Figure 1 , a road sign generation method provided by an embodiment of this application includes:

[0040] Step S110, obtaining multiple frames of target images and respectively marking timestamps for the multiple frames of target images, where each target image includes a target road sign, and the timestamp is a preset duration earlier than the corresponding image capture time.

[0041] In this step, the target images can be obtained by using the image acquisition device on the surveying vehicle, and the image acquisition device can be a monocular camera. During the driving of the surveying vehicle, the surrounding environment can be photographed in real time. Among them, the shooting frame rates of different performance monocular cameras may be different. For example, the shooting frame rate of the monocular camera can be 20 frames per second to 30 frames per second, such as 20 frames per second, 25 frames per second, 30 frames per second. In this step, among the numerous images captured by the monocular camera, multiple frames of target images can be sequentially selected in chronological order, where the target images include the target road sign.

[0042] During the process of a monocular camera capturing images, the captured images are sent to the in-vehicle intelligent device in real time to mark time stamps, so as to associate relevant geographical location information through the time stamps subsequently. Since it takes a certain amount of time to transmit the captured images to the in-vehicle intelligent device, taking the shooting frame rate of the camera as 25 frames per second as an example, the interval time between two adjacent frames is 40 ms, that is, the time for the image to be transmitted to the in-vehicle intelligent device is delayed by about 40 ms. Therefore, in this embodiment, for the same frame of target image, the marking time of the marked time stamp is advanced by a preset duration compared with the time of sending it to the vehicle intelligent device, for example, advanced by 40 ms, so that the time stamp can mark the real shooting time, enabling more accurate geographical location information to be associated in subsequent steps when associating relevant geographical location information according to the time stamp, thereby reducing errors.

[0043] Step S120, determine the camera heading angle corresponding to each target image according to the optimized vehicle geographical locations corresponding to each time stamp.

[0044] It can be understood that as the vehicle moves, the geographical location information of the vehicle will change at different times. In this step, after determining the shooting time corresponding to the selected target image through the time stamp, the geographical location of the vehicle at the current time stamp can be obtained accordingly. It can be understood that the geographical location of the vehicle is generally characterized by a GPS positioning device installed on the vehicle body, that is, the geographical location of the GPS positioning device is the geographical location of the vehicle. However, as Figure 2 and Figure 3 shown, the installation position of the camera on the vehicle body is different from that of the GPS positioning device. On the one hand, there may be a distance between them, for example, the distance is about one meter, and on the other hand, they are not on the same horizontal line, for example, there is an included angle β between them. Therefore, when calculating the camera heading angle of the camera when shooting the target image, if the geographical location of the GPS positioning device is directly used, errors will occur, and the accumulation of errors will affect the accuracy of the finally calculated camera heading angle. Therefore, in this step, the geographical location of the GPS positioning device is not directly used, but after converting and optimizing the geographical location of the GPS positioning device according to their real relative installation positions on the vehicle body, the corresponding optimized vehicle geographical location is obtained, and based on this, the camera heading angle corresponding to each target image when shooting is determined according to relevant technologies. That is to say, for different target images, due to the different shooting positions of the camera, the camera heading angle changes accordingly.

[0045] Step S130, match the feature points of two adjacent frames of target images to obtain the optimized three-dimensional coordinates of the corner points of the target road sign in the target image in the camera coordinate system.

[0046] Among them, the number of frames of the target image is at least 3 frames. In this step, according to the chronological order corresponding to the timestamps, the feature points of the first frame can be matched with the feature points of the second frame, and the feature points of the second frame can be matched with the feature points of the third frame. Since the shape of the road sign itself is generally a regular shape such as a triangle or a quadrilateral and has multiple included angles, the feature points in the target image also include the corner points of the road sign.

[0047] In this step, by pairwise matching the feature points of the target image and then using epipolar geometry, triangulation, and some optimization methods, the optimized three-dimensional coordinates of each feature point in the camera coordinate system can be obtained. It can be understood that while obtaining the optimized three-dimensional coordinates of each feature point, the optimized three-dimensional coordinates of the corner points of the road sign in the camera coordinate system are also obtained.

[0048] Step S140: Determine the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

[0049] Among them, the camera parameters include parameters such as the pre-calibrated camera internal parameter matrix and the camera external parameter matrix. According to the related technology, according to the optimized geographical location of the vehicle and the camera heading angle corresponding to each frame of the target image, the world geographical coordinates of the corner points in each frame of the target image can be calculated. By averaging the world geographical coordinates in each frame of the target image, the unique world geographical coordinates of the same angle can be obtained.

[0050] From this example, it can be seen that by optimizing the timestamps, the optimized geographical location of the vehicle, the camera heading angle, and the optimized three-dimensional coordinates of the corner points associated with the target image in different links, the measurement error is gradually reduced, the total cumulative error becomes smaller, and then the accuracy of the world geographical coordinates finally obtained for the corner points is improved, so that the road sign can be positioned and displayed in the high-precision map according to the more accurate world geographical coordinates, improving the user experience of using the high-precision map.

[0051] Figure 4 It is another schematic flowchart of the road sign generation method shown in the embodiments of the present application.

[0052] See Figure 4 A road sign generation method provided by an embodiment of the present application includes:

[0053] Step S210: Sequentially obtain multiple frames of target images at a preset interval distance and mark timestamps for the multiple frames of target images respectively. The number of target images is greater than or equal to 3 frames, each target image includes a target road sign, and the timestamp is a preset duration ahead of the corresponding image shooting time.

[0054] In this step, the monocular camera mounted on the survey vehicle can be used to take pictures of the road along the way. In the captured images, according to the driving path and the shooting time, for example, starting from the first frame of the target image containing the target road sign, every 3 meters to 5 meters, the second frame, the third frame, the fourth frame... the Nth frame of the target image can be sequentially selected, where N is a natural number. It can be understood that the geographical positions of the cameras corresponding to each selected frame of the target image are different from each other. By selecting the target images taken by the camera at different positions, the mis-matching of subsequent feature points caused by too large a distance between two frames of target images is avoided. Through an appropriate interval distance, more target images are obtained, making the content of the obtained images richer, avoiding the image content being too single, and improving the matching accuracy of feature points in the subsequent steps.

[0055] Correspondingly, according to the determined target images, the timestamp corresponding to each frame of the target image is advanced by a preset duration, that is, the time marked by the in-vehicle intelligent device is advanced by the preset duration, so as to overcome the error caused by the delay in the data transmission process.

[0056] Step S220: Obtain the vehicle longitude and latitude coordinates corresponding to each timestamp respectively; determine the longitude and latitude coordinates of the adjacent points before and after the current vehicle longitude and latitude coordinates; according to the longitude and latitude coordinates of the adjacent points, determine the camera heading angle corresponding to the target image.

[0057] It should be clear that the vehicle longitude and latitude coordinates and the longitude and latitude coordinates of the adjacent points used below are the optimized geographical positions of the vehicle after conversion and optimization according to the true relative installation positions of the GPS positioning device and the camera. That is, the optimized geographical position of the vehicle is the longitude and latitude coordinates obtained by converting the true longitude and latitude coordinates measured by the GPS positioning device, and the longitude and latitude coordinates directly measured by the GPS positioning device will not be used.

[0058] In this step, according to the number of frames of the obtained target images, there are corresponding timestamps. Correspondingly, the vehicle longitude and latitude coordinates corresponding to the timestamps can be obtained. Taking one frame of the target image as an example, after obtaining the vehicle longitude and latitude coordinates of the current point corresponding to the timestamp of the target image, then obtain the longitude and latitude coordinates of the adjacent points at a preset distance before and after the current point. The preset distance can be more than 3 meters, such as 3 meters, 4 meters, etc. Then, obtain the longitude and latitude coordinates of the adjacent points at the preset distance (such as 3 meters each) before and after the current point respectively. Subsequently, calculate the camera heading angle corresponding to the current target image according to the longitude and latitude coordinates of the two adjacent points.

[0059] As Figure 5 shown, in the figure, point O is the optical center of the monocular camera installed on the vehicle body, N and E represent the due north and due east directions, and the vehicle moves continuously along the driving direction D according to the time line, and the GPS coordinates of each passing point along the moving trajectory can be recorded. Taking the current point P among them curTaking a certain frame of target image collected as an example, this target image is one frame among all the images taken along the timeline, and the corresponding timestamp marked in the figure is (2022-01-15 10:48:19). After obtaining the longitude and latitude coordinates of the current point P corresponding to this timestamp cur of the vehicle, the longitude and latitude coordinates of adjacent points P1 and P2 can be obtained at a preset interval greater than or equal to 3 meters. After determining the positions of the specific adjacent points, the angle between the line connecting adjacent points P1 and P2 and the due north is the camera heading angle α.

[0060] Step S230: Match the feature points of two adjacent frames of target images to obtain the optimized three-dimensional coordinates of the corner points of the target road signs in the target images in the camera coordinate system.

[0061] In this step S230 and step S220, they can be carried out in any order or synchronously. In a specific implementation manner, in this step S230, according to the following steps S231 to S235, in N frames of target images, (N - 1) groups of optimized three-dimensional coordinates corresponding to each feature point can be obtained, including the optimized three-dimensional coordinates of the corner points of the target road signs. The specific steps are as follows:

[0062] Step S231: Extract the feature points in each target image respectively, and match the feature points in two adjacent frames of target images to obtain the relative pose between the two adjacent frames of target images and the initial three-dimensional coordinates corresponding to the feature points.

[0063] In this step, after extracting the feature points in each frame of target image according to the relevant algorithm, in the target road sign area of the target image, the corresponding feature points are deleted according to the preset pixel distance. It should be understood that based on the image characteristics of the target road sign, there will be a situation where the local feature points are too dense in the area where it is located, resulting in uneven distribution of the feature points, which will affect the calculation results and cause local deviations. In this step, some feature points in the area of the target road sign in a single frame of target image can be deleted in advance. For example, when the interval distance between two feature points is less than 5 pixels, then one of the feature points is deleted, thereby reducing the density of the local feature points and making the overall distribution tend to be uniform.

[0064] Further, match each pair of adjacent target images. For example, match the feature points in the first frame and the second frame respectively, match the feature points in the second frame and the third frame, match the feature points in the third frame and the fourth frame, and so on. It can be understood that the objects included in each target image may be partially the same and partially different. Therefore, by using relevant algorithms to automatically perform pairwise matching of adjacent two-frame images in each target image, and based on the camera internal parameter matrix and the epipolar geometry constraint, the relative pose of the cameras for each pair of adjacent target images can be obtained. Further, calculate according to the matrix of the relative pose of the cameras and triangulation to obtain the initial three-dimensional coordinates (X c , Y c , Z c ) corresponding to the successfully matched feature points in the camera coordinate system of each target image, where Z c is the depth value corresponding to this feature point in the camera coordinate system.

[0065] Further, after obtaining the relative pose of each pair of adjacent target images and the initial three-dimensional coordinates corresponding to the feature points, in one embodiment, screen the feature points according to the depth values of the three-dimensional coordinates, and delete the feature points with negative depth values. Among them, if the calculated depth value of a pair of feature points after matching is negative, it means that the matching of this group of feature points fails. That is to say, the depth values of the same object in the camera coordinate systems corresponding to different images should all be positive. Therefore, in this step, in a single target image, delete the feature points with negative depth values in its three-dimensional coordinates, so as to reduce the calculation error in the subsequent steps.

[0066] It can be understood that after respectively extracting the feature points of the target image, the corner points corresponding to the target road sign and the initial three-dimensional coordinates of the corner points can be determined among the feature points. That is to say, there are a large number of feature points in the target image, but based on the special shape of the target road sign, the corresponding corner points of the target road sign and the corresponding initial three-dimensional coordinates can be obtained among the feature points according to relevant image recognition algorithms, which is convenient for the calculation of subsequent steps. For example, when the road sign is a quadrilateral, 4 corresponding corner points and their three-dimensional coordinates can be obtained; when the road sign is a triangle, 3 corresponding corner points and their initial three-dimensional coordinates can be obtained.

[0067] Step S232, determine the absolute poses of the remaining target images respectively according to the preset absolute pose of one of the target images and each relative pose.

[0068] To obtain the absolute poses of each target image, in this step, one of the target images can be selected. For example, the first target image can be selected as the reference, and its absolute pose can be preset in advance. For example, the preset absolute pose can be represented by a rotation matrix. Then, based on the preset absolute pose of the selected target image, the absolute poses of the cameras corresponding to the remaining target images are calculated respectively according to the relative poses of every two adjacent target images.

[0069] Step S233: Perform graph optimization based on the absolute pose, the initial 3D coordinates and pixel coordinates of the feature points to obtain the optimized absolute pose.

[0070] In this step, the optimized absolute poses of each target image can be obtained according to the specific implementation manners of the following steps S2331 to S2335.

[0071] S2331: Determine the common feature points between different target images, and use the mean value of the corresponding initial 3D coordinates as the mean 3D coordinates of the common feature points.

[0072] Among them, the co-visibility relationship between different target images can be found through the feature points, that is, the feature points of the same object in different target images, that is, the successfully matched feature points are used as the common feature points. As Figure 6 shown in the figure, C1 in the figure represents the absolute pose of the first target image, C2 represents the absolute pose of the second target image, and C3 represents the absolute pose of the third target image. P1 to P3 represent the feature points in the first target image, P2 to P4 represent the feature points in the second target image, and P3 to P6 represent the feature points in the third target image. The reprojection error e11 represents the position deviation of the pixel coordinates between the original projection position and the adjusted projection position of the feature point P1 in the first target image; e12 represents the position deviation of the pixel coordinates between the original projection position and the adjusted projection position of the feature point P2 in the first target image, and so on.

[0073] As Figure 6 shown in the figure, the feature points P2 and P3 are the common feature points of the first and second target images, the feature point P3 is the common feature point of the first to third target images, and P4 is the common feature point of the second and third target images. In this step, for the common feature points, the average value can be calculated according to the initial 3D coordinates corresponding to each of them in the original target images, so that the common feature points have unified mean 3D coordinates. In the subsequent steps, when a certain feature point is a common feature point, the mean 3D coordinates are used to participate in the relevant calculations. It can be understood that when a certain feature point is not a common feature point, the initial 3D coordinates of this single feature point are also the mean 3D coordinates.

[0074] S2332. Obtain the mean three-dimensional coordinates of each feature point including the common feature points, the absolute poses of each target image, and the correspondence between the pixel coordinates of each feature point in the corresponding target image.

[0075] For the convenience of subsequent graph optimization processing, in this step, establish the correspondence among the mean three-dimensional coordinates of each feature point in the camera coordinate system, the pixel coordinates in the pixel coordinate system, and the absolute pose of the target image where the feature point is located. For the convenience of quickly obtaining data during calculation, IDs can be used for identification respectively, so that the above correspondences can be associated according to the IDs.

[0076] As shown in Table 1 below, combined with Figure 6 , taking 3 frames of target images as an example, for the convenience of obtaining data during subsequent calculations, ID labels can be set for the absolute poses corresponding to each frame of target image, such as 1, 2, 3. Each feature point in each frame of target image has its own corresponding ID. When the feature point is a common feature point, the same ID is displayed in different target images. For example, the feature point P2 appears in the first and second frames of target images at the same time, so the corresponding ID in both frames of target images is 2. The reprojection error of each feature point in the corresponding target image is represented according to its own ID. For example, the reprojection error of the first feature point P1 in the first frame of target image is e11, and the reprojection error of the third feature point P3 in the second frame of target image is e23, and so on. Thus, all feature points can form correspondences based on the target images where they are located, the corresponding absolute poses, mean three-dimensional coordinates, and pixel coordinates. It can be understood that the pixel coordinates can be obtained by conversion from the mean three-dimensional coordinates of the feature points according to related technologies, and then the pixel coordinates of each feature point in the corresponding target image can be obtained.

[0077] Table 1

[0078]

[0079]

[0080] S2333. Using the absolute pose and the mean three-dimensional coordinates of the feature points as vertices, and the correspondence as the connecting edges, through a preset graph optimization model, determine the reprojection error of each common feature point in the corresponding target image.

[0081] In this embodiment, according to relevant algorithms, taking the g2o graph optimization model as an example, input the absolute pose and the mean three-dimensional coordinates of the feature points into the g2o graph optimization model and set the corresponding IDs. Then add the correspondence in step S2332 above as connecting edges to the model. The model can automatically calculate the corresponding reprojection error and synchronously adjust the absolute pose and three-dimensional coordinates during the iteration process to make the reprojection error smaller and smaller.

[0082] S2334, delete the connection edges with reprojection errors greater than a preset value, and perform iteration through a preset graph optimization model to adjust the absolute poses of each target image and the mean three-dimensional coordinates of the feature points until a preset iteration termination condition is reached.

[0083] In one embodiment, when the reprojection error is greater than 1, the corresponding connection edges are deleted to reduce the influence of data with large errors on the calculation. After the initial 1 - 2 iterations of optimization, the connection edges with large errors can be removed, and the remaining connection edges are added to the model for continued iterative optimization, making the absolute poses and the mean three-dimensional coordinates of the feature points more and more accurate.

[0084] In one embodiment, the preset iteration termination condition includes a preset number of iterations and / or a preset total reprojection error threshold. For example, the preset number of iterations is 15 - 20 times. For example, the preset total reprojection error threshold is 100. The total reprojection error is the sum of the reprojection errors corresponding to the pixel coordinates of all feature points. When the iteration reaches one of the above preset iteration termination conditions, the iteration can be stopped.

[0085] S2335, take the absolute pose after the iteration termination as the optimized absolute pose.

[0086] It can be understood that after the iteration is completed, the absolute pose obtained from the last optimization is the optimized absolute pose. This optimized absolute pose is more accurate than the initial absolute pose.

[0087] In summary, in this step S233, an optimized graph can be constructed according to the absolute pose and the three-dimensional coordinates of each feature point through related technologies. For example, through the g20 graph optimization model, the absolute pose and the three-dimensional coordinates of each feature point are used as vertices, and then the corresponding relationship between each absolute pose, the three-dimensional coordinates of the feature points, and the corresponding pixel coordinates is used as the connection edges to construct the optimized graph. In the form of graph optimization, the absolute pose and each three-dimensional coordinate are iteratively adjusted to make the reprojection error smaller and smaller, and finally the qualified absolute pose is obtained as the optimized absolute pose.

[0088] Step S234, obtain the optimized relative poses corresponding to each target image according to the optimized absolute pose.

[0089] It can be understood that after obtaining the optimized absolute pose of each frame of the target image, the corresponding calculation and conversion can be performed to obtain the optimized relative poses of the cameras corresponding to every two adjacent frames of the target images.

[0090] Step S235, perform triangulation according to the optimized relative pose and the pixel coordinates corresponding to the corner points of the target road sign to obtain the optimized three-dimensional coordinates of the corner points in the camera coordinate system.

[0091] In this step, triangulation can be performed according to the optimized relative pose and the pixel coordinates of each corner point, and the corresponding optimized three-dimensional coordinates in the camera coordinate system can be obtained through relevant technologies. Obviously, the optimized three-dimensional coordinates are more accurate than the corresponding initial three-dimensional coordinates.

[0092] Step S240: Determine the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

[0093] Among them, the camera parameters include but are not limited to the pre-calibrated camera internal parameters and camera external parameters, etc. The optimized geographical location of the vehicle in this step is the longitude and latitude coordinates after conversion when the camera captures each frame of the target image. Combining Figure 7 , the specific process of calculating the world geographical coordinates of the corner points is as follows:

[0094] During positioning and measurement, a rectangular coordinate system is established with the optical center O of the monocular camera as the center and the direction of the vehicle heading angle as the Z-axis, and the right side is set as the X-axis and the bottom as the Y-axis. See specifically Figure 7 as shown. The longitude and latitude coordinates of the camera optical center O can obtain the corresponding optimized geographical location according to the timestamp of the target image. According to the Mercator hypothesis, the longitude and latitude coordinates of the camera optical center O are converted into utm coordinates, denoted by O utm (x ou ,you,z ou ). Subsequently, it is necessary to calculate the camera heading angle α corresponding to this timestamp, that is, the camera heading angle α corresponding to each frame of the target image obtained according to the above step S220. Assume that the optimized three-dimensional coordinates of a corner point P in the target object, such as a target road sign, in the camera coordinate system after triangulation are (x, y, z), then its utm(x pu , y pu , z pu ) coordinates can be calculated by formulas (1) to (4):

[0095]

[0096]

[0097] z pu = z ou - y(3)

[0098] In the above formula, θ is the angle between the line connecting the optical center O and the corner point P and the Z-axis, and the angle θ can be calculated by the following formula (4).

[0099] θ = 180 * arctan(xz) * π -1 (4)

[0100] Similarly, in the target images where each frame contains point P, the utm coordinates corresponding to point P, the corner point in each frame of the target image, can be calculated according to the above formula. By calculating the average value of each utm coordinate and then performing reverse conversion, the longitude and latitude coordinates of point P can be obtained, that is, the unique world geographical coordinates of corner point P. By calculating the world geographical coordinates of each corner point of the road sign, the display position of the road sign in the high-precision map and the display area of the sign face can be determined, and a road sign with the corresponding size ratio can be generated in the high-precision map through rendering.

[0101] From this example, it can be seen that the road sign generation method of the present application starts setting relevant conditions from the target image selection stage, selectively filters the target images, and improves the measurement accuracy from the source by adjusting the timestamp; the calculation method of the camera heading angle is no longer based on a single current point, but is calculated based on two adjacent points before and after the current point, further reducing the measurement error; at the same time, after obtaining the optimized absolute coordinates of each frame of the target image through iterative processing, the optimized relative pose is then obtained, and then a more accurate optimized three-dimensional coordinate can be obtained by combining the pixel coordinates of the feature points; finally, based on a series of optimized data, that is, the optimized geographical location of the vehicle, the camera heading angle, and the optimized three-dimensional coordinates, the world geographical coordinates of the corner points can be quickly calculated. Such a design can effectively improve the robustness of the system calculation results, and compared with the initial three-dimensional coordinates with larger errors, the world geographical coordinates of the corresponding corner points can be obtained according to the optimized three-dimensional coordinates according to related technologies, and a more accurate spatial form of the road sign corner points and their display positions in the high-precision map can be obtained.

[0102] Corresponding to the foregoing application function implementation method embodiments, the present application also provides a road sign generation device, an electronic device, and corresponding embodiments.

[0103] Figure 8 It is a schematic structural diagram of the road sign generation device shown in the embodiments of the present application.

[0104] See Figure 8 , a road sign generation device provided in an embodiment of the present application, the device includes an acquisition module 810, an angle determination module 820, a coordinate optimization module 830, and a processing module 840, wherein:

[0105] The acquisition module 810 is configured to acquire multiple frames of target images and mark timestamps for the multiple frames of target images respectively, wherein each target image contains a target road sign, and the timestamp is a preset duration earlier than the corresponding image capture time.

[0106] The angle determination module 820 is configured to determine the camera heading angle corresponding to each target image according to the optimized geographical location of the vehicle corresponding to each timestamp.

[0107] The coordinate optimization module 830 is used to match the feature points of two adjacent frames of target images, and obtain the optimized three-dimensional coordinates of the corner points of the target road signs in the target images in the camera coordinate system.

[0108] The processing module 840 is used to determine the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

[0109] Specifically, the acquisition module 810 is used to sequentially acquire multiple frames of target images at a preset interval distance, and the number of target images is greater than or equal to 3 frames. The angle determination module 820 is used to respectively acquire the longitude and latitude coordinates of the vehicle corresponding to each timestamp; determine the longitude and latitude coordinates of the adjacent points before and after the current vehicle longitude and latitude coordinates; and determine the camera heading angle corresponding to the target image according to the longitude and latitude coordinates of the adjacent points. Among them, the longitude and latitude coordinates of the vehicle corresponding to each timestamp and the longitude and latitude coordinates of the adjacent points are the longitude and latitude coordinates of the optimized geographical location of the vehicle converted and optimized according to the GPS positioning device and the true relative installation position of the camera.

[0110] The coordinate optimization module 830 is used to respectively extract the feature points in each target image, match the feature points in two adjacent frames of target images, and obtain the relative pose between the two adjacent frames of target images and the initial three-dimensional coordinates corresponding to the feature points; determine the absolute poses of the remaining target images respectively according to the preset absolute pose of one of the target images and each relative pose; perform graph optimization according to the absolute pose, the initial three-dimensional coordinates of the feature points, and the pixel coordinates to obtain the optimized absolute pose; obtain the optimized relative pose corresponding to each target image according to the optimized absolute pose; and perform triangulation according to the optimized relative pose and the pixel coordinates corresponding to the corner points of the target road sign to obtain the optimized three-dimensional coordinates of the corner points in the camera coordinate system.

[0111] Among them, performing graph optimization according to the absolute pose, the initial three-dimensional coordinates of the feature points, and the pixel coordinates to obtain the optimized absolute pose includes: determining the common feature points between different target images, and using the mean value of the corresponding initial three-dimensional coordinates as the mean three-dimensional coordinates of the common feature points; obtaining the corresponding relationship between the mean three-dimensional coordinates of each feature point including the common feature points, the absolute poses of each target image, and the pixel coordinates of each feature point in the corresponding target image; using the absolute pose and the mean three-dimensional coordinates of the feature points as vertices, and using the corresponding relationship as the connecting edges, and determining the reprojection error of each common feature point in the corresponding target image through a preset graph optimization model; deleting the connecting edges with reprojection errors greater than the preset value, and performing iteration through the preset graph optimization model to adjust the absolute poses of each target image and the mean three-dimensional coordinates of the feature points until the preset iteration termination condition is reached; where the preset iteration termination condition includes a preset number of iterations and / or a preset total reprojection error threshold. Taking the absolute pose after the iteration termination as the optimized absolute pose.

[0112] The coordinate optimization module 830 is further configured to delete corresponding feature points in the target road sign area of the target image at a preset pixel distance; and / or the coordinate optimization module 830 is further configured to screen the feature points according to the depth value of the three-dimensional coordinates, and delete the feature points with negative depth values.

[0113] In summary, the road sign generation device of the present application can improve the coordinate accuracy of the world geographical location of the road sign corners, and has high robustness to handle various road signs, facilitating more accurate generation of road signs at accurate positions in the high-precision map.

[0114] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0115] Figure 9 It is a schematic structural diagram of an electronic device shown in an embodiment of the present application.

[0116] See Figure 9 , the electronic device 1000 includes a memory 1010 and a processor 1020.

[0117] The processor 1020 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0118] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. Among them, the ROM can store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device can be a readable and writable storage device. The permanent storage device can be a non-volatile storage device that does not lose the stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a mass storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In some other embodiments, the permanent storage device can be a removable storage device (such as a floppy disk, optical drive). The system memory can be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory can store some or all of the instructions and data required by the processor during operation. In addition, the memory 1010 can include any combination of computer-readable storage media, including various types of semiconductor storage chips (such as DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks can also be used. In some embodiments, the memory 1010 can include a removable storage device that is readable and / or writable, such as a compact disc (CD), read-only digital versatile disc (such as DVD-ROM, dual-layer DVD-ROM), read-only Blu-ray disc, super density disc, flash memory card (such as SD card, min SD card, Micro-SD card, etc.), magnetic floppy disk, etc. The computer-readable storage medium does not include carrier waves and instantaneous electronic signals transmitted wirelessly or by wire.

[0119] Executable code is stored on the memory 1010, and when the executable code is processed by the processor 1020, it can cause the processor 1020 to execute some or all of the methods described above.

[0120] In addition, the method according to the present application can also be implemented as a computer program or a computer program product, and the computer program or the computer program product includes computer program code instructions for executing some or all of the steps in the above method of the present application.

[0121] Alternatively, the present application can also be implemented as a computer-readable storage medium (or a non-transitory machine-readable storage medium or a machine-readable storage medium), on which executable code (or a computer program or computer instruction code) is stored. When the executable code (or the computer program or computer instruction code) is executed by a processor of an electronic device (or a server, etc.), it causes the processor to execute some or all of the steps of the above method according to the present application.

[0122] The embodiments of the present application have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for generating a road sign, characterized in that, Including: Obtaining multiple frames of target images and respectively marking timestamps for the multiple frames of target images, where each of the target images includes a target road sign, and the timestamp is a preset duration ahead of the corresponding image capture time; Determining the camera heading angle corresponding to each of the target images according to the optimized geographical location of the vehicle corresponding to each of the timestamps; Respectively extracting feature points in each of the target images, and matching the feature points of two adjacent frames of the target images to obtain the relative pose between two adjacent frames of the target images and the initial three-dimensional coordinates corresponding to the feature points. According to the preset absolute pose of one of the target images and each of the relative poses, respectively determining the absolute poses of the remaining target images; Determining the common feature points between different target images, and using the mean value of the corresponding initial three-dimensional coordinates as the mean three-dimensional coordinates of the common feature points; obtaining the corresponding relationship between the mean three-dimensional coordinates of each feature point including the common feature points, the absolute poses of each target image, and the pixel coordinates of each feature point in the corresponding target image; using the absolute pose and the mean three-dimensional coordinates of the feature points as vertices, and using the corresponding relationship as connecting edges, through a preset graph optimization model, determining the reprojection error of each common feature point in the corresponding target image; deleting the connecting edges with reprojection error greater than a preset value, and performing iteration through the preset graph optimization model to adjust the absolute poses of each target image and the mean three-dimensional coordinates of the feature points until a preset iteration termination condition is reached; where the preset iteration termination condition includes a preset number of iterations and / or a preset total reprojection error threshold; using the absolute pose after iteration termination as the optimized absolute pose; Obtaining the optimized three-dimensional coordinates of the corner points of the target road sign in the target image in the camera coordinate system based on the optimized absolute pose; Determining the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

2. The method according to claim 1, characterized in that, The obtaining of multiple frames of target images includes: Sequentially obtaining multiple frames of target images at a preset interval distance, and the number of the target images is greater than or equal to 3 frames.

3. The method according to claim 1, characterized in that, The determining of the camera heading angle corresponding to each of the target images according to the optimized geographical location of the vehicle corresponding to each of the timestamps includes: Respectively obtaining the longitude and latitude coordinates of the vehicle corresponding to each timestamp; Determining the longitude and latitude coordinates of the adjacent points before and after the current vehicle longitude and latitude coordinates; Determining the camera heading angle corresponding to the target image according to the longitude and latitude coordinates of the adjacent points.

4. The method according to claim 1, characterized in that, The obtaining of the optimized three-dimensional coordinates of the corner points of the target road sign in the target image in the camera coordinate system based on the optimized absolute pose includes: Obtaining the optimized relative pose corresponding to each of the target images according to the optimized absolute pose; Performing triangulation according to the optimized relative pose and the pixel coordinates corresponding to the corner points of the target road sign to obtain the optimized three-dimensional coordinates of the corner points in the camera coordinate system.

5. The method according to claim 1, characterized in that: After respectively extracting the feature points in each of the target images, it further includes: Deleting the corresponding feature points in the target road sign area of the target image at a preset pixel distance. After matching the feature points in two adjacent frames of the target images to obtain the relative pose between two adjacent frames of the target images and the initial three-dimensional coordinates corresponding to the feature points, the method further includes: Filtering the feature points according to the depth values of the three-dimensional coordinates, and deleting the feature points with negative depth values.

6. A road sign generation device, characterized in that, It includes: An acquisition module, configured to acquire multiple frames of target images and respectively mark timestamps for the multiple frames of target images. Each of the target images includes a target road sign, and the timestamp is a preset duration earlier than the corresponding image capture time; An angle determination module, configured to determine the camera heading angle corresponding to each of the target images according to the optimized geographical location of the vehicle corresponding to each of the timestamps; A coordinate optimization module, configured to respectively extract the feature points in each of the target images, match the feature points in two adjacent frames of the target images to obtain the relative pose between two adjacent frames of the target images and the initial three-dimensional coordinates corresponding to the feature points, and respectively determine the absolute poses of the remaining target images according to the preset absolute pose of one of the target images and each of the relative poses; determine the common feature points between different target images, and use the mean value of the corresponding initial three-dimensional coordinates as the mean three-dimensional coordinates of the common feature points; obtain the corresponding relationship between the mean three-dimensional coordinates of each feature point including the common feature points, the absolute poses of each target image, and the pixel coordinates of each feature point in the corresponding target image; use the absolute pose and the mean three-dimensional coordinates of the feature points as vertices, and use the corresponding relationship as connecting edges, and through a preset graph optimization model, determine the reprojection error of each common feature point in the corresponding target image; delete the connecting edges with reprojection errors greater than a preset value, and perform iteration through the preset graph optimization model to adjust the absolute poses of each target image and the mean three-dimensional coordinates of the feature points until a preset iteration termination condition is reached; wherein, the preset iteration termination condition includes a preset number of iterations and / or a preset total reprojection error threshold; use the absolute pose after iteration termination as the optimized absolute pose; based on the optimized absolute pose, obtain the optimized three-dimensional coordinates of the corner points of the target road sign in the target image in the camera coordinate system; A processing module, configured to determine the world geographical coordinates corresponding to the optimized three-dimensional coordinates of the corner points according to the camera parameters, the optimized geographical location of the vehicle, and the camera heading angle.

7. The device according to claim 6, characterized in that, The coordinate optimization module is further configured to delete the corresponding feature points in the target road sign area of the target image at a preset pixel distance; and / or The coordinate optimization module is further configured to filter the feature points according to the depth values of the three-dimensional coordinates, and delete the feature points with negative depth values.

8. An electronic device, characterized in that, It includes: A processor; And A memory, storing executable code thereon, which when executed by the processor, causes the processor to execute the method according to any one of claims 1-5.

9. A computer-readable storage medium, characterized in that,Storing executable code thereon, which when executed by the processor of an electronic device, causes the processor to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Unmanned aerial vehicle (UAV) multispectral image fast splicing method

    CN107274380A

  • Method and device for generating high-precision map guideboard and server

    CN113536854A