A positioning method and device for a 100-meter stake sign

Through image recognition and camera parameter conversion technology, the real position of the 100-meter pile sign is accurately positioned, which solves the problem of inaccurate positioning of the drone caused by the position deviation of the 100-meter pile sign, and improves accident handling efficiency and traffic safety.

CN119152024BActive Publication Date: 2025-08-01INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS +2
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410753197.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-12
Publication Date
2025-08-01
Estimated Expiration
2044-06-12

AI Technical Summary

Technical Problem

The installation location of the existing 100-meter pile signs is deviated from the indicated location, which makes it impossible for the drone to accurately locate the accident site, reducing the accident handling efficiency and the speed of highway disease treatment.

Method used

By obtaining the image of the 100-meter pile sign taken by the camera, using the trained image recognition model to identify the sign information and key point pixel coordinates, and combining the camera parameters to convert the coordinates to establish a latitude and longitude coordinate correlation, and accurately locate the real position of the 100-meter pile sign.

Benefits of technology

The accurate positioning of 100-meter pile signs has been achieved, the accident handling efficiency and traffic safety have been improved, and the drone can quickly and accurately reach the accident site.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119152024B_ABST
    Figure CN119152024B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of transportation, and particularly to a positioning method and device for a hundred-meter post sign. By obtaining a target image containing the hundred-meter post sign captured by a camera, then inputting the target image into a trained image recognition model, identifying the sign information on the hundred-meter post sign in the target image and the key point pixel coordinates corresponding to the hundred-meter post sign in the target image, and finally based on the key point pixel coordinates and the camera parameters of the camera, performing coordinate conversion on the key point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establishing and recording the association relationship between the longitude and latitude coordinates and the sign information. This application can accurately locate the true position of the hundred-meter post sign to further ensure traffic safety and improve the efficiency of accident handling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of transportation, and in particular to a positioning method and device for a 100-meter stake sign. Background Art

[0002] In modern traffic management systems, when a safety incident occurs on a highway, drones can be dispatched to the scene to quickly address it. Based on the on-site situation, a response plan can be prepared before arriving at the scene, allowing for more timely and swift handling. Drones are also commonly used to identify road hazards, such as potholes and cracks. Maintenance personnel also need to obtain the latitude and longitude of these hazards and convert them into pile numbers so they can promptly repair them on-site.

[0003] In such situations, accurately locating the accident site is crucial for dispatching drones to the scene. Currently, the main method for determining the accident site is based on 100-meter stake signs on highways. However, there can be discrepancies between the actual location of the 100-meter stake signs installed by workers and the indicated location. This can lead to discrepancies in the location of the accident site, preventing drones from reaching the accident site immediately and effectively, reducing the efficiency of handling the accident. When handling an accident, every second counts, and the sooner the "roadside illnesses" on highways are addressed, the better. Therefore, a method for locating 100-meter stake signs is urgently needed to accurately locate their true locations, further ensuring traffic safety and improving accident handling efficiency. Summary of the Invention

[0004] The present invention describes a positioning method and device for 100-meter stake signs, which can accurately locate the real position of the 100-meter stake signs to further ensure traffic safety and improve accident handling efficiency.

[0005] According to a first aspect, the present invention provides a method for positioning a 100-meter stake sign, comprising:

[0006] Get the target image containing the 100-meter stake sign captured by the camera;

[0007] Input the target image into a trained image recognition model to identify the signage information on the 100-meter stake signage in the target image and the pixel coordinates of key points corresponding to the 100-meter stake signage in the target image;

[0008] Obtain the camera parameters of the camera, perform coordinate conversion on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters, obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establish and record the association between the longitude and latitude coordinates and the sign information.

[0009] According to a second aspect, the present invention provides a positioning device for a hundred - meter stake sign, comprising:

[0010] An acquisition unit, configured to acquire a target image containing a hundred - meter stake sign captured by a camera;

[0011] An identification unit, configured to input the target image into a trained image recognition model, and identify the sign information on the hundred - meter stake sign in the target image and the key - point pixel coordinates corresponding to the hundred - meter stake sign in the target image;

[0012] A positioning unit, configured to acquire the camera parameters of the camera, based on the key - point pixel coordinates and the camera parameters, perform coordinate conversion on the key - point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key - point pixel coordinates, and establish and record the association relationship between the longitude and latitude coordinates and the sign information.

[0013] According to the positioning method and device for a hundred - meter stake sign provided by the present invention, by acquiring a target image containing a hundred - meter stake sign captured by a camera, then inputting the target image into a trained image recognition model, identifying the sign information on the hundred - meter stake sign in the target image and the key - point pixel coordinates corresponding to the hundred - meter stake sign in the target image, and finally, based on the key - point pixel coordinates and the camera parameters of the camera, performing coordinate conversion on the key - point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key - point pixel coordinates, and establishing and recording the association relationship between the longitude and latitude coordinates and the sign information. This application can accurately locate the true position of the hundred - meter stake sign to further ensure traffic safety and improve the efficiency of accident handling. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 Shows a schematic flowchart of a positioning method for a hundred - meter stake sign according to an embodiment;

[0016] Figure 2 Shows a schematic diagram of the position of key - point pixels of a hundred - meter stake sign according to an embodiment;

[0017] Figure 3 Shows a schematic diagram of the division of a hundred - meter stake sign according to an embodiment

[0018] Figure 4Schematic block diagram of a positioning device for a 100-meter stake sign according to an embodiment. Detailed implementation manners

[0019] The following describes the solution provided by the present invention with reference to the accompanying drawings.

[0020] Figure 1 The flowchart of a positioning method for a 100-meter stake sign according to an embodiment is shown. It can be understood that this method can be executed by any device, equipment, platform, or equipment cluster with computing and processing capabilities. For example Figure 1 As shown, this method includes:

[0021] Step 101, obtain a target image containing a 100-meter stake sign captured by a camera.

[0022] In this embodiment, along the real highway, a monocular camera (such as a mobile phone or a monocular camera) can be used to take a video of the 100-meter stake sign beside the highway, and then the image frame containing the 100-meter stake sign in the video is intercepted as the target image in step 101.

[0023] Step 102, input the target image into a trained image recognition model, and identify the sign information on the 100-meter stake sign in the target image and the key point pixel coordinates corresponding to the 100-meter stake sign in the target image. [[ID=?]]

[0024] In this embodiment, the sign information on the 100-meter stake sign includes the 100-meter number and the kilometer number, which characterize the position indicated by the 100-meter stake sign. For example Figure 2 in the 100-meter stake sign, the number 0 above refers to the 100-meter number being 0m, and the number 1524 below refers to the kilometer number of the 100-meter stake sign being 1524 km. The position characterized by this 100-meter stake sign is 1524 km away from the highway entrance.

[0025] The key point pixel coordinates corresponding to the 100-meter stake sign in the target image refer to the coordinates of the key point pixels corresponding to the 100-meter stake sign in the target image. Here, the key point pixels refer to, for example Figure 2 the pixel point P at the center of the 100-meter stake sign as shown.

[0026] Step 103, obtain the camera parameters of the camera, based on the key point pixel coordinates and the camera parameters, perform coordinate conversion on the key point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establish and record the association relationship between the longitude and latitude coordinates and the sign information.

[0027] In this embodiment, the camera parameters of the camera at least include the focal length f of the camera, the camera internal parameters, and the coordinates of the center point of the camera in the three-dimensional world. Among them, the camera internal parameters at least include the unit physical lengths of the horizontal pixels and vertical pixels of the image captured by the camera on the photosensitive plate of the target camera. The coordinates of the center point of the camera in the three-dimensional world, that is, the longitude and latitude coordinates corresponding to the target camera when capturing the target image, are found from the GPS data of the camera according to the time stamp corresponding to when the target image is captured, which are the longitude and latitude coordinates corresponding to the target camera when capturing the target image.

[0028] Based on the key point pixel coordinates and the camera parameters, the coordinate transformation of the key point pixel coordinates is to convert the key point pixel coordinates corresponding to the hundred-meter stake sign into real coordinates in the three-dimensional world through the camera parameters, and then, according to the longitude and latitude coordinates of the target camera, convert the real coordinates in the three-dimensional world corresponding to the hundred-meter stake sign into longitude and latitude coordinates.

[0029] Among them, the GPS of the camera when capturing the target image can be directly found from the GPS data of the camera according to the time stamp.

[0030] Thus far, the Figure 1 description of the shown process is completed.

[0031] Through Figure 1 the shown embodiment, in this embodiment, by obtaining a target image containing a hundred-meter stake sign captured by the camera, then inputting the target image into a trained image recognition model, identifying the sign information on the hundred-meter stake sign in the target image and the key point pixel coordinates corresponding to the hundred-meter stake sign in the target image, and finally, based on the key point pixel coordinates and the camera parameters of the camera, performing coordinate transformation on the key point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establishing and recording the association relationship between the longitude and latitude coordinates and the sign information. This application can accurately locate the real position of the hundred-meter stake sign to further ensure traffic safety and improve the efficiency of accident handling.

[0032] Regarding step 102,

[0033] As a preferred implementation manner, the image recognition model in step 102 may include a sign extraction module and a sign recognition module.

[0034] Among them, the sign extraction module is used to extract the image block of the 100-meter stake sign corresponding to the target image, determine the key point pixels in the image block of the 100-meter stake sign, and obtain the coordinate of the key point pixels corresponding to the 100-meter stake sign. In this embodiment, to extract the image block of the 100-meter stake sign from the target image, it is first necessary to identify the area where the image of the 100-meter stake sign is located in the target image, determine the bounding box of the image of the 100-meter stake sign, and then extract the image of the 100-meter stake sign from the target image according to the bounding box of the image of the 100-meter stake sign. In specific implementation, the sign extraction module can be implemented by pre-training a corresponding model, which will not be elaborated here for the moment. After describing Figure 1 the process shown above, the training process of the model for implementing the sign extraction module will be described in detail.

[0035] The sign recognition module is used to recognize the image block of the 100-meter stake sign and obtain the sign information on the 100-meter stake sign.

[0036] As an embodiment, the sign recognition module can directly recognize the image block of the 100-meter stake sign based on an image recognition model to obtain the sign information on the 100-meter stake sign.

[0037] Preferably, the above-mentioned sign recognition module for recognizing the image block of the 100-meter stake sign may further include the following methods:

[0038] Divide the image block of the 100-meter stake sign according to a preset partitioning rule to obtain a 100-meter area image block and a kilometer area image block;

[0039] Determine the classification corresponding to the 100-meter area image block through the classification model in the sign recognition module, and obtain the 100-meter number in the 100-meter stake sign based on the classification corresponding to the 100-meter area image block; different classifications in the classification model correspond to different 100-meter numbers;

[0040] Recognize the kilometer area image block through the character recognition sub-module in the sign recognition module to obtain the kilometer number in the 100-meter stake sign; the sign information includes the 100-meter number and the kilometer number in the 100-meter stake sign.

[0041] The 100-meter stake sign contains the 100-meter number and the kilometer number. The above-mentioned division of the image block of the 100-meter stake sign according to a preset partitioning rule means that according to the areas where the 100-meter number and the kilometer number are respectively located, the area containing only the 100-meter number is divided into a 100-meter area image block, and the area containing only the kilometer number is divided into a kilometer area image block. As Figure 3 shown, the image block of the 100-meter stake sign can be divided into the 100-meter area image block and the kilometer area image block as shown in Figure 3 shown, where Figure 3 the area a in the purple line frame in is the 100-meter area image block, and the area b in the orange line frame is the kilometer area image block.

[0042] The reason for dividing the image block of the hundred-meter post sign in this embodiment is that the numerical value of the hundred-meter number in the image block of the hundred-meter post sign is larger and easier to recognize. The model trained by machine learning can recognize the hundred-meter number faster and basically without mistakes. Preferably, in this embodiment, a classification model with faster recognition is also used to recognize the hundred-meter number. Optionally, the classification model divides the image of the hundred-meter area into 0-9 categories according to the hundred-meter number. The hundred-meter number 0 corresponds to the 0th category, the hundred-meter number 1 corresponds to the 1st category, and so on. The hundred-meter number 9 corresponds to the 9th category. In this embodiment, if the classification model divides the image block of the hundred-meter area into the 9th category, it indicates that the hundred-meter number corresponding to the image block of the hundred-meter area is 9. However, the kilometer number in the image block of the hundred-meter post sign is relatively small. Especially when the image block of the hundred-meter post sign is a small piece of image extracted from the target image, if the kilometer number is recognized by the model trained by machine learning, mistakes may easily occur. Therefore, it is more accurate to recognize the kilometer number through a dedicated character recognition sub-module.

[0043] It should be noted that in this embodiment, the image block of the hundred-meter post sign may not be divided, and other recognition models with better computing power and higher accuracy may also be selected to recognize the image block of the hundred-meter post sign. Moreover, after dividing the image block of the hundred-meter post sign, not only the classification model can be used to recognize the image block of the hundred-meter area, but also other lightweight recognition models can be used. This application does not limit this. This embodiment only provides a preferred method. In addition, this application does not limit the dedicated character recognition sub-module used to recognize the image block of the kilometer area. For example, OCR character recognition can be used.

[0044] Preferably, the image recognition model further includes an image processing module:

[0045] The image processing module is used to convert the image format of the image block of the hundred-meter post sign into a preset format, filter the pixel blocks that meet the filtering rules in the image block of the hundred-meter post sign in the preset format, and obtain the filtered image block of the hundred-meter post sign; the filtering rules are used to enhance the difference between the pixels corresponding to the hundred-meter post sign and other pixels in the image block of the hundred-meter post sign.

[0046] In this embodiment, the preset format can be set to the HSV format, and the filtering rules can be correspondingly set to filter out the pixel blocks with the V value or S value lower than the preset value. Among them, the V value is the saturation and the S value is the brightness. The purpose of setting the filtering rules in this embodiment is to help segment the pixel blocks that may be included in the image block of the hundred-meter post sign but do not belong to the image of the hundred-meter post sign through the filtering rules, such as the asphalt background. By setting the filtering rules, it helps to remove the random noise in the image block of the hundred-meter post sign.

[0047] Regarding step 103,

[0048] As a preferred embodiment, in step 103, based on the key point pixel coordinates and the camera parameters, coordinate conversion is performed on the key point pixel coordinates, which specifically includes the following steps:

[0049] Establish a first coordinate system, a second coordinate system, and a third coordinate system.

[0050] Among them, the first coordinate system takes the pixel point at the upper left corner of the target image as the origin, the width of the target image as the u-axis, and the height of the target image as the v-axis.

[0051] The second coordinate system takes the center of the target image as the origin, the width of the target image as the x-axis, and the height of the target image as the y-axis.

[0052] The third coordinate system takes the center of the target image as the origin, the axis parallel to the x-axis as the x c axis, the axis parallel to the y-axis as the y c axis, and the axis coinciding with the optical axis of the camera as the z c axis.

[0053] Based on the key point pixel coordinates, obtain the first coordinates (u, v) of the key point pixel in the first coordinate system.

[0054] According to the preset coordinate conversion formula and the camera parameters, convert the first coordinates (u, v) into the second coordinates (x, y) of the key point pixel in the second coordinate system, and convert the second coordinates (x, y) into the third coordinates (x c , y c , z c ) in the third coordinate system;

[0055] Based on the obtained longitude and latitude coordinates corresponding to the target image, convert the third coordinates (x c , y c , z c ) into the longitude and latitude coordinates corresponding to the key point pixel coordinates.

[0056] As an embodiment, obtaining the first coordinates (u, v) of the key point pixel in the first coordinate system based on the key point pixel coordinates specifically includes: scaling the key point pixel coordinates to the coordinates corresponding to the original image size, and determining the coordinates corresponding to the original image size as the first coordinates (u, v).

[0057] In this example, the pixel coordinates of the key point can be directly obtained based on the pixel distribution on the target image. In order to avoid displacement of the pixel coordinates of the key point on the target image, when determining the first coordinate (u, v), the pixel coordinates of the key point need to be scaled to the coordinates corresponding to the size of the original image.

[0058] The above-mentioned conversion of the first coordinate (u, v) into the second coordinate (x, y) of the key point pixel in the second coordinate system includes:

[0059] The first coordinate (u, v) is converted to the second coordinate (x, y) of the key point pixel in the second coordinate system according to the following formula:

[0060]

[0061] Wherein, (u0, v0) is the coordinate of the center of the target image in the first coordinate system, dx is the unit physical length of the horizontal pixel on the photosensitive plate, and dy is the unit physical length of the vertical pixel on the photosensitive plate.

[0062] The above converts the second coordinate (x, y) into the third coordinate (x c ,y c , z c ),include:

[0063] The second coordinate (x, y) is converted to the third coordinate (x c ,y c , z c ), the formula is:

[0064]

[0065] Among them, f is the focal length of the camera, the third coordinate (x c ,y c , z c ) c The method is determined based on a scale plate carrying a scale placed when the camera captures the target image.

[0066] In specific implementation, when using a target camera to shoot the 100-meter pile sign, a scale plate can be set below the target camera to determine the depth of the 100-meter pile sign captured by the camera (i.e., z c ). The scale plate can be set to be located below the camera and parallel to the ground, with the scale origin on the scale plate located on the optical axis of the camera and the scale direction aligned with the x-axis of the third coordinate system. c Axis parallel.

[0067] Then zc It can be calculated by the following formula:

[0068]

[0069] Wherein, z1 is the scale value corresponding to the intersection point between the straight line connecting the center of the target image and the 100-meter stake sign in the target image and the scale of the scale board, h is the height between the camera and the ground, and h1 is the height between the camera and the scale board.

[0070] Regarding step 103,

[0071] Based on the target image being an image frame obtained from a video, the present application also has another preferred embodiment. In this embodiment, based on the key point pixel coordinates and the camera parameters, coordinate conversion is performed on the key point pixel coordinates, which specifically includes the following steps:

[0072] Obtain at least one image frame containing the 100-meter stake sign adjacent to the target image as a reference image;

[0073] Input the target image into the trained depth recognition model to identify the first three-dimensional coordinates corresponding to the key point pixel coordinates;

[0074] Input the reference image into the depth recognition model to identify the second three-dimensional coordinates corresponding to the key point pixel coordinates of the 100-meter stake sign in the reference image;

[0075] Based on the first three-dimensional coordinates and the second three-dimensional coordinates, calculate the camera position deviation when the camera shoots the target image and the reference image, calculate the error between the camera position deviation and the longitude and latitude deviation when the camera shoots the target image and the reference image in the camera parameters. If the error is less than the first preset error threshold, calculate the longitude and latitude coordinates corresponding to the 100-meter stake sign according to the first three-dimensional coordinates and the second three-dimensional coordinates. If the error is not less than the first preset error threshold, execute the steps of establishing the first coordinate system, the second coordinate system, and the third coordinate system.

[0076] When the target image is an image frame obtained from a video, the image frames adjacent to the target image generally also include the image of the hundred-meter stake sign. Locating the hundred-meter stake sign through multiple images is conducive to accurately positioning the position of the hundred-meter stake sign. Therefore, in this embodiment, the image frames adjacent to the target image and including the hundred-meter stake sign are determined as reference images. Then, a first three-dimensional coordinate of a hundred-meter stake sign is determined based on the target image, and at least one second three-dimensional coordinate of the hundred-meter stake sign is determined based on the reference image. The camera position deviation when the camera captures the target image and the reference image is calculated through the first three-dimensional coordinate and the second three-dimensional coordinate, and whether the first three-dimensional coordinate and the second three-dimensional coordinate calculated by the depth recognition model are accurate is determined through the camera position deviation.

[0077] The first three-dimensional coordinate in this embodiment is in the coordinate system established with the center point of the camera when capturing the target image as the origin, and the second three-dimensional coordinate is in the coordinate system established with the center point of the camera when capturing the reference image as the origin. Since the camera is moving when shooting the video, the positions of the center point of the camera when shooting the target image and the center point of the camera when shooting the reference image in the real world are different (i.e., the latitudes and longitudes are different), which will also result in different first three-dimensional coordinates and second three-dimensional coordinates. Based on the first three-dimensional coordinate and the second three-dimensional coordinate recognized by the depth recognition model, and geometric principles, this embodiment inversely calculates the distance that the camera moves when shooting the target image and the reference image (i.e., the camera position deviation), and then compares the calculated camera position deviation with the actual latitude and longitude deviation when the camera captures the target image and the reference image, so as to determine whether the first three-dimensional coordinate and the second three-dimensional coordinate are accurate.

[0078] In this embodiment, if the error between the above camera position deviation and the latitude and longitude deviation is less than the first preset error threshold, it is determined that the first three-dimensional coordinate and the second three-dimensional coordinate are accurate, and then the first three-dimensional coordinate or the second three-dimensional coordinate can be directly converted into the latitude and longitude coordinates of the hundred-meter stake sign. Or, the first three-dimensional coordinate and the second three-dimensional coordinate can be respectively converted into latitude and longitude coordinates, and then the coordinate values of the converted latitude and longitude coordinates are respectively added and averaged, and the averaged coordinate values are determined as the latitude and longitude coordinates of the hundred-meter stake sign.

[0079] If the error between the above camera position deviation and the latitude and longitude deviation is not less than the first preset error threshold, it is determined that the first three-dimensional coordinate and the second three-dimensional coordinate are inaccurate, and then the latitude and longitude coordinates of the hundred-meter stake sign can be obtained by the method of establishing the first coordinate system, the second coordinate system and the third coordinate system in another embodiment. Among them, the first preset error threshold can be determined according to the error probability of the depth recognition model.

[0080] As an embodiment, the steps of the above depth recognition model recognizing the first three-dimensional coordinate or the second three-dimensional coordinate include:

[0081] Translate each pixel point in the image to be recognized according to different translation lengths, and upsample the multiple images obtained after translation to obtain multiple feature maps corresponding to the image to be recognized and a disparity map corresponding to each feature map; the image to be recognized is the target image or the reference image, and the disparity map is used to represent the disparity between the pixels in the feature map and the pixels in the image to be recognized;

[0082] Extract the pixel points in a specific area of each feature map, and form pixel point pairs by the adjacent pixel points extracted; each pixel point in the image to be recognized has a corresponding pixel point among the pixel points extracted from each feature map;

[0083] Select the pixel points representing the same object according to the disparity map to perform accumulation respectively, and adjust the distance between the pixel points obtained after accumulation based on the pixel point pairs, to obtain a combined view corresponding to the image to be recognized;

[0084] Match the pixel points of the image to be recognized and the pixel points of the combined view, and calculate the three-dimensional coordinates corresponding to the key point pixel coordinates of the hundred-meter stake image in the image to be recognized based on the pixel difference between the matched pixel points.

[0085] In this embodiment, the depth recognition model can be trained based on a training set of a small number of real views. Among them, the above steps are only the recognition algorithm logic inside the depth recognition model in this application. The combined view is a view generated by the depth recognition model based on the image to be recognized and having the same image content as the image to be recognized but different perspectives. The different translation lengths, the specific areas in each feature map, and the pixel points to be selected according to the disparity map in the above embodiments are all parameters obtained by training the depth recognition model.

[0086] In this embodiment, adjusting the distance between the pixel points obtained after accumulation based on the pixel point pairs is to adjust the distance between the pixel points obtained after accumulation according to the correlation features between the adjacent pixel points in the pixel point pairs, so as to avoid distortion of the image content in the combined view.

[0087] In this embodiment, calculating the three-dimensional coordinates corresponding to the key point pixel coordinates of the hundred-meter stake image in the image to be recognized based on the pixel difference between the matched pixel points. The specific manner of calculating the three-dimensional coordinates in this step can refer to the process of binocular image depth calculation in related technologies, which will not be elaborated here.

[0088] Optionally, as a preferred embodiment, if the error between the third coordinate obtained according to the first coordinate system, the second coordinate system, and the third coordinate system and the first three-dimensional coordinate is greater than a second preset error threshold, the camera is used to re-take a target image including the hundred-meter stake sign, and the hundred-meter stake sign is positioned based on the re-taken target image.

[0089] In this embodiment, if the error between the third coordinate obtained according to the first coordinate system, the second coordinate system, and the third coordinate system and the first three-dimensional coordinate is greater than the second preset error threshold, it indicates that data recording errors may occur when calculating the third coordinate using the scale board. At this time, the accurate coordinates of the hundred-meter stake sign cannot be determined, and it is necessary to use the camera to re-take a target image including the hundred-meter stake sign, and then position the hundred-meter stake sign based on the re-taken target image. Among them, the second preset error threshold is the maximum error between the recognition result of the depth recognition model and the true result when testing the depth model. c As an embodiment, the training process of the image recognition model is described below. The following training process is only an example provided by this application, and this application does not limit the training process of the image recognition model.

[0090] First, collect image frames captured from the camera in a real road scene that include the hundred-meter stake sign, and construct a hundred-meter stake sign image dataset. The hundred-meter stake sign image dataset includes samples in different seasons, different light intensities, different angles, and different times of the day. According to the environmental information corresponding to different samples and the sign information on the hundred-meter stake sign, the hundred-meter stake sign image dataset is manually annotated to obtain the category of the hundred-meter stake sign and the pixel coordinates of the bounding box.

[0091] The above-obtained hundred-meter stake sign image dataset is expanded by data augmentation. The data augmentation means can include at least one of random scaling and cropping, random horizontal flipping and rotation, mixing, Gaussian noise addition, HSV transformation, and copy-paste.

[0092] According to the pixel coordinates of the bounding box of the hundred-meter stake sign, the manually annotated bounding box of the hundred-meter stake sign is clustered (that is, the bounding boxes with close pixel coordinates are clustered into one category), and multiple prior boxes of different sizes are generated according to the clustered bounding boxes of the hundred-meter stake sign.

[0093]

[0094] ​Set the initial image recognition model (for example, the YOLOv5 neural network model can be used). Randomly shuffle the dataset of the 100-meter stake sign images after data augmentation and expansion above, and divide it into a training set, a validation set, and a test set. Iteratively train the training set and the validation set to update the parameters of the initial image recognition model, retain the training weights of the round with the highest average precision, and then evaluate the performance of the updated image recognition model through the test set. If the precision rate of the detected image recognition model reaches above the preset precision rate threshold and the recall rate reaches above the preset recall rate threshold, it is regarded as the completion of the training of the image recognition model.

[0095] Among them, during the training process, the image recognition model can be used to perform object detection on the 100-meter stake signs in the 100-meter stake sign images. The image recognition model can predict the actual positions of the signs in the 100-meter stake sign images of the test set based on multiple prior boxes of different sizes, and then calculate the normalized offset between the actual position of the sign and the prior box according to the validation set, and then adjust the parameters of the image recognition model for iteration until the bounding box information of each 100-meter stake sign finally output by the image recognition model passes the test of the test set.

[0096] According to an embodiment of another aspect, the present invention provides a positioning device for 100-meter stake signs. Figure 4 The schematic block diagram of the positioning device for 100-meter stake signs according to an embodiment is shown. It can be understood that the device can be implemented by any device, equipment, platform, and device cluster with computing and processing capabilities. As Figure 4 shown, the device includes:

[0097] An acquisition unit 401, configured to acquire a target image including a 100-meter stake sign captured by a camera;

[0098] An identification unit 402, configured to input the target image into a trained image recognition model to identify the sign information on the 100-meter stake sign in the target image and the key point pixel coordinates corresponding to the 100-meter stake sign in the target image;

[0099] A positioning unit 403, configured to acquire the camera parameters of the camera, perform coordinate transformation on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters to obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establish and record the association relationship between the longitude and latitude coordinates and the sign information.

[0100] As a preferred implementation manner, the image recognition model includes a sign extraction module and a sign recognition module;

[0101] The sign extraction module is used to extract the sign image block corresponding to the 100-meter stake sign from the target image, determine the key point pixels in the sign image block of the 100-meter stake, and obtain the key point pixel coordinates corresponding to the 100-meter stake sign;

[0102] The sign recognition module is used to recognize the sign image block of the 100-meter stake to obtain the sign information on the 100-meter stake sign.

[0103] As a preferred embodiment, the sign recognition module recognizes the sign image block of the 100-meter stake, including:

[0104] Divide the sign image block of the 100-meter stake according to a preset zoning rule to obtain a 100-meter area image block and a kilometer area image block;

[0105] Determine the classification corresponding to the 100-meter area image block through the classification model in the sign recognition module, and obtain the 100-meter number in the 100-meter stake sign based on the classification corresponding to the 100-meter area image block; different classifications in the classification model correspond to different 100-meter numbers;

[0106] Recognize the kilometer area image block through the character recognition sub-module in the sign recognition module to obtain the kilometer number in the 100-meter stake sign; the sign information includes the 100-meter number and the kilometer number in the 100-meter stake sign.

[0107] As a preferred embodiment, the image recognition model further includes an image processing module:

[0108] The image processing module is used to convert the image format of the sign image block of the 100-meter stake into a preset format, filter the pixel blocks that meet the filtering rules in the sign image block of the preset format to obtain a filtered sign image block of the 100-meter stake; the filtering rules are used to enhance the difference between the pixels corresponding to the 100-meter stake sign and other pixels in the sign image block of the 100-meter stake.

[0109] As a preferred embodiment, the positioning unit 403 performs coordinate conversion on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters, including:

[0110] Establish a first coordinate system, a second coordinate system, and a third coordinate system; the first coordinate system takes the pixel point at the upper left corner of the target image as the origin, the width of the target image as the u-axis, and the height of the target image as the v-axis, the second coordinate system takes the center of the target image as the origin, the width of the target image as the x-axis, and the height of the target image as the y-axis, and the third coordinate system takes the center of the target image as the origin, the axis parallel to the x c axis as the x cThe axis with the axis coinciding with the optical axis of the camera as the z c axis;

[0111] Based on the key point pixel coordinates, obtain the first coordinates (u, v) of the key point pixel in the first coordinate system;

[0112] According to the preset coordinate conversion formula and the camera parameters, convert the first coordinates (u, v) to the second coordinates (x, y) of the key point pixel in the second coordinate system, and convert the second coordinates (x, y) to the third coordinates (x c , y c , z c );

[0113] Based on the obtained longitude and latitude coordinates corresponding to the target image, convert the third coordinates (x c , y c , z c ) to the longitude and latitude coordinates corresponding to the key point pixel coordinates.

[0114] As a preferred embodiment, the positioning unit 403 obtains the first coordinates (u, v) of the key point pixel in the first coordinate system based on the key point pixel coordinates, including:

[0115] Scale the key point pixel coordinates to the coordinates corresponding to the size of the original image, and the coordinates corresponding to the size of the original image are the first coordinates (u, v).

[0116] As a preferred embodiment, the camera parameters include the unit physical lengths of the horizontal pixels and the vertical pixels on the photosensitive plate of the target camera;

[0117] The positioning unit 403 converts the first coordinates (u, v) to the second coordinates (x, y) of the key point pixel in the second coordinate system, including:

[0118] Convert the first coordinates (u, v) to the second coordinates (x, y) of the key point pixel in the second coordinate system according to the following formula, and the formula is:

[0119]

[0120] where (u0, v0) are the coordinates of the center of the target image in the first coordinate system, dx is the unit physical length of the horizontal pixels on the photosensitive plate, and dy is the unit physical length of the vertical pixels on the photosensitive plate.

[0121] As a preferred embodiment, the third coordinates (x c , y c , zc z in c Determined based on a scale plate with scales placed when the camera captures the target image;

[0122] Convert the second coordinate (x, y) of the positioning unit 403 into the third coordinate (x c , y c , z c ) in the third coordinate system, including:

[0123] Convert the second coordinate (x, y) into the third coordinate (x c , y c , z c ) in the third coordinate system according to the following formula, and the formula is:

[0124]

[0125] where f is the focal length of the camera.

[0126] As a preferred embodiment, the scale plate is located below the camera and parallel to the ground, the origin of the scales on the scale plate is on the optical axis of the camera, and the scale direction is parallel to the x c axis of the third coordinate system;

[0127] The z c is calculated by the following formula:

[0128]

[0129] where z1 is the scale value corresponding to the intersection point between the straight line connecting the center of the target image and the 100 - meter stake sign in the target image and the scales on the scale plate, h is the height between the camera and the ground, and h1 is the height between the camera and the scale plate.

[0130] As a preferred embodiment, the target image is an image frame obtained from a video;

[0131] The coordinate conversion of the key - point pixel coordinates based on the key - point pixel coordinates and the camera parameters includes:

[0132] Obtain at least one image frame containing the 100 - meter stake sign adjacent to the target image as a reference image;

[0133] Input the target image into a trained depth recognition model to identify the first three - dimensional coordinate corresponding to the key - point pixel coordinates;

[0134] Input the reference image into the depth recognition model to identify the second three-dimensional coordinates corresponding to the key point pixel coordinates of the 100-meter stake sign in the reference image;

[0135] Based on the first three-dimensional coordinates and the second three-dimensional coordinates, calculate the camera position deviation when the camera captures the target image and the reference image, calculate the error between the camera position deviation and the longitude and latitude deviation in the camera parameters when the camera captures the target image and the reference image. If the error is less than the first preset error threshold, calculate the longitude and latitude coordinates corresponding to the 100-meter stake sign according to the first three-dimensional coordinates and the second three-dimensional coordinates. If the error is not less than the first preset error threshold, execute the steps of establishing the first coordinate system, the second coordinate system, and the third coordinate system.

[0136] As a preferred embodiment, the steps for the depth recognition model to identify the first three-dimensional coordinates or the second three-dimensional coordinates include:

[0137] Translate each pixel point in the image to be recognized according to different translation lengths, and perform upsampling on the multiple images obtained after translation to obtain multiple feature maps corresponding to the image to be recognized and the disparity map corresponding to each feature map; the image to be recognized is the target image or the reference image, and the disparity map is used to characterize the disparity between the pixels in the feature map and the pixels in the image to be recognized;

[0138] Extract the pixel points in a specific area of each feature map, and form pixel point pairs by combining adjacent pixel points extracted; each pixel point in the image to be recognized has a corresponding pixel point among the pixel points extracted from each feature map;

[0139] Select the pixel points representing the same object according to the disparity map for accumulation respectively, and adjust the distance between the pixel points obtained after accumulation based on the pixel point pairs to obtain the combined view corresponding to the image to be recognized;

[0140] Match the pixel points of the image to be recognized with the pixel points of the combined view, and calculate the three-dimensional coordinates corresponding to the key point pixel coordinates of the 100-meter stake image in the image to be recognized based on the pixel difference between the matched pixel points.

[0141] As a preferred embodiment, if the error between the third coordinate obtained according to the first coordinate system, the second coordinate system, and the third coordinate system and the first three-dimensional coordinate is greater than the second preset error threshold, re-capture the target image including the 100-meter stake sign by the camera, and position the 100-meter stake sign based on the re-captured target image.

[0142] Each embodiment in the present invention is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for relevant content.

[0143] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0144] The specific embodiments described above further elaborate on the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention shall be included in the protection scope of the present invention.

Claims

1. A positioning method for a hundred-meter stake sign, characterized in that, Including: Obtain a target image captured by a camera and containing a hundred-meter stake signboard; Input the target image into a trained image recognition model to identify the signboard information on the hundred-meter stake signboard in the target image and the key point pixel coordinates corresponding to the hundred-meter stake signboard in the target image; Obtain the camera parameters of the camera, and based on the key point pixel coordinates and the camera parameters, perform coordinate transformation on the key point pixel coordinates to obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establish and record the association relationship between the longitude and latitude coordinates and the signboard information; The performing coordinate transformation on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters includes: Establish a first coordinate system, a second coordinate system, and a third coordinate system; the first coordinate system takes the pixel point at the upper left corner of the target image as the origin, the width of the target image as the u-axis, and the height of the target image as the v-axis, the second coordinate system takes the center of the target image as the origin, the width of the target image as the x-axis, and the height of the target image as the y-axis, and the third coordinate system takes the center of the target image as the origin, the axis parallel to the x-axis as the x c axis, the axis parallel to the y-axis as the y c axis, and the axis coinciding with the optical axis of the camera as the z c axis; Obtain the first coordinates (u, v) of the key point pixel in the first coordinate system based on the key point pixel coordinates; According to the preset coordinate transformation formula and the camera parameters, convert the first coordinate (u, v) into the second coordinate (x, y) of the key point pixel in the second coordinate system, and convert the second coordinate (x, y) into the third coordinate (x c , y c , z c ) Based on the obtained longitude and latitude coordinates corresponding to the target image, convert the third coordinate (x c , y c , z c ) into the longitude and latitude coordinates corresponding to the key point pixel coordinates; The target image is an image frame obtained from a video; The performing coordinate transformation on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters includes: Obtain at least one image frame adjacent to the target image and containing the hundred-meter stake signboard as a reference image; Input the target image into a trained depth recognition model to identify the first three-dimensional coordinates corresponding to the key point pixel coordinates; Input the reference image into the depth recognition model to identify the second three-dimensional coordinates corresponding to the key point pixel coordinates of the hundred-meter stake signboard in the reference image; Based on the first three-dimensional coordinates and the second three-dimensional coordinates, calculate the camera position deviation when the camera captures the target image and the reference image, calculate the error between the camera position deviation and the longitude and latitude deviation when the camera captures the target image and the reference image in the camera parameters. If the error is less than the first preset error threshold, calculate the longitude and latitude coordinates corresponding to the hundred-meter stake signboard according to the first three-dimensional coordinates and the second three-dimensional coordinates. If the error is not less than the first preset error threshold, perform the steps of establishing the first coordinate system, the second coordinate system, and the third coordinate system.

2. The method according to claim 1, characterized in that, The image recognition model includes a signboard extraction module and a signboard recognition module; The signboard extraction module is used to extract the hundred-meter stake signboard image block corresponding to the hundred-meter stake signboard from the target image, determine the key point pixels in the hundred-meter stake signboard image block, and obtain the key point pixel coordinates corresponding to the hundred-meter stake signboard; The signboard recognition module is used to recognize the hundred-meter stake signboard image block to obtain the signboard information on the hundred-meter stake signboard.

3. The method according to claim 2, wherein The signboard recognition module recognizing the hundred-meter stake signboard image block includes: Divide the hundred-meter stake signboard image block according to a preset partitioning rule to obtain a hundred-meter area image block and a kilometer area image block; Determine the classification corresponding to the hundred-meter area image block through the classification model in the signboard recognition module, and obtain the number of hundreds in the hundred-meter stake signboard based on the classification corresponding to the hundred-meter area image block; different classifications in the classification model correspond to different numbers of hundreds; The kilometer number in the hundred - meter stake signboard is obtained by recognizing the kilometer - area image block through the character recognition sub - module in the signboard recognition module; the signboard information includes the hundred - meter number and the kilometer number in the hundred - meter stake signboard.

4. The method according to claim 1, wherein The camera parameters include the unit physical lengths of the horizontal pixels and the vertical pixels on the photosensitive plate of the target camera. The conversion of the first coordinate (u, v) to the second coordinate (x, y) of the key - point pixel in the second coordinate system includes: The first coordinate (u, v) is converted to the second coordinate (x, y) of the key - point pixel in the second coordinate system according to the following formula: where (u0, v0) is the coordinate of the center of the target image in the first coordinate system, dx is the unit physical length of the horizontal pixels on the photosensitive plate, and dy is the unit physical length of the vertical pixels on the photosensitive plate.

5. The method according to claim 1, wherein The z coordinate in the third coordinate (x c , y c , z c ) is determined based on a scale plate with scales placed when the camera captures the target image; c ​ Convert the second coordinate (x, y) into a third coordinate (x c , y c , z c ) in the third coordinate system, including: Convert the second coordinate (x, y) to the third coordinate (x c , y c , z c ) in the third coordinate system according to the following formula: where f is the focal length of the camera. The scale plate is located below the camera and parallel to the ground, and the origin of the scale on the scale plate is located on the optical axis of the camera, and the scale direction is parallel to the x-axis of the third coordinate system; c axis; The said z c is calculated by the following formula: where z1 is the scale value corresponding to the intersection point between the straight line connecting the center of the target image and the hundred - meter stake signboard in the target image and the scale of the scale board, h is the height between the camera and the ground, and h1 is the height between the camera and the scale board.

6. The method according to claim 1, wherein The steps for the depth recognition model to recognize the first three - dimensional coordinate or the second three - dimensional coordinate include: Translating each pixel point in the image to be recognized according to different translation lengths, and up - sampling the multiple images obtained after translation to obtain multiple feature maps corresponding to the image to be recognized and the disparity map corresponding to each feature map; the image to be recognized is the target image or the reference image, and the disparity map is used to represent the disparity between the pixels in the feature map and the pixels in the image to be recognized. Extracting the pixel points in a specific area of each feature map, and forming pixel - point pairs by the adjacent pixel points extracted; each pixel point in the image to be recognized has a corresponding pixel point among the pixel points extracted from each feature map. Selectively accumulating the pixel points representing the same object according to the disparity map, and adjusting the distances between the pixel points obtained after accumulation based on the pixel - point pairs to obtain the combined view corresponding to the image to be recognized. Matching the pixel points of the image to be recognized with the pixel points of the combined view, and calculating the three - dimensional coordinates corresponding to the key - point pixel coordinates of the hundred - meter stake image in the image to be recognized based on the pixel differences between the matched pixel points.

7. The method according to claim 1, wherein if the error between the third coordinate obtained according to the first coordinate system, the second coordinate system and the third coordinate system and the first three - dimensional coordinate is greater than the second preset error threshold, a target image including the hundred - meter stake signboard is re - photographed by the camera, and the hundred - meter stake signboard is positioned based on the re - photographed target image.

8. A positioning device for a hundred-meter stake sign, characterized in that, Based on the method according to any one of claims 1 - 7, it includes: An acquisition unit for acquiring a target image containing a hundred - meter stake signboard photographed by a camera. An identification unit, configured to input the target image into a trained image recognition model, and identify the sign information on the 100-meter stake sign in the target image and the key point pixel coordinates corresponding to the 100-meter stake sign in the target image; A positioning unit, configured to obtain the camera parameters of the camera, perform coordinate conversion on the key point pixel coordinates based on the key point pixel coordinates and the camera parameters, obtain the longitude and latitude coordinates corresponding to the key point pixel coordinates, and establish and record the association relationship between the longitude and latitude coordinates and the sign information.

Citation Information

Patent Citations

  • Traffic sign automatic extraction method and system based on binocular CCD camera

    CN112101299A

  • Road sign generation method and device for high-precision map, electronic equipment and storage medium

    CN114677458A

  • Guideboard generation method, device and equipment

    CN114820783A

  • Highway stake mark matching method based on binocular stereoscopic vision

    CN116721408A