Rod recognition method and device in image, computer device and storage medium

By performing rod detection and epipolar line search on multiple consecutive images, the same rod can be identified and located, solving the problem of recognition accuracy under the influence of changes in lighting in outdoor environments, and achieving high-precision rod recognition and map updating.

CN117011739BActive Publication Date: 2025-11-28TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211163449.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-23
Publication Date
2025-11-28
Estimated Expiration
2042-09-23

AI Technical Summary

Technical Problem

When existing technologies identify rod-shaped objects in outdoor environments, they are greatly affected by changes in lighting, resulting in low identification accuracy.

Method used

By performing rod detection on multiple consecutive images, at least two consecutive images are identified, including a reference image and the previous image. Based on the initial rod recognition result of the reference image and the projection of the detection points in the previous image, the same rod is identified. The accuracy of recognition is improved by localization processing through epipolar search and matching points of fitted straight lines.

Benefits of technology

It effectively reduces the impact of lighting changes on recognition, improves the recognition accuracy of rod-shaped objects in outdoor environments, and is suitable for high-precision map updates and advanced driver assistance systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011739B_ABST
    Figure CN117011739B_ABST
Patent Text Reader

Abstract

The application relates to a rod-shaped object recognition method and device in an image, computer equipment, a storage medium and a computer program product. The method can be applied to the field of maps and comprises the following steps: based on initial recognition results of rod-shaped objects in multiple continuous images, determining at least two continuous images in which rod-shaped objects exist in the multiple continuous images; based on a fitting straight line of the initial recognition results of the rod-shaped objects in a reference image and the projection of a detection point closest to a reference surface in a previous image on the reference image, recognizing the same rod-shaped object in the reference image and the previous image, obtaining a matching result of the reference image and the previous image; projecting the detection point of the same rod-shaped object in the previous image to the reference image through epipolar line searching to obtain an epipolar line; and based on the matching points of the epipolar line and the fitting straight line, performing positioning processing on the rod-shaped object in the reference image to obtain the recognition result of the rod-shaped object in the reference image. The application can improve the accuracy of rod-shaped object recognition and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a rod-shaped object recognition method and device in image, computer device, storage medium and computer program product. BACKGROUND

[0002] With the development of computer technology and artificial intelligence technology, computer vision technology (Computer Vision, CV) appears, which is a science of how to make machines "see". Further, it refers to using cameras and computers to replace human eyes to identify, follow and measure targets, and further process images to make them more suitable for human observation or transmission to instruments for detection. Target detection and positioning is an application of computer vision technology. For example, the current road-side rod-shaped objects and lane lines can be detected to assist in building high-precision maps.

[0003] At present, for rod-shaped objects on the road, a line segment matching can be performed by calculating a descriptor, so as to realize the matching of rod-shaped objects. However, this method is easily affected by factors such as light, and can only be applied to indoor and structured objects. The structure of outdoor rod-shaped objects is similar, and the light changes greatly. Therefore, the accuracy of rod-shaped object matching using the calculation descriptor is low. SUMMARY

[0004] Therefore, it is necessary to provide a rod-shaped object recognition method and device in image, computer device, computer readable storage medium and computer program product, which can effectively improve the accuracy of rod-shaped object recognition and matching.

[0005] In a first aspect, the present application provides a rod-shaped object recognition method in image. The method comprises:

[0006] Based on the initial rod-shaped object recognition result obtained by detecting rod-shaped objects in a plurality of continuous images, at least two continuous images in which rod-shaped objects exist in the plurality of continuous images are determined, and the at least two continuous images include a reference image and a previous image of the reference image;

[0007] Based on the fitting straight line of the initial rod-shaped object recognition result of the reference image, and the projection of the detection point closest to the reference surface in the rod-shaped object recognition result of the previous image in the reference image, the same rod-shaped object in the reference image and the previous image is recognized.

[0008] The detection point closest to the reference surface of the same rod-shaped object in the previous image is projected to the reference image by epipolar search to obtain an epipolar line.

[0009] Based on the epipolar line and the matching point of the fitting straight line, the rod-shaped object in the reference image is positioned and processed to obtain a recognition result of the rod-shaped object in the reference image.

[0010] In a second aspect, the present application further provides an image rod-shaped object recognition device. The device comprises:

[0011] A target recognition module is configured to determine at least two continuous images in which rod-shaped objects exist in a plurality of continuous images based on initial recognition results of rod-shaped objects obtained by performing rod-shaped object detection on the plurality of continuous images, wherein the at least two continuous images comprise a reference image and a previous image of the reference image.

[0012] A rod-shaped object matching module is configured to identify a same rod-shaped object in the reference image and the previous image based on a fitting straight line of the initial recognition result of the rod-shaped object in the reference image and a projection of a detection point in the previous image that is closest to a reference surface in the reference image.

[0013] An epipolar line searching module is configured to project the detection point of the same rod-shaped object in the previous image that is closest to the reference surface to the reference image by epipolar line searching to obtain an epipolar line.

[0014] A rod-shaped object recognition module is configured to perform positioning and processing on the rod-shaped object in the reference image based on the epipolar line and the matching point of the fitting straight line to obtain a recognition result of the rod-shaped object in the reference image.

[0015] In a third aspect, the present application further provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:

[0016] Based on initial recognition results of rod-shaped objects obtained by performing rod-shaped object detection on a plurality of continuous images, at least two continuous images in which rod-shaped objects exist in the plurality of continuous images are determined, wherein the at least two continuous images comprise a reference image and a previous image of the reference image.

[0017] Based on a fitting straight line of the initial recognition result of the rod-shaped object in the reference image and a projection of a detection point in the previous image that is closest to a reference surface in the reference image, a same rod-shaped object in the reference image and the previous image is identified.

[0018] The detection point of the same rod-shaped object in the previous image that is closest to the reference surface is projected to the reference image by epipolar line searching to obtain an epipolar line.

[0019] Based on the epipolar line and the matching point of the fitting straight line, a positioning process is performed on the rod-shaped object in the reference image to obtain a recognition result of the rod-shaped object in the reference image.

[0020] In a fourth aspect, the present application further provides a computer readable storage medium. The computer readable storage medium has a computer program stored thereon, and the computer program, when executed by a processor, implements the following steps:

[0021] Based on the initial recognition result of the rod-shaped object obtained by performing rod-shaped object detection on the plurality of continuous images, at least two continuous images in which the rod-shaped object exists are determined from the plurality of continuous images, and the at least two continuous images include the reference image and a previous image of the reference image.

[0022] Based on the fitting straight line of the initial recognition result of the rod-shaped object of the reference image and the projection of the detection point closest to the reference plane in the recognition result of the rod-shaped object of the previous image on the reference image, the same rod-shaped object in the reference image and the previous image is recognized.

[0023] The detection point of the same rod-shaped object closest to the reference plane in the previous image is projected onto the reference image by epipolar line search to obtain an epipolar line.

[0024] Based on the epipolar line and the matching point of the fitting straight line, a positioning process is performed on the rod-shaped object in the reference image to obtain a recognition result of the rod-shaped object in the reference image.

[0025] In a fifth aspect, the present application further provides a computer program product. The computer program product includes a computer program, and the computer program, when executed by a processor, implements the following steps:

[0026] Based on the initial recognition result of the rod-shaped object obtained by performing rod-shaped object detection on the plurality of continuous images, at least two continuous images in which the rod-shaped object exists are determined from the plurality of continuous images, and the at least two continuous images include the reference image and a previous image of the reference image.

[0027] Based on the fitting straight line of the initial recognition result of the rod-shaped object of the reference image and the projection of the detection point closest to the reference plane in the recognition result of the rod-shaped object of the previous image on the reference image, the same rod-shaped object in the reference image and the previous image is recognized.

[0028] The detection point of the same rod-shaped object closest to the reference plane in the previous image is projected onto the reference image by epipolar line search to obtain an epipolar line.

[0029] Based on the epipolar line and the matching point of the fitting straight line, a positioning process is performed on the rod-shaped object in the reference image to obtain a recognition result of the rod-shaped object in the reference image.

[0030] The image rod recognition method, device, computer device, storage medium and computer program product determine at least two continuous images in which rods exist in the plurality of continuous images based on the initial rod recognition result obtained by rod detection on the plurality of continuous images, so as to determine the two images that need to be processed. The projection of the detection point in the previous frame image that is closest to the reference plane in the initial rod recognition result of the reference image is identified as the same rod in the reference image and the previous frame image based on the fitting straight line of the initial rod recognition result of the reference image and the detection point in the rod recognition result of the previous frame image that is closest to the reference plane, so as to obtain the matching result of the reference image and the previous frame image. The detection point in the previous frame image that is closest to the reference plane of the same rod is projected to the reference image through the epipolar search, so as to obtain the epipolar line. Thus, the error of the projection is reduced through the epipolar line, and the accurate position of each rod in the reference image is determined through the matching point of the epipolar line and the fitting straight line, so as to obtain the recognition result of the rods in the reference image. BRIEF DESCRIPTION OF DRAWINGS

[0031] Figure 1 An application environment diagram of the image rod recognition method in an embodiment;

[0032] Figure 2 A flowchart of the image rod recognition method in an embodiment;

[0033] Figure 3 A schematic diagram of the initial rod recognition result in an embodiment;

[0034] Figure 4 A schematic diagram of the rod beside the road in an embodiment;

[0035] Figure 5 A schematic diagram of the high-precision map vectorization result in an embodiment;

[0036] Figure 6 A schematic diagram of the high-precision map vectorization result in another embodiment;

[0037] Figure 7 A schematic diagram of the projection process of the homography matrix in an embodiment;

[0038] Figure 8 A schematic diagram of the detection point projection result of the straight line process and the turning process in an embodiment;

[0039] Figure 9 A schematic diagram of the rod projection result in an embodiment;

[0040] Figure 10 A schematic diagram of the epipolar search result in an embodiment;

[0041] Figure 11 A schematic diagram of high-precision mapping results in an embodiment;

[0042] Figure 12 A schematic diagram of a sliding window in an embodiment;

[0043] Figure 13 A block diagram of a structure of a rod recognition device in an image in an embodiment;

[0044] Figure 14 An internal structure diagram of a computer device in an embodiment. DETAILED DESCRIPTION

[0045] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0046] In this document, it should be understood that the terms involved:

[0047] Visual Inertial Odometry (VIO): also called Visual-Inertial System (VINS), is an algorithm that fuses camera and Inertial Measurement Unit (IMU) data to realize Simultaneous Localization and Mapping (SLAM).

[0048] PreIntegration: under the condition of knowing the IMU state quantity (attitude and velocity, displacement) at the last time, the linear acceleration and angular velocity measured by the IMU are used to do integration operation to obtain the state quantity at the current time.

[0049] Sliding-Window algorithm: operating on a string or array of a specific size, rather than operating on the entire string and array, which reduces the complexity of the problem, thereby also reducing the nesting depth of the loop. In the present application, it refers to following the rod-shaped object by moving the time-sequentially continuous image frames, rather than following in adjacent frames, which effectively avoids the influence of single image missed detection on vectorization, thereby ensuring more reliable rod matching.

[0050] Essential matrix: also called E matrix, reflecting the relationship between the representations of the image point of a point P in space in the camera coordinate system under different viewing angles of the camera.

[0051] Fundamental matrix: also called F matrix, the point in image coordinate system realizes the matching between frames through F matrix, and is similar to E matrix, generally uses F matrix to do polar line search.

[0052] Homography matrix: also called H matrix, is the perspective transformation of a plane in real world and corresponding image thereof; the image is transformed from one view to another view through the perspective transformation, and the specific introduction can refer to the following blog.

[0053] Normalized plane: the normalized plane is simultaneously divided by Z (depth direction) of the three-dimensional point in the camera coordinate system.

[0054] Triangulation: triangulation, also called triangulation, refers to the angle of a feature point in a three-dimensional space observed from different positions, so as to measure the depth value of the point.

[0055] Bundle adjustment: simply called BA, the pose of the camera and the three-dimensional coordinates of the measured point are taken as unknown parameters, the feature point coordinates detected on the image for the forward intersection are taken as observation data, so as to adjust and obtain the optimal camera parameters and world point coordinates.

[0056] The rod recognition method in the image provided by the embodiments of the present application can be applied to, for example Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be integrated on the server 104, or placed on the cloud or other servers. When the terminal 102 collects multiple frames of continuous images, in order to locate the rod-shaped object existing in the image, the multiple frames of continuous images and the corresponding pose data (the pose of the image acquisition device for collecting the image) and position data can be sent to the server 104, and the server 104 is used to realize the accurate positioning of the rod-shaped object in the image. After the server 104 obtains multiple frames of continuous images, the initial recognition result of the rod-shaped object obtained by detecting the rod-shaped object in the multiple frames of continuous images is used to determine at least two frames of continuous images in which the rod-shaped object exists in the multiple frames of continuous images, and the at least two frames of continuous images include a reference image and a previous frame image of the reference image; based on the projection of the detection point closest to the reference surface in the rod-shaped object recognition result of the previous frame image in the fitting straight line of the rod-shaped object initial recognition result of the reference image, the same rod-shaped object in the reference image and the previous frame image is recognized; the detection point closest to the reference surface of the same rod-shaped object in the previous frame image is projected to the reference image through epipolar search to obtain an epipolar line; based on the matching points of the epipolar line and the fitting straight line, the rod-shaped object in the reference image is positioned and processed to obtain the recognition result of the rod-shaped object in the reference image. Among them, the terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers.

[0057] In one embodiment, as shown in Figure 2 , an image rod-shaped object recognition method is provided, which can be applied to a terminal or a server. The following will be described by taking the case that the method is applied to the server 104 in Figure 1 , including the following steps:

[0058] Step 201, based on the initial recognition result of the rod-shaped object obtained by detecting the rod-shaped object in multiple frames of continuous images, determine at least two frames of continuous images in which the rod-shaped object exists in the multiple frames of continuous images, and the at least two frames of continuous images include a reference image and a previous frame image of the reference image.

[0059] Among them, multi-frame continuous images refer to a series of images continuously acquired by the same image acquisition device. These images are ordered according to the shooting time and are also the detection target of this application. In these multi-frame continuous images, the same rod-shaped object may exist in consecutive images. Therefore, this application uses the existence of the same rod-shaped object in consecutive images to achieve accurate positioning. Rod-shaped object detection refers to the identification of rod-shaped objects in an image through image recognition technology. The initial rod-shaped object recognition result refers to the rod-shaped object detection result identified by image recognition technology. The initial rod-shaped object recognition result is as follows: Figure 3 As shown, the system consists of multiple sets of detection points contained in each image. Each set of detection points contains multiple detection points, and the detection points in the same set are aligned in a straight line, representing a rod-shaped object. "At least two consecutive images" does not mean that only two consecutive images in the multiple consecutive images contain the rod-shaped object; rather, it means that two or more images in the multiple consecutive images contain the rod-shaped object. However, the solution in this application only requires two images to achieve rod-shaped object localization. The reference image refers to the image used as the standard for rod-shaped object localization; rod-shaped object localization refers to locating the rod-shaped objects present in the reference image.

[0060] Specifically, the solution of this application mainly obtains the specific location of the rod-shaped objects in the image by identifying and locating them. When the user of terminal 102 needs to perform positioning processing, they can provide server 104 with a series of continuously collected images, poses, and position data, and server 104 will then perform positioning processing on the rod-shaped objects in the images. After obtaining this data, server 104 will first perform rod detection on multiple consecutive images to determine the initial recognition result of the rod-shaped objects in each frame. Thus, the specific location of the rod-shaped objects in each image is initially determined. In one embodiment, the solution of this application can be applied to the drawing of high-precision maps, such as... Figure 4 As shown, pole-shaped objects in roads are one of the most important semantic features of high-precision maps, which can be used for semantic localization in lane-level navigation and advanced autonomous driving assistance. However, data collected by surveying vehicles is insufficient for updating high-precision maps, especially in urban roads where data changes are measured weekly or even daily. Timely detection of changes (missing or added) in pole-shaped objects is crucial for the freshness of high-precision maps. Therefore, the solution in this application can perform pole-shaped object recognition processing on multiple consecutive frames of images captured by vehicle-mounted cameras, thereby updating the high-precision map in real time based on the pole-shaped object recognition results. Figure 5 and Figure 6The high-precision map vectorization result shown has different types of rod-shaped objects distributed on both sides of the road, which can help the vehicle to perform high-level assisted driving. The high-precision map is generally laid by a high-precision collection vehicle, which generally needs to be operated by professional equipment and collection personnel, and the vehicle is few and the cost is high, so the update is slow. The scheme of the present application can realize the crowdsourcing update of the rod-shaped objects in the high-precision map, and the general vehicle or mobile phone can complete the data difference, so as to ensure the freshness of the high-precision data. Therefore, after the vehicle terminal submits a plurality of continuous images, the server 104 can perform rod-shaped object detection on the plurality of continuous images to obtain a rod-shaped object initial recognition result. Further high-precision rod-shaped object positioning recognition is performed based on the rod-shaped object initial recognition result. Since the rod-shaped object in the image needs to be positioned, the rod-shaped object in the image needs to be present to perform positioning. In the rod-shaped object detection on the plurality of continuous images to obtain the rod-shaped object initial recognition result, a plurality of pictures containing rod-shaped objects are first selected, and then the rod-shaped object positioning is realized with the aid of these pictures. According to the order corresponding to the pictures, one of them can be selected as a reference picture (the first picture cannot be selected), and the previous picture of the reference picture is obtained to assist in realizing the rod-shaped object recognition. In one embodiment, the scheme of the present application can be applied to the drawing of a high-precision map. At this time, the plurality of continuous images can be a series of images taken by a vehicle-mounted camera. Assuming that there is only one rod-shaped object within 20 meters, the vehicle speed is generally 10-30 m / s, and the frame rate of the vehicle-mounted camera is generally 5-10 Hz. Therefore, 5 to 10 pictures can be selected as a plurality of continuous images, and if the shooting is normal, the same rod-shaped object is contained in these images. Therefore, these can be used as at least two continuous images for rod-shaped object recognition, for example, the picture closest to the shooting time point is selected as the reference picture, and the picture second closest to the shooting time point is selected as the previous picture of the reference picture.

[0061] In step 203, the same rod-shaped object in the reference image and the previous image is recognized based on the fitting straight line of the rod-shaped object initial recognition result of the reference image and the projection of the detection point closest to the reference surface in the rod-shaped object recognition result of the previous image on the reference image.

[0062] The reference plane refers to a plane as a contrast plane in the two pictures, and the ground can be selected as the reference plane because the ground does not change relative to the rod-shaped object in the two pictures. The projection refers to a shadow of a figure projected onto a plane or a line. In this application, the detection point corresponding to at least one rod-shaped object in the previous frame image is projected onto the reference image to obtain the position of the detection point in the reference image. The fitted straight line refers to a plurality of straight lines determined based on the initial identification result of the rod-shaped object. Since the initial identification result of the rod-shaped object includes a plurality of detection points, each group of detection points is on the same straight line. Therefore, a straight line can be fitted for each group of detection points. It is noted that in this process, if the number of detection points in the initial identification result of the rod-shaped object is less than 3, it can be directly omitted to avoid the influence of no detection and outliers. The same rod-shaped object refers to the same rod-shaped object existing in the reference image and the previous frame image, but due to the difference in shooting machines, it exists in the two images at the same time.

[0063] Specifically, since the same rod-shaped object exists in the two continuous frames of images, in order to match these same rod-shaped objects, the detection point closest to the reference plane in the rod-shaped object identification result of the previous frame image can be projected onto the reference image through the projection technology. In a specific embodiment, since the shooting device may have moved during image shooting, the projection can be performed based on the pose data corresponding to the reference image and the previous frame image of the reference image. In the specific projection, a homography matrix can be constructed based on the reference image and the previous frame image of the reference image, and then the projection is realized through the homography matrix. After the detection point closest to the reference plane in the rod-shaped object identification result of the previous frame image is projected onto the reference image, the approximate distance between the rod-shaped object corresponding to the detection point in the previous frame image and the rod-shaped object corresponding to the fitted straight line of the initial identification result of the rod-shaped object can be determined through the projection of the detection point and the fitted straight line of the initial identification result of the rod-shaped object, so as to determine whether the two rod-shaped objects are the same rod-shaped object. In a specific embodiment, since the shooting machines of the previous and subsequent images in the plurality of continuous frames of images are separated by a short interval, the same rod-shaped object can exist, and the scheme of the present application is to identify the same rod-shaped object to perform positioning identification. Therefore, after the projection is completed, the straight line distance between the projection of the detection point and the fitted straight line of the initial identification result of the rod-shaped object is determined first, and then it is determined whether the rod-shaped objects corresponding to the two are the same rod-shaped object. If yes, subsequent rod-shaped object identification can be performed.

[0064] In step 205, the detection point closest to the reference plane of the same rod-shaped object in the previous frame image is projected onto the reference image through the epipolar search to obtain an epipolar line.

[0065] The epipolar search refers to determining the epipolar line in the reference image. The epipolar line is a concept in epipolar geometry, and refers to the intersection line of the epipolar plane and the image. The epipolar geometry describes the internal projective relationship between two views, which is independent of the external scene and only depends on the internal parameters of the image acquisition device and the relative pose between the two views.

[0066] Specifically, the scheme of the present application determines two images of the reference image and the previous frame image, takes the two images as two views of epipolar geometry, and then performs epipolar search to determine the epipolar line in the reference image. The projection here is different from the projection process in step 203, which is only an approximate projection. Here, the projection is performed by the way of epipolar search, and the epipolar line in the reference image can be accurately searched based on the determined same rod-shaped object. When the relative pose between the two images is known as R, t. The formula of the epipolar search is specifically:

[0067] l=K -T t x RK -1 x

[0068] wherein K, t x and R represent the internal parameters of the camera, the skew-symmetric matrix of the translation vector, and the rotation matrix, respectively.

[0069] In step 207, the rod-shaped object in the reference image is located based on the matching points on the epipolar line and the fitting straight line, and the recognition result of the rod-shaped object in the reference image is obtained.

[0070] The matching points refer to the intersection points of the epipolar line and the fitting rod in the two-dimensional reference image. The location processing specifically refers to the process of determining the specific position of the rod-shaped object corresponding to the matching points in the world coordinate system. The specific position determined is the recognition result of the rod-shaped object in the reference image.

[0071] Specifically, by converting the two-dimensional matching points into triangulation, the position of the rod-shaped object corresponding to the matching points in the reference image in the camera coordinate system can be obtained by positioning. Then, by coordinate system conversion, the position in the camera coordinate system can be converted into the world coordinate system, so as to determine the accurate position of the rod-shaped object in the world coordinate system, complete the positioning, and obtain the recognition result of the rod-shaped object in the reference image. In one of the embodiments, the scheme of the present application can be applied to the drawing of a high-precision map, at this time, the rod-shaped object position determined can be used to realize the rod-shaped object position updating in the high-precision map, and the rod-shaped object can also be used as an anchor point of the road to realize the positioning of the absolute accuracy of the lane line. In another embodiment, the scheme of the present application is suitable for visual matching of lane-level navigation, at this time, the determined rod-shaped object position and the high-precision master database data are matched to determine the absolute position of the vehicle in the world coordinate system. In addition, the scheme of the present application can also be applied to high-level assisted driving, and the absolute position of the vehicle in the world coordinate system can also be determined through visual matching.

[0072] The above rod-shaped object recognition method in the image is based on the initial recognition result of the rod-shaped object obtained by detecting the rod-shaped object in the plurality of continuous images, to determine at least two continuous images in which the rod-shaped object exists in the plurality of continuous images; so as to determine the two images to be processed. Then, the projection of the detection point closest to the reference surface in the rod-shaped object recognition result of the previous image on the reference image is identified, the same rod-shaped object in the reference image and the previous image is recognized, the matching result of the reference image and the previous image is obtained, and the detection point closest to the reference surface of the same rod-shaped object in the previous image is projected onto the reference image through epipolar search to obtain an epipolar line; so as to reduce the error of the projection through the epipolar line, and then based on the matching points of the epipolar line and the fitting straight line, the rod-shaped object in the reference image is positioned and processed to determine the accurate position of each rod-shaped object in the reference image, and the recognition result of the rod-shaped object in the reference image is obtained.

[0073] In one embodiment, before step 201, the method further comprises: performing rod-shaped object recognition processing on the plurality of continuous images through the object detection network to obtain a plurality of groups of detection points, and taking the plurality of groups of detection points as the initial recognition result of the rod-shaped object in the plurality of continuous images.

[0074] The object detection network specifically refers to a network mechanism similar to lanenet, which is generally used for lane line detection at present. It can extract the position of the lane line in the road image through the photographed road image. The present application realizes the detection of the rod-shaped object through the object detection network, and can detect the relatively vertical rod-shaped object existing in the image. The detection point refers to a point on the detected rod-shaped object, and each group of detection points corresponds to a detected rod-shaped object.

[0075] Specifically, the preliminary pole-shaped object detection in the present application can be implemented by a lane net-like object detection network. The pole-shaped object can be taken as a detection target to train the initial object detection network. Then, each frame of image in a plurality of continuous images is input into the trained object detection network, respectively, to realize the detection and labeling of the pole-shaped object in the image by the object detection network. The labeling refers to labeling the pole-shaped object existing in the image by detection points. Each pole-shaped object corresponds to a group of detection points connected in a line. Finally, the initial recognition result of the pole-shaped object is obtained according to the detection result. In one embodiment, the initial recognition result of the pole-shaped object can be screened again after being obtained to remove the pole-shaped object with less than 3 detection points, so as to avoid the influence of no detection and outliers on the pole-shaped object recognition. In one embodiment, the scheme of the present application can be applied to the drawing of a high-precision map. At this time, the pole-shaped objects to be recognized specifically include road signs, street lamps, traffic signal lights and the like beside the road. The labeling pictures containing the road signs, street lamps and traffic signal lights can be taken as training data to train the object detection network, so as to realize the recognition of the pole-shaped objects in the road. In the embodiment, the initial recognition of the pole-shaped object is realized by the object detection network, which can effectively ensure the accuracy of the pole-shaped object recognition.

[0076] In one embodiment, before step 203, the method further includes: constructing a homography matrix based on the pose data of the at least two continuous images; and projecting the detection point closest to the reference plane in the pole-shaped object recognition result of the previous frame of image to the reference image by the homography matrix.

[0077] The pose data refers to the position and attitude. The attitude is described by attaching a coordinate system to the object and then giving the description of the coordinate system relative to the reference system, i.e. by a rotation matrix to describe the direction. The pose data of the at least two continuous images specifically refers to the position and attitude of the image acquisition device when acquiring the two images. The homography matrix, also called H matrix, is a perspective transformation of a plane in the real world and its corresponding image. The perspective transformation is used to realize the transformation of the image from one view to another view. In the scheme of the present application, the detection point is transformed from the view of the previous frame of image to the view of the reference image. The reference plane refers to an objective plane existing at the same time as the at least two continuous images, which can exist as a reference plane. In specific embodiments, the ground can be selected as the reference plane.

[0078] Specifically, the scheme of the present application can project the rod-shaped object reference points in the previous frame image onto the reference image through a homography matrix, so that the homography matrix can be constructed by the pose data of the image acquisition device corresponding to the two frames of images before the projection. In one embodiment, the projection process specifically projects the detection points closest to the reference plane in the rod-shaped object recognition result of the previous frame image to the image normalization plane through the homography matrix, and then converts the coordinates on the image normalization plane into coordinates on the reference image in combination with the intrinsic parameters of the image acquisition device. The image normalization plane is a plane obtained by dividing the three-dimensional points of the image acquisition device coordinate system by Z (the value in the depth direction). In this embodiment, the homography matrix is constructed by the pose data, which can accurately project the detection points representing the rod-shaped object recognition result of the previous frame image to the reference image, thereby ensuring the accuracy of rod-shaped object recognition.

[0079] In one embodiment, before constructing the homography matrix based on the pose data of at least two consecutive frames of images, the method further comprises: obtaining speed data of the image acquisition device when acquiring a plurality of consecutive frames of images and initial pose data of the image acquisition device corresponding to the plurality of consecutive frames of images; and pre-integrating the speed data to obtain the pose data corresponding to each frame of image in the at least two consecutive frames of images.

[0080] The speed data specifically includes the measured values of the acceleration and angular velocity of the image acquisition device corresponding to the at least two consecutive frames of images, the bias of the acceleration and the bias of the angular velocity, and can also include the measured values of the speed and the bias of the speed. For the acquisition process of the speed data, for image acquisition devices without wheel speed such as smart phones, the speed is obtained by integrating the acceleration, and for vehicle-mounted cameras with wheel speed, the speed can be directly measured. The pre-integration uses the linear acceleration and angular velocity measured by the inertial measurement unit to obtain the state quantity at the current time by integrating the known state quantity (attitude and speed, displacement) at the previous time.

[0081] Specifically, since the homographic matrix needs to be constructed by combining the pose data of at least two continuous images, the pose data corresponding to each of the two images needs to be estimated first in the scheme of the present application. Therefore, the speed data and initial pose data of the image acquisition device when acquiring a plurality of continuous images can be obtained first, and then the pre-integration calculation is performed based on the initial pose data and the speed data corresponding to each of the plurality of continuous images, the pose data corresponding to the first image is determined based on the initial pose data, and then the pose data corresponding to the second image is determined based on the pose data of the first image. The calculation is continued in this way until the pose data of each image is determined, and then the pose data corresponding to each of the at least two continuous images is determined. In the embodiment, the pre-integration calculation is performed based on the initial pose data of the image acquisition device and the speed data corresponding to each of the images, which can effectively determine the pose data corresponding to each of the at least two continuous images, thereby accurately establishing the homographic matrix and ensuring the accuracy of the rod recognition.

[0082] In one embodiment, the speed data includes a measured value of acceleration, a measured value of angular velocity, a bias of acceleration, and a bias of angular velocity, and the pose data includes position data, speed data, and angle data. The pre-integration processing of the speed data to obtain the pose data corresponding to each of the at least two continuous images includes: based on the measured value of acceleration and the bias of acceleration, determining the position data and the speed data corresponding to each of the at least two continuous images; and based on the measured value of angular velocity and the bias of angular velocity, determining the angle data corresponding to each of the at least two continuous images.

[0083] Specifically, the pose data includes position data, speed data, and angle data, so when calculating, the three data need to be calculated respectively, and the final pose data is obtained by comprehensively considering the three data. For the position data and the speed data, the two data can be determined by the measured value of acceleration and the bias of acceleration. For the angle data, it is determined based on the measured value of angular velocity and the bias of angular velocity. The specific calculation formula of the position data is:

[0084]

[0085] The specific calculation formula of the speed data is:

[0086] v←v+(R(a m -a b )+g)Δt

[0087] Wherein, p and v before the arrow respectively represent the position data and the speed data in the previous frame image pose, and p and v after the arrow represent the position data and the speed data of the current frame image.a m and a b respectively represent the measured value data of the acceleration and the bias of the acceleration. R represents the rotation matrix between adjacent frame images, and g is the gravity acceleration. Δt represents the motion time between two frames.

[0088] For the angle data, the specific calculation formula is:

[0089]

[0090] Wherein, q before the arrow represents the angle data in the previous frame image pose, and q after the arrow represents the angle data of the current frame image.w m and w b respectively represent the measured value data of the angular velocity and the bias of the angular velocity. In this embodiment, by determining the related data of the acceleration, the position data and the speed data corresponding to each frame image in the plurality of continuous images can be effectively calculated, and by the angular velocity related data, the angle data corresponding to each frame image in the plurality of continuous images can be effectively calculated, so as to obtain the pose data corresponding to each frame image.

[0091] In one of the embodiments, constructing the homography matrix based on the pose data of the at least two continuous images comprises: determining the rotation data and the displacement data of the reference image and the previous frame image relative to the reference plane based on the pose data of the at least two continuous images; and constructing the homography matrix based on the rotation data and the displacement data.

[0092] Wherein, the rotation data and the displacement data are data used for constructing the homography matrix, and the two data can be determined by comparing the pose changes between the previous frame and the next frame. The rotation data represents the rotation angle change of the image acquisition device relative to the reference plane between the previous frame and the next frame, and the displacement data represents the position change of the image acquisition device relative to the reference plane between the previous frame and the next frame.

[0093] Specifically, after obtaining the pose data corresponding to the reference image and the previous frame image respectively, the homography matrix can be constructed by the pose change between the two. Specifically, the rotation data and the displacement data of the reference image and the previous frame image relative to the reference plane can be determined based on the pose data of the at least two continuous images. Then the homography matrix is constructed based on the rotation data and the displacement data. In one of the specific embodiments, the scheme of the present application can be applied to the drawing of high-precision map, at which time the continuous images can be specifically the images collected by the vehicle-mounted camera. For the projection process of the homography matrix, please refer to Figure 7As shown in the figure, where the plane on the right lower side of the figure represents the ground, the line segment through the plane above it represents the rod, the quadrilateral above the ground represents the second plane parallel to the ground and passing through the lowest point of the rod detection, and the arrow represents the normal vector of the ground. The left side is the vehicle body and the camera, the direction of the vehicle body and the imaging plane of the camera are basically coincident (the camera has an installation angle on the vehicle body), and the rotation and translation between the two frames are known with respect to the ground plane, so that the homography matrix can be constructed, and the formula is as follows:

[0094]

[0095] where R represents the rotation matrix between the two frames, t represents the translation between the two frames, n represents the normal vector of the plane, and in the image coordinate system, n is approximately [0, 1, 0] T , and d represents the distance from the camera optical center to the plane. Therefore, the points on the normalized plane of the adjacent frames satisfy the following relationship:

[0096]

[0097] For the points passing through the second plane, the following homography matrix can also be constructed:

[0098]

[0099] where λ is a scale factor, and the value range satisfies: 0<λ<1, and for the points on the ground, λ=1, then the points on the normalized plane between adjacent frames satisfy the following relationship:

[0100]

[0101] where s is a scale factor, so the above formula can also be written as:

[0102]

[0103] And by dividing the first two rows of the matrix by the third row, we get:

[0104]

[0105]

[0106] where u2 and v2 are the theoretical values of the projection of the ground homography matrix to the image normalized plane coordinates, and u′2 and v′2 are the theoretical values of the projection of the intersection of the rod to the image normalized plane coordinates of the current frame. For a car, except for lane changing and uphill, the straight direction is much larger than the up and down and left and right movement, so it satisfies: t z >>t x and t z >>t yTherefore, when the vehicle is driving straight, u'2=u2, v'2=v2. When the vehicle is changing lanes or climbing, the straight direction is generally larger than the up and down or left and right directions, and u'2≈u2, v'2≈v2. Thus, the detected points on the previous frame image can be projected to the image normalized plane coordinates. After obtaining the image normalized plane coordinates, the positions of the detected points corresponding to the reference image can be projected to the reference image in combination with the camera internal parameters. Figure 8 In the above embodiment, the left image in the upper part represents the reference image, the right image in the upper part represents the previous frame image, and the vehicle is basically driving straight. The detected points on the left image are the results of the detected points on the right image projected by the homography matrix. The left image in the lower part represents the reference image, and the right image in the lower part represents the previous frame image. The vehicle is basically turning. The detected points on the left image are the results of the detected points on the right image projected by the homography matrix. In the embodiment, the homography matrix is constructed by the change of the pose data, which can effectively ensure the accuracy of the constructed homography matrix, thereby ensuring the accuracy of the pole-shaped object recognition and positioning.

[0107] In one of the embodiments, before the same pole-shaped object in the reference image and the previous frame image is recognized, the fitting straight line of the initial recognition result of the pole-shaped object in the reference image and the projection of the detected point closest to the reference plane in the initial recognition result of the pole-shaped object in the previous frame image on the reference image further include: fitting a straight line to the initial recognition result of the pole-shaped object in the reference image by a first-order function to obtain the fitting straight line corresponding to each initial recognition result of the pole-shaped object.

[0108] Specifically, the straight line fitting can be realized by a first-order function x=ky+b. The first-order function corresponding to the pole-shaped object can be determined by substituting each detected point coordinate in the initial recognition result of the pole-shaped object into the first-order function, and the processing procedure of the straight line fitting can be realized. Through the straight line fitting, the straight line equation corresponding to the initial recognition result of the pole-shaped object can be effectively determined, thereby determining the distance between the coordinates of the projection of the detected point and the straight line equation corresponding to the initial recognition result of the pole-shaped object, and ensuring the accuracy of the pole-shaped object matching.

[0109] In one of the embodiments, the step 205 includes: determining the distance data between the fitting straight line of the initial recognition result of the pole-shaped object in the reference image and the projection of the detected point closest to the reference plane in the initial recognition result of the pole-shaped object in the previous frame image on the reference image; screening the distance data by a preset distance threshold to obtain a matching group composed of the fitting straight line and the projection of the detected point with the distance less than the preset distance threshold; and determining that the initial recognition result of the pole-shaped object corresponding to the fitting straight line in the matching group and the initial recognition result of the pole-shaped object corresponding to the detected point are the same pole-shaped object in the reference image and the previous frame image.

[0110] The distance data between the projection of the detection point and the fitting straight line is specifically that for each detection point, the distance between the detection point and all the fitting straight lines corresponding to the rod-shaped objects in the reference image is calculated. The preset distance threshold is a screening criterion for screening the matching rod-shaped objects, and can be set according to the speed of the image acquisition device and the shooting interval between the previous frame image and the reference image. For each matching group, each matching group includes one detection point and one fitting straight line, and if there are multiple fitting straight lines less than the preset distance threshold for one detection point, the fitting straight line with the smallest distance can be selected to construct the matching group.

[0111] Specifically, the matching relationship between the detection point and the rod-shaped object in the reference image can be determined by the distance between the detection point and the fitting straight line less than the preset distance threshold, that is, the matching relationship between the rod-shaped object in the previous frame image and the rod-shaped object in the reference image is determined. Therefore, the distance data between the projection of the detection point and the fitting straight line is calculated first, and after the projection of the detection point to the reference image, the coordinate corresponding to the detection point in the reference image can be determined, and the distance between the coordinate and each fitting straight line can be calculated through analytic geometry. Then, the matching group composed of the fitting straight line and the projection of the detection point with the distance less than the preset distance threshold can be obtained by screening the distance data in combination with the preset distance threshold. The data constituting the matching group includes the projection of one detection point and one fitting straight line, but in essence, the two represent the rod-shaped object in the previous frame image and the rod-shaped object in the reference image, respectively. Therefore, the initial identification result of the rod-shaped object corresponding to the fitting straight line in the matching group and the initial identification result of the rod-shaped object corresponding to the detection point can be determined based on the matching group, and the final result of the rod-shaped object matching can be obtained, which is the same rod-shaped object in the reference image and the previous frame image. As shown in FIG. 6, the left side represents the reference image, the right side represents the previous frame image, and the points on the left side represent the results obtained by projecting the rod-shaped object in the right image through the homography matrix. In this embodiment, the same rod-shaped object in the two images is screened through the distance between the projection and the fitting straight line, so that the accuracy of the rod-shaped object identification between the two frame images can be effectively ensured. Figure 9

[0112] In one of the embodiments, based on the matching point of the epipolar line and the fitting straight line, the step 207 includes: determining the matching point where the epipolar line and the fitting straight line intersect; performing triangulation processing on the matching point to obtain the image acquisition device coordinate system position of the matching point; and transferring the image acquisition device coordinate system position of the matching point to the world coordinate system to obtain the identification result of the rod-shaped object in the reference image.

[0113] ​The matching point is specifically a point where the epipolar line and the fitting straight line intersect. Triangulation, also called triangulation, refers to measuring the depth value of a point by observing the angle of a feature point in a three-dimensional space from different positions. The image acquisition device coordinate system is a three-dimensional rectangular coordinate system with the focal center of the image acquisition device as the origin and the optical axis as the Z axis. The world coordinate system refers to selecting a reference coordinate system in the environment to describe the positions of the image acquisition device and the object, which is called the world coordinate system. The relationship between the image acquisition device coordinate system and the world coordinate system can be described by a rotation matrix R and a translation vector t.

[0114] Specifically, the feature point in the present application refers to the identified matching point. By triangulating the matching point, the depth value corresponding to the matching point can be determined. Then, the image acquisition device coordinate system position of the matching point is obtained by combining the depth value of the point with the coordinates of the point on the reference image. After obtaining the image acquisition device coordinate system position, it can be converted into the world coordinate system. The position of the rod-shaped object in the reference image is accurately described by the coordinates in the world coordinate system. In one embodiment, the scheme of the present application can be applied to high-precision map drawing. At this time, the image acquisition device is specifically a camera device on the vehicle body. At this time, the position of the matching point in the camera coordinate system can be determined first. According to the coordinate transformation, it is sequentially converted to the vehicle body coordinate system and the world coordinate system, and the position of the rod-shaped object in the world coordinate system is obtained. After obtaining the world coordinate system position of the rod-shaped object, the position can be compared with the parent library data to determine whether there is a newly added rod-shaped object. If there is, the difference is completed. If not, it means that the data of the parent library is not timely and needs to be updated, thereby realizing the drawing and updating operation of the high-precision map. As shown in Figure 10 The left side represents the reference image, and the right side represents the previous frame image of the reference image. The line in the left image is obtained by epipolar search using the line in the right image, and the intersection point of the line and the fitting line is the matching point of the left and right two frame images. As shown in Figure 11 When applied to a high-precision mapping task, the display effect of the high-precision map, the small fine line outside the road is the result of rod vectorization. By matching with the high-precision map, it can be used for data difference or high-level auxiliary driving. The points represent the trajectory points of the vehicle, and the lines on both sides of the road represent the vectorized lane lines. In this embodiment, the coordinate system conversion of the matching point can be effectively realized by triangulation processing, so as to effectively identify the accurate position of the rod-shaped object in the world coordinate system, and ensure the accuracy of the rod-shaped object recognition and positioning.

[0115] In one of the embodiments, after obtaining the recognition result of the rod-shaped object in the reference image based on the matching points of the epipolar line and the fitting straight line, the method further comprises: constructing a sliding window based on at least two continuous images; determining the recognition result of the rod-shaped object corresponding to each image in the sliding window except the first image; constructing a loss function based on the recognition result of the rod-shaped object; optimizing the pose data corresponding to each image in the sliding window through the loss function to obtain optimized pose data; and optimizing the recognition result of the rod-shaped object corresponding to each image based on the optimized pose data.

[0116] The sliding window refers to a sliding window algorithm, which originally operates on a specific size of string or array, rather than on the entire string and array, thereby reducing the complexity of the problem and the nesting depth of the loop. In this application, the sliding window refers to the movement of the continuous image frames in time sequence to follow the rod-shaped object, rather than following the rod-shaped object in adjacent frames, thereby effectively avoiding the influence of single image missing detection on vectorization, and ensuring more reliable rod matching.

[0117] Specifically, the scheme of the present application can position the rod-shaped object in the reference image based on the reference image and the previous image. Thus, it can be extended to other images of at least two continuous images. By constructing a sliding window, the recognition result of the rod-shaped object corresponding to each image in the sliding window except the first image can be determined in sequence based on the above-mentioned recognition method. Then, the rod-shaped object in the sliding window can be subjected to a bundle adjustment method constraint processing, i.e., a loss function is constructed based on the consistency of the position and orientation of the rod-shaped object in the three-dimensional space to constrain the pose of the vehicle, thereby obtaining more accurate local positioning. The optimization function is specifically as follows:

[0118]

[0119] and represent the constraints of pre-integration and vision respectively. The optimal pose can be obtained by minimizing the errors of vision and pre-integration. The sliding window in the present embodiment can refer to FIG. 1 shown in the description. Figure 12 As time goes on, the first frame is discarded and the sixth frame is added to form a new sliding window. In the present embodiment, the coordinate system conversion of the matching points can be effectively realized through triangulation processing, thereby effectively identifying the accurate position of the rod-shaped object in the world coordinate system and ensuring the accuracy of the recognition and positioning of the rod-shaped object. ​

[0120] The application further provides an application scenario of the rod-shaped object recognition method in an image. Specifically, the rod-shaped object recognition method in an image is applied as follows in the application scenario:

[0121] When a user needs to draw a high-precision map to realize automatic driving through the map, the image pole recognition of the application can assist in drawing and updating the poles in the high-precision map. First, the user can find some continuous images collected by a vehicle-mounted camera in a corresponding section of the map. Then, the poles are drawn based on the map, and in the subsequent, the poles in the corresponding section can be updated according to the new images. Specifically, first, on the server side, the pole recognition processing is performed on the multiple continuous images by the object detection network to obtain multiple sets of detection points, and the multiple sets of detection points are taken as the initial recognition results of the poles in the multiple continuous images. Then, based on the initial recognition results of the poles, at least two continuous images in which the poles exist in the multiple continuous images are determined, and the at least two continuous images include a reference image and a previous image of the reference image. At the same time, the pose data corresponding to each image can be determined. First, the speed data of the image collection device when collecting the multiple continuous images and the initial pose data of the image collection device corresponding to the multiple continuous images are obtained. The speed data is pre-integrated to obtain the pose data corresponding to each image in the at least two continuous images. Specifically, based on the measured value of the acceleration and the bias of the acceleration, the position data and the speed data corresponding to each image in the at least two continuous images can be determined. Based on the measured value of the angular velocity and the bias of the angular velocity, the angle data corresponding to each image in the at least two continuous images can be determined. Then, a homography matrix is constructed based on the pose data of the at least two continuous images. The detection point closest to the reference plane in the pole recognition result of the previous image is projected to the reference image through the homography matrix. For the process of constructing the homography matrix, the rotation data and the displacement data of the reference image and the previous image relative to the reference plane can be determined based on the pose data of the at least two continuous images. Then, the homography matrix is constructed based on the rotation data and the displacement data. Then, a first-order function is used to perform straight line fitting on the initial recognition result of the pole in the reference image to obtain a fitting straight line corresponding to each initial recognition result of the pole. Then, the distance data between the projection of the detection point and the fitting straight line is determined. The distance data is filtered through a preset distance threshold to obtain a matching group composed of the fitting straight line and the projection of the detection point with a distance less than the preset distance threshold. It is determined that the fitting straight line corresponding to the initial recognition result of the pole in the matching group and the initial recognition result of the pole corresponding to the detection point are the same pole in the reference image and the previous image. After the same pole is obtained, the detection point closest to the reference plane of the same pole in the previous image is projected to the reference image through epipolar search to obtain an epipolar line. Then, the matching point intersected by the epipolar line and the fitting straight line can be determined. The matching point is triangularized to obtain the image collection device coordinate system position of the matching point. The image collection device coordinate system position of the matching point is transferred to the world coordinate system to obtain the recognition result of the pole in the reference image.Meanwhile, in order to further optimize the processing, a sliding window can be constructed based on at least two continuous images; a rod recognition result corresponding to each image in the sliding window except the first image is determined; a loss function is constructed based on the rod recognition result; pose data corresponding to each image in the sliding window is optimized through the loss function, to obtain optimized pose data; and the rod recognition result corresponding to each image is optimized based on the optimized pose data.

[0122] It should be understood that, although each step in the flowchart involved in each of the above embodiments is shown in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in the flowchart involved in each of the above embodiments can include multiple steps or stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least part of other steps or stages.

[0123] Based on the same inventive concept, the embodiments of the present application also provide an image rod recognition device for implementing the above-mentioned image rod recognition method. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more image rod recognition device embodiments provided below can refer to the limitations of the image rod recognition method in the above text, and will not be repeated here.

[0124] In one embodiment, as shown in Figure 13 An image rod recognition device is provided, comprising:

[0125] The target recognition module 1302 is configured to determine, based on the initial rod recognition result obtained by performing rod detection on the plurality of continuous images, at least two continuous images in which rods exist in the plurality of continuous images, the at least two continuous images including a reference image and a previous image of the reference image.

[0126] The rod matching module 1304 is configured to identify the same rod in the reference image and the previous image based on the fitting straight line of the initial rod recognition result of the reference image and the projection of the detection point closest to the reference surface in the rod recognition result of the previous image to the reference image.

[0127] The epipolar line searching module 1306 is configured to project the detection point closest to the reference surface of the same rod in the previous image to the reference image through epipolar line searching to obtain an epipolar line.

[0128] The rod recognition module 1308 is configured to perform positioning processing on the rods in the reference image based on the epipolar lines and the matching points of the fitted straight lines, to obtain a recognition result of the rods in the reference image.

[0129] In an embodiment, the initial detection module is further configured to: perform rod recognition processing on the plurality of continuous images by using the object detection network to obtain a plurality of groups of detection points, and take the plurality of groups of detection points as initial recognition results of the rods in the plurality of continuous images.

[0130] In an embodiment, the detection point projection module is further configured to: construct a homography matrix based on the pose data of the at least two continuous images; and project, by using the homography matrix, the detection point closest to the reference plane in the rod recognition result of the previous image to the reference image.

[0131] In an embodiment, the pose calculation module is further configured to: obtain speed data of the image acquisition device when the plurality of continuous images are acquired, and initial pose data of the image acquisition device corresponding to the plurality of continuous images; and perform pre-integration processing on the speed data to obtain the pose data corresponding to each of the at least two continuous images.

[0132] In an embodiment, the speed data includes a measured value of acceleration, a measured value of angular velocity, a bias of acceleration, and a bias of angular velocity, and the pose data includes position data, speed data, and angle data; and the pose calculation module is specifically configured to: determine the position data and the speed data corresponding to each of the at least two continuous images based on the measured value of acceleration and the bias of acceleration; and determine the angle data corresponding to each of the at least two continuous images based on the measured value of angular velocity and the bias of angular velocity.

[0133] In an embodiment, the pose calculation module is further configured to: determine rotation data and displacement data of the reference image and the previous image relative to the reference plane based on the pose data of the at least two continuous images; and construct a homography matrix based on the rotation data and the displacement data.

[0134] In an embodiment, the straight line fitting module is further configured to: perform straight line fitting on the initial recognition results of the rods in the reference image by using a first-order function to obtain a fitted straight line corresponding to each of the initial recognition results of the rods.

[0135] In an embodiment, the rod matching module 1304 is specifically configured to: determine distance data between the projection of the detection point and the fitted straight line; and perform screening on the distance data by using a preset distance threshold to obtain a matching group composed of the fitted straight line and the projection of the detection point with a distance less than the preset distance threshold, and determine that the initial recognition result of the fitted straight line and the initial recognition result of the detection point in the matching group are initial recognition results of the same rod in the reference image and the previous image.

[0136] In one embodiment, the rod object recognition module 1308 is specifically configured to: determine the epipolar line and the matching points of the intersection of the fitted straight line; triangulate the matching points to obtain the image acquisition device coordinate system positions of the matching points; and transfer the image acquisition device coordinate system positions of the matching points to the world coordinate system to obtain the recognition result of the rod object in the reference image.

[0137] In one embodiment, the position optimization module is further configured to: construct a sliding window based on at least two continuous images; determine the rod object recognition result corresponding to each image in the sliding window except the first image; construct a loss function based on the rod object recognition result; optimize the pose data corresponding to each image in the sliding window through the loss function to obtain optimized pose data; and optimize the rod object recognition result corresponding to each image based on the optimized pose data.

[0138] The above-mentioned modules of the rod object recognition device in the image can be realized by software, hardware, or a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to the above-mentioned modules.

[0139] In one embodiment, a computer device is provided, which can be a server, and the internal structure diagram thereof can be as shown in Figure 14 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store rod object recognition related data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through a network connection. The computer program is executed by the processor to implement an image rod object recognition method.

[0140] Those skilled in the art can understand that Figure 14The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0141] In an embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method embodiments when executing the computer program.

[0142] In an embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program implements the steps in the above method embodiments when executed by a processor.

[0143] In an embodiment, a computer program product or computer program is provided, including computer instructions stored in a computer readable storage medium. A processor of a computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in the above method embodiments.

[0144] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0145] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.

[0146] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0147] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method of identifying a rod in an image, characterized by, The method comprises: determining at least two frames of continuous images in which a rod exists in the plurality of frames of continuous images based on a rod initial identification result obtained by performing rod detection on the plurality of frames of continuous images, the at least two frames of continuous images comprising a reference image and a previous frame image of the reference image; identifying a same rod in the reference image and the previous frame image based on a projection of a detection point in the previous frame image that is closest to a reference surface in a rod identification result of the previous frame image and a fitting straight line of a rod initial identification result of the reference image; projecting the detection point of the same rod in the previous frame image that is closest to the reference surface to the reference image by epipolar search to obtain an epipolar line; performing positioning processing on a rod in the reference image based on a matching point of the epipolar line and the fitting straight line to obtain an identification result of the rod in the reference image.

2. The method of claim 1, wherein, Before the determining at least two frames of continuous images in which a rod exists in the plurality of frames of continuous images based on a rod initial identification result obtained by performing rod detection on the plurality of frames of continuous images, the method further comprises: performing rod identification processing on the plurality of frames of continuous images by an object detection network to obtain a plurality of groups of detection points, and taking the plurality of groups of detection points as rod initial identification results in the plurality of frames of continuous images.

3. The method of claim 1, wherein, Before the identifying a same rod in the reference image and the previous frame image based on a projection of a detection point in the previous frame image that is closest to a reference surface in a rod identification result of the previous frame image and a fitting straight line of a rod initial identification result of the reference image, the method further comprises: constructing a homography matrix based on pose data of the at least two frames of continuous images; projecting the detection point in the previous frame image that is closest to the reference surface to the reference image by the homography matrix.

4. The method of claim 3, wherein, Before the constructing a homography matrix based on pose data of the at least two frames of continuous images, the method further comprises: obtaining speed data of an image acquisition device when the image acquisition device acquires the plurality of frames of continuous images and initial pose data of the image acquisition device corresponding to the plurality of frames of continuous images; performing pre-integration processing on the speed data to obtain pose data corresponding to each frame of image in the at least two frames of continuous images.

5. The method of claim 4, wherein, The speed data comprises a measurement value of acceleration, a measurement value of angular velocity, a bias of acceleration and a bias of angular velocity, and the pose data comprises position data, speed data and angle data; The pre-integration processing on the speed data to obtain pose data corresponding to each frame of image in the at least two frames of continuous images comprises: determining position data and speed data corresponding to each frame of image in the at least two frames of continuous images based on the measurement value of acceleration and the bias of acceleration; determining angle data corresponding to each frame of image in the at least two frames of continuous images based on the measurement value of angular velocity and the bias of angular velocity.

6. The method of claim 3, wherein, The constructing a homography matrix based on pose data of the at least two frames of continuous images comprises: determining rotation data and displacement data of the reference image and the previous frame image relative to a reference plane based on the pose data of the at least two frames of continuous images; Construct a homography matrix based on the rotation data and the displacement data.

7. The method of claim 1, wherein, Before identifying the same rod-shaped object in the reference image and the previous frame image, the fitting straight line based on the rod-shaped object initial identification result of the reference image and the projection of the detection point closest to the reference plane in the rod-shaped object identification result of the previous frame image in the reference image further include: Perform straight line fitting on the rod-shaped object initial identification result in the reference image through a first-order function to obtain a fitting straight line corresponding to each rod-shaped object initial identification result.

8. The method of claim 1, wherein, Before identifying the same rod-shaped object in the reference image and the previous frame image, the fitting straight line based on the rod-shaped object initial identification result of the reference image and the projection of the detection point closest to the reference plane in the rod-shaped object identification result of the previous frame image in the reference image include: Determine the distance data between the fitting straight line of the rod-shaped object initial identification result of the reference image and the projection of the detection point closest to the reference plane in the rod-shaped object identification result of the previous frame image in the reference image; Filter the distance data through a preset distance threshold to obtain a matching group composed of the projection of the fitting straight line and the detection point with a distance less than the preset distance threshold, and determine that the rod-shaped object initial identification result corresponding to the fitting straight line and the rod-shaped object initial identification result corresponding to the detection point in the matching group are the same rod-shaped object in the reference image and the previous frame image.

9. The method of claim 1, wherein, The matching point based on the epipolar line and the fitting straight line, the positioning processing of the rod-shaped object in the reference image to obtain the identification result of the rod-shaped object in the image include: Determine the matching point intersected by the epipolar line and the fitting straight line; Perform triangulation processing on the matching point to obtain the image acquisition device coordinate system position of the matching point; Transfer the image acquisition device coordinate system position of the matching point to the world coordinate system to obtain the identification result of the rod-shaped object in the reference image.

10. The method according to any one of claims 1 to 9, characterized in that, After obtaining the identification result of the rod-shaped object in the reference image based on the matching point of the epipolar line and the fitting straight line, further include: Construct a sliding window based on the at least two continuous images; Determine the rod-shaped object identification result corresponding to each frame image in the sliding window except the first frame image; Construct a loss function based on the rod-shaped object identification result; Optimize the pose data corresponding to each frame image in the sliding window through the loss function to obtain pose optimization data; Optimize the rod-shaped object identification result corresponding to each frame image based on the pose optimization data.

11. A device for recognizing rod-shaped objects in an image, characterized in that, The device includes: A target identification module configured to determine at least two continuous images in which a rod-shaped object exists in a plurality of continuous images based on a rod-shaped object initial identification result obtained by detecting the rod-shaped object in the plurality of continuous images, the at least two continuous images including a reference image and a previous frame image of the reference image; A rod-shaped object matching module configured to identify the same rod-shaped object in the reference image and the previous frame image based on a projection of a fitting straight line of a rod-shaped object initial identification result of the reference image and a detection point closest to a reference plane in a rod-shaped object identification result of the previous frame image in the reference image; An epipolar line searching module is configured to project the closest point of the same rod-like object to the reference plane in the previous frame image to the reference image by epipolar line searching to obtain an epipolar line; A rod-like object recognizing module is configured to locate the rod-like object in the reference image based on the matching point of the epipolar line and the fitting straight line to obtain a recognition result of the rod-like object in the reference image.

12. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to implement the steps of the method in any one of claims 1 to 10.

13. A computer readable storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 10.

14. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 10. The computer program is executed by the processor to implement the steps of the method in any one of claims 1 to 10.

Citation Information

Patent Citations

  • Airport foreign matter identification method and device, computer equipment and storage medium

    CN108764202A

  • Rod-shaped object fusion method and system based on road structure, server and medium

    CN112508111A