Method for geolocating and qualifying signaling infrastructure devices
Patent Information
- Application Number
- DE602020056403
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-12-02
- Filing Date
- 2020-12-01
- Publication Date
- 2025-08-13
- Estimated Expiration
- 2040-12-01
AI Technical Summary
Current methods for generating high-definition 3D maps are tedious, require significant data processing, and struggle with the complex recognition and geolocation of discrete objects like road signs, especially in the absence of GPS/GNSS signals.
A method using single panoramic images from a spherical camera system combined with neural networks and triangulation algorithms to geolocate discrete objects by determining their position and dimensions, reducing data processing complexity and improving object recognition.
Enables efficient and accurate geolocation of discrete objects like road signs by simplifying data processing and reducing human intervention, while providing reliable 3D map data for autonomous vehicles.
Description
[0001] The present invention relates to a method for geolocation and qualification of signaling infrastructure devices comprising detection and classification of discrete objects such as signaling infrastructure devices and in particular road signs with a view to integrating this data into a high-resolution map. Technical field
[0002] The invention relates to the field of producing high-definition 3D maps for the geolocation of autonomous vehicles by combining GPS / GNSS data, a high-definition map and movement data, in particular using an inertial unit. Prior art
[0003] An autonomous vehicle must know its position relative to its surroundings in order to navigate. Its location is possible using GPS / GNSS. However, in the absence of GPS / GNSS and in the case of disrupted satellite signals, other methods are needed to help a vehicle determine its position. A high-definition (HD) on-board map fills in the information gaps while providing a reliable 3D description of the vehicle's surroundings.
[0004] Current solutions for generating high-definition 3D maps and covering large areas are tedious, require significant work time, and the volume of data required for these maps remains high. At the same time, automated object recognition in LIDAR images and point clouds has progressed in recent years.
[0005] Mapping systems incorporate at least three types of data: geolocation data, 3D spatial measurements, and 360° visualization with panoramic optical images. High-quality geolocation measurements are the foundation for creating faithful and accurate maps. To achieve this, LIDAR scanners provide dense, high-precision 3D point clouds that can represent the morphology of the environment, and panoramic images provide additional information, textures, signage, text, and other visual details.
[0006] Together, these data help create dense, highly detailed 3D maps. Mobile mapping systems generate over 1 gigabyte of data per 50m, making data management and processing complex, requiring large-capacity data storage platforms, and powerful computing and video processing resources for data processing, visualization, navigation, and manipulation.
[0007] Processing and analyzing mobile mapping data to create high-definition maps is delicate and time-consuming for several reasons: (1) the heterogeneous and unstructured nature of 3D point clouds, which are a relatively new technology, makes them difficult to analyze, (2) 2D images when taken alone lack spatial information, (3) most classification and pattern extraction methods require significant human analyst intervention and, (4) stand-alone software used to make HD maps are inherently designed as vertical solutions (5) the recognition of discrete objects such as signaling devices and their geolocation is complex.
[0008] To extract certain objects such as discrete objects, particularly signs or other signaling elements, it is preferable to use images rather than a point cloud, particularly because point clouds are not suitable for sign recognition.
[0009] This involves recognizing these discreet objects and their geolocation.
[0010] For geolocation, it is known, for example, from document US2018 / 188060 A1, to combine camera images and depth maps or point clouds from detection and distance measurement sensors such as lidars. Such a technique requires processing a large amount of data.
[0011] Document WO2018 / 140656 A1 describes combining panoramic images and distance data obtained by distance detection devices, for example lidar or radar, to geolocate objects.
[0012] Finally, the paper FOOLADGAR FAHIMEH ET AL: "Geometrical Analysis of Localization Error in Stereo Vision Systems", IEEE SENSORS JOURNAL, IEEE SERVICE CENTER, NEW YORK, NY, US, vol. 13, no. 11, November 1, 2013 (2013-11-01), pages 4236-4246, XP011527941 proposes methods for calculating target positions using stereo images. Statement of the invention
[0013] The present application deals more particularly with the geolocation of signaling devices and in particular of road signs in images and in particular in panoramic images and proposes a method of geolocation in global references of discrete objects from images originating from a spherical and non-stereo imaging system with positioning and geolocation sensors on board a moving vehicle.
[0014] The method of the invention uses single panoramic images, i.e. non-stereo, geolocated spherical or semi-spherical, taken successively and combines a part of classification of objects based on neural networks and an algorithmic part of calculation of position and dimensions of the classified objects in order to simplify the shooting and reduce the mass of data to be processed.
[0015] More specifically, the present invention provides a method according to claim 1.
[0016] The camera device carried by the mobile vehicle comprising several cameras distributed in a spherical manner adapted to take synchronized, dated and localized photos by localization and dating means such as a GPS / GNSS system, the method may comprise taking photos periodically by the cameras of said camera device during the movement of the vehicle carrying them, said synchronized, dated and geolocated photos being processed in software adapted to construct spherical panoramic images from said photos.
[0017] The camera device may comprise five cameras pointing in directions spaced 72° apart in a horizontal plane around a vertical axis relative to the rolling plane of the vehicle and one camera pointing upwards on the vertical axis, the simultaneous photos taken by the 6 cameras are connected to produce a 360° spherical panoramic photo in a horizontal plane.
[0018] The method may comprise fixing an origin point with coordinates (0, 0) at the level of the panoramic images with a width of W pixels and a height of H pixels, a step of recognizing and classifying said objects by means of deep learning models based on neural networks producing a selection frame around a subset of pixels which contains the object detected in the panoramic image as well as its classification code and a step of memorizing: of said selection frame of each detected object P, the height h and the width w of this frame and, the coordinates i (vertical), j (horizontal) in number of pixels of a position marker of the selection frame containing said object associated with the panoramic photos relative to said point of origin.
[0019] The method may include defining the origin point of coordinates (0, 0) of the panoramic image as the top left point of the panoramic image and defining the position mark i, j relative to the origin (0, 0) of the selection frame as the top left point of the selection frame.
[0020] The triangulation method may comprise, for a series of panoramic images, a step of determining the angular position of the objects detected in each panoramic image of said series in which an object classification is present relative to the shooting center which comprises, from the dimensions of the selection frame: a calculation of the position of the estimated center A of each object in the panoramic image; a calculation of an estimated radius r of said object such that r is equal to half of the smallest of the dimensions of width w and height h of the selection frame in the panoramic image and the calculation of the position of a point B on the selection frame and closest to point A such that AB=r; a calculation of the angular position of the center A of the image of said object and of point B relative to the center O of the group of cameras 100 in the spherical camera coordinates, the panoramic image being assimilated to a sphere of center O then, the calculation of a vector u A joining the center of the cameras O and the center A of the object, the calculation of a vector u B joining the center of the cameras O and point B and the calculation of a cone of axis u A and vertex angle α such that α = cos − 1 u → A ⋅ u → B in Cartesian camera coordinates; a transformation into the global Cartesian terrestrial coordinates and a storage of the coordinates of the vectors u A of the center O of the cameras and the angle α for each object of each panoramic image.
[0021] The step of determining the position of the detected objects can in particular use: the position parameters of the cameras, which include the geolocated Cartesian coordinates (x 0 , y 0 , z 0 ) of the center O of the camera device and the orientation (Ω, Φ, K) of each of the cameras of the camera device; the height and width of the original panoramic image H, W; the borders of the bounding boxes containing objects detected in the panoramic images.
[0022] The panoramic image being considered as the upper part of the surface of a sphere with the camera device at the center of this sphere, the triangulation method preferably comprises the determination of the position of a particular detected object by means of a projection of a cone directed towards the center A of the selection frame of said object on said upper part of said surface of a sphere, said cone having as its apex the center of the camera device and the angle opening α and comprises, for a given particular detected object, a calculation of intersections of pairs of cones generated in at least two distinct panoramic images comprising said object and taken at at least two different locations U 0 , V 0 , so as to calculate by triangulation a spatial distance between the particular detected object and the positions of the centers of the camera device at said at least two different locations and to calculate a geolocated position of said particular detected object.
[0023] Preferably, only cone pairs whose distance between their apices is limited by a distance parameter D are taken into account, the distance parameter D being expressed in meters and chosen according to the spatial frequency of photo taking to reduce the number of false detections.
[0024] The method advantageously comprises a validity analysis of 3D IJ intersections of said cone projections in order to reduce the number of detected object candidates and identify the uniqueness of the objects and determine whether intersections are sufficiently close, a validity condition of a candidate object being the fact that the direction vectors u A and v A of the pairs of cones are distant by a distance less than a parameter d at the IJ intersection of the cones, the method further comprising an extraction as valid objects of representatives satisfying said validity analysis, the storage of their position in the global coordinate system, the storage of their predicted classification and the storage of their dimension.
[0025] The validity analysis advantageously comprises a search for intersection nodes and a calculation of proximity between the axes of the cones to determine whether intersections are sufficiently close, by means of the generation of a KD tree of points I and J and the selection from said KD tree of pairs of cones having a minimum distance less than a parameter d at the intersection IJ of the cones, the value in meters of said parameter d being chosen according to the size of the classes of objects detected, and comprising, for the condition that the intersections are sufficiently close, an algorithm for identifying points I and J of closest distance from the vectors u A and v A direction of the pairs of intersecting cones, an algorithm for constructing spheres of influence P,Q whose radius is the radius of the cone along a perpendicular to the directing axis of the cone around points I and J called attachment points and an algorithm for validating the condition if points I and J are mutually contained in the sphere of influence of the opposite cone the axes of the cones then being contained in the cone opposite the intersection.,
[0026] The method may comprise, for all pairs of cones whose intersections are sufficiently close and whose direction vectors are mutually contained in the cones and when two or more attachment points have been found, an analysis for detecting parasitic intersections, said analysis comprising the creation of a graph: for which each attachment point constitutes a node, for which two nodes are connected if they satisfy the conditions of sufficiently close intersections and direction vectors mutually contained in the cones, said analysis further comprising for each connected component of the graph, a choice of the node having the maximum degree of connections as representative of a single object and a sorting by decreasing degrees of connectivity of said representatives then an iterative process of validation and subtraction of the highest order representative and of the cones converging towards the latter, of analysis of convergences of remaining cones and of deletion of the representatives finding themselves without converging cones, said iterative process being repeated until all the representatives have been validated or invalidated. Brief description of the drawings
[0027] Other characteristics, details and advantages of the invention will appear on reading the detailed description below, and on analyzing the attached drawings, in which: [ Fig. 1 ] shows an example of panoramic photography; [ Fig. 2] shows a diagram of a camera reference system in spherical coordinates; [ Fig. 3 ] shows a diagram of a global reference system in Cartesian coordinates; [ Fig. 4 ] shows a diagram of angular camera tracking; [ Fig. 5 ] shows a diagram of an image projected onto a sphere centered on the cameras; [ Fig. 6 ] shows a panel treatment according to one aspect of the application; [ Fig. 7 ] shows a representation of two solid angles converging at a triangulated point; [ Fig. 8 ] shows an example of object detection on a trajectory; [ Fig. 9 ] shows a flowchart of the steps of a first process of the application; [ Fig. 10 ] shows a flowchart of the steps of a second application process; [ Fig. 11 ] is a schematic representation of an image-taking vehicle [ Fig. 12 ] shows a graph corresponding to the object detection of the figure 8 . Description of the embodiments
[0028] The drawings and description below describe non-limiting exemplary embodiments useful for understanding the invention.
[0029] For HD map generation, it is desirable to perform detection and classification of roads and traffic signs from panoramic images.
[0030] The panoramic images 11 as represented in figure 1 are taken from mobile vehicles 1 as shown schematically in figure 11equipped with panoramic cameras 2a, 2b, for example an assembly of cameras, called a panoramic camera device, which may in particular comprise five cameras 2a pointing in directions spaced 72° apart in a horizontal plane around a vertical axis relative to the rolling plane of the vehicle and a camera 2b pointing upwards on the vertical axis. The cameras take synchronized photos, dated and located by location and dating means such as a GPS / GNSS system 3 and an inertial unit. The GNSS / GPS antenna is normally positioned as close as possible to the inertial unit and in the vertical axis of the inertial unit and is in the immediate vicinity of the camera block. The vehicle may also be equipped with a LIDAR device for scanning the environment to generate point clouds for producing 3D maps.
[0031] The 6 photos taken by the cameras are connected to produce a 360° panoramic photo in a horizontal plane.
[0032] Vertically, the lower part of the photo has no information due to the cameras' field of view but is used to complete the image by 180°. The image is referenced to have a point with coordinates 0, 0 at the top left point and has a width of W pixels by a height of H pixels, for example 1920 pixels in a horizontal direction and 1080 pixels in a vertical direction.
[0033] The cameras take pictures synchronously as the vehicle carrying them moves, for example every two meters. The single images are processed in software adapted to construct a panoramic image from the six simultaneous images.
[0034] The problems associated with recognizing and positioning objects such as signs on a map from panoramic images are, on the one hand, knowing whether the images of signs that are repeated in several panoramic images correspond to the same sign or to different signs and, on the other hand, the impossibility of knowing the real position of a sign from its image on a panoramic image.
[0035] The following description takes the example of road signs, but the method described applies to any discrete object that one wishes to geolocate, such as traffic lights, bus stop shelters or other discrete objects whose geolocation is desired.
[0036] The method has also been tested in railway environments to geolocate terminals and ground-based devices. In addition, specific ground markings could be geolocated using this method, such as squares on "give way" lines, bicycle symbols on cycle lanes, and manhole covers. However, the larger the objects to be detected, the less accurate the geolocation.
[0037] Generally speaking, the information required for a panoramic image containing at least one object to be geolocated is: the predicted class of the detected object, the pixel coordinates of a selection frame of a sub-image containing the object to be geolocated, the h, w dimensions of the selection frame of the sub-image, the name of the panoramic image, its orientation and its geolocation.
[0038] A step prior to the method of the present application is the creation of a model for recognizing MRP objects, such as panels 400, in the form of a neural network using a learning base which contains examples of the classes of interest of these objects. This panel recognition model 400 is then used to recognize the panels in panoramic images to be processed.
[0039] Object classes are groups of object types to be found. Depending on the size of the database and the number of samples, a class can contain an object type but also a family of objects. For example, for signs, a class can contain the type stop sign, but also a type of sign such as direction signs. The granularity of the class can vary depending on the amount of data available or to be processed.
[0040] According to the present application, the method begins with a succession of steps of taking panoramic images 11 by one or more mobile vehicles 1 carrying panoramic camera devices discussed above. An example of a panoramic image is given in figure 1 . In this image, the coordinates i, j in pixels of a point P are Cartesian coordinates from an origin point (0, 0) at the top left of the image.
[0041] The image is also W pixels wide, which corresponds to 360°, and H pixels high, over 180°, with a lower black band of 10 on the part of the image hidden by the vehicle.
[0042] Once the images are taken and saved in a database or a 500 file, depending on the figure 9, inference steps, by the model on the new images to be processed, perform recognition and classification of objects 410 in the panoramic images by means of the neural network. This makes it possible to recognize the classes of the detected objects, for example classes of signs, the traffic light class, etc. Each object detected in each image will be associated with the detected class. Similarly, a selection frame as represented in figure 6 discussed below will give the maximum dimensions of the object in the corresponding image.
[0043] A database or 300 file with the objects classified in the images and photos and camera positions is then created.
[0044] Then an important part for the positioning of detected and classified objects from mobile mapping images in a 3D map is to determine their geolocated coordinates. To do this, the present application proposes to perform a triangulation of the objects from two or more detections of these objects.
[0045] The triangulation performed uses the principle of parallax. A minimum of two images are required for this method to work. For each detected sign or object, its position relative to the center of the camera group, relative to the extrinsic camera parameters, and its position in the panoramic image must be determined.
[0046] Specific information associated with each panoramic image is therefore necessary to extract the geolocation of classified objects. The camera position parameters, which include the geolocated Cartesian coordinates of the center O of the camera group 100 (x 0 , y 0 , z 0 ) according to the figure 3 and the orientation of each of the six cameras (Ω, Φ, K) in the group of 6 cameras 100 according to the figure 4 ; The height and width of the original panoramic image (H, W) of the figure 1 ; the boundaries of all selection frames 13 containing the objects detected in the panoramic images such as panel P shown in figure 6 . For geolocation, several transformations are necessary: image pixels must be located in the camera spherical coordinates (ρ, θ, φ) according to the figure 2 where the cameras 100 are arranged at the origin point O of the reference frame; the spherical camera coordinates must be translated into Cartesian camera coordinates (X, Y, Z) according to the figure 3because the coordinates used to locate the vehicle, and therefore the center of the shooting device 100, are Cartesian coordinates, and finally; the Cartesian camera coordinates must be translated into global coordinates (x 0 , y 0 , z 0 ) according to the figure 3 (e.g. WGS-84 geodetic coordinates).
[0047] The expected result is the determination of the geolocation of a position P of the objects classified in global coordinates (x, y, z).
[0048] In the following we will consider a traffic sign.
[0049] To geolocate the panels, it is necessary to determine the position of each detected panel relative to the center of the camera and relative to the panoramic image. Solid angles 110 are projected from the center of the camera 100 called the apex, towards the center A of the selection frame 13 of the detected panel called the sub-image as shown in Figure 5 .
[0050] A "sub-image" is a subset of the pixels of the panoramic image including the object to be geolocated, the selection frame framing the sub-image.
[0051] The position of the sub-image of the traffic sign relative to the center of the camera can be considered as a part of the surface of a sphere 120 (the panoramic image) with the camera device 100 at the center of this sphere.
[0052] For each sub-image, solid angles, the cones 110, are generated as follows. Conservatively, a solid angle opening can then be defined by the radial distance r between the central coordinates, A, of the bounding box and thus of the traffic sign, to the nearest edge of the bounding box of the figure 6 according to the relationship: A = i + h 2 , j + w 2 with : r = min w h 2
[0053] Here, (i, j) are the coordinates of the topmost pixel on the left of the sub-image 13 enclosing the panel, h and w are respectively the height and width of the sub-image. The coordinates of point A in the image are vertically A i = i + h / 2 and horizontally A j = j + w / 2. This also makes it possible to define a point B of the selection frame closest to A on the perimeter defined by the radius r.
[0054] The transformation of the pixel coordinates of the central point A of the sub-image with coordinates (A i , A j ) into spherical camera coordinates (partial transformation, because there is no depth dimension) is then carried out as follows:
[0055] In the horizontal plane the angle φ is: φ = π − 2 A j π w modulo 2 π
[0056] For this formula, since the center of the image corresponds to an angle φ 0 =0, we must subtract 2A j π / W from π to obtain the value of the angle φ.
[0057] In the vertical plane the angle θ is: θ = A i π H modulo π
[0058] In this calculation the image width corresponds to W in pixels and 2π in spherical coordinates and the image height corresponds to H in pixels and π in spherical coordinates.
[0059] For each sub-image, solid angles (cones) are generated by assigning a value of 1 to the parameter ρ of the vector defining the axis of the cone in the spherical camera coordinates.
[0060] To place oneself in the global coordinate system, several transformations are carried out.
[0061] The Cartesian camera coordinates are expressed in spherical camera coordinates in the form: X = ρ sin θ cos φ , Y = ρ sin θ sin φ , Z = ρ cos θ with ρ=1 two vectors u A and u B are defined such that using the angles determined for points A and B we have: u → A = X A Y A Z A u → B = X B Y B Z B with an origin at (0, 0, 0) center of the camera set.
[0062] The angle between these two unit vectors allows us to calculate the opening of the solid angle from the scalar product of these two vectors such that: α = cos − 1 u A → ⋅ u B →
[0063] Finally, the cone vector, u A is transformed into global coordinates: x A y A z A = X A Y A Z A . R T
[0064] Where R is the rotation matrix using the camera's extrinsic parameters: R = R z K R y Φ R x Ω either also: R = cos K − sin K 0 sin K cos K 0 0 0 1 ⋅ cos Φ 0 sin Φ 0 1 0 − sin Φ 0 cos Φ ⋅ 1 0 0 0 cos Ω − sin Ω 0 sin Ω cos Ω and RT< the transposed matrix.
[0065] Thus, for each panoramic image, vectors, starting from the center of the camera group (x 0 ,y 0 ,z 0 ) of the figure 3 in global coordinates at the time of taking this image and pointing towards the center of the panel (x A ,y A ,z A ) in global coordinates as seen in the panoramic image, create lines oriented in space. Each panoramic image thus produces a cone centered on the line going from the center of this image to point A of an object in this image and of opening angle α.
[0066] For a given class of objects resulting from the classification, a search for intersections of the cones, derived from a series of panoramic images and trajectory data for the capture points of the images containing detected objects, is then carried out. This makes it possible to calculate a spatial distance between the detected object(s) and the positions of the center of the camera at the time of the photos taken by triangulation. To reduce the number of false detections, only pairs of cones whose apices are separated by a distance 204 according to the figure 8less than a distance parameter D are taken into account. This search can be carried out using a KD tree on the trajectory points which makes it possible to accelerate the search for pairs of cones which are close to each other. Such a KD tree is not mandatory, but it lightens the calculation when the quantity of data increases. D is expressed in meters and its choice depends on the maximum distance at which it is considered that an object can no longer be detected on an image. The parameter D will be adapted by the operator according to the spatial frequency of the shots. For example, for a class of panel objects, a vehicle taking photos every 2m, D can be defined at 100m.
[0067] The frequency of taking photos is preferably determined based on the distance traveled by the vehicle, for example, photos are taken every meter or every two meters traveled. In some configurations, the frequency can be determined based on a time frequency, with the vehicle speed determining the interval in meters between photos. For a vehicle traveling at 50 km / h and cameras taking photos at 14 fps, in this case, one photo would be obtained per meter and a backup of one photo per 2 or 3 meters could be made.
[0068] There figure 7represents two vectors u A and v A central respectively to the cones 110a, 110b and corresponding to the detection of a panel in two images taken with the cameras in position U 0 and in position V 0 . To reduce the number of detected candidates and identify the uniqueness of an object, the projections of the cones 110a, 110b are analyzed to determine the intersection points in 3D which represent the geolocated position of the objects in space.
[0069] Because cone vectors are three-dimensional and due to measurement inaccuracies, cone vectors may not intersect perfectly, creating multiple close intersections. Intersections are considered valid if several conditions are met: i - They must be sufficiently close; ii - The direction vectors are mutually contained in the cones and; iii - The intersection is not a parasitic intersection.
[0070] For condition i, to determine whether intersections are close enough, a KD-Tree is generated and only pairs of cones whose minimum distance 130 between their axes defined by vectors u A and v A is less than a parameter d are considered to satisfy this condition. The value of d in meters is chosen based on the size of objects of an object class detected in the images.
[0071] For condition ii, according to which the direction vectors are mutually contained in the cones, we identify the points I and J of closest distance along the direction axes of the pairs of intersecting cones and we construct spheres of influence P, Q whose radius is the radius of the cone along a perpendicular to the direction axis of the cone around these points called attachment points. The condition is satisfied if the points I and J are mutually contained in the spheres of influence P, Q taken into account and these points I and J are said to be mutually connected.
[0072] For all pairs of cones that satisfy both conditions i and ii, when two or more attachment points have been found, a position in global coordinates is calculated for each attachment point.
[0073] A line segment IJ orthogonal to the two lines formed by the vectors u A and v A, directions of each cone, connects the two lines formed by said vectors.
[0074] The problem is to minimize the distance ∥ J - I ∥ 2< of the line segment IJ. It is necessary to deduce the position of the attachment points in global coordinates through the equation according to which the scalar product of the two perpendicular vectors is zero.
[0075] A general equation expresses the position of 3D points along their respective vector: x y z = M 0 + t M A − M 0
[0076] Since solid cones are defined by their apex M 0 and a directional point MA , we have: M 0 = x 0 y 0 z 0 M A = x A y A z A
[0077] A particular value of the variable t which defines the distance of the points on the line defined by the directional vector of the cone will define the position respectively of the attachment point I or J.
[0078] Since the scalar product of two mutually orthogonal directional vectors is zero, we have for u A and v A: I − J ⋅ U A − U 0 = 0 I − J ⋅ V A − V 0 = 0
[0079] By rewriting the scalar products with the general vector line equation where: I − J = U 0 − V 0 + t A U A − U 0 − t B V A − V 0
[0080] Then by evaluating the scalar product equations with the known points of the vectors u A and v A , and by realizing an equality between the equations, it is possible to solve the equations first for t A , then for t B in order to obtain the global coordinates of the attachment points I and J.
[0081] Once the parameters t A and t B have been obtained, the radii of the spheres for each vector can be calculated with: sin α A = R S / t A R S = t A sin α A
[0082] The radius RS is used to define the sphere of influence at point I, a similar calculation with t B allows the sphere of influence at point J to be calculated. The radius Rs is then the dimensional analogue of r in the equation r=(min (w,h) / 2 above.
[0083] Once the intersections of cones at points I and J have been determined, i.e. when the intersection is such that the pairs of spheres for which the distance between their centers is less than d and these centers are mutually contained in the opposite sphere, it is necessary to verify condition iii - The intersection is not a parasitic intersection or a ghost panel.
[0084] To do this, a subsequent analysis is carried out to detect unique or parasitic intersections, i.e. false intersections as shown in figure 8representing a simplified case where on trajectory 200 a vehicle takes photos at positions 200a, 200b, 200c, 200d, 200e, of two panels 201 and 202, panel 201 being visible on the photos taken at points 200b, 200c, 200d, 200e while panel 202 is visible on the photos taken at points 200a, 200b, 200c. We see that in this case, a vector from a photo of panel 201 in position 200b and a vector from a photo of panel 202 taken at point 200c intersect at 203. In this case, the intersection at point 203 corresponds to a ghost panel. This comes from the fact that in reality, a cone can have an intersection with several other cones which generates an attachment point representing the possible position of an object. In such a case it is necessary to delete the parasitic intersections to keep only the valid intersections at points 201 and 202. To delete the points corresponding to a ghost panel, a graph 600 as represented in figure 12 for which each attachment point Ap constitutes a node is realized. The branches Br of the graph are the connections between a node and all the other nodes forming together a connected component. For each connected component, the node Rp with the maximum degree is chosen as the representative of a single object. The representatives are then sorted by decreasing degrees of connectivity.
[0085] To filter out spurious points, the highest order representatives and the cones converging to them are validated and removed. The remaining cone convergences are analyzed and, if one or more representatives are found without converging cones, these representatives are identified as a projection of a ghost or duplicate panel called a ghost projection and invalidated. The operation is repeated until all representatives have been validated or invalidated. All representatives that pass the last test are then extracted as valid objects, their position in the global coordinate system is stored as well as their predicted classification and dimension. It should be noted that storing the measured dimension then allows to automatically check whether the size of the detected objects is consistent with their theoretical size inherent to their class.
[0086] For example in the figure 12 , point 201 of the figure 8corresponds to an attachment point A1 with a representative Rp1 on graph 600, point 202 to an attachment bridge A2 with a representative Rp2 and point 203 to an attachment point A3 with a representative Rp3. The representative Rp1 having 11 connections is validated and removed with its cones then the representative Rp2 having 5 connections is removed with its cones. This leaves the representative Rp3 which no longer has any cones and is therefore classified as a phantom projection and therefore invalidated.
[0087] Furthermore, the radius Rs of the highest order representative allows us to evaluate the spatial dimension and size of the object.
[0088] To facilitate the elimination of ghost signs due to cone projections intersecting at much further distances or intersecting with cones of other traffic signs, it is possible to adapt D the maximum distance between cone apices considered according to the density of traffic signs and to reduce D inside cities where the density of signs is high compared to scenes outside cities where the density of signs is lower.
[0089] After the road sign extraction algorithm has finished processing a panoramic image, sub-images that could not be triangulated and localized are documented in a separate log for manual verification.
[0090] The traffic sign localization algorithm of the present application depends on fundamental geometric transformations and a constrained three-dimensional triangulation method.
[0091] The output data are the geolocated coordinates (x, y, z) of the objects and in particular the signs in global coordinates, their predicted classification and a minimum dimension, r.
[0092] Another point addressed by the parasite point filtering described above is the extraction of ghost objects when several cones have an intersection with a single cone. This happens due to the sensitivity of the method to parallax at small angles coupled with a profusion of cones which add background noise. A sufficient distance between the viewpoints (images taken) is required to clearly distinguish the intersections of the cone vectors during triangulation.
[0093] Furthermore, to avoid computing intersections with cones that are too far apart, the parameter D, which reduces interference between cones for instances of signs of the same type that are close together, can be adapted according to the location of the signs. D can be reduced in cities where the density of signs is higher and increased outside cities. Since the cones are projected to infinity, according to the trajectory of the vehicle, particularly in curves, roundabouts and intersections, reducing the search radius reduces the background noise. The value of D depends on the size of the objects and the frame capture frequency.
[0094] The aforementioned distance parameters D and d are introduced when implementing the method on a set of images corresponding to a surface whose panels or other objects are to be geolocated.
[0095] All of these operations are synthesized in figure 10which describes, from the database of images with classified objects 300, the steps 305 of reading the selection frames of the objects and 310 reading the trajectories of the cameras having taken the images, the step 320 of constructing the cones 320 in parallel with the step 330 of creating the KD trees of proximity of the intersections, the step 340 of selecting the objects of a class to carry out the step 350 of searching for the pairs of distance cones <D par rapport à l'origine de la photo en prenant en compte le résultat de l'étape de création de l'arbre K-D des points de trajectoire puis l'étape 355 de recherche des nœuds d'intersection et le calcul de la proximité entre les axes des cônes pour définir les nœuds correspondant à des objets probables.
[0096] These steps are followed by a step 360 of calculating the connected components with the generation of a graph 600 of all the node connections then a step 365 of selecting the node of maximum degree for each connected component followed by a step 370 of ordering the representative nodes by descending degrees to carry out the step 375 of filtering the dummy representatives by deleting the vectors already referenced for higher order nodes and deleting the remaining single vector nodes.
[0097] At the end of this operation, the output data are the objects for which the coordinates of the center of the object in global coordinates and the radius of the object are stored in step 380. To summarize,
[0098] A - The triangulation method is performed with the criteria: 1) at least two images where the model has detected the same object, 2) the position in global coordinates of the center of the camera block is known 3) the camera orientation angles are known which allows to calculate the position of the detected object in the image in a global frame. B - The triangulation method projects the directional vectors, with their apex at the center of the camera block and searches for intersections in 3D space as potential positions of the object. C - The determination of the uniqueness and validity of the object as well as the execution time performance are ensured by the introduction of two parameters: D, a search distance for the pairs of intersecting vectors which makes the algorithm more efficient in execution time and d, a distance threshold determining that at the intersection, the vectors are close enough to reduce the candidate intersection points.D - The method provides a method of filtering the crossings by a voting system to identify among the candidates the most probable position of the objects. E - After calculating the number of votes (degree of connectivity of a position) the points are ordered in rank. The projected vectors are assigned to the points by their decreasing degree of connectivity and the points no longer having associated vectors are considered as false positives and not real positions.
[0099] Once all existing object classes in the objects have been processed, the result is a database of 390 objects that can be integrated into a 3D map.
[0100] The present application thus proposes a highly automated method for the recognition and geolocation of discrete objects such as signs, traffic lights or other infrastructure objects from panoramic images taken by one or more vehicles moving on traffic lanes. The vehicles may in particular be motor vehicles moving on road traffic lanes or railway vehicles for geolocating railway signaling devices.
Claims
1. Method for geolocating discrete objects from a series of unitary geolocated panoramic images (11) taken by a 360° panoramic camera device of a mobile vehicle on the route of said vehicle, which comprises: a. - at least one succession of steps (510) of taking said unitary semispherical or spherical panoramic images by a mobile vehicle carrying a panoramic camera device and carrying a geolocation system for which a periodicity of taking the photographs producing the unitary panoramic images is determined according to a distance travelled by the vehicle or according to a time frequency, the vehicle speed then determining the interval in meters between said photographs, b. - a succession of steps (410) of recognising and classifying objects, in said unitary panoramic images, performed by means of neural network-based deep learning models producing detection and classification of objects in the panoramic images and producing definition of selection frames around subsets of pixels that contain the objects detected in the panoramic images and, for each object in each panoramic image, a classification code and the coordinates i vertical, j horizontal in number of pixels of a position location of the width selection frame w and height selection h frame containing said object in said panoramic image; c. - a succession of steps (305 to 380) of geolocating the classified objects by a method of triangulating said objects from at least two of said separate unitary panoramic images containing said classified objects and the positions of the panoramic cameras when taking said panoramic images containing said classified objects, said geolocation steps including, for a series of unitary panoramic images and for each detected object: i.- a step of determining, in each unitary panoramic image of said series in which an object classification is present, the angular position with respect to the photographing centre of said detected object, ii.- at least one change of reference of the coordinates of said object in the images in terrestrial Cartesian coordinates according to the coordinates of the cameras at the time of the shootings, iii.- a projection of cones (110) made from an angle α directed towards the centre A of the selection frame of said object, said projection having as apex the centre of the camera device and an opening angle defined by an estimated radius r, per image, equal to half of the smallest of the width w and height h dimensions of the selection frame, iv.- calculating intersections of pairs of cones (110a, 110b) generated in at least two unitary panoramic images including said object and taken at at least two different locations (100a, 100b), so as to calculate by triangulation a spatial distance between detected objects and the positions of the centres of the camera device at said at least two different locations and then to calculate a geolocated position of said detected objects.
2. Geolocation method according to claim 1 wherein, the camera device carried by the mobile vehicle(s) (1) including several cameras (2a, 2b) distributed spherically adapted to take photographs synchronised, dated and located by location and dating means (3) such as a GPS / GNSS system, the method includes taking photos periodically by the cameras of said camera device when the vehicle carrying them is moved, said synchronised, dated and geolocated photographs being processed in software adapted to construct spherical panoramic images from said photographs.
3. Geolocation method according to claim 2, wherein the camera device including five cameras (2a) pointing in directions spaced 72° apart in a horizontal plane about a vertical axis with respect to the plane of travel of the vehicle and a camera (2b) pointing upward on the vertical axis, the simultaneous photographs taken by the 6 cameras are connected to produce a 360° spherical panoramic photo in a horizontal plane.
4. Geolocation method according to any one of the preceding claims including fixing a point of origin of coordinates (0, 0) on the panoramic images with a width of W pixels over a height of H pixels, a step (410) recognition and classification of said objects by means of deep learning models based on neural networks producing a selection frame around a subset of pixels that contains the object detected in the panoramic image along with its classification code and a step of memorising: - said selection frame (13) of each detected object P, - the height h and the width w of this frame and, - coordinates i vertical, j horizontal in number of pixels of a position mark of the selection frame containing said object associated with the panoramic photographs relative to said point of origin.
5. Geolocation method according to claim 4 including defining the point of origin of coordinates (0, 0) of the panoramic image as the top left point of the panoramic image and defining the position marker i, j (13a) relative to the origin (0, 0) of the selection frame as the top left point of the selection frame.
6. Geolocation method according to claim 4 or 5, wherein the triangulation method includes, for a series of panoramic images, a step of determining the angular position of the objects detected in each panoramic image of said series wherein an object classification is present with respect to the photographing centre which comprises, from the dimensions of the selection frame: - a calculation of the position of the estimated centre A of each object in the panoramic image; - a calculation of an estimated radius r of said object such that r is equal to half of the smallest of the width w and height h dimensions of the selection frame in the panoramic image and calculating the position of a point B on the selection frame closest to the point A such that AB=r; - a calculation of the angular position of the centre A of the image of said object and of the point B with respect to the centre O of the group of cameras 100 in the camera spherical coordinates, the panoramic image being assimilated to a sphere of centre O, then, - the calculation of a vector uA joining the centre of the cameras O and the centre A of the object, calculating a vector uB joining the centre of the cameras O and the point B and calculating a cone of axis uA and angle at the vertex α such that α = cos − 1 u → A ⋅ u → B in camera Cartesian coordinates; - a transformation in the global Cartesian terrestrial coordinates and a memorisation of the coordinates of the vectors uA of the centre O of the cameras and of the angle α for each object of each panoramic image.
7. The geolocation method according to claim 6 wherein the step of determining the position of the detected objects uses: - the position parameters of the cameras, which comprise the geolocated Cartesian coordinates (x0, y0, z0) of the centre O of the camera device and the orientation (Ω, Φ, K) of each of the cameras (2a, 2b) of the camera device; - the height and width of the original panoramic image H, W; - the borders of the selection frames (13) containing objects detected in the panoramic images.
8. Geolocation method according to claim 6 or 7 wherein, the panoramic image being considered as the upper part of the surface of a sphere (120) with the camera device (100) at the centre of this sphere, the triangulation method includes determining the position of a particular detected object by means of a projection of a cone (110) directed towards the centre A of the selection frame of said object on said upper part of said surface of a sphere, said cone having as its apex the centre of the camera device and the opening of angle α and comprises, for a given particular object detected, a calculation of intersections of pairs of cones (110a, 110b) generated in at least two distinct panoramic images including said object and taken at at least two different locations U0, V0, so as to calculate by triangulation a spatial distance between the particular object detected and the positions of the centres of the camera device at said at least two different locations and to calculate a geolocated position of said particular object detected.
9. Geolocation method according to claim 8, wherein only the pairs of cones where the distance between their apexes is limited by a distance parameter D are taken into account, the distance parameter D being expressed in metres and selected according to the spatial frequency of photographing to reduce the number of false detections.
10. Geolocation method according to claim 8 or 9, wherein the method includes a validity analysis of 3D intersections I-J of said cone projections in order to reduce the number of object candidates detected and to identify the uniqueness of the objects and to determine whether intersections are sufficiently close, a validity condition of an object candidate being the fact that the directing vectors uA and vA of the cone pairs (110a, 110b) are distant by a distance less than a parameter d from the intersection I-J of the cones, the method further comprising extracting as valid objects representatives satisfying said validity analysis, storing their position in the global coordinate system, storing their predicted classification and storing their dimension.
11. Geolocation method according to claim 10, wherein the validity analysis includes a search (355) for intersection nodes and a proximity calculation between the axes of the cones to determine whether intersections are sufficiently close, by means of generating a K-D tree of the points I and J and selecting from said K-D tree pairs of cones having a minimum distance (130) less than a parameter d at the intersection I-J of the cones, the value in metres of said parameter d being selected according to the size of the classes of objects detected, and including (355), for the condition according to which the intersections are sufficiently close, an algorithm for identifying points I and J closest to the distance of the directing vectors uA and vAof the intersecting pairs of cones, an algorithm for constructing spheres of influence P, Q whose radius is the radius of the cone along a perpendicular to the directing axis of the cone around the points I and J called attachment points and an algorithm validating the condition if the points I and J are mutually contained in the sphere of influence of the opposite cone, the axes of the cones being then contained in the cone opposite to the intersection.
12. Geolocation method according to claim 11 including, for all pairs of cones whose intersections are sufficiently close and whose directing vectors are mutually contained in the cones opposite the intersection and when two or more attachment points have been found, an analysis for detecting parasitic intersections, said analysis including a step (360) of producing a graph (600): - for which each attachment point constitutes a node, - for which two nodes are connected if they meet the conditions of intersections sufficiently close and directing vectors mutually contained in the cones, said analysis furthermore including, for each related component (A1, A2, A3) of the graph (600), a choice (365) of the node (Rp1, Rp2, Rp3) having the maximum degree of connections as representing a single object and a sorting (370) by decreasing degrees of connectivity of said representatives and then an iterative process of validating and subtracting representative with the highest order and the cones converging thereto, analysing convergences of remaining cones and removing representatives found without convergent cones (375), said iterative process being repeated until all representatives have been validated or invalidated.