Method for locating individuals in a three-dimensional frame of reference from two-dimensional images captured by a fixed camera and associated locating device
The method and device enhance three-dimensional localization of individuals using geometric calibration and neural networks to process two-dimensional images, addressing precision and cost issues in complex environments.
Patent Information
- Application Number
- FR2023015483
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2043-12-29
AI Technical Summary
Existing two-dimensional camera-based localization methods struggle to provide precise three-dimensional localization of individuals in complex environments like airports and train stations, while existing technologies such as beacons, GPS, cameras, Lidar, and RFID chips are costly, intrusive, or limited in precision.
A method and device that utilize a fixed camera to geometrically calibrate two-dimensional images, detect identifiable body elements, and estimate an individual's absolute position on a three-dimensional ground plane using a predetermined average height, incorporating a neural network for image processing.
Enables precise, economical, and non-intrusive three-dimensional localization of individuals, even in occluded conditions, by leveraging conventional cameras with improved accuracy through geometric calibration and average height estimation.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for locating individuals in a three-dimensional frame of reference from two-dimensional images captured by a fixed camera and associated locating device
[0001] The present invention relates to a method for locating individuals in a three-dimensional frame of reference from two-dimensional images captured by a fixed camera. The present invention also relates to a locating device associated with such a method.
[0002] In various fields such as security, robotics or navigation, it is necessary to be able to locate an individual located in different environments. Airports, train stations or even warehouses are sensitive places and need to be monitored thanks to three-dimensional localization of individuals who pass through them despite disturbances due to several constraints such as crowds or occultations.
[0003] To effectively locate individuals in these places, different methods are generally used depending on the needs.
[0004] The state-of-the-art localization methods include all types of technologies including: beacons, GPS, cameras, Lidar or RFID chips.
[0005] All these technologies are more or less used in the field of security and have advantages and disadvantages always depending on the needs.
[0006] Overall, beacons, GPS, cameras, Lidar and all other types of sensors still pose many problems of cost and applicability in complex environments while RFID chips but also beacons, involve intrusion and precision that are still too limiting.
[0007] Cameras remain the most widely used technology for locating individuals in specific environments.
[0008] However, it is known that cameras are often limited to two-dimensional surveillance and have still limited performance in complex environments such as train stations or airports.
[0009] There is therefore a need for the localization of individuals in a three-dimensional reference system with good precision which is less expensive, less intrusive and more suited to complex environments.
[0010] For this purpose, the present description relates to a method for locating individuals in a three-dimensional frame of reference, comprising a ground plane, from two-dimensional images captured by a fixed camera, the camera having a field of vision of a three-dimensional surveillance space, each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional frame of reference of the image. This method comprises the following steps: - geometric calibration of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space; - processing of two-dimensional images, captured by the camera, for detection of at least one identifiable element of an individual's body, each identifiable element having an associated position in the two-dimensional reference system;
[0011] estimation of an absolute position of the individual on the ground plane of the three-dimensional reference frame as a function of at least one identifiable element, a predetermined average height of the individual and said geometric calibration. Advantageously, the estimation of the absolute position of the individual as a function of at least one identifiable element of the body of an individual and the predetermined average height of the individual makes it possible to improve the positioning accuracy in the three-dimensional reference frame from two-dimensional images.
[0012] According to other advantageous aspects of the invention, the method comprises one or more of the following characteristics taken in isolation or in all possible combinations.
[0013] The method further comprises a reconstruction step comprising a determination of a silhouette of the individual in three dimensions from the estimation of an absolute position of the individual on the ground plane of the three-dimensional reference frame and of said predetermined average height of the individual.
[0014] The processing step includes detection of at least 4 identifiable elements.
[0015] The processing step comprises detecting a number of identifiable elements configurable.
[0016] The step of estimating the absolute position of the individual on the ground plane of the three-dimensional reference frame comprises a calculation of an average of estimated positions of the individual in the three-dimensional reference frame.
[0017] The two-dimensional image processing step is implemented by a neural network.
[0018] The identifiable elements of an individual's body in a two-dimensional image are part of a set including the feet, head, hips, shoulders, eyes or nose.
[0019] The step of estimating an absolute position of the individual on the ground plane of the three-dimensional reference frame implements a calibration matrix.
[0020] The invention also relates to a device for locating individuals in a three-dimensional frame of reference comprising a ground plane, from two-dimensional images captured by a fixed camera, the device being connected or integrated into said camera, the camera having a field of vision of a three-dimensional surveillance space, each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional frame of reference of the image. This device comprises a processor configured to implement: - a geometric calibration module of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space; - a module for processing two-dimensional images, captured by the camera, [for detecting at least one identifiable element of an individual's body, each identifiable element having an associated position in the two-dimensional frame of reference; - a module for estimating an absolute position of the individual on the ground plane of the three-dimensional reference system as a function of at least one identifiable element, a predetermined average height of the individual and said geometric calibration.
[0021] According to another aspect, the invention relates to a system for locating individuals in a three-dimensional frame of reference, comprising a fixed camera and a device for locating individuals in a three-dimensional frame of reference comprising a ground plane, from two-dimensional images captured by said fixed camera, the locating device being as briefly described above.
[0022] According to another aspect, the invention relates to a computer program comprising software instructions, which, when executed by computer, implement a method of locating individuals in a three-dimensional reference frame as briefly described above.
[0023] The invention will appear more clearly on reading the description which follows, given solely by way of non-limiting example and made with reference to the drawings in which:
[0024] [Fig-1] [Fig.l] is a schematic representation of a localization system of an individual in a three-dimensional surveillance space comprising a ground plane, said surveillance being implemented by a location device associated with a fixed camera;
[0025] [Fig.2] [Fig.2] is a flowchart of the steps of the localization method by the location device of [Fig.l];
[0026] [Fig.3] [Fig.3] is a schematic representation of a determination of a silhouette of the individual in the three-dimensional surveillance space.
[0027] In [Fig.l], a three-dimensional surveillance space 5 comprising an individual 10 is monitored using two-dimensional images in a two-dimensional frame of reference with perpendicular axes (U, V) provided with a reference frame with transverse axis V and longitudinal axis U, the reference frame being centered on the lower left edge of the images captured by a camera 15, for example an optical camera.
[0028] The monitoring space 5 is included in a three-dimensional reference frame comprising a longitudinal axis X, a transverse axis Y and an anteroposterior axis Z. For example, the three-dimensional reference frame is the terrestrial reference frame (or world reference frame).
[0029] The camera 15 is for example a video surveillance camera, configured to capture successions of images at a predetermined capture frequency. The camera 15 is spatially positioned at a predetermined position. In other words, the camera 15 has a fixed position in space.
[0030] Advantageously, this makes it possible to calculate a reliable geometric calibration between the three-dimensional reference frame and the two-dimensional reference frame.
[0031] The three-dimensional surveillance space 5 is in a direct field of vision of the camera 15.
[0032] Each captured two-dimensional image is represented by a matrix of pixels, each pixel having a position in the two-dimensional reference system associated with the image.
[0033] In the remainder of the description, it is considered that all of the pixels making up the two-dimensional images correspond to Cartesian positions in the three-dimensional reference frame.
[0034] The camera 15 comprises or is connected to a location device 40. The location device 40 and the camera 15 form a location system 60 of an individual in the three-dimensional frame of reference.
[0035] The location device 40 is a programmable electronic device, e.g. a computer, configured to communicate via a communication link, wired or wireless, with the camera 15.
[0036] According to a variant, the location device 40 is a processing unit integrated into the camera 15.
[0037] The device 40 for locating individuals in a three-dimensional frame of reference implements a detection of a plurality of identifiable elements 20A, 20B, 20C, 20D of the body of said individual 10, and an estimation of a plurality of estimated positions 22A, 22B, 22C, 22D associated with each of the identifiable elements 20A, 20B, 20C, 20D and of an absolute position 24.
[0038] The plurality of identifiable elements 20A, 20B, 20C, 20D is a set of physical parts characteristic of an individual 10 being chosen, for example, from the group consisting of: feet, head, hips, shoulders, eyes or nose.
[0039] In addition, each identifiable element 20A, 20B, 20C, 20D then has a two-dimensional position in the two-dimensional frame of reference of the image.
[0040] The plurality of estimated positions 22A, 22B, 22C, 22D are projections in the three-dimensional reference frame, on the ground plane, along the Y axis of the two-dimensional positions of the plurality of identifiable elements 20A, 20B, 20C, 20D in the two-dimensional reference frame, the projections using a predetermined average height of the individual.
[0041] The location device 40 is configured to estimate an absolute position 24 of the individual 10 considered as a point of the feet of said individual 10 in the three-dimensional reference frame.
[0042] The location device 40 comprises an information processing unit 48, formed for example by a processor 50, and an electronic memory 52 associated with the processor 50.
[0043] The location device 40 comprises a geometric calibration module 42, an image processing module 44 and a position estimation module 46.
[0044] The geometric calibration module 42 is configured to associate each of the pixels of the two-dimensional images captured by the camera 15 with a Cartesian position in the three-dimensional reference frame, the pixels being projected onto the ground of the surveillance space 5.
[0045] The image processing module 44 is configured to detect at least one identifiable element 20A, 20B, 20C, 20D of the body of an individual 10, each identifiable element 20A, 20B, 20C, 20D having an associated position in the two-dimensional reference frame.
[0046] The position estimation module 46 is configured to estimate an absolute position 24 of the individual on the ground plane of the three-dimensional reference system as a function of the at least one identifiable element 20A, 20B, 20C, 20D, of a predetermined average height of the individual and of the geometric calibration.
[0047] In one embodiment, the geometric calibration module 42, the image processing module 44 and the position estimation module 46 are produced in the form of software, or a software brick, executable by the processor 50. The memory 52 is then capable of storing calibration, image processing and estimation software.
[0048] In a variant not shown, the geometric calibration module 42, the image processing module 44 and the position estimation module 46 are produced in the form of a programmable logic component, such as an FPGA (Field Programmable Gate Array), or in the form of an integrated circuit, such as an ASIC (Application Specified Integrated Circuit).
[0049] In one embodiment, the image processing module 44 implements a previously trained neural network to recognize from two-dimensional images and position in the two-dimensional reference frame of each image, at least one identifiable element of the individual's body.
[0050] Generally speaking, the neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.
[0051] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.
[0052] Alternatively, more complex neural network structures can be envisaged with a layer that can be connected to a layer further away than the immediately preceding layer.
[0053] Each neuron is also associated with an operation, i.e. a type of processing, to be carried out by said neuron within the corresponding processing layer.
[0054] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a connection between two neurons. It is often a real number, which takes both positive and negative values. In some cases, the synaptic weight is a complex number.
[0055] Each neuron is capable of performing a weighted sum of the value(s) received from the neurons of the previous layer, each value then being multiplied by the respective synaptic weight of each synapse, or connection, between said neuron and the neurons of the previous layer, then applying an activation function, typically a non-linear function, to said weighted sum, and delivering at the output of said neuron, in particular to the neurons of the following layer which are connected to it, the value resulting from the application of the activation function. The activation function makes it possible to introduce a non-linearity into the processing carried out by each neuron. The sigmoid function, the hyperbolic tangent function, the Heaviside function are examples of activation functions.
[0056] As an optional addition, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.
[0057] Such a neural network is trained on a database comprising images of individuals having identifiable elements 20A, 20B, 20C, 20D as defined previously.
[0058] The operation of the location device 40 is now described with reference to [Fig.2] which illustrates the main steps of a location method in one embodiment.
[0059] The localization method comprises a first geometric calibration step 100 associating each of the pixels of the two-dimensional images captured by the video device 15 with a Cartesian position in the three-dimensional reference frame.
[0060] For example, the geometric calibration module 42 performs a geometric calibration from calibration matrices, according to the following relationship for a point P with Cartesian coordinates (X, Y, Z) in the three-dimensional reference frame whose image is with coordinates (uv) in the 2D image:
[0061] [Math.l] Z \ 1 /
[0062] With 'MiraMext; projection matrix associating each pixel with a Cartesian position in the three-dimensional reference frame, - Mint; so-called intrinsic matrix, whose coefficients are the intrinsic parameters of the camera - so-called extrinsic matrix, whose coefficients are parameters extrinsic parameters that can vary depending on the position of the camera, these parameters representing a translation and a rotation
[0063] Thus, based on a matrix of pixels of the captured image, a calibration matrix is determined associating each point of the three-dimensional reference frame with a pixel of the matrix of pixels.
[0064] Geometric calibration is implemented manually by an operator for greater precision.
[0065] In addition, during geometric calibration, possible geometric distortions are taken into account using various methods such as the checkerboard pattern method.
[0066] Alternatively, the geometric calibration is implemented semi-automatically or automatically for easy deployment, the automatic calibration method comprising a search for the horizon line or vanishing points from structuring elements of the image or from a checkerboard.
[0067] In a two-dimensional image processing step 200, the image processing module 44 implements a detection of at least one identifiable element 20A, 20B, 20C, 20D each identifiable element 20A, 20B, 20C, 20D having an associated estimated position 22A, 22B, 22C, 22D.
[0068] As explained previously, preferably the image processing step 200 implements a previously trained neural network to detect said identifiable elements 20A, 20B, 20C, 20D. For example, the image processing step 200 implements a deep learning algorithm, among the following algorithms: CMU open-pose, Mask RCNN, AlphaPose.
[0069] Preferably, the image processing step 200 comprises a detection of at least 4 identifiable elements 20A, 20B, 20C, 20D.
[0070] Alternatively, the image processing step 200 comprises a detection of a configurable number of identifiable elements 20A, 20B, 20C, 20D.
[0071] Finally, in a position estimation step 300, the position estimation module 46 calculates an absolute position 24 of the individual 10 in the three-dimensional monitoring space 5 from the geometric calibration, from the at least one identifiable element 20A, 20B, 20C, 20D of the body of the individual 10 and from a predetermined average height of the individual 10.
[0072] The average height of the individual is configurable, for example it is set to lm75.
[0073] From the predetermined average height of the individual, a predetermined average height of each identifiable element of the individual is estimated.
[0074] Thus, considering an average height of the individual fixed at lm75, the predetermined height of the hips of the individual 10 is estimated at 88cm, the predetermined height of the shoulders and the neck of the individual 10 is estimated at lm52.
[0075] Knowing the position of each identifiable element of the body of the individual 10 and the average height of each element Ye, using the projection matrix, the position:
[0076] [Math.2] Pose^= Mext, Pu, Pv, Fe)
[0077] With: - P°sest(i): estimated position of the identifiable element i; - Pu: two-dimensional position of the identifiable element i along the axis u; - Py: two-dimensional position of the identifiable element i along the v axis; - Ye ; average height of element i - f: projection function, which allows to solve [Mathl], and which from the projection matrices Mint and Mexh of the position of the pixel (PU,PV) and the height Ye estimates X,Z.
[0078] In one embodiment, the absolute position 24 of the individual in the ground plane of the three-dimensional reference frame is deduced from the average of the estimated positions on the ground plane 22A, 22B, 22C, 22D associated with each of the identifiable elements 20A, 20B, 20C, 20D and from the geometric calibration such that:
[0079] [Math.3] Posahs =---—-----
[0080] With: - P°sabs: absolute position 24 in the three-dimensional frame of reference, - E: number of identifiable elements,
[0081] As a variant, each identifiable element 20A, 20B, 20C, 20D is associated with a reliability weight making it possible to determine the absolute position 24 according to a weighted average of the estimated positions 22A, 22B, 22C, 22D of each of the identifiable elements 20A, 20B, 20C, 20D associated with said reliability weight such that:
[0082] [Math.4] +- ■ -+
[0083] With: - A5; reliability weight of the identifiable element i.
[0084] As a variant and with reference to [Fig.3], during a reconstruction step 400, a digital silhouette of the individual 10, in the three-dimensional frame of reference, is determined from the estimation of the absolute position 24 of the individual 10 and at least one predetermined average height of the individual.
[0085] In addition, an identifier is associated with each silhouette thus reconstructed and makes it possible to follow an individual 10 in a crowd.
[0086] It is then understood that the localization method, and the associated localization device 40, make it possible to locate an individual 10 precisely from images captured by conventional cameras, making this localization economical, flexible and non-intrusive.
[0087] The proposed method is particularly useful for obtaining a more precise localization of individuals in the presence of occlusions in the camera's field of view, since the absolute position takes into account an estimated average height based on the predetermined average height of the individual for each identified element of the individual's body.
Claims
Claims
1. Method for locating individuals in a three-dimensional frame of reference, comprising a ground plane, from two-dimensional images captured by a fixed camera, the camera having a field of vision of a three-dimensional surveillance space (5), each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional frame of reference of the image, the method being characterized in that it comprises the following steps: - geometric calibration (100) of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional frame of reference of the surveillance space (5); - processing two-dimensional images (200), captured by the camera, for detection of at least one identifiable element (20A, 20B, 20C, 20D) of the body of an individual (10), each identifiable element (20A, 20B, 20C, 20D) having an associated position in the two-dimensional frame of reference;- estimation (300) of an absolute position (24) of the individual on the ground plane of the three-dimensional reference system as a function of at least one identifiable element (20A, 20B, 20C, 20D), of a predetermined average height of the individual and of said geometric calibration.;
2. Method according to claim 1, further comprising a reconstruction step (400) comprising a determination of a silhouette of the individual (10) in three dimensions from the estimation of an absolute position (24) of the individual (10) on the ground plane of the three-dimensional reference frame and of said predetermined average height of the individual.
3. A method according to any one of claims 1 or 2, wherein the processing step (200) comprises detecting at least 4 identifiable elements (20A, 20B, 20C, 20D).
4. A method according to any one of claims 1 to 3, wherein the processing step (200) comprises detecting a configurable number of identifiable elements (20A, 20B, 20C, 20D).
5. Method according to any one of claims 1 to 4, in which the step of estimating (300) the absolute position (24) of the individual on the ground plane of the three-dimensional reference frame comprises a calculation of an average of estimated positions (22A, 22B, 22C, 22D) of the individual (10) in the three-dimensional reference frame.
6. A method according to any one of claims 1 to 5, wherein the two-dimensional image processing step (200) is implemented by a neural network.
7. A method according to any one of claims 1 to 6, wherein the identifiable elements (20A, 20B, 20C, 20D) of the body of an individual (10) in a two-dimensional image are part of a set comprising the feet, the head, the hips, the shoulders, the eyes or the nose.
8. Method according to any one of claims 1 to 7, in which the step of estimating (300) an absolute position (24) of the individual (10) on the ground plane of the three-dimensional reference frame implements a calibration matrix.
9. A computer program comprising software instructions which, when executed by a computer, implement a location method according to any one of claims 1 to 4.
10. O. Device (40) for locating individuals in a three-dimensional frame of reference comprising a ground plane, from two-dimensional images captured by a fixed camera, the device being connected or integrated into said camera, the camera having a field of vision of a three-dimensional surveillance space (5), each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional frame of reference of the image, the device being characterized in that it comprises a processor configured to implement: - a geometric calibration module (42) of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional frame of reference of the surveillance space (5);- a module (44) for processing two-dimensional images, captured by the camera, [for detecting at least one identifiable element (20A, 20B, 20C, 20D) of the body of an individual (10), each identifiable element (20A, 20B, 20C, 20D) having an associated position in the two-dimensional frame of reference; - a module (46) for estimating an absolute position (24) of the individual on the ground plane of the three-dimensional frame of reference as a function; of at least one identifiable element (20A, 20B, 20C, 20D), of a predetermined average height of the individual and of said geometric calibration.
11. System (60) for locating individuals in a three-dimensional frame of reference, characterized in that it comprises a fixed camera (15) and a device (40) for locating individuals in a three-dimensional frame of reference comprising a ground plane, from two-dimensional images captured by said fixed camera (15), the locating device (40) being in accordance with claim 10.
Citation Information
Patent Citations
Image processing apparatus for estimating three-dimensional position of object and method therefor
US20160292533A1
Three dimensional position estimation mechanism
US20190130602A1