Method for locating individuals in a three-dimensional reference frame from two-dimensional images captured by a fixed camera and associated localization device
The method and device enhance three-dimensional localization of individuals using geometric calibration and neural networks to process two-dimensional images from fixed cameras, addressing limitations of existing two-dimensional methods in complex environments.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Patents
- Current Assignee / Owner
- THALES SA
- Filing Date
- 2023-12-29
- Publication Date
- 2026-05-08
AI Technical Summary
Existing two-dimensional camera-based localization methods are limited in complex environments like airports and train stations, requiring more accurate, cost-effective, and less intrusive three-dimensional localization of individuals.
A method and device that utilize a fixed camera to geometrically calibrate two-dimensional images, detect identifiable body elements, and estimate absolute positions in a three-dimensional reference frame using a predetermined average height, incorporating a neural network for image processing.
Enables precise and economical three-dimensional localization of individuals in complex environments, even with occlusions, by leveraging conventional cameras with improved accuracy and flexibility.
Abstract
Description
Title of the invention: Method for locating individuals in a three-dimensional reference frame from two-dimensional images captured by a fixed camera and associated localization device
[0001] The present invention relates to a method for locating individuals in a three-dimensional reference frame from two-dimensional images captured by a fixed camera. The present invention also relates to a localization device associated with such a method.
[0002] In various fields such as security, robotics, and navigation, it is necessary to be able to locate an individual in different environments. Airports, train stations, and warehouses are sensitive locations and need to be monitored using three-dimensional localization of individuals passing through them, despite disruptions caused by various constraints such as crowds or obstructions.
[0003] To efficiently locate individuals in these places, different methods are generally employed depending on the needs.
[0004] Prior art location methods include all types of technologies including: beacons, GPS, cameras, Lidar or RFID chips.
[0005] All these technologies are used to varying degrees in the field of security and have advantages and disadvantages depending on the needs.
[0006] Overall, beacons, GPS, cameras, Lidar and all other types of sensors still pose many problems of cost and applicability in complex environments, while RFID chips, but also beacons, involve an intrusion and precision that are still too limited.
[0007] Cameras therefore remain the most used technology for locating individuals in specific environments.
[0008] However, it is known that cameras are often limited to two-dimensional surveillance and have still limited performance in complex environments such as train stations or airports.
[0009] There is therefore a need for a localization of individuals in a three-dimensional reference frame of good accuracy which is less expensive, less intrusive and more adapted to complex environments.
[0010] To this end, the present description relates to a method for locating individuals in a three-dimensional reference frame, including a ground plane, from two-dimensional images captured by a fixed camera, the camera having a field of view of a In a three-dimensional surveillance space, each captured image is represented by a pixel matrix, with each pixel having a position in a two-dimensional image reference frame. This process includes the following steps: - geometric calibration of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space; - processing of two-dimensional images, captured by the camera, for the detection of at least one identifiable element of an individual's body, each identifiable element having an associated position in the two-dimensional reference frame;
[0011] Estimating the absolute position of an individual on the ground plane of the three-dimensional reference frame as a function of at least one identifiable element, a predetermined average height of the individual, and said geometric calibration. Advantageously, estimating the absolute position of the individual as a function of at least one identifiable element of the individual's body and the predetermined average height of the individual makes it possible to improve the positioning accuracy in the three-dimensional reference frame from two-dimensional images.
[0012] According to other advantageous aspects of the invention, the method comprises one or more of the following features taken individually or in all possible combinations.
[0013] The method further comprises a reconstruction step involving the determination of a three-dimensional silhouette of the individual from the estimation of an absolute position of the individual on the ground plane of the three-dimensional reference frame and of said predetermined average height of the individual.
[0014] The processing step includes the detection of at least 4 identifiable elements.
[0015] The processing step includes the detection of a number of identifiable elements configurable.
[0016] The step of estimating the absolute position of the individual on the ground plane of the three-dimensional reference frame includes a calculation of an average of estimated positions of the individual in the three-dimensional reference frame.
[0017] The two-dimensional image processing step is implemented by a neural network.
[0018] The identifiable body features of an individual in a two-dimensional image are part of a set including the feet, head, hips, shoulders, eyes or nose.
[0019] The step of estimating an absolute position of the individual on the ground plane of the three-dimensional reference frame implements a calibration matrix.
[0020] The invention also relates to a device for locating individuals in a three-dimensional reference frame including a ground plane, based on two-dimensional images captured by a fixed camera. The device is connected to or integrated into said camera. The camera has a field of view of a three-dimensional surveillance space. Each captured image is represented by a pixel matrix, each pixel having a position in a two-dimensional image reference frame. This device includes a processor configured to implement: - a geometric calibration module for the camera associating each pixel of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space; - a two-dimensional image processing module, captured by the camera, [for the detection of at least one identifiable element of an individual's body, each identifiable element having an associated position in the two-dimensional reference frame; - a module for estimating an absolute position of the individual on the ground plane of the three-dimensional reference frame as a function of at least one identifiable element, a predetermined average height of the individual and said geometric calibration.
[0021] According to another aspect, the invention relates to a system for locating individuals in a three-dimensional reference frame, comprising a fixed camera and a device for locating individuals in a three-dimensional reference frame including a ground plane, from two-dimensional images captured by said fixed camera, the locating device being as briefly described above.
[0022] According to another aspect, the invention relates to a computer program comprising software instructions which, when executed by computer, implement a method for locating individuals in a three-dimensional reference frame as briefly described above.
[0023] The invention will become clearer upon reading the following description, given solely by way of non-limiting example and with reference to the drawings in which:
[0024] [Fig-1] [Fig.1] is a schematic representation of a localization system of an individual in a three-dimensional surveillance space including a ground plane, said surveillance being implemented by a location device associated with a fixed camera;
[0025] [Fig.2] [Fig.2] is a flowchart of the steps in the localization process by the location device of the [Fig.l];
[0026] [Fig.3] [Fig.3] is a schematic representation of a determination of an individual's silhouette in the three-dimensional surveillance space.
[0027] In [Fig.1], a three-dimensional surveillance space 5 comprising an individual 10 is monitored using two-dimensional images in a two-dimensional reference frame with perpendicular axes (U, V) equipped with a frame of transverse axis V and longitudinal axis U, the frame being centered on the lower left edge of the images captured by a camera 15, for example an optical camera.
[0028] The surveillance space 5 is included in a three-dimensional reference frame comprising a longitudinal axis X, a transverse axis Y and an anteroposterior axis Z. For example, the three-dimensional reference frame is the terrestrial reference frame (or world reference frame).
[0029] The camera 15 is, for example, a video surveillance camera, configured to capture sequences of images at a predetermined capture frequency. The camera 15 is spatially positioned at a predetermined location. In other words, the camera 15 has a fixed position in space.
[0030] Advantageously, this allows for the calculation of a reliable geometric calibration between the three-dimensional reference frame and the two-dimensional reference frame.
[0031] The three-dimensional surveillance space 5 is in a direct field of vision of the camera 15.
[0032] Each two-dimensional image captured is represented by a matrix of pixels, each pixel having a position in the two-dimensional reference frame associated with the image.
[0033] In the following description, it is considered that the set of pixels composing the two-dimensional images correspond to Cartesian positions in the three-dimensional reference frame.
[0034] The camera 15 includes or is connected to a location device 40. The location device 40 and the camera 15 form a location system 60 of an individual in the three-dimensional reference frame.
[0035] The location device 40 is a programmable electronic device, e.g. a computer, configured to communicate via a wired or wireless communication link with the camera 15.
[0036] According to one variant, the localization device 40 is a processing unit integrated into the camera 15.
[0037] The localization device 40 of individuals in a three-dimensional reference frame implements a detection of a plurality of identifiable elements 20A, 20B, 20C, 20D of the body of said individual 10, and an estimation of a plurality of estimated positions 22A, 22B, 22C, 22D associated with each of the identifiable elements 20A, 20B, 20C, 20D and of an absolute position 24.
[0038] The plurality of identifiable elements 20A, 20B, 20C, 20D is a set of physical parts characteristic of an individual 10 being chosen, for example, from the group consisting of: feet, head, hips, shoulders, eyes or nose.
[0039] In addition, each identifiable element 20A, 20B, 20C, 20D then has a two-dimensional position in the two-dimensional reference frame of the image.
[0040] The plurality of estimated positions 22A, 22B, 22C, 22D are projections in the three-dimensional reference frame, on the ground plane, along the Y axis of the two-dimensional positions of the plurality of identifiable elements 20A, 20B, 20C, 20D in the two-dimensional reference frame, the projections using a predetermined average height of the individual.
[0041] The localization device 40 is configured to estimate an absolute position 24 of the individual 10 considered as a point of the feet of said individual 10 in the three-dimensional reference frame.
[0042] The localization device 40 includes an information processing unit 48, formed for example of a processor 50, and an electronic memory 52 associated with the processor 50.
[0043] The localization device 40 includes a geometric calibration module 42, an image processing module 44 and a position estimation module 46.
[0044] The geometric calibration module 42 is configured to associate each of the pixels of the two-dimensional images captured by the camera 15 with a Cartesian position in the three-dimensional reference frame, the pixels being projected onto the floor of the surveillance space 5.
[0045] The image processing module 44 is configured to detect at least one identifiable element 20A, 20B, 20C, 20D of the body of an individual 10, each identifiable element 20A, 20B, 20C, 20D having an associated position in the two-dimensional reference frame.
[0046] The position estimation module 46 is configured to estimate an absolute position 24 of the individual on the ground plane of the three-dimensional reference frame as a function of at least one identifiable element 20A, 20B, 20C, 20D, a predetermined average height of the individual and the geometric calibration.
[0047] In one embodiment, the geometric calibration module 42, the image processing module 44 and the position estimation module 46 are implemented as software, or a software block, executable by the processor 50. The memory 52 is then capable of storing calibration, image processing and estimation software.
[0048] In an alternative not shown, the geometric calibration module 42, the image processing module 44 and the position estimation module 46 are implemented as a programmable logic component, such as an FPGA (Field Processing Gate Array). Programmable Gate Array), or in the form of an integrated circuit, such as an ASIC (Application Specifies Integrated Circuit).
[0049] In one embodiment, the image processing module 44 implements a neural network previously trained to recognize from two-dimensional images and position in the two-dimensional reference frame of each image, at least one identifiable element of the individual's body.
[0050] In general, the neural network comprises an ordered succession of layers of neurons, each of which takes its inputs from the outputs of the previous layer.
[0051] More precisely, each layer comprises neurons taking their inputs from the outputs of the neurons of the previous layer, or from the input variables for the first layer.
[0052] Alternatively, more complex neural network structures can be envisaged with a layer that can be linked to a layer further away than the immediately preceding layer.
[0053] Each neuron is also associated with an operation, that is to say a type of processing, to be carried out by said neuron within the corresponding processing layer.
[0054] Each layer is connected to the other layers by a plurality of synapses. A synaptic weight is associated with each synapse, and each synapse forms a link between two neurons. This is often a real number, which takes on both positive and negative values. In some cases, the synaptic weight is a complex number.
[0055] Each neuron is designed to perform a weighted sum of the value(s) received from the neurons of the preceding layer, each value being multiplied by the respective synaptic weight of each synapse, or connection, between said neuron and the neurons of the preceding layer, and then to apply an activation function, typically a non-linear function, to said weighted sum, and to deliver at the output of said neuron, in particular to the neurons of the next layer connected to it, the value resulting from the application of the activation function. The activation function introduces non-linearity into the processing performed by each neuron. The sigmoid function, the hyperbolic tangent function, and the Heaviside function are examples of activation functions.
[0056] As an optional complement, each neuron is also capable of applying, in addition, a multiplicative factor, also called bias, to the output of the activation function, and the value delivered at the output of said neuron is then the product of the bias value and the value from the activation function.
[0057] Such a neural network is trained on a database containing images of individuals possessing identifiable elements 20A, 20B, 20C, 20D as defined above.
[0058] The operation of the localization device 40 is now described with reference to [Fig.2] which illustrates the main steps of a localization process in one embodiment.
[0059] The localization process includes a first geometric calibration step 100 associating each of the pixels of the two-dimensional images captured by the video device 15 with a Cartesian position in the three-dimensional reference frame.
[0060] For example, the geometric calibration module 42 performs a geometric calibration from calibration matrices, according to the following relation for a point P with Cartesian coordinates (X, Y, Z) in the three-dimensional reference frame whose image has coordinates (uv) in the 2D image:
[0061] [Math.l] Z \ 1 /
[0062] With 'MiraMext'; projection matrix associating each pixel with a Cartesian position in the three-dimensional reference frame, - Mint; so-called intrinsic matrix, whose coefficients are the intrinsic parameters of the camera - an extrinsic matrix, whose coefficients are parameters extrinsic parameters that can vary depending on the camera's position, these parameters representing a translation and a rotation
[0063] Thus, depending on a pixel matrix of the captured image, a calibration matrix is determined associating each point of the three-dimensional reference frame with a pixel of the pixel matrix.
[0064] Geometric calibration is implemented manually by an operator for greater accuracy.
[0065] In addition, during geometric calibration, possible geometric distortions are taken into account using various methods such as the checkerboard target method.
[0066] Alternatively, geometric calibration is implemented semi-automatically or automatically for easier deployment, the automatic calibration method including a search for the horizon line or vanishing points from structuring elements of the image or from a checkerboard.
[0067] In a two-dimensional image processing step 200, the image processing module 44 implements the detection of at least one identifiable element 20A, 20B, 20C, 20D each identifiable element 20A, 20B, 20C, 20D having an associated estimated position 22A, 22B, 22C, 22D.
[0068] As explained previously, preferably the image processing step 200 implements a neural network previously trained to detect said identifiable elements 20A, 20B, 20C, 20D. For example, the image processing step 200 implements a deep learning algorithm, among the following algorithms: CMU open-pose, Mask RCNN, AlphaPose.
[0069] Preferably, the image processing step 200 includes the detection of at least 4 identifiable elements 20A, 20B, 20C, 20D.
[0070] Alternatively, the image processing step 200 includes the detection of a configurable number of identifiable elements 20A, 20B, 20C, 20D.
[0071] Finally, in a position estimation step 300, the position estimation module 46 calculates an absolute position 24 of the individual 10 in the three-dimensional monitoring space 5 from the geometric calibration, of at least one identifiable element 20A, 20B, 20C, 20D of the body of the individual 10 and of a predetermined average height of the individual 10.
[0072] The average height of the individual is configurable, it is for example fixed at lm75.
[0073] From the predetermined average height of the individual, a predetermined average height of each identifiable element of the individual is estimated.
[0074] Thus, considering an average height of the individual fixed at 1m75, the predetermined height of the hips of individual 10 is estimated at 88cm, the predetermined height of the shoulders and neck of individual 10 is estimated at 1m52.
[0075] Knowing the position of each identifiable element of the body of individual 10 and the average height of each element Ye, using the projection matrix, the position:
[0076] [Math.2] Pose^= Mext, Pu, Pv, Fe)
[0077] With: - P°sest(i): estimated position of the identifiable element i; - Pu: two-dimensional position of the identifiable element i along the u axis; - Py: two-dimensional position of the identifiable element i along the v axis; - Ye; average height of element i - f: projection function, which allows to solve [Mathl], and which from the projection matrices Mint and Mexh of the position of the pixel (PU,PV) and the height Ye estimates X,Z.
[0078] In one embodiment, the absolute position 24 of the individual in the ground plane of the three-dimensional reference frame is deduced from the average of the estimated positions on the ground plane 22A, 22B, 22C, 22D associated with each of the identifiable elements 20A, 20B, 20C, 20D and the geometric calibration such that:
[0079] [Math.3] Posahs =---—-----
[0080] With: - P°sabs: absolute position 24 in the three-dimensional reference frame, - E: number of identifiable elements,
[0081] Alternatively, each identifiable element 20A, 20B, 20C, 20D is associated with a reliability weight allowing the absolute position 24 to be determined according to a weighted average of the estimated positions 22A, 22B, 22C, 22D of each of the identifiable elements 20A, 20B, 20C, 20D associated with said reliability weight such that:
[0082] [Math.4] +- ■ -+
[0083] With: - A5; reliability weight of identifiable element i.
[0084] Alternatively and with reference to [Fig.3], during a reconstruction step 400, a digital silhouette of the individual 10, in the three-dimensional reference frame, is determined from the estimation of the absolute position 24 of the individual 10 and at least one predetermined average height of the individual.
[0085] In addition, an identifier is associated with each silhouette thus reconstructed and makes it possible to track an individual 10 in a crowd.
[0086] It is therefore conceivable that the localization process, and the associated localization device 40, make it possible to locate an individual 10 precisely from images captured by conventional cameras, making this localization economical, flexible and non-intrusive.
[0087] The proposed method is particularly useful for obtaining a more precise localization of individuals in the presence of occlusions in the field of view of the camera, since the absolute position takes into account an average height estimated according to the predetermined average height of the individual for each identified element of the individual's body.
Claims
Demands
1. A method for locating individuals in a three-dimensional reference frame, including a ground plane, from two-dimensional images captured by a fixed camera, the camera having a field of view of a three-dimensional surveillance space (5), each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional reference frame of the image, the method being characterized in that it comprises the following steps: - geometric calibration (100) of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space (5); - processing of two-dimensional images (200) captured by the camera, for the detection of at least one identifiable element (20A, 20B, 20C, 20D) of the body of an individual (10), each identifiable element (20A, 20B, 20C, 20D) having an associated position in the two-dimensional reference frame;- estimation (300) of an absolute position (24) of the individual on the ground plane of the three-dimensional reference frame as a function of at least one identifiable element (20A, 20B, 20C, 20D), a predetermined average height of the individual and said geometric calibration.;
2. A method according to claim 1, further comprising a reconstruction step (400) comprising determining a three-dimensional silhouette of the individual (10) from the estimation of an absolute position (24) of the individual (10) on the ground plane of the three-dimensional reference frame and of said predetermined average height of the individual.
3. A method according to any one of claims 1 or 2, wherein the processing step (200) comprises a detection of at least 4 identifiable elements (20A, 20B, 20C, 20D).
4. A method according to any one of claims 1 to 3, wherein the processing step (200) comprises a detection of a configurable number of identifiable elements (20A, 20B, 20C, 20D).
5. A method according to any one of claims 1 to 4, wherein the estimation step (300) of the absolute position (24) of the individual on the ground plane of the three-dimensional reference frame comprises a calculation of an average of estimated positions (22A, 22B, 22C, 22D) of the individual (10) in the three-dimensional reference frame.
6. A method according to any one of claims 1 to 5, wherein the two-dimensional image processing step (200) is implemented by a neural network.
7. A method according to any one of claims 1 to 6, wherein the identifiable elements (20A, 20B, 20C, 20D) of the body of an individual (10) in a two-dimensional image are part of an ensemble comprising the feet, head, hips, shoulders, eyes or nose.
8. A method according to any one of claims 1 to 7, wherein the estimation step (300) of an absolute position (24) of the individual (10) on the ground plane of the three-dimensional reference frame implements a calibration matrix.
9. A computer program comprising software instructions which, when executed by a computer, implement a localization method according to any one of claims 1 to Q
10. O. A device for locating individuals in a three-dimensional reference frame including a ground plane, from two-dimensional images captured by a fixed camera, the device being connected to or integrated into said camera, the camera having a field of view of a three-dimensional surveillance space (5), each captured image being represented by a matrix of pixels, each pixel having a position in a two-dimensional reference frame of the image, the device being characterized in that it includes a processor configured to implement: - a geometric calibration module (42) of the camera associating each of the pixels of the two-dimensional images captured by said camera with a Cartesian position in the three-dimensional reference frame of the surveillance space (5);- a two-dimensional image processing module, captured by the camera, [for detection of at least one identifiable element (20A, 20B, 20C, 20D) of the body of an individual (10), each identifiable element (20A, 20B, 20C, 20D) having an associated position in the two-dimensional reference frame; - a module (46) for estimating an absolute position (24) of the individual on the ground plane of the three-dimensional reference frame as a function; of at least one identifiable element (20A, 20B, 20C, 20D), of a predetermined average height of the individual and of said geometric calibration.
11. A localization system (60) for individuals in a three-dimensional reference frame characterized in that it comprises a fixed camera (15) and a localization device (40) for individuals in a three-dimensional reference frame comprising a ground plane, from two-dimensional images captured by said fixed camera (15), the localization device (40) being in accordance with claim 10.