Method and device for capturing images of moving subjects

JP2026148536APending Publication Date: 2026-09-17アイデミアパブリックセキュリティフランス
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026035447
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-05
Filing Date
2026-03-05
Publication Date
2026-09-17

Smart Images

  • Figure 2026148536000001_ABST
    Figure 2026148536000001_ABST
Patent Text Reader

Abstract

This invention provides a method and device for capturing images of moving subjects. [Solution] The present invention relates to a method for capturing an image of a region of interest of a moving subject (103), the method comprising: calculation (repeated at computational frequency) of the three-dimensional position of the region of interest in a capture volume; estimation of a three-dimensional path of the region of interest from the computational position, the path being defined by a parametric path model combining periodic and linear time components; and at least one future time t from the estimated path. capture Determining the location of the region of interest; time t capture This includes capturing an image of the region of interest using a capture camera that is oriented in the direction of the location where the region of interest is determined.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This invention relates to the field of capturing images of a region of interest localized on a moving subject. Specifically, this invention relates to tracking and acquiring clear images of a region of interest localized on a moving person.

[0002] In the first example, this could be the problem of recognizing the iris of a person passing in front of an image capture system in the case of biometric recognition of a moving person. In the second example, this could also be the problem of reading information such as a two-dimensional code or QR code on a badge worn by a moving person. Other examples of technical contexts may benefit from the present invention.

[0003] Systems already exist that allow a region of interest to be captured while the subject is stationary. The capture volume is the spatial region considered when capturing an image of the region of interest. This is the region in which the subject can move while the image of the region can be captured. Once the subject leaves the capture volume, it is no longer visible to the image capture system.

[0004] When the subject is stationary, the image capture system can operate according to the following principle: In the first step, the subject is detected and localized within the capture volume. Next, the image capture system waits for the subject to stabilize. Once the subject is stable, the precise step of localizing the region of interest allows for the acquisition of its three-dimensional coordinates. These coordinates then allow a capture camera, equipped with a motor that enables its orientation, to be positioned precisely in the direction of the target region of interest. Next, a series of images are captured until the acquired quality is sufficient for the application (typically iris recognition or information reading). Note that "this assumed application requires that high-resolution images be captured, regardless of whether they are for iris recognition or information reading."

[0005] When the subject is moving, capturing a clear image of the region of interest becomes more difficult. Specifically, the subject's movement can produce a blurred image. In addition, due to the required resolution and sensitivity, the captured image has a very shallow depth of field. At this time, inaccuracies in the distance of the region of interest also result in the capture of a blurred image. Furthermore, it has been observed that when capturing an iris image at short distances (especially less than 1 meter), aiming becomes difficult because the person's movement can no longer be compensated for by the camera's field of view, and the target moves out of the capture camera's field of view. Similarly, at short distances, it has been observed that the time delay between eye detection and the control of the capture camera's movement by its motor can significantly affect the accuracy of the estimated position. These shortcomings can lead to failures when attempting to track at short distances or at high walking speeds.

[0006] The present invention aims to solve all or some of these problems by proposing a method for capturing images of a region of interest of a moving subject, including when the subject is walking at high speed (e.g., more than 1 m / s, and especially up to 1.5 m / s), and particularly at short distances (e.g., less than 1 meter from the capture camera, and especially up to 70 cm), enabling reliable acquisition of clear images of the target region of interest of the moving subject.

[0007] According to one aspect of the present invention, a method is proposed for capturing an image of a region of interest of a moving subject within a capture volume, the method comprising the following steps: - A calculation process (repeated at a computation frequency) for the three-dimensional position of a region of interest, wherein the calculation process outputs a series of three-dimensional positions of the region of interest (on which the instantaneous observable position signal depends) from a series of acquired images in which the region of interest is detected and localized; - The process of extracting a periodic observable signal (at the applicable frequency) from an instantaneous observable position signal; - An estimation step of a three-dimensional path of a region of interest within a captured volume from a computational location, wherein the three-dimensional path is defined in at least one dimension by a parametric path model combining periodic and linear time components; - At least one future time t from the estimated three-dimensional path of the region of interest capture The process of determining the location of the region of interest; -Image of the region of interest captured by the camera (time t) capture The capture process in which the capture camera is at time t capture A capture process in which the location of the determination of the area of ​​interest is oriented in the direction of the area.

[0008] This method enables the capture of a clear image of a moving subject's region of interest (including at short distances) through a short tracking time after the detection and localization of a series of three-dimensional positions of the region of interest. Furthermore, this method allows for latency compensation. Consequently, the number of images of the subject captured by the capture camera is reduced, and even capturing a single image of the subject by the capture camera is sufficient. Due to the image quality required for the application (especially in biometric applications), capturing a series of images of the subject by the capture camera is no longer necessarily required. Moreover, this method allows for the parallel tracking of multiple subjects.

[0009] In particular, since the subject's movement is walking, especially at speeds of less than 2 m / s and especially between 1 and 1.5 m / s, this method is suitable not only for average walking speed but also for brisk walking speed.

[0010] Advantageously, the acquired image is an image of the detection volume, also called a context image, which is acquired by the context camera system. This enables tracking by the context camera, during which the three-dimensional position of a series of regions of interest within the detection volume is tracked and acquired after detection and localization. Similarly, if multiple subjects are within the detection volume and appear in the context image, one region of interest will be tracked for each subject.

[0011] Advantageously, if the acquisition is performed at a given acquisition frequency, the calculation frequency is particularly less than or equal to the acquisition frequency.

[0012] Advantageously, the estimation is performed at an estimation frequency preferably equal to the calculation frequency.

[0013] Advantageously, the image is captured by the capture camera in the capture volume at the capture frequency.

[0014] Advantageously, the detection volume (corresponding to a segment of the field of view of the context camera system where the acquired image quality is sufficient to perform subject detection (e.g., at a depth greater than 2 meters from the context camera system, and preferably at a depth of 2.50 meters)) and the capture volume (corresponding to a segment of the field of view of the capture camera where the captured image quality is sufficient to achieve recognition of the region of interest) have an intersection, which allows working with a detection volume having a greater depth than the depth of the capture volume, which is particularly constrained by the high resolution required for iris acquisition in the context of biometrics (e.g., a depth of 1.50 meters to 0.50 meters from the capture camera in the case of iris recognition). The large depth of the detection volume allows the model to track so that, given that initialization occurs, for example, in a half-step motion (a motion of about 50 centimeters), the subject has already converged before entering the capture volume.

[0015] Preferably, the instantaneously observable position signal consists of a series of positions calculated with respect to the region of interest.

[0016] Advantageously, the periodic time component of the path model includes amplitude, phase, and angular frequency parameters, allowing for the representation of the periodic sinusoidal time component. Frequency decomposition is possible in a manner equivalent to this spectral angular frequency decomposition.

[0017] Advantageously, the linear time component of the path model includes a rate of change parameter and an intercept parameter.

[0018] Advantageously, one of the at least one of the aforementioned dimensions is the vertical dimension (y), which allows for precise modeling of the height of the region of interest of a moving subject, particularly by taking into account its up-and-down rhythm, which is especially important for orienting the capture camera in the direction of the determined location of the region of interest, especially when the subject is walking.

[0019] Advantageously, one of the at least one of the dimensions is the lateral dimension (x), which allows for precise modeling of the horizontal lateral swaying motion (motion in the x-direction) of the region of interest of the moving subject (especially when the subject is walking).

[0020] Similarly, the at least one dimension may specify both of the aforementioned dimensions, since stepping (i.e., walking) involves not only the up-and-down vertical movement (in the y-direction) of the subject but also the left-and-right swaying movement (in the x-direction).

[0021] Advantageously, one of the at least one of the aforementioned dimensions is the longitudinal dimension (z), which allows for precise modeling of the horizontal longitudinal depthwise motion (motion in the z direction) of the region of interest of a moving subject, especially when the subject is walking. Specifically, the path in the depth dimension (z) can also be modeled in this way, to a lesser extent, which provides a better estimation than linear approximation alone, since the subject's walking follows a specific rhythm (which varies depending on the supporting leg, for example), and the resulting motion is not necessarily symmetrical, especially when a person is dragging their feet, generating a back-and-forth swaying effect with each step.

[0022] Therefore, the at least one dimension can favorably specify all three of the aforementioned dimensions so that the location of the region of interest in the three dimensions can be precisely estimated.

[0023] To your advantage, pointing the capture camera is time t capture This involves adjusting the focus of the capture camera according to the distance between the capture camera and the determined position of the region of interest, which allows for improved sharpness of the acquired image (especially if the depth of field of the capture camera lens is not sufficiently wide). Advantageously, adjustment and aiming can be performed simultaneously.

[0024] To the advantage, the focus of the capture camera is "time t capture The focal tilt is adjusted depending on the distance between the capture camera and the determination position of the region of interest, plus a focal coefficient that varies in set or variable steps within a fixed interval (especially a focal coefficient centered on zero), thereby enabling the realization of a focal tilt.

[0025] Conveniently, the focal coefficient varies linearly between the extrema at the zero-center interval.

[0026] Conveniently, the focal coefficient varies sinusoidally between the extrema at zero center intervals.

[0027] Advantageously, the focal coefficient varies in increments corresponding to the depth of field between the extremes of the zero-center distance.

[0028] In one embodiment, the subject is a person, and the region of interest is a part of the face of a moving subject including the eyes (preferably the iris), or a visual representation (emoji) such as a two-dimensional barcode attached by the subject.

[0029] Advantageously, the capture frequency of the capture camera is higher than the acquisition frequency of the context camera system.

[0030] In one embodiment, the path estimation step includes implementing an extended Kalman filter on the parametric path model in at least one dimension of the detection region of interest of each subject, which includes an initialization step of the extended Kalman filter, a prediction step by the extended Kalman filter, and an update step of the extended Kalman filter at the sampling frequency. The extended Kalman filter enables the local linearization of the nonlinear components of the parametric path model in the region of interest. The extended Kalman filter enables the simultaneous creation of predictions (future values) and denoising the model by a nonlinear equation of sinusoidal and measurement noise (something particularly useful if it is conceivable that the instantaneous observable position signal can be decomposed into a combination of linear functions). In particular, if multiple subjects are detected, the path estimation step includes implementing an extended Kalman filter on the parametric path model in at least one dimension with respect to the detection region of interest of each subject.

[0031] The calculated or determined position and the estimated path are preferably located within or converted to a reference system belonging to the capture camera.

[0032] In particular, at least one future time t from the path estimated via the implementation of the extended Kalman filter capture The determination of the location of the region of interest is performed only when the Kalman filter converges.

[0033] Advantageously, the path estimation frequency is the same as the sampling frequency of the extended Kalman filter, also known as the update frequency of the extended Kalman filter, and is less than or equal to the frequency of context image acquisition by the context camera system, and is preferably greater than or equal to the calculation frequency (i.e., the frequency of calculating the three-dimensional position of the region of interest from a series of acquired context images).

[0034] Advantageously, the three-dimensional path of the region of interest is estimated at a frequency equal to the frequency of the calculation of the three-dimensional position of the region of interest, thereby enabling the calculation to be synchronized. In a variant form, the sampling frequency of the extended Kalman filter can be variable (for example, in the case of a context camera system that includes many sensors whose data are fused and have various frequencies), so that the position calculation does not necessarily occur at the synchronous time, and the filter is preferably updated with each new calculated position, and such devices and methods in particular make it possible to limit concealment within the detection volume thanks to the fusion of sensor data.

[0035] Advantageously, the frequency of position calculations fluctuates over time, thereby allowing it to adapt to the availability of sensor measurements, particularly if the target is momentarily hidden, for example in the context of iris acquisition, a person may turn around, which would remove the person's eyes from the field of view of the acquisition, or if changes in the exposure time of the context camera system may occur depending on the lighting conditions. Similarly, this also allows for mitigation of image loss caused by overloading the central processing unit of the data processing device implementing the invention. For example, in the case of non-detection of the region of interest in the image, the frequency of position calculations can be reduced. Thus, the update frequency of the extended Kalman filter is variable depending on whether the region of interest is detected in the acquired image, and the extended Kalman filter is not updated during calculation cycles in which the region of interest was not detected in the acquired image during the cycle, thereby avoiding calculations.

[0036] In one embodiment, the method includes the step of applying the band-pass filter in at least one dimension to an instantaneous observable position signal at the application frequency for the purpose of extracting a periodic observable signal.

[0037] In particular, the periodic observable signal is centered at zero, which is a result of applying a band-pass filter and enables compatibility with the periodic time component of the zero-centered path model.

[0038] Advantageously, the update frequency of the extended Kalman filter is equal to the application frequency of the band-pass filter.

[0039] In particular, the application frequency of the band-pass filter is variable, which makes it possible to address cases of subject obscuration (i.e., when the three-dimensional position cannot be calculated in several acquired images).

[0040] Advantageously, the band-pass filter is second-order, allowing for customized and robust embodiments. Advantageously, the coefficients of the band-pass filter are determined experimentally based on tests performed on samples of subjects whose mean vibration frequency (in the y-direction) varies preferably between 1 and 2 Hz by spectral analysis of the data thus acquired (especially at varying walking speeds).

[0041] In one embodiment, the extended Kalman filter is based on a measurement vector that includes at least two observers, one of which has an instantaneous observable position signal (consisting of a series of calculated positions in the region of interest) as a first observer and the other having a periodic observable signal as a second observer.

[0042] In one embodiment, the state vector of the extended Kalman filter includes the following five states: the average position of the region of interest in at least one dimension, the rate of change of the average position of the region of interest in at least one dimension, the amplitude of the oscillation of the region of interest centered on the average position, the phase of the oscillation centered on the average position, and the angular frequency of the oscillation of the average position of the region of interest. For example, in the case of an application to the iris, in the case of a parametric path model in only dimension y, the states listed above, for example, correspond to the height of the person's eyes, the rate of change of this height (excluding any oscillations ("linear" components)), the amplitude of the oscillation of the eye height, the phase of the sine wave modeling this oscillation centered on the average height of the eyes, and the rate of change of this phase (i.e., the angular frequency of the oscillation resulting in particular from the pace of the person's steps). This composition of states allows for improved accuracy by increasing fidelity to the parameters describing the gait. Equivalently, angular frequency and phase can be substituted with frequency and phase.

[0043] In one embodiment, the initialization step of the extended Kalman filter includes determining an initial state vector by determining initialization parameters including amplitude parameters; phase parameters; angular frequency parameters; rate of change parameters; and intercept parameters, the initialization parameters being determined from all or part of a series of calculated three-dimensional positions of a region of interest. In particular, with respect to the region of interest, a single initialization step is sufficient, including the case of temporary occlusion of subjects in one or more images.

[0044] Advantageously, the portion of a series of calculated three-dimensional positions includes at least two positions calculated based on images continuously acquired by a context camera system, and, for example, the first n calculated three-dimensional positions of a region of interest, where n is in particular 4 or greater, and especially 8 with respect to a subject whose step frequency is substantially 1 Hz.

[0045] In one embodiment, the portion of a series of calculated three-dimensional positions includes positions calculated with respect to the region of interest based on sequentially acquired images, the first of which is the image in which the region of interest is first detected, and subsequent images are images in which the region of interest is detected until at least two local extrema are detected, which allows for reaching the end of a step motion and for initialization to occur with respect to at least one half-step motion, with the center of the step motion considered to correspond to a maximal and the end of the step motion considered to correspond to a minimum (and vice versa). Thus, this embodiment enables a robust initialization that is appropriate for multidimensional path models.

[0046] In one embodiment, the covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and at least one of: the variance of the rate of change of the average position of the region of interest in said at least one dimension, the variance of the angular frequency of oscillation of the region of interest centered on the average position of the region of interest, or the variance of the amplitude of oscillation of the region of interest centered on the average position of the region of interest. Accordingly, the covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and at least one variance among a representative velocity of the moving object, a period of oscillation representing the motion of the moving object, or an amplitude of oscillation representing the motion of the moving object. Said variance is preferably empirically calibrated in advance (e.g., based on statistical prior studies) and / or depends on the position of the region of interest calculated before the prediction step.

[0047] Preferably, the covariance matrix of the process noise of the extended Kalman filter depends on the sampling period of the extended Kalman filter and all three variances.

[0048] In one embodiment, for any future time t tracking , the estimation of the path of the region of interest at said any future time t tracking is obtained by calculating the tangent to the modeled path, and the parameters thereof result from the update of the extended Kalman filter based on images acquired prior to the estimation; in particular, the modeled path is written in terms of position pos and velocity spd as follows: pos(t tracking )=P average path (t measure position )+drift average path *(t tracking -t measure position )+Amplitude*sin(Φ(t measure position )+ω*(t tracking -t measure position )) spd(t tracking )=drift average path +Amplitude*ω*cos(Φ(t measure position )+ω*(t tracking -t measure position )), where, tmeasure position is the time when the final context image used to estimate the path was acquired, and the tangent to the parameter model is written as follows: pos(t) = pos(t tracking )+spd(t tracking )*(tt tracking ), This is in the future time t capture The determination of the location of the region of interest is obtained by t=t capture This can be made possible by applying [the relevant principle].

[0049] Any future time t tracking This is advantageous because it allows for the estimation of the path reception time by a real-time coprocessor, which enables the consideration of latency between real time and the time of context image acquisition.

[0050] In one embodiment, the capture camera is directed towards the position of the region of interest, which is determined based on the three-dimensional path of the region of interest at a frequency higher than the calculation frequency, where the aiming direction changes at a rate higher than the position calculation frequency and real-time tracking is performed. tracking This allows the process to be performed based on the latest estimation of the path in the region of interest.

[0051] Advantageously, the context camera system includes two cameras that generate stereo images.

[0052] Advantageously, context camera systems include time-of-flight (TOFF) cameras.

[0053] According to another aspect of the present invention, a computer program is proposed that includes instructions configured to perform each step of the process of the method according to the present invention when the computer program is executed on a computer.

[0054] According to another aspect of the present invention, a removable or non-removable information storage means that is partially or fully readable by a computer or microprocessor is proposed, the means including a code instruction of a computer program for performing each step of the steps of the method according to the present invention.

[0055] According to another aspect of the present invention, a device is proposed for capturing an image of a region of interest of a moving subject in a capture volume, the device being - Context camera system; - Capture camera; - Includes a processor configured for the following processes: - A computational process (repeated at computational frequency) for the three-dimensional position of a region of interest, which outputs a series of three-dimensional positions of the region of interest (on which the instantaneous observable position signal depends) from a series of acquired images obtained by a context camera, on which the region of interest is detected and localized; - The process of extracting a periodic observable signal (at the applicable frequency) from an instantaneous observable position signal; - An estimation step of a three-dimensional path of a region of interest within a captured volume from a computational location, wherein the three-dimensional path is defined in at least one dimension by a parametric path model combining periodic and linear time components; - At least one future time t from the estimated three-dimensional path of the region of interest capture The process of determining the location of the region of interest; -time t capture Image of the region of interest taken by a capture camera facing the direction of the location of the determination of the region of interest (time t capture The capture process in which the capture camera takes place at time t capture The region of interest is directed within the direction of the judgment location, and the image is time t capture A capture step that enables capture by a capture camera and has the same advantages as the method according to the present invention.

[0056] In one embodiment, the capture camera is rotatable so as to face the direction of the determination position of the region of interest in at least one dimension.

[0057] Advantageously, the context camera system includes two cameras that generate stereo images.

[0058] Advantageously, context camera systems include Tof cameras.

[0059] The present invention will be better understood through the following description, which refers to embodiments and variations of the present invention given by non-limiting examples and described with reference to the accompanying drawings. [Brief explanation of the drawing]

[0060] [Figure 1] This document shows an image capture device according to one embodiment of the present invention in the context of an application. [Figure 2] This diagram shows the architecture of an image acquisition device according to one embodiment of the present invention. [Figure 3] This diagram shows the hardware architecture of an image acquisition device according to one embodiment of the present invention. [Figure 4] This document shows a software architecture used for path estimation and tracking in one embodiment of the present invention. [Figure 5] The principle of path estimation in one embodiment of the present invention is shown. [Figure 6] This is an example of a data processing device layout that may be used to implement one or more embodiments of the present invention.

[0061] The same reference was used in all attached drawings to specify elements that are identical or similar in form or function.

[0062] The present invention is applicable to a variety of contexts. The present invention may relate to the recognition of the iris of a moving person 103 passing in front of an image-capturing device 100, or even to the recognition of a badge worn by the moving person 103 passing in front of the image-capturing device 100, for example, to allow or prevent the person from accessing and / or to timestamp their entry. The subject may, for example, pass through a security gate (in an airport, virus laboratory, or nuclear power plant, etc.), or simply move around in a free space (such as a public space) and be tracked by the image-capturing device. Another example of an embodiment of the present invention relates to the capture of a two-dimensional code or QR code (QR stands for Quick Response) placed on an object being moved around by a moving person (particularly a person walking, especially during delivery or processing). This example involves the problem of capturing an image of the code on an object while the object is being processed (usually in a warehouse). In all cases, this is the problem of capturing a clear image of a region of interest localized on or thereby worn on a moving subject. The embodiment shown in Figure 1 relates to the context of iris recognition of a moving person 103 passing in front of an image capture device 100 according to one embodiment of the present invention, which is placed in an airport. Acquisition of the person 103's iris in motion allows the person to be recognized and authenticated upstream of the device 100, which allows, for example, if the person is recognized as permitted to enter, a security gate that would otherwise prevent access to be opened, and if necessary, a record of their entry is maintained, making entry smooth and rapid without the person having to stop in particular, and allowing the person to maintain their walking speed.

[0063] Figure 2 shows the architecture of device 100 according to one embodiment of the present invention. This figure shows a capture volume 101 equipped with a capture camera 102. The capture camera 102 is advantageous in that it can capture an image of a subject 103 having a region of interest 104. To do this, the capture camera 102 is typically motorized to allow it to be oriented in space. In this embodiment, the capture camera 102 is equipped with two motors, one of which enables rotation in the horizontal plane and the other enables rotation in the vertical plane. The movement of these two motors allows the capture camera to be oriented in any direction within the capture volume 101. A third motor adjusts the focus of the capture camera, i.e., adapts the image capture to the distance of the subject from the capture camera (more specifically to the distance of the region of interest 104 of the subject 103 from the capture camera), where this distance corresponds to depth.

[0064] The image acquisition device 100 is controlled by a data processing device 106, which enables the movement of these motors to be controlled. Typically, but not always, the image acquisition device 100 is also a data processing device 106 that receives images from the capture camera 102 and processes the received images. The data processing device is typically a computer, tablet, smartphone, or any other device that enables a computer program responsible for controlling the camera and acquiring images, and various steps of the method according to the present invention to be performed.

[0065] In one reference embodiment, the capture camera 102 is captured by a context camera system (particularly including one or more context cameras 105), and the distance from the subject is probably calculated by triangulation using at least two context cameras 105. This context camera system enables capturing images of the detection volume covering all or part of the capture volume, then recognizing the subject 103 in the acquired images before localizing them, and then localizing the region of interest 104 of the subject. In some embodiments, the context camera system is also controlled by a data processing device 106. Alternatively, a dedicated device may be responsible for controlling the context camera system. In this case, this device communicates with device 106 for the purpose of capturing images. The context camera system may, in variant forms, or in addition to the context cameras, include a ToF (Time of Flight) sensor, radar, or any other sensor that can contribute to localizing the subject. Context cameras and their optional attached sensors enable the calculation of the three-dimensional position of the subject 103, more specifically the three-dimensional position of the subject's region of interest 104 in space, and in particular, the three-dimensional position within the detection and capture volume.

[0066] This estimation of the region of interest's position is used to position the capture camera and to capture and adjust the image using the capture camera. However, the acquisition of images by the context camera 105, their analysis to determine the three-dimensional position of the region of interest, the calculation of motor control that enables the capture camera 102 to be oriented in the calculated position direction, and the acquisition of images by the capture camera all take a certain amount of time. Therefore, there is a delay between the position calculation and the acquisition of an image of the region of interest. Due to the continuous movement of the subject, these delays, and inherent inaccuracies in the positioning calculation, acquiring a clear captured image of the region of interest is challenging.

[0067] According to the present invention, measures are taken to control the capture camera not by relying on the location of the region of interest calculated by the context camera, but by relying on an estimation of where this region of interest will be when an image of it is actually captured by the capture camera. Thus, the location estimation is an estimation of the future location of the region of interest at the time of estimation. This measure is based on the calculated actual location, which allows the path of the subject 103 to be estimated and thus allows their future locations to be determined at the moment the capture camera captures the image. What is meant by path is location as a function of time, modeling motion as a function of time: in other words, the path of a point is a geometric curve that describes it as it passes through a reference frame, and thus provides information about its position and velocity over time in the reference frame. To determine this future location, the three-dimensional path of the region of interest 104 is estimated from the calculated location, and the three-dimensional path is defined by a parametric path model that combines periodic and linear time components in at least one dimension. Therefore, it becomes clear that the future position can be determined with appropriate accuracy for the subject's movement thanks to a combination of two components, regardless of whether the subject is, for example, walking or in a wheelchair. Specifically, if the person is in a wheelchair, the periodic component of the parameter model is estimated to be zero.

[0068] Where appropriate, focusing may be performed, and the determined three-dimensional position is then corrected by a focus coefficient in a dimension representing the distance of the region of interest from the capture camera. This focus coefficient represents the sum of values ​​that vary linearly between a negative minimum and a positive maximum, centered around a value of zero, and changes between each image capture. Slightly adjusting the focus of the image capture in this way increases the chance of obtaining at least one sharp image within a set of captured images.

[0069] Advantageously, the motors equipped in the capture camera are brushless motors, and moreover, direct-drive motors (enabling continuous motor motion). Brushless motors allow the capture camera to smoothly and continuously track the motion of the subject without compromising image quality. When the motors are direct-drive motors, the absence of the reduction gears inherent in these motors eliminates associated mechanical play. This would not be possible with stepper motors controlled using step control methods as used in known systems.

[0070] Figure 3 shows the hardware architecture of an image acquisition device 100 according to an example embodiment of the present invention. The image acquisition device 100 is primarily controlled by a main processor 201, which is preferably (but not necessarily) hosted in a data processing device 106. This processor 201 is typically operated by a conventional operating system such as Linux®, Windows®, or MacOS®. These operating systems are not strictly real-time and therefore may affect the accuracy of the estimation, so this must be taken into consideration.

[0071] The image acquisition device 100 also preferably (but not necessarily) includes a coprocessor 202 hosted in a data processing device 106. The main feature of this coprocessor 202 is that it is real-time in a guaranteed manner (i.e., it can perform several tasks in a precise given time). The real-time coprocessor 202 is responsible for triggering the context camera 105 (211), triggering the capture camera 102 (212), controlling a motor 205 for focusing the capture camera (213), controlling an aiming motor 206 that enables the capture camera to be oriented (214), and finally controlling the illumination of the region of interest 207 in synchronization with the image acquisition by the capture camera (215). The image 209 generated by the context camera 105 and the image 210 generated by the capture camera 102 are sent to the main processor 201 for processing.

[0072] The real-time coprocessor 202 is responsible for real-time tasks. The real-time coprocessor 202 triggers image acquisition by the context camera system 105 at regular intervals (preferably at a frequency of 7 Hz or higher, up to 15 Hz in this case, for example, in the embodiments described). The real-time coprocessor 202 controls the motion of the aiming motor, for example, via servo control of position and / or velocity. The real-time coprocessor 202 controls the motion of the focusing motor. The real-time coprocessor 202 triggers image acquisition by the capture camera 102 at appropriate times and activates illumination of the region of interest in synchronization with the acquisition by the capture camera 102.

[0073] The main processor 201 receives a stream of images acquired by the context camera 105. Within these acquired images (also called context images), the main processor 201 detects the presence or absence of the subject 103 and localizes the region of interest 104 within the detected volume in three dimensions. Based on these three-dimensional coordinates, the main processor estimates the path of the region of interest and then transmits this path to the coprocessor 202.

[0074] The coprocessor 202 servo-controls the position and speed of the aiming motor 206 to reach the path transmitted by the main processor 201 as quickly as possible.

[0075] The aiming motor, at time t capture When the coprocessor 202 "locks on" to the ideal path and point in the direction of the determined location in the region of interest, time t capture At this point, the capture camera 102 triggers the capture of an image of the region of interest, and this capture is probably triggered, for example, at time t capture If the subject is hidden and not visible, the process is repeated (specifically, periodically).

[0076] More frequently, coprocessor 202 receives route updates as it gets closer to the actual path in the region of interest.

[0077] The advantage of synchronization between the main processor and the coprocessor is that instruction transmission by the main processor is not subject to real-time constraints, thereby enabling the use of a multitasking non-real-time operating system. In other alternative embodiments, a single real-time processor handles all tasks that are shared between the central processing unit and the real-time coprocessor.

[0078] Figure 4 shows a software architecture used for path estimation and tracking in an example of an embodiment of the present invention.

[0079] An example of this embodiment is based on the use of two context cameras 105 used in stereo mode to enable three-dimensional localization of objects detected in the image. Note that other embodiments may use other techniques as an alternative to or in addition to stereo imaging (such as a Toff camera), or even use systems for realizing three-dimensional vision based on structured light.

[0080] The main processor 201 executes a first module 301 responsible for localizing the region of interest (e.g., the subject's eye in the case of iris recognition in a context image). The recognition and localization algorithms used are known in themselves and are therefore not the subject of this specification.

[0081] A second module 302, executed by the main processor 201, is responsible for calculating the three-dimensional position of the region of interest positioned by the first module. Preferably, the frequency at which the three-dimensional position of the region of interest is calculated is less than or equal to the acquisition frequency by the context camera (e.g., 15 Hz). The advantage of calculating the position at a frequency lower than the image acquisition frequency is that it reduces the processor load during high demand, and therefore the calculation frequency can be varied depending on the load, which is advantageous. These positions are stored in memory, or at least the most recent positions are stored in memory. At least two positions are stored in memory for the purpose of estimating a path defined in at least one dimension by a parametric path model that combines periodic and linear time components. Preferably, the constituent positions of the subject in full steps are used in the initialization stage of an extended Kalman filter employed to estimate the path.

[0082] A third module 303, executed by the main processor 201, is responsible for estimating the path in the region of interest. The path is estimated as a function of time. To do this, the received context image is time-marked with a timestamp when it is received. These timestamps allow the time index n to be associated with the computed three-dimensional position. Thus, the time index n corresponding to time t corresponds to the three-dimensional position P n =( X n ,Y n ,Z n ) corresponds to a series of indexed positions P i This is how it is obtained. Route model P eyes (t) includes the following: - Amplitude parameter, phase parameter Φ0, and angular frequency parameter ω walkPeriodic time component P that depends on periodic (t). Frequency decomposition is equivalent to this angular frequency spectral decomposition. - Rate of change parameter drift average path and intercept parameter P average path Linear time component P that depends on linear (t). Parametric pathway model of the eye in three dimensions P eyes (t) is written as follows: P eyes (t)=P linear (t) + P periodic (t), here, P linear (t)=P average path (t), in this example, corresponds to the average position of the eyes excluding any vibrations over time: that is, the average position of the two eyes (in this case, in three dimensions) as a function of time excluding any vibrations, where, P average path (t)=P average path +drift average path *(t-t0); and, P periodic (t) = Amplitude * sin(Φ(t)), This corresponds to the oscillation of the average position of the eyes over time in this example: that is, the oscillation of the average position of two eyes as a function of time (in this case, in three dimensions), where, Φ(t) = Φ0 + ω walk *(t-t0).

[0083] To separate the linear component from the periodic component of the instantaneous observable position signal, a band-pass filter (which may vary depending on the dimension) is applied to the instantaneous observable position signal, which is constructed from the calculated eye position over time.

[0084] The extended Kalman filter is used to estimate this parametric path model. The extended Kalman filter is based on the following: -The first observer is an instantaneously observable position signal consisting of a series of computed eye positions.

number

number

number

number

number

number

number

number

[0085] For each new subject, their eye pathways P eyes (t) is estimated by a parameter model via the implementation of an extended Kalman filter on the parametric path model in each dimension, which includes an initialization phase of the extended Kalman filter, a prediction phase by the extended Kalman filter, and an update phase of the extended Kalman filter at the sampling frequency. The following diagram is used to illustrate in more detail how this path estimation process is carried out. In the case presented here, since the modeled path is the path of the midpoint between the two eyes, only a single eye path is modeled for each person, and the field of view of the capture camera allows both eyes to be acquired simultaneously, as the capture camera favorably includes two sensors with two lenses sharing the same motor; nevertheless, it will be possible to compute the path for each eye and update the path for each eye in parallel.

[0086] Advantageously, in cases where the acquired image contains multiple regions of interest, for example, when multiple subjects (each with one region of interest) are detected in the acquired image, the paths of the regions of interest (especially including filter updates) can be estimated in parallel, and in this case, images of each region of interest are probably captured sequentially (especially if there is only one capture camera 102).

[0087] The path estimation module 303 transmits a continuous estimation to the path tracking module 304, which is executed by the real-time coprocessor 202. The path tracking module can determine the location of the region of interest at any time based on the final path estimation transmitted by module 303, in order to ensure continuous optimized aiming. In practice, the frequency of the transmission of the continuous estimation is usually higher than the frequency of the capture of the context image. In this example, the path estimation is transmitted at a frequency of 15 Hz, and then, in the real-time coprocessor, the path tracking module is responsible for determining the location of the region of interest at a frequency of 1 kHz based on the final path estimation transmitted by module 303, which ensures continuity in directing the capture camera toward disengaging from the path, even if the capture frequency is 15 Hz or variable in this case. Specifically, if the captured image is suitable for biometric authentication of the subject, no further capture is necessary.

[0088] This ability to determine the location of the region of interest at any given time allows for the control of module 305, which controls the motor used to aim the capture camera 102. This is how the capture camera 102 continuously tracks the region of interest.

[0089] Module 306 also receives the determined position of the region of interest. This module 306 is responsible for calculating a focal factor that is added to the depth coordinate (Z) of the determined position (i.e., to the distance between the capture camera and the region of interest). The focal factor varies linearly, for example, between a minimum negative value and a maximum positive value. The distance between the capture camera and the determined position of the region of interest, corrected by the focal factor, allows module 307, which controls the motor used to focus the capture camera 102, to be controlled. Modules 306 and 307 are optional and are therefore depicted with dashed lines. Device 100 can operate without additional focusing, for example, if an EDOF (EDOF stands for Extended Depth of Focus) sensor is used, up to a few centimeters, provided that the depth of field of the capture camera lens is sufficiently wide compared to the error in position prediction.

[0090] Optionally, a third module 303 is responsible for additional estimation of the path in the region of interest. In the prediction phase with the extended Kalman filter, the extended Kalman filter is considered to have converged if its prediction error |Pk / k-1 -Pk / k| related to the calculation of the actual position falls below a convergence threshold (e.g., statistically or dynamically determined). Advantageously, additional estimation is used, especially during the initialization phase, unless the extended Kalman filter has converged. This additional estimation is usually performed via a linear approximation of the storage positions. In the simplest embodiment, only the coordinates of a straight line passing through the last two storage positions are estimated. In this case, a local straight-line path is estimated, which may suffice. In more complex embodiments, it is possible to compute a polynomial model of the path based on three or more positions. Such a polynomial model allows for obtaining a more precise path estimation than the local straight-line path. The path is estimated as a function of time. A series of indexed positions P i This allows for the calculation of velocity associated with a given location, provided that at least two locations are stored. This sequence may have indices for which no associated locations exist. This can be due to problems associated with transmitting context images, or, for example, the inability to recognize regions of interest in some context images. These missing locations must be taken into consideration when estimating the path, and especially in evaluating velocity. For example, local linear path estimation P based on two stored locations. n-1 =( X n-1 ,Y n-1 ,Z n-1 ) and P n =( X n ,Y n ,Z n Regarding this, the speed of the subject can be estimated at time t corresponding to the time index n as follows: V n =(VX n ,VY n ,VZ n )=(X n -X n-1 ,Y n -Y n-1 ,Z n -Zn-1 ). Position P n-1 is missing, position P can be obtained in a similar manner n-2 can be used, but at this time, in order to take the missing position into account, the obtained velocity needs to be divided by 2. By applying the velocity over the entire time period (T-t), it is possible to estimate the position at time T after the current time t from the current position at time t.

[0091] Figure 5 shows the principle of path estimation in an example embodiment of the present invention. For clarity, the present description focuses exclusively on path P in only the y-dimension, which characterizes the eye height of a subject that is most affected by periodic fluctuations eyes (t) focuses exclusively on parametric modeling of .

[0092] Eye dynamics is modeled using the system state equation by any of the following: [Formula] Via Euler discretization, k represents the following time index: [Formula] where, x k actual state u k input command, zero given that there is no command herein; w k covariance matrix Q k transient noise which is a centered Gaussian distribution of z k measurement v k covariance matrix R k measurement noise which is a centered Gaussian distribution of ; Alternatively, P eyes (t) is P average path (t)=P average path +drift average path *(t-t0) and P linear (t)=P periodicConsidering that (t) is constructed as Amplitude*sin(Φ(t)), here Φ(t)=Φ0+ω walk *(t-t0), in the case of variation in the eye position along the y-axis, the following equation can be written:

number

[0093] By applying the extended Kalman filter to the two-observer model, the state vector X(t) and the measurement vector Z(t) can be written as follows:

number

[0094] Module 303 for estimating the eye's path receives as input the three-dimensional position of the eye, calculated at the calculation frequency based on one or more images captured by the context camera system. Each received three-dimensional position is instantaneously observable position signal

number

[0095] The instantaneously observable position signal is its component along the Y-axis.

number

number

number

number

number

[0096] The initialization step 303b of the extended Kalman filter for the parametric path model includes determining the initial state vector X(t0) by determining the initialization parameters, which include the following: - Amplitude parameter; -Phase parameter Φ0; -Angular frequency parameter ω; - Rate of change parameter: drift; -Intercept parameter P0; The initialization parameters are determined from all or part of a series of calculated three-dimensional positions of the region of interest of the detected volume. Preferably, the initialization parameters forming the initial state vector X(t0) are determined from instantaneous observable position signals in such a way that part of this series covers the captured "1 / 2 first step" of the subject in the image acquired by the context camera system 105. To achieve this objective, part of the series of positions includes the calculated positions of the region of interest based on images continuously acquired by the context camera system, the first image being the image in which the region of interest is first detected, and subsequent images being images in which the region of interest is detected until at least two local extreme values ​​are detected.

[0097] This makes it possible to write the following equation for the extended Kalman filter: - Observation function:

number

number

number

[0098] At each current time step, a prediction stage 303c of the state X(t+1) in the subsequent time occurs, and similarly, (at the current time step

number

[0099] In other words, during the prediction phase (also called the convergence phase), the state estimated in the previous time is used to generate an estimate of the current state:

number

number

[0100] Constant frequency 1 / T s Regarding path estimation in this context, transient noise is the frequency (q drift or q angular frequency ) relies almost entirely on noise:

number

[0101] The initial covariance (uncertainty) matrix Q of the model's transient noise.

number

[0102] In prediction step 303c using the extended Kalman filter, if the prediction error of the extended Kalman filter for the calculation of the actual position |Pk / k-1 -Pk / k| is greater than the convergence threshold, the extended Kalman filter is not considered to have converged. This check also ensures that the filter does not diverge at each time step. The convergence threshold is determined favorably statistically or dynamically. - The convergence threshold is, for example, around 1 meter, or less than 1 meter, and especially 0.2 m. Independent of the filter convergence, the prediction stage 303c continues with the update stage 303d.

[0103] During the update phase, observations of the current time are used to refine the predicted state in order to obtain a more precise estimate:

number

[0104] Matrix P is updated in each computation cycle corresponding to a new location calculation and gives the confidence of the path. Preferably, in the case of loss (hidden) of the region of interest, the filter is not updated, and therefore when the region of interest is detected again during a subsequent acquisition, the filter is updated by a sampling period corresponding to the length of time between the time when the region of interest was last acquired (and effectively when its location was calculated) and the time when the region of interest is acquired again.

[0105] When the extended Kalman filter converges, the future time t tracking The path estimation in the region of interest is performed using the time t of the last acquisition, which is used to update the extended Kalman filter. measure positionBased on the acquisitions made up to this point, the tangent to the modeled path is calculated (i.e., including the region of interest), and in particular, the modeled path is written in the form of position pos and velocity spd as follows: pos(t tracking )=P average path (t measure position )+drift average path *( t tracking -t measure position ) + Amplitude * sin(Φ(t measure position )+ω*(t tracking -t measure position spd(t tracking )=drift average path +Amplitude*ω*cos(Φ(t measure position )+ω*(t tracking -t measure position .

[0106] t tracking If is selected to be the estimated time of receiving the path estimated by the real-time coprocessor 202, this allows for compensation of latency between real time and the time of the last acquisition used to update the extended Kalman filter.

[0107] A predetermined time t of reception by the route tracking module 304 tracking The y-dimensional path estimation of the eye in is obtained by calculating the tangent to the modeled path and is written as follows: pos(t) = pos(t tracking )+spd(t tracking )*(tt tracking )

[0108] The paths in the other two dimensions can be obtained, for example, by using the model described in relation to additional path estimation.

[0109] Therefore, at present time, the capture camera is at time t capture By applying the preceding formula to future time t capture The estimated location of the region of interest regarding pos(t) captureThe camera is commanded to orient itself in the coordinate direction of ), thereby enabling compensation for the response time of the aiming system 305 and / or focusing system 307 of the capture camera.

[0110] To obtain reliable path estimations, it is necessary to sample position calculations at a sufficient frequency. This frequency depends on the speed of the subject's motion and therefore on the intended application. For example, in the case of iris recognition of a walking person, a frequency of 15 context images per second and therefore 15 computational position estimates per second has been proven to be sufficient, and the same applies to path estimation.

[0111] Figure 6 shows an example of a data processing device 106 for carrying out one or more embodiments of the present invention. The data processing device 106 typically includes one or more central processing units (CPUs) 601 and / or one or more graphics processing units (GPUs) 605, a physical communications module (NET) 604, one or more physical input / output modules 607 for exchanging data with external devices (such as a context camera system and a capture camera), a temporary storage medium 602 such as random access memory (RAM), a non-temporary recording medium 603 (FLASH), and a communications bus (not shown) for transferring data between internal components of the data processing device 106.

[0112] The data processing device 106 may be used to execute one or more program modules 301, 302, 303, 304, 305, 306, 307, which contain instructions that cause the data processing device 106 to execute the method according to the present invention when a program module or group of modules is executed. The program module or group of modules may be written in any compiled or interpreted programming language. The program module or group of modules may form part of a software solution (i.e., a collection of executable instructions, code, scripts, etc., and / or databases).

[0113] The data processing device 106 includes the following elements that are connected to each other via a communication bus: - A central processing unit (CPU) 601 such as a microprocessor, and more particularly a central processing unit (CPU) 601 including an internal clock; -Temporary memory 602 for storing executable code for a method to carry out the present invention, and registers configured to record variables and parameters required to carry out a method according to an embodiment of the present invention, wherein the memory capacity of the data processing device is preferably supplemented by an optional random access memory 602 connected to an expansion port, for example; - A non-temporary memory 603 for storing a computer program and calibration data required to carry out embodiments of the present invention; the stored computer program particularly includes a computer program which includes instructions configured to carry out all or some of the steps of the method according to the present invention when the program is executed on the processing device 106, and the non-temporary memory 603 is an example of a (removable or non-removable) non-temporary information storage means; -A communication module 604 including a network interface 604 connected to a communication network on which digital data to be processed is transmitted or received; the network interface 604 may consist of a single network interface or a set of various network interfaces (e.g., wireless and wireless interfaces or various types of wired or wireless interfaces). Data packets, upon receipt under the control of a software application running on a processor 601, are transmitted to or read from the network interface for transmission. - A user interface (HMI), in particular a user interface (HMI) including a graphics processor 605 for receiving input from the user or for displaying information (in particular (visual and / or audio) guidance information) to the user; - Input / output module 607 for receiving data from / to external peripheral devices (especially hard disks and removable storage media).

[0114] The executable code may be stored in non-temporary memory 603 (e.g., flash memory or read-only memory) or on a removable digital medium such as a disk. In one variant, the program's executable code may be received via a communication network through a network interface 604 to be stored in one of the storage means of a data processing device 106, such as memory 603, before execution.

[0115] The central processing unit 601 is configured to control and guide the execution of instructions (stored in one of the aforementioned storage means, such as non-temporary memory 603) or segments of software code of a program or group of programs according to one embodiment of the present invention. After being turned on, the CPU 601 can execute instructions from temporary RAM memory 602 relating to a software application. When such software is executed by the processor 601, it enables the execution of the method according to the present invention.

[0116] In one embodiment, the device is a programmable device that uses software to carry out the present invention. In a modified form, the present invention may be carried out in hardware form (for example, in the form of an application-specific integrated circuit (ASIC) or a field-programmable gate array (FPGA)).

[0117] According to one embodiment, the data processing device 106 is exclusively hosted locally, or in a variant form located outside terminal 1, or even distributed, and includes a plurality of processing subunits (in particular, at least partially located externally and communicating with each other via the network interface 604). Similarly, depending particularly on the nature of the terminal, all or part of the memory may be physically distant and hosted, for example, on a remote server. [Explanation of Symbols]

[0118] 100 Image Acquisition Devices 101 Capture Volume 102 Capture Camera 103 Subject 104 Areas of Interest 105 Context Camera 106 Data Processing Devices 201 Processor 202 coprocessors 205 Motor 206 Aiming Motor 207 Lighting 209 images 210 images 211 Trigger 212 Trigger 213 Control 214 Control 215 Control 301 First Program Module 302 Second Program Module 303 Path Estimation Module 303a Application Stage 303b Initialization stage 303c Prediction stage 303d Update Stage 304 Route Tracking Module 305 Aiming System 306 Program Modules 307 Focus system 601 Central Processing Unit 602 Temporary storage media 603 Non-temporary recording media 604 Physical Communication Module 604 Network Interface 605 Graphics Processing Unit 607 Physical Input / Output Module

Claims

1. A method for capturing an image of a region of interest of a moving subject (103) within a capture volume, wherein the method is: - A calculation process repeated at the calculation frequency of the three-dimensional position of the region of interest, wherein the instantaneous observable position signal is obtained from a series of acquired images in which the region of interest is detected and localized. [Math 1] A computation process that delivers a series of three-dimensional positions of the region of interest on which it depends as output; - The aforementioned instantaneous observable position signal [Math 2] Periodic observable signals from [Math 3] Extraction process at the applicable frequency; - Three-dimensional path of the region of interest within the capture volume from the calculation position. [Math 4] The estimation process, wherein the path has a periodic time component (P periodic (t)) and the linear time component (P linear Estimation steps defined in at least one dimension by a parametric path model combining (t); - At least one future time t from the estimated three-dimensional path of the region of interest capture Steps for determining the location of the region of interest; - The time t of the image of the region of interest captured by the capture camera (102) capture The capture process in which the capture camera (102) is at time t capture A capture process that is oriented in the direction of the judgment position in the region of interest. A method that includes this.

2. The method according to claim 1, wherein the subject (103) is a person, and the region of interest is a part of the face of the moving subject (particularly the eyes, preferably the iris), or a visual representation such as a two-dimensional barcode attached to the subject.

3. The aforementioned path (P eyes The method according to claim 1 or 2, wherein the estimation step of (t)) includes performing an extended Kalman filter on the parametric path model in at least one dimension of the detection region of interest of the subject, the extended Kalman filter comprising an initialization step (303b), a prediction step by the extended Kalman filter (303c), and an update step of the extended Kalman filter at the sampling frequency (303d).

4. The aforementioned periodically observable signal [Math 5] The purpose of extracting the instantaneously observable position signal [Math 6] The method according to claim 1 or 2, comprising the step (303a) of applying a band-pass filter in at least one dimension to the application frequency.

5. The extended Kalman filter comprises a measurement vector consisting of at least two observers, one having the instantaneous observable position signal as a first observer and the other having the periodic observable signal as a second observer, which are composed of a series of calculated positions in the region of interest. [Numerical 7] The method according to claim 4, based on the present invention.

6. The state vector of the extended Kalman filter (X(t) = [Number 8] There are five states: at least one dimension (P average path (t)) the average position of the region of interest, the rate of change of the average position of the region of interest in at least one dimension [Number 9] , the amplitude of the vibration in the region of interest centered on the average position (Amplitude), the phase of the vibration centered on the average position (Φ(t)), and the angular frequency of the vibration at the average position in the region of interest. [Number 10] The method according to claim 3, including the method described in claim 3.

7. The initialization step of the filter (303b) involves determining the initialization parameters and then determining the initial state vector (X(t) 0 The initialization parameters include determining the following: - Amplitude parameter; -phase parameter (Φ 0 ); - Angular frequency parameter (ω); - Rate of change parameter (drift); - Intercept parameter (P 0 ); Includes, The method according to claim 3, wherein the initialization parameter is determined from all or part of the three-dimensional positions of the series of calculated regions of interest.

8. The method according to claim 7, wherein the portion of the series of positions includes positions calculated with respect to the region of interest based on sequentially acquired images, the first image being the image in which the region of interest is first detected, and the subsequent images being images in which the region of interest is detected until at least two local extreme values ​​are detected.

9. The covariance matrix of the process noise (Q) of the extended Kalman filter is the sampling period (T) of the extended Kalman filter. s ) and the rate of change of the mean position of the region of interest in at least one dimension. [Math 11] Variance of (q drift ), the oscillation of the region of interest centered on the average position of the region of interest [Number 12] The dispersion of the angular frequency (q angular frequency ), or the variance (q) of the amplitude (Amplitude) of the vibration of the region of interest centered on the average position of the region of interest. a The method according to claim 3, which depends on at least one of the following.

10. Any future time t tracking The estimation of the path in the region of interest is performed at any future time t. tracking The parameters are obtained by calculating the tangent to the modeling path, and the parameters arise from the update of the extended Kalman filter based on the image obtained prior to the estimation, in particular the modeling path is pos(t tracking )=P average path (t measure position )+drift average path *(t tracking -t measure position )+Amplitude*sin(Φ(t measure position )+ω*(t tracking -t measure position ) spd(t tracking )=drift average path +Amplitude*ω*cos(Φ(t measure position )+ω*(t tracking -t measure position ) It is written in terms of position pos and velocity spd, as shown above. Check measure position is the time when the final context image used to estimate the path was acquired, and the tangent to the parameter model is, pos(t)=pos(t tracking )+spd(t tracking )*(t-t tracking ) The method according to claim 3, as described above.

11. The method according to claim 1 or 2, wherein the capture camera (102) is directed in the direction of the position of the region of interest, which is determined based on the three-dimensional path of the region of interest at a frequency higher than the calculation frequency.

12. A computer program comprising instructions configured to perform each step of the process according to claim 1 or 2 when the program is executed on a computer.

13. Detachable or non-detachable information storage means that is partially or entirely readable by a computer or microprocessor, comprising code instructions for a computer program for performing each step of the steps of the method according to claim 1 or 2.

14. A device for capturing an image of a region of interest of a moving subject within a capture volume, wherein the device is - Context camera system (105); - Capture camera (102); - Includes a processor, the processor is - A calculation process repeated at the calculation frequency of the three-dimensional position of the region of interest, wherein an instantaneous observable position signal is obtained from a series of acquired images acquired by the context camera system in which the region of interest is detected and localized. [Number 13] A computation process that delivers as output the three-dimensional positions of a series of the region of interest on which it depends; - The aforementioned instantaneous observable position signal [Number 14] Periodic observable signals from [Number 15] Extraction process at the applicable frequency; - The three-dimensional path of the region of interest within the capture volume from the calculation position (P eyes (t)) estimation step, wherein the path is a periodic time component (P periodic (t)) and the linear time component (P linear Estimation steps defined in at least one dimension by a parametric path model combining (t); - At least one future time t from the estimated three-dimensional path of the region of interest capture Steps for determining the location of the region of interest; - The aforementioned time t capture At time t capture Capture process in A device configured for this purpose.

15. The device according to claim 14, wherein the capture camera (102) is rotatable to face the direction of the determination position of the region of interest in at least one dimension.