Method and system for estimating trajectories
The RGB-D camera and machine learning-based method effectively reduces false alarms and enhances collision prevention in construction zones by accurately predicting element trajectories and triggering alerts only when necessary.
Patent Information
- Application Number
- PCT/ES2025/070218
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-22
- Filing Date
- 2025-04-21
- Publication Date
- 2025-10-30
AI Technical Summary
Existing people detection systems in construction zones struggle with excessive false collision alarms and inadequate real-time monitoring of moving and stationary elements, failing to effectively predict trajectories and prevent collisions.
A method using an RGB-D camera and machine learning models, such as YOLO v5, to identify and predict the trajectories of potentially mobile elements, determine collision risks, and trigger alerts only when necessary, while adapting to varying lighting and visibility conditions.
Reduces false alarms and enhances real-time monitoring by accurately predicting element trajectories and preventing collisions, ensuring safety in construction zones with adaptive collision risk assessment.
Smart Images

Figure 00000022_0000 
Figure 00000023_0000 
Figure 00000024_0000
Abstract
Description
[0001] DESCRIPTION
[0002] METHOD AND SYSTEM FOR TRAJECTORY ESTIMATION
[0003] FIELD OF INVENTION
[0004] The present invention belongs to the field of monitoring people and objects using cameras and predicting their trajectories.
[0005] PRIOR STATE OF THE ART
[0006] People detection in construction zones using cameras typically involves systems that employ real-time image analysis algorithms to identify the presence of people within the monitored area. These people may be moving or stationary.
[0007] Furthermore, it is important to be able to monitor areas of interest, such as construction sites, people, and other objects present in the environment, like vehicles, machinery, or building materials. This helps ensure safety in the environment and prevent accidents.
[0008] People detection systems in construction areas can also be integrated with larger video surveillance systems, allowing real-time remote monitoring and automatic generation of alerts upon detection of suspicious activities or situations that pose a risk to safety in the area of interest.
[0009] In particular, the field of video surveillance is characterized by advanced systems that combine image analysis algorithms and the ability to integrate with existing surveillance systems to provide effective and secure surveillance in construction zones.
[0010] DESCRIPTION OF THE INVENTION
[0011] For all the cases presented above, a method for detecting elements in an area, such as a construction site, is highly advantageous in order to predict their trajectories, generate risk or danger alarms, and prevent collisions between moving elements and / or between moving and static elements. In a first aspect of the invention, a method for preventing collisions of people and / or objects is provided, comprising the steps of: capturing a plurality of images, said images formed by pixels, with an RGB-D camera attached to a first element; obtaining a distance to each pixel in the plurality of images; identifying one or more second elements in the plurality of images; and discriminating, from among the one or more second elements, whether said second elements are potentially mobile, based on the distance to each pixel in the plurality of images and the one or more potentially mobile second elements identified.predict a trajectory of at least one of the second potentially mobile elements identified; and determine a collision risk between said at least one of the second potentially mobile elements identified whose trajectory has been predicted and another element, where the other element is either said first element or a second element of the second elements identified in the plurality of images.
[0012] In embodiments of the invention, the one or more secondary subjects in the images are at least one person and / or at least one object. In equivalent embodiments of the invention, the plurality of images simultaneously captures a plurality of persons and / or a plurality of objects. The secondary subjects, at the time of photography, may be either in motion at a certain speed relative to the first subject or at rest. In a typical embodiment of the invention, the objects of the identified secondary subjects may comprise materials such as, for example, construction materials, vehicles, or elements characteristic of construction sites or worksites. The prevention method described in this disclosure is also compatible with secondary subjects other than those characteristic of construction sites or worksites.
[0013] In embodiments of the invention, the first element to which the RGB-D camera is attached is in relative motion with respect to at least one of the second identified elements whose trajectory has been predicted. Relative motion is understood to mean the movement of one element with respect to another. For example, the movement of the first element with respect to the second element, or the movement of the second element with respect to the second element.
[0014] In possible embodiments of the invention, the first element and the coupled camera are at rest, and one or more second elements are also at rest. In this case, the distance obtained by the camera to the pixels representing the second element is constant over time.
[0015] In possible embodiments of the invention, the first element and the attached camera are at rest, while one or more second elements are in motion. In this case, the distance obtained by the camera to the pixels representing the second element varies over time depending on the magnitude and direction of the velocity of the second element. In other possible embodiments of the invention, the first element and the attached camera are in motion, while one or more second elements are at rest. In this case, the distance obtained by the camera to the pixels representing the second element varies over time depending on the magnitude and direction of the velocity of the first element.
[0016] In possible embodiments of the invention, the first element and the attached camera are in motion, and one or more second elements are also in motion. In this case, the distance obtained by the camera to the pixels representing the second element varies over time depending on the magnitude and direction of the velocities.
[0017] Knowing the speed of the first element and the camera attached to it, it is possible to determine the absolute movement between the first and second elements, since the change in the distance obtained by the camera to the pixels representing the second element can be deduced by applying the concept of speed addition.
[0018] The state of rest of an element is defined as a state of velocity v = 0 with respect to the ground.
[0019] In preferred embodiments of the invention, a machine learning model identifies one or more second elements in a plurality of images. Furthermore, the machine learning model identifies the same one or more second elements in a series of successive images. That is, the machine learning model is capable of identifying a second element in a first image, possibly in an image with a plurality of second elements, and uniquely and unambiguously identifying or distinguishing the same second element in successive images, allowing the measurement of the evolution of the second element's position and velocity over time.
[0020] In preferred embodiments of the invention, the machine learning model discriminates whether these secondary elements are potentially mobile. That is, the model is trained with training data that allows it to distinguish secondary elements of interest from other objects that are either unidentifiable or identifiable but considered of no interest.
[0021] In preferred embodiments of the invention, the machine learning model can be a deep learning model. In particular, the deep learning model can be a YOLO (You Only Look Once) model, such as a YOLO v5 model.
[0022] The machine learning model can be trained with images that include the positions and displacements of secondary elements. Furthermore, these training images allow the model to discern and recognize secondary elements of interest. Specifically, the machine learning model can be trained with images that include the positions and displacements of secondary elements under different lighting and visibility conditions. For example, the training images include images taken during the day, at night, and in situations of low visibility due to dust, snow, fog, rain, or other factors that reduce visibility.
[0023] In a preferred embodiment of the invention, the trajectory prediction of the second elements, identified as potentially mobile, is calculated using one of the following methods: linear regression, polynomial regression, Kalman filter, or uniformly accelerated rectilinear motion. Other known methods for calculating the trajectories of bodies from known initial conditions may be used to predict the trajectory of the second elements.
[0024] The predicted trajectory of at least one second potentially moving element appearing in a first image is calculated using images taken prior to that first image. The initial conditions for calculating the predicted trajectory of an element identified in an image are the images immediately preceding the image in which the second element was identified. The preceding time period comprises at least the time required to acquire at least two images, preferably at least twelve images, or even more preferably at least thirty-six images.
[0025] The number of previous images used to predict the trajectory of the second element will ultimately depend on the number of frames per second at which the RGB-D camera operates. For example, using an RGB-D camera at 12 frames per second, 12 previous images correspond to the images taken in the second before the image on which the trajectory is calculated, and 36 images correspond to the images taken three seconds earlier.
[0026] The trajectory prediction of at least one second potentially moving element is calculated for a period of time subsequent to the first image. Using the trajectory calculation methods mentioned above, the trajectory of an element identified in an image is predicted for the few seconds following the first image. Preferably, this subsequent period of time is configurable. By configurable period, it is meant that this time can be reduced or increased depending on the specific needs of the trajectory prediction. For example, a subsequent period of time could be 1 second. A person skilled in the art will understand that the shorter the predicted time for the trajectory, the lower the uncertainty of the trajectory.
[0027] In an alternative implementation, the machine learning model associates a centroid with each second element. This measure saves computational resources when calculating the trajectories of the second elements. The centroid can be defined as a point within the identified second element, from which all prediction calculations are performed. For example, the centroid of an element could correspond to the pixel of the second element closest to the camera or to the pixel at the center of mass of the 2D image representing the second element. Therefore, the trajectory of each potentially moving second element is calculated by tracing the trajectory of the centroid of that second element.
[0028] In a preferred embodiment of the invention, the collision risk is assigned based on the predicted collision risk level between the trajectories of at least one of the first and / or second potentially moving elements identified. Collision risk is understood to mean that the predicted trajectory of a second element comes into direct contact with the first element, or that it approaches within a distance defined as the danger distance. A collision risk can be assigned to a situation in which two elements have predicted trajectories that come into contact or approach each other, increasing the risk of collision. Furthermore, in various scenarios, the trained machine learning module can consider a situation to be of low risk despite having a second element moving at close range to the first element.This case could be understood, for example, as a vehicle traveling parallel to the first element where the camera is located, since there would be no probability of collision.
[0029] If a predetermined trajectory poses a collision risk, a signal is triggered if the predicted trajectory of a component indicates a collision between a first and a second element. This signal can take various forms, such as acoustic, visual, tactile, digital, or virtual. Another type of alert issued is a direct, individual notification to an agent responsible for managing a second element at risk of collision. The signal can be collective, such as a siren sounding upon detecting a collision risk, or individual, for example, acting solely on the second element causing the risk.An example of signal transmission could be a bracelet worn by a person within the area monitored by the camera. This bracelet, connected via Bluetooth or another type of connectivity that allows for remote connection, could transmit the alarm generated by the system, particularly, for example, by the computing component. Transmission methods could include sound, vibration, or light.
[0030] Other objects present in each image, different from the potentially moving elements identified in the images (preferably by the Machine Learning model), are identified using a point cloud. Examples of elements represented by these pixels could include, but are not limited to, trees, signs, posts, or other features.
[0031] Furthermore, this unidentified point cloud comprises both static and moving points. Initially, the points in the point cloud are treated interchangeably. If movement is detected in the point cloud, it is associated with a second, static element, unidentified by the Machine Learning module, that has begun to move. These pixels are monitored by the Machine Learning module to determine if their movement could potentially cause a collision with the first element or another element. Specifically, a collision risk is assessed between these moving points in the point cloud and the first element.
[0032] In some implementations, the method uses multiple synchronized RGB-D cameras to capture images. These synchronized cameras are arranged to minimize blind spots in the monitored area. A blind spot is, for example, an area of the monitored region that is not within the camera's field of view. By strategically positioning multiple cameras at different angles, it is possible to overlap their individual fields of view, creating a collective field of view around the cameras where there are no blind spots.
[0033] In a second aspect of the invention, a system for detecting and tracking elements is provided, comprising an RGB-D camera, configured to obtain a plurality of images, coupled to a first element, and a computing unit; wherein the camera is configured to: capture a plurality of images, said images formed by pixels, and wherein the computing unit is configured to obtain a distance to each pixel of the plurality of images; identify one or more second elements in the plurality of images; discriminate, from the one or more second elements, whether said second elements are potentially mobile, from the distance to each pixel of the plurality of images and from the one or more potentially mobile second elements identified; and predict a trajectory of at least one of the potentially mobile second elements identified.and determine a collision risk between said at least one of the second potentially mobile elements identified whose trajectory has been predicted and another element, where the other element is either said first element or a second element of the second elements identified in the plurality of images.;
[0034] In preferred embodiments of the invention, the computing unit further comprises a machine learning system configured for detecting, predicting, and tracking the movement of objects in the vicinity of the camera. The computing unit is, for example, a PC with a GPU, capable of reading and executing programs. The computing unit is equivalent to a processor and may be integrated into a system that also includes the camera, or alternatively, form part of a separate system, communicating with the camera either wired or wirelessly. The computing unit may be, for example, located remotely or within a monitoring unit.
[0035] Specifically, the first element is a mobile vehicle or a stationary object in the vicinity of the first and / or second elements. For example, this mobile vehicle could be a construction machine to which the camera is attached, monitoring the machine's surroundings and identifying and distinguishing nearby elements to prevent hazardous situations due to proximity or collisions.
[0036] The system, or the system described in this disclosure in general, can be mounted on a plane at the same height as the elements it monitors. For example, it can be mounted on a vehicle, whether stationary or in motion, surrounding the elements and / or people being monitored.
[0037] In alternative embodiments, the device or system described in this disclosure may be mounted on a plane higher than the height of the objects it monitors. For example, it may be mounted on a vertical structure, such as a pole or column.
[0038] In alternative embodiments, the device or system described in this disclosure may be arranged on a plane lower than the height of the objects it monitors.
[0039] In a third aspect of the invention, a computer program product is provided comprising instructions such that, when the program is executed by at least one computing device, it causes at least one computing unit to carry out the steps of the method of the first aspect of the invention.
[0040] BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To complement the description and to aid in a better understanding of the characteristics of the invention, in accordance with some examples of practical embodiments of the invention, a set of figures is included as an integral part of the description, in which, for illustrative and non-limiting purposes, the following has been represented:
[0042] Figure 1 shows the position of a plurality of identified elements, their trajectory followed and the prediction of their trajectory, in addition to the risk and danger zones around the camera, according to embodiments of the invention.
[0043] Figure 2 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a low risk situation, according to embodiments of the invention.
[0044] Figure 3 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a medium risk situation, according to embodiments of the invention.
[0045] Figure 4 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a high risk situation, according to embodiments of the invention.
[0046] Figure 5 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a medium risk situation in which the identified object moves at high speed, according to embodiments of the invention.
[0047] Figure 6 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a high-risk situation in which the identified object moves at high speed, according to embodiments of the invention.
[0048] Figure 7 shows the position of an identified element, its trajectory followed and the prediction of its trajectory, and of the risk and danger zones around the camera, in a low risk situation in which the identified object has greatly reduced its speed, according to embodiments of the invention.
[0049] Figure 8 shows the steps of a method for detecting and tracking moving elements according to embodiments of the invention. DESCRIPTION OF A PREFERRED EMBODIMENT OF THE INVENTION
[0050] The description of the possible preferred embodiments of the invention requires providing numerous details to facilitate a better understanding of the invention. Even so, it will be apparent to someone skilled in the art that the invention can be implemented without these specific details. Furthermore, well-known features have not been described in detail to avoid unnecessarily complicating the description.
[0051] Figure 1 represents an embodiment of a possible typical trajectory prediction scene (100) according to the invention. In this example, the camera is mounted on or carried on the first element. In other embodiments, the camera may be mounted directly on the first element or may be positioned detached from the first element. In Figure 1, the first element (101) carries the camera, which is an RGB-D camera that obtains images of its surroundings.
[0052] An RGB-D camera is a camera that combines color image capture (RGB) with depth measurement capabilities (D). This allows it to capture the color information of a scene, as well as measure the distance between the camera and objects in the scene, providing three-dimensional information. Each pixel in the RGB image is assigned a measurement of the distance between the object represented by that pixel and the depth camera. Typically, the RGB sensor and the D sensor are positioned side by side, separated by the smallest possible distance to assign precise depth measurements to each pixel in the RGB image. In this case, intrinsic calibration between the RGB and D sensors is possible to obtain spatially accurate information from each sensor.Other RGB-D cameras include RGB image measurement and depth (D) measurement on the same sensor, thus associating a more accurate depth measurement by eliminating the parallax error that occurs when the RGB and D sensors are separated by a distance.
[0053] The images captured by the RGB-D camera are provided to the computing unit, specifically the Machine Learning module. A virtual hazard zone (102) and a risk zone (103) are defined in front of the camera sensor. An object is located at position (104), having moved from a certain position and traveled along a path (107). Using the Machine Learning module, the trajectory is predicted based on previously acquired images of the same object, and the path (106) is predicted up to point (105). Another object (108) is in the scene, having moved from a different position and heading towards position (109). A third object (110) is stationary in the scene. The camera (111) has a blind spot area on both sides of the sensor.
[0054] The camera's blind spot can be eliminated by deploying additional cameras. Each additional camera is positioned in a direction that minimizes the overall blind spot. These cameras can be deployed together and facing different directions, or separately and pointing in different directions. When cameras are deployed together and facing different directions, as many hazard zones are created as there are cameras deployed. These hazard zones can be combined into a single hazard zone. However, the created hazard zones may overlap between different cameras, generating a general hazard zone around the object being monitored. When cameras are deployed separately and pointing in different directions, as many hazard zones are created as there are cameras deployed.
[0055] RGB-D cameras typically use technologies such as structured infrared light or time-of-flight to measure depth. For a preferred embodiment of the present invention, any system that allows the calculation of the distance between the pixels that make up an image and the object they represent may be used.
[0056] Figure 2 shows a typical no-risk situation (200), with the camera (201), a danger zone in front (202), and a risk zone (203). An element (204), coming from (205) and with a predicted trajectory (206), moves without entering any zone. In this case, no alarm is triggered.
[0057] Figure 3 shows the same situation (300) as Figure 2, but with the object (301) having a predicted trajectory (303). Being within the risk zone, an alarm is triggered to prevent a collision. Figure 4 shows the same situation (400) as Figure 3, but with the object (401) having a predicted trajectory (403). Being within the danger zone, an alarm is triggered to prevent a collision.
[0058] The capabilities described in this disclosure also allow for the adaptation of monitoring people and / or objects in environments with adverse visibility conditions. The Machine Learning module can adapt to different lighting conditions and adverse environments, ensuring accurate and reliable detection in diverse situations.
[0059] The predicted trajectory of a second element depends on its speed relative to the first element. In particular, an element approaching at high speed will result in a more reactive trajectory prediction and potentially generate a greater risk of collision. Figure 5 shows an example of a person running in front of the first element carrying the camera. The figure shows a second element (501), coming from point (502) and having traveled along trajectory (503) (used for the subsequent prediction). Its predicted trajectory (504) would enter the risk zone (203) and end at point (505). In this case, the trajectory prediction is reactive to the element's speed, and the camera would emit a danger signal to the operator of the machine or the person walking.
[0060] Figure 6 shows the continuation of Figure 5, where a second element (601), originating from point (602) and having traveled along path (603) (and used for the subsequent prediction), has a predicted path (604), which would enter the danger zone (202) ending at point (605). In this case, the predicted path is reactive to the element's speed, and the system would issue a collision warning signal to the operator in charge of the machine or to a person on foot.
[0061] Finally, Figure 7 shows the same second element (701) from Figures 5 and 6 and the evolution of its predicted trajectory after sudden braking (704). Because the speed has decreased so considerably, the predicted trajectory (704) is no longer within the risk area (203), becoming a non-immediate risk and making the issuance of the risk alert unnecessary. To assign a risk level, the predicted distance between the two elements or the relative speed between the objects is taken into account.
[0062] The hazard zones defined in figures (203) and (202) are for illustrative purposes only. The shapes may change and be adapted to the specific needs of the environment and / or the elements to be identified. Generally, the risk zone (103) includes the hazard zone (102).
[0063] When calculating the trajectory of detected objects, the relative movement of the camera with respect to the other detected elements is taken into account. This information prevents false alarms from triggering for objects that appear to be moving, but in reality, it is the camera itself that is in motion. In the case of pixels detected as a point cloud, the alarm is only triggered if the predicted trajectory collides with the first element the camera is positioned over, specifically when the camera is directly above that element.
[0064] Some non-limiting examples of potentially movable secondary elements include: people, bicycles, scooters, wheelbarrows, tow trucks, stationary vehicles, moving vehicles, vibrating vehicles, part of a work team wearing PPE and another part of the team without PPE. Other examples are scenes in a work situation with people in different positions (standing, squatting, kneeling, etc.), situations with one or more people lying down, workers using work tools (shovels, picks, etc.), workers transporting different materials (construction bags, profiles, etc.), workers at different distances (less than 3.5m, between 3.5 and 8m, more than 8m), operators moving around the vehicle approaching and moving away from the machinery, workers moving around the vehicle moving parallel, usual working situation in the rain, usual working situation with the asphalt giving off vapors, workers wearing clothing similar to the environment (clothing and background of the same color), operators in front of and behind a guardrail or on top of the vehicle, machinery on a banked surface or on a ramp with part of the work equipment, vehicles that could enter a construction site (cars, machinery, scooters, bicycles, ...) or any type of animals.
[0065] In a preferred embodiment of the invention, the YOLO network identifies elements and calculates their trajectory. The remaining unidentified pixels are treated as point clouds (if other objects are present in the scene), and their distance from the camera is evaluated to detect potential collisions. If no objects are present in the scene, no point cloud is defined. In that case, a hazard or danger alarm would be triggered. It is not necessary to calculate the trajectory of the point cloud; only its distance from the camera needs to be monitored. This solves the problem of unintentionally triggering false collision alarms due to stationary objects. The ground is considered a detected point cloud, which can be removed from the calculations, for example, using segmentation techniques.
[0066] Preferably, two models based on Y0L0v5-m are used in the proposed solution. One is for detecting people and vehicles, and the other is for face detection to determine if people are looking towards the machinery. The YOLOv5 structure consists of:
[0067] • Backbone: A CSPDarknet53 backbone containing several convolutional layers and residual modules.
[0068] • Neck: The neck uses PANet (Path Aggregation Network) to improve feature propagation at multiple scale levels, essential for detecting objects of different sizes.
[0069] • Head: The head of the model, where the final predictions are made,
[0070] The logic behind the YOLO network application can change the predicted detection time for a detected person depending on their level of awareness of their surroundings. For example, it detects if a person is using a mobile phone, has their back turned, is lying down, is facing forward, or is in a place with reduced visibility. It is important to note that the alarm level is self-adjusting depending on the level of awareness of the people in the situation or other factors that affect the reaction time to the trajectories of objects and / or people.
[0071] The proposed invention aims to solve the problem of excessive false collision alarms by emitting an alarm signal only when necessary. This problem exists, for example, in construction sites or work environments where the proximity of personnel to machinery is practically unavoidable. It is essential to be able to discern when a collision is imminent. To achieve this, AI is used to discriminate when a collision might actually occur, anticipate the movement of people or objects, and determine whether a dangerous interaction is truly likely.
[0072] The signals emitted can be sent only to the first element at risk of collision, to the general system, or to subsequent elements. All alarms can also be saved for later review.
[0073] The height level at which the camera is located can be raised or lower than the rest of the elements, to monitor the trajectories while interfering with them as little as possible.
[0074] The camera can also detect clearance range collisions, in which the machine may collide with objects at a different height.
[0075] The trajectory of the elements is calculated taking their dimensions into account. For example, a first element can change dimensions, such as an excavator with its arm retracted or extended, and this will be taken into account when calculating the collision.
[0076] The risk of collision is considered nonexistent for detected objects that are not potentially mobile. In alternative situations, such potentially non-mobile objects may generate a risk of collision if the trajectory of the first element carrying the camera system has a collision path with these non-mobile objects.
[0077] The collision can be between a person and the first element (in this case it could be a vehicle carrying the camera), between two people, between a vehicle and the first element (in this case it could be a vehicle carrying the camera), between two vehicles, between a person and a vehicle, or any combination of elements identified by the camera.
[0078] The Deep Learning model detects people and vehicles in 2D images. Using this information, along with object tracking, depth information from the RGBD image, and the application of Kalman filters and linear regression, it can predict the future location of a person or vehicle. The preferred embodiment of the invention allows the method and system to detect and prevent collisions between the first element (which may include the camera system) and second elements (moving objects), the camera, and stationary objects that suddenly move toward the camera, as well as collisions between second elements.
Claims
Claims 1. Method (800) for the prevention of collisions of people and / or objects, comprising the steps of: • capture (801) a plurality of images, said images being formed by pixels, with an RGB-D camera coupled to a first element, • obtain (802) a distance to each pixel of the plurality of images, • identify (803) one or more second elements in the plurality of images, • discriminate (804), of the one or more second elements, whether said second elements are potentially mobile, • from the distance to each pixel of the plurality of images and from the one or more second potentially moving elements identified, predict (805) a trajectory of at least one of the second potentially moving elements identified, and • determine (806) a collision risk between said at least one of the second potentially moving elements identified whose trajectory has been predicted and another element, wherein the other element is either said first element or a second element of the second elements identified in the plurality of images.
2. The method according to claim 1, wherein the second one or more elements of the images are at least one person and / or at least one object.
3. The method in accordance with any of the preceding claims, wherein the first element is in relative motion with respect to at least one of the second identified elements whose trajectory has been predicted.
4. The method in accordance with any of the above claims, wherein a Machine Learning model identifies the one or more second elements in the plurality of images.
5. The method in accordance with any of the preceding claims, wherein The Machine Learning model identifies the same one or more elements in a set of successive images.
6. The method in accordance with any of the above claims, wherein the Machine Learning model discriminates whether said second elements are potentially mobile.
7. The method according to claim 6, wherein the Machine Learning model is a Deep Learning model.
8. The method in accordance with claim 7, wherein the Deep Learning model is a YOLO model.
9. The method according to claims 1-6, wherein the Machine Learning model has been trained with images comprising positions of second elements and their displacements.
10. The method according to claim 9, wherein the Machine Learning model has been trained with images comprising positions of second elements and their displacements under different lighting and visibility conditions.
11. The method in accordance with any of the preceding claims, wherein the trajectory prediction is calculated by applying one of the following methods: linear regression, polynomial regression, Kalman filter, or uniformly accelerated rectilinear motion.
12. The method according to any of the preceding claims, wherein the prediction of the trajectory of at least one second potentially moving element appearing in a first image is calculated using images taken in a period prior to that first image.
13. The method according to claim 12, wherein the above period comprises at least the time required to acquire at least 2 images, preferably at least 12 images, or more preferably at least 36 images. images.
14. The method according to any of the preceding claims, wherein the prediction of the trajectory of at least one second potentially movable element is calculated for a later period of time.
15. The method according to claim 14, wherein the subsequent time period is configurable.
16. The method in accordance with any of the above claims, wherein the Machine Learning model associates a centroid to every second element.
17. The method according to claim 16, wherein the calculation of the trajectory of each second potentially movable element is performed by calculating the trajectory of the centroid of each second element.
18. The method in accordance with any of the preceding claims, wherein the collision risk is assigned depending on the predicted collision risk level between the trajectories of at least one of the first and / or second potentially mobile elements identified.
19. The method according to any of the preceding claims, wherein the risk of collision triggers the emission of a signal if the prediction of the trajectory of a component predicts a collision between a first element and a second element.
20. The method in accordance with any of the preceding claims, wherein the emission of the signal generates an acoustic, visual, tactile, digital or virtual signal.
21. The method in accordance with any of the preceding claims, wherein the unidentified pixels of the plurality of images form a cloud of unidentified points.
22. The method according to any of the preceding claims, wherein said unidentified point cloud comprises static points and points in motion.
23. The method in accordance with any of the preceding claims, wherein a collision risk is determined between said moving points of the point cloud and the first element.
24. System for the detection and tracking of elements, comprising • an RGB-D camera, configured to obtain a plurality of images, coupled to a first element, and • a computing unit, wherein the camera is configured to capture a plurality of images, said images being formed by pixels, and wherein the computing unit is configured to • obtain a distance to each pixel of the plurality of images, • Identify one or more secondary elements in the plurality of images, • to discriminate, from one or more second elements, whether said second elements are potentially mobile, • from the distance to each pixel of the plurality of images and the one or more second potentially moving elements identified, predict a trajectory of at least one of the second potentially moving elements identified, and • determine a collision risk between said at least one of the second potentially moving elements identified whose trajectory has been predicted and another element, where the other element is either said first element or a second element of the second elements identified in the plurality of images.
25. The system according to claim 24, wherein the computing unit further comprises a Machine Learning system configured for the detection and prediction of the movement of elements in the vicinity of the camera.
26. The system according to claims 24 and 25, wherein the first element is a mobile vehicle or a stationary object in the vicinity of the first and / or second elements. l. The system according to claims 24 - 26, wherein the system is at the same height level as the first and / or second elements.
28. The system according to claims 24-27, wherein the system is located at a higher level with respect to the first and / or second elements.
29. The system according to claims 24-27, wherein the system is located at a lower level with respect to the first and / or second elements.
30. A computer program product comprising instructions such that, when the program is executed by at least one computer system, it causes at least one computing unit to carry out the steps of the method of any of claims 1-23.