Method and system for determining 3D location of objects detected in video

By optimizing image parameters and employing a two-point calibration method, the system achieves precise 3D location determination of objects in video streams, addressing the inaccuracies and inefficiencies of existing methods.

WO2025133642A1PCT designated stage expired Publication Date: 2025-06-263VISIOND SIA +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/HR2023/000013
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing methods for determining the 3D location of objects in video streams are inaccurate and time-consuming, particularly when dealing with objects moving in 3D space or requiring geo-coordinate precision.

Method used

The method involves creating an undistorted image by optimizing camera, image, and lens parameters, followed by a unique two-point calibration process that determines the camera's height and geo-location, as well as the height and location of the camera's field of view center, allowing for precise geo-coordinate calculation.

Benefits of technology

This approach enables high-precision and accurate determination of object locations in both indoor and outdoor settings, even for objects moving in 3D space, by transforming 2D camera images into accurate 3D real-space coordinates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure HR2023000013_26062025_PF_FP_ABST
    Figure HR2023000013_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The method and system for determining the location of an object in real world detected inside the image based on it's location in pixels by using image calibration, camera calibration and map calibration is disclosed, where camera calibration implies the definition of two calibration points in real space - the camera location and the center of the camera's field of view
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Method and system for determining 3D location of objects detected in video

[0002] The subject invention relates to methods and systems for determining the exact location of objects in the real world, which are recognized in a video stream based on camera calibration and map / image calibration. If the object is recognized in the outdoor space, then its location should be expressed in geo coordinates (longitudinal and latitudinal angle), while for objects that are recognized in indoor spaces (buildings, warehouses,...) their location should be expressed in x,y coordinates in meters.

[0003] Today's applications that use this technical solution are applications that are monitoring objects on a flat surface, for example, monitoring and managing traffic and parking lots or monitoring people inside one floor of a shopping center.

[0004] Another technical problem that the present invention solves is the tracking of objects that move in 3D space, whether indoor or outdoor. The solution to the mentioned technical problem requires that, in addition to the location of the object, the height at which it is located is also detected. In other words, it is necessary to determine the difference between the height of the camera and the current height at which the detected object is located.

[0005] The solution to the mentioned technical problems is applicable regardless of the type of camera whose video recording is analyzed, which in itself represents the next (additonal) technical problem that is solved by the subject invention.

[0006] Following approaches have been taken when solving the mentioned technical problems:

[0007] 1. camera calibration using a chessboard pattern or similar patterns of regular geometric shapes;

[0008] 2. calibration of systems using cameras by using a moving object whose location is known;

[0009] 3. creation and use of detailed 3D maps within which the system, as a rule, the autonomous driving system searches for landmarks in order to determine its own position or geo location based on the mutual distances between the landmarks;

[0010] 4. using principles of artificial intelligence algorithms.

[0011] Document US2015154753 uses a moving object that has a chessboard pattern on it and that transmits its current geo-location. The vision system in one or more cameras recognizes this pattern and associates the current position of the object withing the image with its current geolocation.

[0012] Document US2022108460 uses specific markers, for example, a person or more wearing clothes of a certain color and transmitting their current geo-location. In this case too, the vision system connects the current position of the marker within the image with the geolocation. Document US11,625,86OB1 also uses a chessboard pattern to calibrate the camera and an additional two patterns to connect the pixels to the measured geo coordinates of the comer points on the pattern.

[0013] The disadvantage of such approaches is that the accuracy of the system depends on the number of points taken as a sample. Greater accuracy requires a greater number of points, i.e. a more complicated and time-consuming system calibration.

[0014] The invention in question uses two reference points and at the same time very precisely and accurately determines the location or geo-coordinates of the object / point for which there is an interest in determining the geo-coordinates.

[0015] Documents WO2019152662, US2023222681 use landmarks with known locations to improve the accuracy of determining their current location based on the arrangement and mutual distances of these landmarks that they recognize in the video stream. In this case, the vision system is used to detect and extract feature points or landmarks from existing digital maps and compare them with the same ones recognized in the video stream. If digital 3D maps do not exist, patent WO2019152662 describes an example of sending a vehicle that contains all the sensors needed to create a 3D map to the required area and thus creates a 3D map of the outdoor space.

[0016] The disadvantage of this approach is the need for continuous creation of 3D maps of the outdoor space because it is constantly changing. The construction of new buildings, parked vehicles, etc. are elements that change the characteristics of the outdoor space within which an autonomous delivery vehicle or drone needs to determine its current location. These methods also need camera calibration and bring all the problems addressed by the patent group from the previously stated patent group.

[0017] The last group of documents (US2023060211 , WO2022238933, US 10,580,164 B2) uses neural networks to calibrate the camera so that the camera can be used in vision systems that place the detected object within the video stream in a real 3D world. Generally speaking, the methods used by these systems are methods that are already known, but given the innovations that Al technology brings, they are used in a new way.

[0018] For example, the intrinsic and extrinsic parameters of the video camera as well as the location of the detected object are calculated by the patent (US 10,580,164 B2) through neural networks using the principle of recognizing the object and its characteristic points. This patent uses the example of vehicles and their unique mutual distances between their characteristic points to, once it recognizes the type of vehicle, based on these distances, intrinsic and extrinsic parameters of the camera are calculated.

[0019] Document US2021019914 uses neural networks to recognize people within video stream. It places a "Bounding box" around each detected person and uses the bottom of the bounding box as a reference for the ground under the assumption that all persons are touching the ground with their feet. By following a large number of people, i.e. based on a large sample, the neural network reconstructs the parameters and position of the camera.

[0020] The disadvantage of methods using neural networks lies in the fact that each network must have a reference model by which it calculates camera parameters, and that by changing a camera parameter such as the camera's field of view or mounting it to a different location, the process of calibrating the camera parameters needs to be repeated from the beginning. After defining the camera parameters, these systems must use the already described calibration methods to determine location or geo coordinates of the object detected in video stream.

[0021] The subject invention solves the mentioned problems using a method for determining the location of an object based on the position where that object is located within the digital camera image, which contains the following steps: a. creating an image without distortion by optimizing camera parameters, image parameters and lens parameters, which is known from G. Bradsky et al: Learning open CV, O'Reilly, 2008; (image correction is partially described in the mentioned book, while additional optimizations used in this invention are described later in the text). b. of camera calibration, which represents a unique solution that is disclosed by the patent application in question; c. image calibration, which after calibration is used as a geogrephic map, which also represents a unique solution that enables high precision and accuracy in determining the exact position of the object / point shown in the camera image.

[0022] The closest state of the art is represented by a document IWA: Calibration & Geolocation with FW 8.471 FW 8.80 and CM 7.70“, Whitepaper, 2023.

[0023] As is well known, camera calibration is necessary in order to place the objects recognized within the 2D image in the real 3D space.

[0024] The closest state of the art defines the parameters that are needed and for which calibration is done:

[0025] • horizontal and vertical camera field of view;

[0026] • pan tilt and roll camera angles;

[0027] • camera height;

[0028] • latitude, longitude, azimuth;

[0029] • determining the position of reference points in the real world.

[0030] These parameters are needed to mathematically calculate the coordinates of the distance of the point in the image from the reference point.

[0031] Camera calibration is the process by which we determine the values of these parameters for a particular camera.

[0032] The first, prima facie, difference between the closest state of the art and the subject invention lies in the fact that calibration according to the closest state of the art does not allow importing a 3D map or a 3D model of the space, i.e. Bosch's calibration method, and therefore the system is not suitable for objects that are not on one plane (flat ground plane), that is, which move along a three-dimensional path, for example, on the stairs.

[0033] In the subject invention (according to the present invention) there are two calibration points:

[0034] 1. camera height and camera mounting location as the first calibration point;

[0035] 2. height and position and / or geo location of the center of the camera field of view as a second calibration point; where the second calibration point is a unique solution brought by this invention in relation to the state of the art and especially to the closest state of the art which enables the 2D camera image to be transformed / converted into 3D real space with the help of the first calibration point. After calibration, i.e. converting the field of view into an existing 3D space, anyone who is a skilled professional can calculate an equivalent point in real space for each point in the camera's field of view.

[0036] Therefore, the position and / or geolocation of the center of the camera's field of view is a new concept used by the subject invention, and it represents the actual point in space to which the center of the video camera's image looks. Through the relationship between the position of this reference point and the position of the height and the location of the camera installation, the pixels in the image are connected to the real space coordinates.

[0037] Furthermore, in order to connect the pixels to the real space coordinates, the pixels in the image must not be distorted. All known calibration methods take the influence of camera parameters such as lens distortion and eliminate it using known methods. However, none of the methods covers the fact that the digital image is influenced by a number of other factors that can introduce distortion into the image. For example, changing the aspect ratio changes the angle of visibility, image resize changes the relationships in the image. Any change in the relationship in the image results in a change in the function in the relationship between the 2D image and the real 3D space, therefore the calibration according to the subject invention also takes into account other elements that can introduce distortion into the image, especially changes in the relationship in the image (e.g. Horizontal and vertical camera field of view transfer the world in front of the camera into an image with horizontal and vertical resolution. If the horizontal resolution of the image is reduced, we have disturbed the relationship between the total number of horizontal pixels and the total horizontal field of view.)

[0038] As one of the novelties, the concept of image calibration is introduced, as a process that is done before calibrating the camera. Through the image calibration process, it is adjusted to the camera parameters and as such represents a real digital twin of the space in front of the camera. The image calibration process is not part of this patent application. Before calibrating the camera, it is necessary to calibrate the image.

[0039] So, as a novelty, the concept of image calibration is introduced, as a process that is done before calibrating the camera. Through the process of image calibration, the image itself is adjusted to the parameters of the camera and as such represents a real digital twin of the space in front of the camera. The image calibration process is not part of this patent application. For example, it is a well-known fact that changing the aspect ratio also changes the angle of visibility of the camera. In this case, from Fig. 1A to Fig. 1C, the horizontal field of view of the camera is changed in such a way that on a Volvo car it is clearly visible where the individual radio aspect cuts the image.

[0040] These images show an example when the horizontal field of view is not changed as a camera parameter, but rather its aspect ratio is changed as an image parameter. Changing the aspect ratio of the image changes the camera horizontal field of view. Furthermore, the image from the camera is usually displayed inside a player. The dimensions of the player and the dimensions of the image are usually not the same, therefore the video player adapts the image to its dimensions by introducing black bars as filling for areas that do not belong to the image. It is important to note here that the origin of the video player 201 and the origin of the image 202 are different. This difference results in different point in pixels where image starts compared with the pixels where video player starts.This has several repercussions for vision systems that process the image, we will mention only one here. In the left figure (Fig. 2a), the field of view of the camera defined by the vertical angle of the camera is shown in the figure as the vertical resolution of the image. The vertical resolution of the player displaying that image is not the same as the vertical resolution of the image, therefore the location of the car recognized in the image of the video player is not the actual location of the car in the camera image and should be recalculated taking into account the thickness of the black bar. The image obtained from such a player needs to be cut (adjusted) to its actual dimensions. Display scaling is also a factor to consider.

[0041] A display with a horizontal resolution of 1920 pixels with a display scaling of 100% shows an image with a horizontal resolution of 1920 pixels without its distortion. Display scaling of 150% reduces the display resolution to 1950 / 1.5=1280, so in this case the image with a horizontal resolution of 1920 is reformatted to a horizontal resolution of 1280. The distortion introduced by display scaling in this case is the reduction of the display of the horizontal field of view from 1920 to 1280 pixels.

[0042] Changing the aspect ratio, displaying video in a letterbox and display scaling are just some of the indicated image deformations that the vision system needs to correct. Image deformations that are corrected by the subject invention are not the exhaustion of the aforementioned deformations.

[0043] In any case, before calibrating the camera, it is necessary to calibrate the image.

[0044] Bosch has three camera calibration methods

[0045] Autocalibration

[0046] Autokalibracija je ogranicena samo na automobile i prostore unutar kojih se krecu automobili. Npr, nadzor automobila na raskrscima i cestama i to je njeni osnovni nedostatak. Nairne, nadzorom raskrsca Al zna da se vozila moraju kretati pravocrtno ili pod kutom od 90 stupnjeva. Pracenjem lokacija velikog uzorka automobila u slici, slika se iskrivljuje kako bi se prilagodila pravocrtnom kretanju automobila. Nema auta - nema ove kalibracije.

[0047] Autocalibration is limited only to cars and spaces within which cars move. For example, monitoring cars at intersections and roads, and that is its main drawback. Namely, by monitoring the intersection, Al knows that vehicles must move in a straight line or at an angle of 90 degrees. By tracking the locations of a large sample of cars in the image, the image is warped to accommodate the straight-line motion of the cars. Drawback of this aproach is no car - no autocalibration.

[0048] Other disadvantages are the need to use Al technology to determine calibration parameters. Furthermore, auto-calibration is limited only to cameras that support this functionality.

[0049] After autocalibration, Bosch recommends that a person check the autocalibration results, which in this case raises the question of the very purpose of autocalibration. What are the benefits of powerful Al scaling when a person needs to measure and control after it?

[0050] The output of this system is neither geolocation nor the display of recognized cars on the map. Autocalibration was made in order for the system to determine the horizontal and vertical camera field of view, as well as the pan, tilt androll angle. Therefore, in order to determine the object's geo-location, another calibration using a map (another bosch method) is required. So, in the case of the closest state of the art, we have autocalibration, human verification and additional calibration with the help of a map.

[0051] Calibration using map

[0052] Here, the calibration is based on markers, as in those patents that used feature points. They calculate the geolocation by linear interpolation between the current position of the tracked object and the known geolocations of the markers.

[0053] The subject invention, in contrast to the above, determines geolocation by exact calculation based on camera, image and lens parameters.

[0054] This calibration method also suffers from the phenomenon of lens distortion. The more distorted the image, the more markers are needed thus making the calibration process more complex.

[0055] This type of calibration also has limitations because it needs a "sufficient" number of markers, so its use is limited by the number of markers. If the system using this calibration method monitors a space with a small number of markers, then they need to be measured and the system must be aligned with the markers, which then makes this calibration complicated and "time consuming"

[0056] A big disadvantage is the need for "alignment" of the map and the field of view of the camera. The map must be placed in the "viewing" direction of the camera. The subject invention does not need this because it uses two points so that the map can be placed in any direction, which is explained later in the text.

[0057] There is also a verification problem here. After placing the markers and everything, time is spent for system verification.

[0058] According to the subject invention, the method for determining the location of an object in the real world based on the location where that object is located within a digital camera image consists of the following steps: a. creating an image without distortion by optimizing camera parameters, image parameters and lens parameters; b. camera calibration (901) or (904) in such a way as to determine the height of the camera and the geo-location of the camera mounting location as the first calibration point, and then to determine the height and actual location of the center of the camera's field of view (400) as the second calibration point, and to the first and second calibration points are entered into the vision system (900) which, based on this data and the 3D map of the space, defines the 3D space within which the camera calibration points are located; c. connecting pixels within the image with coordinates in the real world in such a way that each pixel in the image is assigned a unique location in real space.

[0059] The term "actual location" is defined either by a 3D Cartesian coordinate system or expressed in terms of latitude, longitude and azimuth. In this connection, the terms "actual location" and "geo-location" are often used interchangeably in the text. The term "actual location" can also be replaced by the term "location" if the term "location" is accompanied by the feature "in the real world" or an equivalent feature that unequivocally indicates that the term "location" is used in the meaning of the term "actual location".

[0060] At the same time, the height of the geo location of the camera mounting location is different from the height of the geo location of the center of the camera's field of view, which enables tracking the location of an object that is located or moves up or down in relation to the plane of the camera mounting height at the location where it is monitored. 3D maps are generated from the coordinates x and y of the distance of the object from the camera with the help of existing 3D maps or Google API services that give us the height for the requested location of the object. The subject invention (subject matter of invention) in one embodiment includes a method of determining the actual location of objects based on their position within the image of a digital camera, and the method itself includes the following steps

[0061] 1. creating an image without distortion by optimizing camera parameters, image parameters and lens parameters;

[0062] 2. Measurements of the actual location of the camera 901 or the camera 904 from the map as well as the measurements of the height of the camera 901 or the camera 904 relative to the ground, wherein the actual location of the camera 901 or the camera 904 and the height of the camera 901 or the camera 904 relative to the ground represent the first calibration points of the camera 901 or cameras 904;

[0063] 3. Measurements of the actual location of the center of the field of view of the camera 400 and its height relative to the ground, where the actual location of the center of the field of view of the camera 400 and its height relative to the ground represent the second calibration point of the camera 901 or the camera 904;

[0064] 4. Connections of the actual location of the calibration points from 1. and 2. with their coordinates in pixels on the image that will be calibrated as a map;

[0065] 5. Representations of all calibrated cameras 901 or cameras 904 that have an actual location within the boundaries defined by the calibration of the image depicting a specific geographic area;

[0066] 6. Determination of the horizontal and vertical angle of the detected object in the image of one of the cameras 901 or the camera 904 in relation to the center of the field of view of the camera 400 from its location expressed in pixels;

[0067] 7. Determining the distance along the horizontal and vertical axis of the image from camera 901 or camera 904 from the height difference and mutual distance of calibration points;

[0068] 8. Determination of the actual location of the detected object from the distance from point 7 in relation to the actual locations of the calibration points;

[0069] 9. Determining the coordinates of the detected object in pixels of the image used as a map from the pixel-to-image ratio and camera calibration points 301 and 302 for the actual coordinates of the detected object;

[0070] 10. Repeats the process for each object detected by either camera 901 or 904.

[0071] In the rest ofthe text, the term "camera 901 or camera 904" will be replaced by the term "camera 901 or 904" with the same meaning.

[0072] In step 10, the plural of the word camera is used, because several cameras can be connected in one vision system by an expert in the field in already known ways, where it does not matter which camera provides the image through which the exact real location of the object is known, if calibration has been made for all cameras and for all areas covered by cameras pictures and cameras according to the subject invention. It is also not important whether all or individual cameras have the ability to process images, but it depends on the needs and requirements of the user, because image processing can be done, apart from the camera, with edge devices, servers or cloud applications. The term "actual coordinates" used in the text refers either to geo-coordinates in the case of external spaces or to Cartesian coordinates in the case of internal spaces. Likewise, the term actual location used in the text refers to the geo-location in the case of outdoor locations or the location expressed in Cartesian coordinates in the case of indoor locations.

[0073] The map from points 2 and 9 refers to a geographic map, or to maps of the Google earth pro or Open street maps type.

[0074] The image from points 4 and 5 refers to the image of a certain geographical area that is shown on the display. The display is defined by the resolution in pixels, therefore the image of the geographical area, which is also defined by pixels, needs to be connected with geo coordinates in order to be able to place camera calibration points expressed in geo coordinates within that image

[0075] The invention in question in its favorable variant implies determining the actual location of objects based on their position within the digital camera image, where the objects are not on the same plane as the camera's height plane, i.e. where the objects move up or down in relation to the camera's height plane.

[0076] Map is a graphic representation, drawn to scale, of a part of the earth and is defined by geo coordinates. The image that comes from the video camera is a visual representation of the world in front of the camera and is defined by pixels. Therefore, when the video camera provides an image of a certain part of the earth, it can be used as a map after applying the procedures according to the subject invention by applying the calibration procedure described below

[0077] Before an image can be used as a map, it needs to be calibrated. Through the calibration process, we connect the pixels in the image with the geo coordinates of the space it shows. SI. 3a shows a top view of the earth when the camera is looking at the earth at a right angle. The upper part of the picture is directed towards the north. Here it can be immediately noted that other views are also possible, but they bring more mathematical operations to calculate the actual location, that's why this one was taken as an example because it requires the least mathematical operations.

[0078] Thus, Figure 3 a represents an image of a part of the earth obtained by the Google Earth Pro program, while it should be remembered that other programs or other images can also be used, for example topographic maps, Open Street map, etc. In Figure 3a, it is necessary to mark two arbitrary points which can be called as Cal. Point 1 301 and Cal. Point 2 302.

[0079] The actual coordinates of these points can be measured with a GNSS receiver or determined from a geo map such as a topographic map. This example will use Google Earth Pro.

[0080] Figure 3b represents the calibration process of an image that will be used as a map after the calibration process.

[0081] Map Point 1 box contains coordinates in pixels of point Cal. Point 1 in the picture (822,189) and its real coordinates in real space. The same applies to Map Point 2 and point Cal. Point 2.

[0082] In this way, the pixels in the image are connected to the real space coordinates. For example, the difference in horizontal pixels between Cal Point 1 and Cal Point 2 is 1257-822=435 pixels. The difference in long coordinates between these points is 0.0023034498. A point that is 10 pixels away from Cal Point 1 is 0.0000529529 degrees away from it in real space. This simple math connects the pixels in the image with the real coordinates of the real space. Through mathematical formulas for calculating the distance and angle between two geocoordinates, the connection between pixels in the image and meters is defined. This is important because in this case the height at which the camera is located can be measured and added as a parameter to the image. In this way, the calibration of the image that will be used as a map is completed.

[0083] Figure 3c shows the mounting location of a particular camera. Let the camera be mounted at a height of 5 meters. The actual coordinates of its mounting location can be measured with a GNSS receiver, but according to the subject invention, the actual coordinates of the camera mounting location can be determined using a method that uses a laser range finder and Google Earth Pro.

[0084] The method according to the invention involves measuring two distances to places that are easily recognizable on the map, in this case they are the two ends of the pier. In the case in Figure 3c, the distances are dl=48.39m, d2=42.89m.

[0085] Pozivajuci se na sliku 3d, iz tocke na kraju jednog gata 303 opisuje se kruznica koja ima radijus jednak izmjerenoj vrijednosti di ili d? (u ovom slucaju d2) dok se iz tocke kraja drugog gata 304 povlaci pravac koji ima drugu izmjerenu vrijednost d (u ovom slucaju di) i ciji kraj zavrsava tocno na kruznici.

[0086] Referring to Figure 3d, from the point at the end of one pier 303, a circle is drawn that has a radius equal to the measured value dl or d2 (in this case d2), while from the end point of the second pier 304, a line is drawn that has another measured value d (in this case case dl) and whose end ends exactly on the circle

[0087] The intersection point between the circle and the straight line is the camera mounting point. A Google Earth placemark is placed at that point. In this way, the actual coordinates of the camera installation site are obtained, which determine the actual location of the camera. We add the measured height of the camera to these real coordinates and thus obtain a unique 3D point.

[0088] The same procedure can be used to find the actual coordinates of the central point of the camera's field of view. Its height also needs to be measured and thus we get a unique 3D point.

[0089] Figure 4 shows an example of calculating geo coordinates from image coordinates. Let's say that on an image with a resolution of 1920x1080 we want to calculate the actual location for the location in the image of 1445x540 pixels. The angle between the axis 401 of the center of the field of view of the camera 400 and the required point is calculated from the parameter of the horizontal field of view.In this case the angle is equal to one quarter of horizontal field ov view (HFoV / 4). Now the distance in meters is calculated from the height of the camera and the angle of the camera to the ground. From the distance and HFoV / 4, the distance of the point from the calibration point 2 is obtained. Based on the known actual coordinates of the calibration points (their mutual distance and angle), the actual location of the required point is calculated.

[0090] This calculation took the angle of the point with respect to the center of the camera's field of view, which was calculated from the coordinates of the image (from the pixels). From this angle comes the distance of the required point from the center of the camera's field of view in meters, and this value is converted into real coordinates. To display this point on an image that is used as a map, we go back to connecting the pixels in the image with real space coordinates. In the example used previously, it was defined that for 435 pixels in the horizontal direction, the difference in long coordinates amounts to 0.0023034498 degrees. Based on these relationships, the coordinate in pixels of the image that we calibrated as a map is determined.

[0091] This is the procedure for calculating the actual location of an object recognized in a video camera image and displaying it in an image that is used as a map.

[0092] The integration of 3D space results from the fact that calibration points are essentially real space points and that they can be integrated into 3D space. In this way, the camera is integrated into the space. The calculation of the 3D location of the object recognized in the image is performed in the same way, except that after the calculated x and y coordinates in meters, the height is corrected. For example, a man recognized on the window of the second floor will be correctly positioned only when the x and y coordinates are corrected for the height at which he was recognized.

[0093] The implementation of the subject invention is shown in FIG. 5, which shows a picture of the part of the city where there are roads, intersections, pedestrian crossings, parking lots, that is, all the elements that make up an urban environment. The users of these urban spaces are people, vehicles, ships, cyclists and they all use the same space, so one of the tasks of managing urban spaces is to enable the most efficient use of these spaces. Efficient use of space requires a realtime review of all space users for the purpose of managing traffic lights, directing traffic, and managing parking lots.

[0094] The technical solution requires the use of computer systems to perform analysis tasks. Information about the condition inside the space is transmitted via video cameras. Scientific tests confirm that the concentration of a person observing multiple displays (operators in a video surveillance center) does not last longer than 30 minutes and that after that period the operators fail to notice events..

[0095] Therefore, there is a strong demand to leave the job of video review to computer systems that will notify operators of anomalies.

[0096] But it turned out that computers are the worst in tasks that are very easy for humans, namely the tasks of recognizing and classifying objects in images (videos).

[0097] A huge effort has been directed in the field of recognition and classification of objects in the image and significant steps have been made in this direction. The results that have been achieved are very good, but also show that more work is needed in this direction.

[0098] The moment an object is recognized in a picture, the first question that arises is, where is that object? This invention relates to the problem of determining the actual location of an object recognized in an image. This invention presents a technical solution that, among other things, enables the vision system to put information on the image of the part of the city under surveillance about traffic conditions, free parking spaces, illegally parked vehicles, places with a higher concentration of people, etc. Figure 6 shows one layer that such a system generates and which is placed on the image of the city so that the operator is informed that the vision system has recognized anomalies and displayed them (most often with a red icon) on the map in real time. Figure 6 shows how to display data collected by a number of video cameras. A higher concentration of vehicles is shown by a thicker line across the road 601. A higher concentration of people is shown by a larger icon showing a person 602. The size and color of the vehicle icon shows parked vehicles 603. Illegally parked vehicles are shown with a larger icon or a different color icon 604. Vision system collects data and, based on the criteria set by the user, displays the relevant data to the operator in graphic form (icons and lines above the folder map).

[0099] In order to be able to generate this and this amount of data, a larger number of cameras need to be installed and calibrated.

[0100] SI. 7 shows the integration of a single camera 901 or 904, designated in this figure by the reference numeral 701, and the center of the field of view of the camera 400.

[0101] The camera, marked as 701 in the figure, is placed on top of the building in the central part of figure 7, while the center of the field of view of the camera in figure 7 is marked as point 400. Camera 701 and the center of the field of view of camera 400 are calibration points whose actual coordinates and heights are known. In addition, the image also shows the actual location of the vehicle 703. In order to find out the actual location of the vehicle, it is determined how far the vehicle is along the x-axis 704 and along the y-axis 705 from the center of the field of view of the camera 400.

[0102] SI. 8 gives an example of the appearance of the image obtained from the video camera 901. The image shows the location in pixels of the detected vehicle 801. It has a pixel difference in the horizontal (dx) and vertical (dy) directions compared to the center of the camera's field of view (400). From the linear relationship between the horizontal field of view and the horizontal resolution, as well as the vertical field of view and the vertical resolution, dx and dy define the angles from the camera field of view center point to the detected object. The specified angles, camera heights and distances between the actual location of the camera mounting and the actual location of the center of the camera's field of view give dx and dy values expressed in meters.

[0103] Using formulas for calculating the distance between two real coordinates and formulas for calculating the angle between two real coordinates, the real location of the object is calculated.

[0104] From these real coordinates, taking into account the calibration parameters of the image we use as a map, the pixel values corresponding to this real location are obtained.

[0105] A series of video cameras will generate a series of geo coordinates that will represent people, vehicles, bicycles and other moving or static objects.

[0106] The vision system that receives geo locations of all detected objects from all calibrated cameras has an overview of all objects within the monitored area. By calibrating the image and calibrating the camera, a digital twin of real space was created, which gives the computer system an image of space and elements in space in real time. With these calibration methods, it is possible to create a layer of detected elements above the spatial map.

[0107] The computer system is now enabled to create predictive models of traffic, pedestrian behavior, etc., so that by comparing these models and the actual situation, it could manage traffic or any other process that takes place in the city / settlement, i.e. the monitored area. The method described so far assumes a two-dimensional space. The implementation of the third dimension refers to the height correction in relation to the center point of the camera's field of view using the camera's height point.

[0108] U navedenom primjeru, korekcija lokacije obzirom na razliku u visinu mjesta montaze kamere i lokacije na kojoj je detektirano vozilo zahtjeva od sustava da nakon izracunate lokacije detektiranog objekta za tu lokaciju, iz 3D geo mape cita iznos visine i usporeduje ga s iznosom visine lokacije mjesta montaze kamere. Ukoliko postoji razlika izmedu visina ovih dviju tocaka visina kamere se korigira u iznosu koji je jednak razlici visine i ponavlja se postupak izracuna geo lokacije detektiranog objekta.

[0109] In the given example, location correction due to the difference in the height of the camera mounting location and the location where the vehicle was detected requires the system to, after calculating the location of the detected object for that location, reads the height amount from the 3D geo map and compare it with the height amount of the camera video stream centre point location. If there is a difference between the heights of these two points, the height of the camera is corrected in an amount equal to the height difference and the process of calculating the geo location of the detected object is repeated.

[0110] This is how height correction is done.

[0111] Another object of this invention is a system through which the method described above is applied to determine the location in Cartesian coordinates, and / or the actual location of objects. The system is shown schematically in the picture 9.

[0112] The system includes at least one video camera, at least one video processing device integrated into or outside the video camera, predictive Al algorithms and control algorithms based on predictive models and real-time real-time data.

[0113] This invention describes a system for obtaining real 3D space coordinates for objects recognized in an image. The block diagram of the system in Figure 9 shows a system that has a larger number of cameras (901 or 904). The videos (902) of each of the cameras are analyzed and objects (people, cars, ships, etc.) are found in them based on user criteria..

[0114] Video analysis (902) of each of the cameras (901 or 904) can be done in three places. Video analysis (902) can be performed by the input part of the system (903) which is an algorithm that searches for objects within the image and is located within the computer, server or cloud application (900). Another place where video analysis can be done is a video camera (904) that contains a video processing algorithm, so instead of video, it sends metadata data (906) about objects and their current locations within the video image. In addition to cameras (901), there are also dedicated edge devices for video processing (905) that find objects in it and pass metadata about their current locations to the module for correcting the coordinates of detected objects (908). The combination of cameras (901) and the input part of the system (903) provide identical metadata as the combination of cameras (901) and the recording device (905). The mentioned combinations can be substituted with cameras (904). The choice of the specific selection depends on the user and his needs.

[0115] Metadata from all sources (eg, camera) is fed into the detected object coordinate correction module (908) along with image correction parameters for each of the cameras. The output from this module is metadata data (909) with corrected coordinates. This metadata together with the calibration parameters (910) of each of the cameras 901 or 904 form the input of the module (911) for calculating the coordinates of the recognized objects in the real world.

[0116] Thus, a system for determining the location of an object that contains at least one video camera 901 or 904; at least one video processing device 902 that can be integrated into the video camera

[0117] 904 or located outside 905 or 903 of the video camera 901 ; predictive Al algorithms and control algorithms based on predictive models and real-time real-time data, wherein the combination of camera 901 and system input 903 provide identical meta data as the combination of camera 901 and recording device 905 and wherein said combination of camera 901 and parts 903 and

[0118] 905 can be replaced by cameras 904 that provide identical meta data 906; and wherein the module for corrected coordinates of detected objects 908 receives said meta data 906 on the one hand together with data on image correction parameters 907 for each of the cameras 901 or 904 on the other hand, and wherein the outputs of said module are meta data 909 s corrected coordinates; wherein the said metadata 909 together with the calibration parameters 910 of each of the cameras 901 or 904 form the input of the module 911 for calculating the coordinates of recognized objects in the real world.

[0119] These coordinates together with 3D maps 912 are the input parameters of the module for height correction 913. The output of this module is the real locations in space 914 of the detected objects that are saved in the database 915. Based on the real data, the Al algorithm 916 makes predictive models according to the user's requirements.

[0120] The described system can be used in parking lot management applications, traffic management applications, city port management applications and other applications. Parking lot management application 918 uses a predictive parking model as well as real-time real-time data for parking management module 917. Traffic management module 919 uses an application 920 that uses a predictive traffic management model as well as real-time real-time traffic condition data. The city port management module 921 uses an application 922 that uses a predictive ship monitoring model as well as real ones.

[0121] Fig. 1A shows the video camera image aspect ratio 4:3.

[0122] Fig. IB shows the image of the video camera aspect ratio 16:9.

[0123] Fig. 1C shows the video camera image aspect ratio 16:10.

[0124] SI. 2a and 2b show the differences in video player and video origin.

[0125] SI. 3 A Displays the image to be calibrated and used as a map

[0126] SI. 3B Displays image calibration points

[0127] SI. 3C Shows the location of the camera installation and two distances that are measured for the purpose of determining the geo location of the installation location of the camera out of the Gogle Earth Pro program.

[0128] SI. 3D Shows the circle and the direction from the measurement points and their intersection, which indicates the location of the camera installation.

[0129] SI. 4 Shows an example of calculating the actual coordinate from the image coordinate SI. 5 It shows a picture that we use as a map of a part of the city

[0130] SI. 6 Shows the image we use as a folder with a layer generated by the system based on the locations of the detected objects collected. Heavy traffic is shown with a thick line, a larger number of pedestrians at an intersection is shown with a larger pedestrian icon, an illegally parked vehicle is shown with a red icon, etc.

[0131] SI. 7 Shows the field of view of one camera on the map as well as the location of the vehicle

[0132] SI. 8 Shows the center of the camera's field of view and the location of the detected vehicle within the image

[0133] SI. 9 Shows the system diagram.

Claims

Claims1. A method for determining the actual location of an obj ect based on the location in pixels where that object is located within the digital camera image (901) or (904), wherein said method contains the following steps: a. a. creating a distortion-free image by optimizing camera parameters, image parameters and lens parameters; b. calibration of the camera (901 ) or (904) in such a way as to determine the height of the camera and the geo-location of the camera mounting location as the first calibration point, and then to determine the height and actual location of the camera's field of view center point (400) as the second calibration point and that the first and second calibration points are entered into the vision system (900) which, based on this data and imported 3D maps of the space, defines the 3D space within which the camera calibration points are located; c. connecting pixels within the image with coordinates in the real world in such a way that each pixel in the image is assigned a unique location in real space.

2. The method according to claim 1, characterized in that it contains the following steps: a. creating an image without distortion by optimizing camera parameters, image parameters and lens parameters; b. Measurements of the actual location of the camera (901) or (904) from the map as well as measurements of the height of the camera (901) or (904) relative to the ground, where the actual location of the camera (901) or (904) and the height of the camera (901) or (904) ) relative to the ground represent the first calibration point of the camera (901) or (904); c. Measurements of the actual location of the center of the camera's field of view (400) and its height relative to the ground, where the actual location of the center of the camera's field of view (400) and its height relative to the ground represent the second camera calibration point (901) or (904); d. Connections of the actual location of the calibration points from a. and b. with their coordinates in pixels on the image that we will calibrate as a map; e. Representations of all calibrated cameras (901) or (904) that have an actual location within the boundaries defined by the calibration of the image showing a specific geographic area; f. Determination of the horizontal and vertical angle of the detected object in the image of one of the cameras (901) or (904) in relation to the center of the field of view of the camera (400) from its location expressed in pixels; g. Determining the distance along the horizontal and vertical axis of the image from the camera (901) or (904) from the height difference and the mutual distance of the calibration points; h. Determining the actual location of the detected object from the distance from point g. in relation to the actual locations of the calibration points; i. Determining the coordinates of the detected object in pixels of the image used as a map from the ratio of pixels to the image and camera calibration points (301) and (302) for the actual coordinates of the detected object;j. Repeating the process for each object detected by any of the cameras (901) or (904).

3.

3. The method according to claim 1 or 2, characterized by the fact that the height of the geo location of the camera mounting location (901) or (904) is different from the height of the geo location of the center of the camera's field of view (400), which enables tracking the location of an object that is moving up or down in relation to the plane of the mounting height of the camera (901) or (904) in the place where it is mounted.

4. A method for determining the location of an object according to claim 1 or 2, characterized by the fact that 3D maps are generated from the x and y coordinates of the distance of the object from the camera (901) or (904) with the help of 3D maps or Google API services that give us the height for the requested location of the object.

5. The method according to claim 3 or 4, characterized by the fact that the distance in meters of the required point is calculated via the angle of the point in relation to the center of the camera field of view (400), which is obtained from the required point coordinates in the image of the video camera (900) or (904) in pixels.

6. A method for determining the location of an object according to any of the preceding claims, characterized in that a digital twin of detected objects from video camera images (901) or (904) is created by image calibration and camera calibration (901) or (904) which together with 3D maps ( 912) provides the computer system with an image of space and detected objects in real space, real space and real time.

7. A system for determining the location of an object containing at least one video camera (901) or (904); at least one video processing device (902) that can be integrated into the video camera (904) or located outside (905 or 903) the video camera (901); predictive Al algorithms and control algorithms based on predictive models and real data in real time, where the combination of cameras (901) and the input part of the system (903) provide the same meta data as the combination of cameras (901) and the recording device (905) and wherein said combinations of camera (901) and parts (903) and (905) can be replaced by cameras (904) that provide identical metadata (906); and wherein the module for corrected coordinates of detected objects (908) receives said meta data (906) on the one hand together with image correction parameter data (907) for each of the cameras (901 or 904) on the other hand, and exits from said meta data module (909) with corrected coordinates; wherein the said metadata (909) together with the calibration parameters (910) of each of the cameras (901) or (904) form the input of the module (911) for calculating the coordinates of recognized objects in the real world.

8. System according to claim 7, characterized in that the height correction module (913) additionally generates real locations of detected objects in space (914) which are stored in the database (915).

9. The system according to claim 8, characterized by the fact that based on the data obtained from the actual locations of detected objects in space (914) and the predictive Al algorithm (916) it manages parking lot monitoring applications (917), traffic management applications (918) and port management (919).

10. A system for determining the location of an object through which a method for determining the location of objects according to any of the requirements is performed 1 - 6.

11. Use of the method according to requirements 1-6 for parking management, traffic management, port management.

12. Use of systems according to requirements 7 - 10 for parking management, traffic management, port management.

Citation Information

Patent Citations

  • Camera Calibration Method

    US20230095500A1

  • Neural network object pose determination

    US20230145701A1