Method for a stationary monocular camera recording frames of a scene
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- DEEPSCENARIO GMBH
- Filing Date
- 2024-07-10
- Publication Date
- 2026-05-20
AI Technical Summary
Monocular cameras provide only 2D representations of objects, leading to inaccurate 2D bounding boxes that cannot achieve centimeter-level accuracy, which is insufficient for detailed applications like traffic analysis, and existing methods to convert 2D to 3D data are complex or require expensive equipment like lasers or stereoscopic systems.
A method for a stationary monocular camera to estimate 3D bounding boxes by obtaining calibration parameters automatically using a 3D model computed from retrieved frames, where pixel pairs are matched between selected and recorded frames to derive extrinsic and intrinsic camera parameters, enabling the generation of 3D replicas of the scene without manual measurement or expensive equipment.
Enables accurate 3D bounding box estimation for objects in the scene, providing detailed 3D representations for applications like automated driving studies without the need for expensive equipment or manual calibration, allowing for efficient and robust computation of calibration parameters.
Smart Images

Figure EP2024069561_16012025_PF_FP_ABST
Abstract
Description
[0001] Method for a stationary monocular camera recording frames of a scene
[0002] The invention relates to stationary monocular cameras, which is used to observe its environment continuously or for a longer period in time. Examples for such stationary monocular cameras are stationary mobile cameras, traffic cameras, surveillance cameras or drones that hover stationary above a fixed location.
[0003] Such cameras often observe vehicular and / or pedestrian traffic on a road. Typically, traffic cameras are arranged inside or outside cities along roads or major roads, such as highways, freeways, expressways or crossroads. Such cameras send their images, also called frames, or videos to a monitoring center, which receives the frames or video in real time and serves as a dispatcher, if there is a disruptive incident such as a traffic collision or road safety issue. For example, traffic cameras also can be used to activate safety equipment remotely based on information provided by the camera.
[0004] However, monocular cameras are inexpensive and widely used. Thus, initial efforts are known to use these cameras, which are available anyway, to understand how people, goods, animals, etc. move through space and time, namely the scene which is recorded. Consequently, a stationary monocular camera can be used to observe its environment for a longer period in time.
[0005] For technical applications, i.e. to provide a detailed analysis, like a traffic analysis, it is key to estimate the position and orientation of all objects in the scene, to estimate the size of their bounding box, namely length, width and height, and further preferably to estimate to which category, for example car, bicycle, pedestrian, the respective object belongs and to assign a consistent track ID to all objects in a scene. However, with normal monocular cameras, only a 2D representation of the objects is possible and common practice.
[0006] Depending on the position of the monocular camera and thus, depending on the perspective by which the monocular camera sees the scene, 2D bounding boxes of objects are thus inaccurate and cannot be used to achieve a detailed, in example centimeter-level, accuracy. However, a detailed representation is necessary for a variety of applications, such as automated driving studies.
[0007] Therefore, it is also known according to the state of the art to convert the 2D video data of a monocular camera into 3D data in a very complex way, whereby however the use of personal or expensive equipment like lasers on a side is necessary to determine parameters for converting. Furthermore, further effort are known to directly provide 3D data of a scene by installation of systems like radar or stereoscopic cameras. Such systems, however, are very expensive compared to monocular cameras and also a high effort has to be made for their installation.
[0008] It is therefore desirable to address at least one of the above problems. In particular, the object of the invention is to use a stationary monocular camera recording at least one frame of a scene to estimate 3D bounding boxes for objects in the scene in order to generate 3D replicas of the scene. A further object of the present invention is to propose an alternative to the state of the art.
[0009] In accordance with the invention, the object is solved in a first aspect by a method for a stationary monocular camera in particular a traffic camera, a surveillance camera, a hovering drone, or a stationary mobile camera, recording at least one frame of a scene according to claim 1 .
[0010] In accordance with the invention, a method for a stationary monocular camera recording at least one frame of a scene comprises to obtain calibration parameters, comprising preferably extrinsic and / or intrinsic parameters. Thus, the method comprises according to the first aspect to obtain calibration parameters for the stationary monocular camera. The method comprises for obtaining calibration parameters of the camera the following steps, which are preferably automated without the use of personal on a side of the scene or in a manual task afterwards.
[0011] In a first step, a plurality of frames recorded from different camera poses of a scene are retrieved. The scene is preferably a location where the identification and / or behavior of objects is to be analyzed. The frames which are retrieved are referred to as retrieved frames. Frames are also referred to as images. Further, a 3D model of the scene is computed using the retrieved frames. The 3D model of the scene is preferably a 3D point cloud or a 3D mesh. The 3D model is computed in such a way that it consists of a plurality of 3D elements.
[0012] To each of the 3D elements, 3D element information are assigned, respectively. For example, in the case the 3D model is a 3D point cloud, a 3D element corresponds to a 3D point. The 3D element information comprise for example the arrangement of the 3D point with respect to a defined reference or to other points. For example, 3D element information of a 3D point correspond to x-, y-, z-coordinates in a Cartesian coordinate system with three dimensions. Further, a 3D element could be a polygon. In a further step of the method, at least one frame of the scene is recorded using the stationary monocular camera. The frame is referred to as recorded frame. The scene of the recorded frame corresponds to the scene of the retrieved frames. The recorded frame comprises a plurality of pixel, wherein 2D information correspond to each pixel of the recorded frame, respectively. For example, each pixel of the recorded frame is assigned to a 2-di- mensional coordinate system and thus, comprises for example a u-coordinate and a v- coordinate.
[0013] Further, the method comprises to select at least one of the retrieved frames, which is referred to as selected frame.
[0014] Moreover, the method comprises to determine a plurality of pixel pairs. Each pixel pair comprises one pixel of the selected frame and one pixel of the recorded frame, corresponding to each other. Corresponding pixel of a selected frame and the recorded frame may for example be determined on a defined criterion.
[0015] Further, the method comprises to assign 3D element information to each of the pixel pairs. Assigning to a pixel pair 3D element information comprises to assign the 3D element information of the 3D element, which corresponds to the pixel of the selected frame in the pair, to the pixel pair. A correspondence between the 3D element and a pixel of the selected frame in the pair can be determined for example based on a further criterion.
[0016] Moreover, the method comprises to assign 2D information to each of the pixel pairs. Assigning 2D information to a pixel pair comprises to assign to the pixel pair the 2D information corresponding to the pixel in the recorded frame, which is assigned to the respective pair.
[0017] As the 3D model with its 3D elements is computed from the retrieved frames and the selected frame is one of the retrieved frames the pixel pairs provide a correspondence between the 3D model and the selected frame.
[0018] Thus, calibration parameters of the stationary monocular camera can be calculated based on the 3D element information and the 2D information assigned to the pairs. Further, preferably extrinsic parameters, like a pose comprising a position and orientation of the stationary monocular camera, or intrinsic parameters, like an intrinsic matrix and / or distortion coefficients, can be calculated based on the 3D element information and the 2D information, which are assigned to each other by the pairs. Different methods for calculating parameters comprising intrinsic and extrinsic parameters based on correspondences between the real world represented by the 3D element information and the image of the scene represented by the 2D information are known and thus, calculating the parameters can be executed by those methods.
[0019] Thus, the idea of the present invention is to compute a 3D model of the scene based on retrieved frames without further detailed information of the scene and the pose and intrinsic parameters of the stationary monocular camera. Consequently, correspondence between the real 3D world represented by the 3D model and the representation in the 2D frame can be used to determine the calibration parameters without manually measure at least one pose of the stationary monocular camera in the scene or manually observe details of predefined features in the scene. Calibration parameters for a stationary monocular camera can thus preferably be derived completely automatically.
[0020] According to a first development of the first aspect, the method further comprises for retrieving a plurality of frames as retrieved frames recording a plurality of frames using a mobile camera. During recording the plurality of frames, the mobile camera is moving through the scene. The mobile camera is, for example, a monocular camera which can be referred to as further monocular camera or mobile monocular camera, because the mobile camera is preferably different from the stationary monocular camera. According to an alternative, the mobile camera corresponds to the stationary camera, wherein in particular the camera is moved through the scene as mobile camera and then mounted stationary at a fixed point as the stationary camera. Preferably, the mobile camera is arranged at a vehicle like a drone, a car, a bicycle, a plane or a human. Further, the method comprises to provide the plurality of frames recorded by the mobile camera for retrieving. Consequently, the retrieved frames correspond to frames, which are recorded by a mobile camera moving through the scene.
[0021] Preferably, the mobile camera comprises a localization device for acquiring poses or positions of the mobile camera. The localization device is for example a receiver for receiving signals of the Global Positioning System (GPS), the Galileo-System, the Glonass-System, the Beidou-System or the IRNSS-System. Preferably, the mobile camera comprises additionally or alternatively to a localization device an inertial measurement unit (IMU) to derive the pose of the mobile camera. Preferably, additionally or alternatively to a localization device and / or to an inertial measurement unit the mobile camera comprises a detection unit for detecting Ground Control Points for acquiring poses or positions of the mobile camera based on the Ground Control Points. Further, the method preferably comprises to assign to some or to each of the frames recorded by the mobile camera a camera pose of the mobile camera taken during recording the respective frame, preferably by using an algorithm like Structure from Motion (SfM) or a Simultaneous Localization and Mapping (SLAM) algorithm
[0022] Providing the retrieved frames by a mobile camera, for example by an autonomous drone, enables to provide all data of the scene, which is needed to generate a 3D model and thus to provide calibration parameters for the stationary camera. All data can thus be obtained automatically without personal deployment.
[0023] According to a further development, recording a plurality of frames with the mobile camera comprises positioning the mobile camera to a defined pose, preferably which is substantially similar to the pose of the stationary monocular camera or which is near or nearest to the pose of the stationary monocular camera. Further, the method comprises recording at least one frame at the defined pose. Further, the method comprises to provide the frame which is recorded by the mobile camera at the defined pose as one of the retrieved frames and select the retrieved frame at the defined pose as selected frame.
[0024] Based on this development, a retrieved frame is provided, which can be used as the selected frame. The retrieved frame can be recorded by the mobile camera by defining the pose for being on a pose most similar to the pose of the stationary monocular camera. Consequently, the step of determining a plurality of pixel pairs can be executed with less effort and determining the plurality of pixel pairs is less computationally intensive. Thus, the method can be executed faster and with less computing power and more robustly.
[0025] According to a further development, selecting at least one of the retrieved frames comprises comparing a plurality of the retrieved frames or each of the retrieved frames with the recorded frame. The plurality of the retrieved frames for comparison are for example selected based on a pose assigned to the retrieved frame or frames, wherein the poses preferably are derived by an algorithm supported by the poses derived from the localization device, the inertial measurement unit or the Ground Control Points. By the comparison step at least one of the retrieved frames can be identified which perspective of the scene is most similar to the perspective of the recoded frame or which shares the most image features with the recorded frame or shows the most similar image features as the recorded frame.
[0026] Consequently, thus a further option is provided, also if the retrieved frames are retrieved from an external source and not from a mobile camera, which can be placed on a defined pose. Consequently, selecting the retrieved frame based on a comparison also enables the method to be executed with less effort and thus to be executed faster in a more generic way.
[0027] According to a further development, computing a 3D model based on the retrieved frames comprises to use an automatic algorithm. The automatic algorithm can be according to a special development an algorithm, which is based on artificial intelligence. For example, a Structure from Motion (SfM) algorithm or a Simultaneous Localization and Mapping (SLAM) algorithm is used to compute the 3D model based on the retrieved frames.
[0028] Preferably, the algorithm uses camera poses or positions assigned to the retrieved frames and / or receives an input of information, for example coordinates of Ground Control Points shown in at least one of the retrieved frames, to compute the 3D model of the scene, in particular to compute for each of the retrieved frames a pose of the mobile camera taken during recording the respective frame.
[0029] Using an algorithm like a SfM or SLAM algorithm for computing a 3D model enables the use of established and optimized methods to achieve a suitable result, namely 3D model, for the further steps. Preferably, the algorithm is used for generating a 3D point cloud or a 3D mesh. Preferably, a further algorithm is used to generate a 3D mesh from a 3D point cloud which is output by a SfM or SLAM algorithm.
[0030] According to a further development, determining a plurality of pixel pairs comprises to determine for each of a number of arbitrary pixel or all pixel in the recorded frame which represent a feature of the scene, a pixel in the selected frame corresponding to the same feature of the scene. Consequently, determining a plurality of pixel pairs comprises to find those pixel of the recorded frame and the selected frame showing or representing the same feature. For example, a feature is a piece of information about the content of an image or frame, which can be obtained by image processing. Features may be specific structures in the image such as points, edges or objects. Further, features can be defined or differentiated from other features by structures, colors or other pixel characteristics. Consequently, determining a plurality of pixel pairs comprises to find features, which are shown in both, in the recorded frame and in the selected frame, and to link or associate corresponding pixel to each other by the pairs.
[0031] Preferably, determining a plurality of pixel pairs is executed manually or using an algorithm like Scale Invariant Feature Transform (SIFT) or Detector-Free Local Feature Matching with Transformers (LoFTR). A further algorithm, which has the capability to determine pixel pairs is known as SuperGlue. Pixel pairs could thus be also referred to as matches or correspondences.
[0032] According to a further development of the method, corresponding 3D element information are assigned to each of the pixel pairs, respectively. Assigning to a pixel pair, for each of the pixel pairs, a corresponding 3D element information comprises to determine for the respective pixel pair the 3D element of the 3D model, which corresponds to the pixel of the selected frame. Each pixel of the selected frame has a corresponding 3D element of the 3D model. Consequently, for each of the pixel pairs, the 3D element is determined, which corresponds to the pixel in the pixel pair. Further, the 3D element information of the determined 3D element are assigned to the pixel pair, respectively.
[0033] For example, thus, a pixel pair is selected, in the next step, the pixel of the selected frame is used to find the corresponding 3D element of the 3D model and then the 3D element information of the found 3D element is assigned to the first pixel pair. Afterwards, the next pixel pair is selected and the steps are executed for the next pixel pair. This continues until 3D element information are assigned to each of the pixel pairs. However, this example describes a reordering sequence in which the pixel pairs are processed one after the other. This is for illustrative purposes only. According to another development, the pixel pairs are be processed in parallel.
[0034] Consequently, according to the development, in a first step, pixel pairs are determined. The number of the pixel pairs depends on the features, which are recognizable in both frames, namely the selected frame and the recorded frame. The more features are recognized the more pixel pairs can be determined. In the next step, 3D element information are assigned to each of the determined pixel pairs.
[0035] Thus, the information of the 3D model can be transferred by the selected frame to the pixel pairs. Thus, correspondences between 3D element information and 2D information can be obtained by the pairs. The 2D information can thus be linked to the 3D model.
[0036] According to a further development, determining for the respective pixel pair the 3D element of the 3D model corresponding to the pixel of the selected frame is executed by using a pinhole camera model and raytracing. Consequently, parameters of the camera which provided the retrieved frames are used. The parameters are for example provided by a manufacturer of the camera and / or by the algorithm for computing a 3D model based on the retrieved frames. Thus, using the parameters a perspective projection of a pixel of the selected frame to the 3D model is possible to find the 3D element of the 3D model corresponding to the pixel.
[0037] A suitable process for finding 3D elements corresponding to the pixel in the pixel pair is thus provided.
[0038] According to a further development, preferably as an alternative to the previous development, the method comprises to determine a plurality of pixel pairs and to find 3D element information, that in a first step for each of a number of arbitrary or all 3D elements of the 3D model a pixel in the selected frame representing the 3D element is identified. For example, the algorithm for computing the 3D model, preferably SfM or SLAM, provides for the 3D elements in the 3D model corresponding 2D information of each pixel in the selected frame which can be used to identify the pixel in the selected frame. Further, in a second step, the method comprises for determining a plurality of pixel pairs determining for each of the identified pixel further representing a feature of the scene, a pixel in the recorded frame corresponding to the same feature of the scene and assigning corresponding pixel to a pair. Determining a plurality of pixel pairs is executed for example by using an algorithm like S2DNet.
[0039] According to the previous development in a further development assigning to a pixel pair, for each of the pixel pairs (36, 38), corresponding 3D element information comprises assigning the 3D element information to the pixel pair (36, 38), which corresponds to the 3D element of the identified pixel of the pixel pair (36, 38).
[0040] According to a further development, determining a plurality of pixel pairs is executed using a neural network or based on artificial intelligence.
[0041] According to a further development, assigning to a pixel pair, for each of the pixel pairs, 3D element information of a 3D element of the 3D model corresponding to the pixel of the selected frame in the pair, is executed using a neural network or based on artificial intelligence.
[0042] According to a further development, calculating parameters and preferably the pose, the intrinsic matrix and / or the distortion coefficients of the stationary monocular camera based on the 3D element information and 2D information assigned to the pairs is executed using a neural network or based on the use of artificial intelligence. However, according to a development of the method, 2D information are assigned to 3D element information by the pairs. The plurality of pairs can thus be used to solve the following equation:
[0043] In this equation the values correspond to the 2D information assigned to the i-th pixel pair. The information y;and z;correspond to the 3D element information assigned to the i-th pair. The parameters K, which corresponds to intrinsic parameters represented by an intrinsic matrix, and the parameters R, t, which correspond to the extrinsic parameters, can be derived from the correspondences of the 2D information and the 3D element information by solving the equation. can thus be calculated, too.
[0044] Consequently, the intrinsic matrix K and the extrinsic parameters R, t namely the matrix / ? and the vector t correspond to the parameters of the stationary monocular camera. Consequently, an exact pose of the stationary camera can be derived.
[0045] According to a development of the first aspect or to a second aspect, the method further comprises providing a stationary monocular camera and preferably calibration parameters of the monocular camera. According to the development or second aspect, the method comprises for providing a 3D bounding box of at least one object in the scene and preferably for assigning a unique identifier to the object across frames further steps, which are preferably executed with a neural network and thus are based on the use of artificial intelligence.
[0046] The steps comprise, for providing the bounding box of an object in a frame of the scene recorded by the stationary monocular camera, a first step to identify an object in the frame of the scene. The object is preferably identified by an image processing algorithm. Preferably an identifier is assigned to the object. Wherein the identifier is preferably a unique track identifier. Further, 2D information of a reference pixel of the identified object is estimated or determined. The 2D information preferably comprise the u- and v-coordinates in the recorded frame. The reference pixel may be one arbitrary pixel or preferably a centre pixel of the object. Further, a depth of the object is determined. The depth corresponds in particular to a distance between the object represented by the reference pixel and the stationary camera. Determining the depth of the object is preferably executed by estimating the depth, preferably using a neural network, or by calculating the depth. Calculating the depth comprises to use a 3D model of the scene. The 3D model corresponds to the 3D model, which is derived from the retrieved frames. Calculating the depth comprises to determine the 3D element information of the 3D model corresponding to the reference pixel, for example by using a pinhole camera model and raytracing. Further calculating comprises to calculate the depth based on the pose of the stationary monocular camera and the 3D element information.
[0047] In a further step, a 3D bounding box surrounding the object is estimated. Preferably estimating the 3D bounding box comprises to estimate a size of the 3D bounding box, wherein the size preferably comprises the length, width and height of the 3D bounding box. Estimating a bounding box surrounding the object can be estimated by defining edges delimiting the object and thus forming the bounding box. Estimating the bounding box can be executed by an image processing algorithm and / or a neural network.
[0048] Further, a yaw angle of the 3D bounding box is estimated. Estimating the yaw angle is preferably based on the camera parameters of the stationary camera and or the neural network.
[0049] In a further step, the roll and pitch angle of the 3D bounding box relative to the camera position is determined. Determining the roll and pitch angle is preferably executed by estimating the roll and pitch angle, preferably using a neural network, or by calculating the roll and pitch angle. Calculating the roll and pitch angle comprises to use a 3D model of the scene. The 3D model corresponds to the 3D model, which can be derived from the retrieved frames.
[0050] Consequently, the 3D model is preferably already available. However, the 3D model is computed based on a plurality of retrieved frames of the scene, which are recorded from different camera poses of the scene. The 3D model preferably is computed by an algorithm, for example an SfM or SLAM algorithm and the 3D model consists of a plurality of 3D elements, wherein 3D element information are assigned to each 3D element.
[0051] Further, calculating the roll and pitch angle comprises to determine 3D element information, in particular 3D polygon information of a 3D element corresponding to the reference pixel. The 3D element is in particular a 3D polygon, in the 3D model. In particular, a 3D polygon corresponds to a 3D element and thus, 3D polygon information corresponds to the 3D element information of the 3D element. The 3D element or 3D polygon corresponding to the reference pixel is preferably the 3D element or 3D polygon which represents the reference pixel in the 3D model. Preferably the 3D element information correspond the normal vector of a polygon or triangle which corresponds to the 3D element.
[0052] Further, the method comprises to calculate the pitch and roll angle of the bounding box based on the determined 3D polygon information or 3D element information, in particular a normal vector of the 3D element and the estimated yaw angle and preferably the camera pose.
[0053] Consequently, a method is provided for defining a 3D bounding box of an object with an exact depth, pitch and roll angle with respect to the real world based on the use of a 3D model.
[0054] The camera parameters are provided for example by the method according to the first aspect. Using intrinsic parameters, comprising an intrinsic matrix and / or distortion coefficients, and extrinsic parameters of the camera can thus be used in the process to find the size, position, comprising the depth, the yaw, roll and pitch angle of the 3D bounding box.
[0055] According to a further development, the 3D element information, preferably x, y and z coordinates, corresponding to the reference pixel is determined using a pinhole camera model and raytracing.
[0056] According to a further development, the 3D model is a polygon mesh model, in particular a triangle mesh model. The 3D model thus comprises a plurality of polygons defined by the 3D elements or correspond to the 3D elements. According to a further development, the 3D polygons are defined by 3D elements, which correspond to 3D points of the 3D element.
[0057] According to a further development, the method further comprises to provide output data comprising a sequence of frames, namely a video sequence of the scene, wherein each frame comprises the identified object represented by the 3D bounding box and preferably an identifier assigned to the identified object and / or the 3D bounding box.
[0058] According to a further development, before identifying an object in the recorded frame, a normalization step is executed. In particular, the normalization step comprises to rectify the recorded frame recorded by the stationary monocular camera using estimated distortion coefficients and / or normalize the focal lengths to the same value across frames. Further preferably, the normalization step can comprise to normalize the extrinsic parameters, for example to compensate drift in the pose of the camera.
[0059] According to a third aspect of the invention, the invention comprises a computer program product, which comprises instructions. The instructions, when executed on a processor, cause the processor to execute the method according to the first and / or the second aspect of the invention.
[0060] Moreover, the invention is directed to a fourth aspect, wherein the fourth aspect comprises output data. Output data according to the invention comprise a sequence of frames of a scene, comprising at least one 3D bounding box, representing an object in the scene, wherein the data is produced by a method according to the first and / or the second aspect of the invention.
[0061] Further advantages, features and details of the invention result from the following description of the preferred embodiments as well as from the drawings, which show in:
[0062] Fig. 1 a scene with a stationary monocular camera,
[0063] Fig. 2 a general step of the method according to an embodiment,
[0064] Fig. 3 steps for obtaining calibration parameters of the camera and
[0065] Fig. 4 steps for providing a bounding box of an object in a scene.
[0066] Fig. 1 shows a scene 10 comprising an intersection of streets. In one corner of the intersection, a stationary monocular camera 12 is positioned to record frames 14, namely images, of the scene 10. The recorded frames 14 are output from the stationary monocular camera 12 to a computer system 16 executing a method for obtaining calibration parameters of the stationary monocular camera 12 and for identifying 3D bounding boxes 18 of objects 20 in the scene 10.
[0067] Further, a mobile camera 22 is shown, which is arranged at a vehicle 24, namely a drone 26. The mobile camera 22 records frames of the scene 10 from different poses during moving through the scene 10 along a path 28. The plurality of frames recorded with the mobile camera 22 are output to the computer system 16 and thus retrieved from the computer system 16. The frames are referred to as retrieved frames 30.
[0068] The computer system 16 calculates based on the retrieved frames 30 a 3D model 32 of the scene 10. Further, at least one of the retrieved frames 30 is selected as a selected frame 34. The selected frame 34 is selected as being the frame of the retrieved frames 30, which shows most features 31 , for example traffic light 33, of the scene 10, which are recorded by the stationary monocular camera 12, also. The embodiment in fig. 1 is directed to select one retrieved frame 30 as the selected frame 34. However, according to a further embodiment a plurality of frames 30 can be selected as selected frames.
[0069] The selected frame 34 and one of the recorded frames 14 are investigated. Thus, pixel pairs 36, 38 are determined. Each of the pixel pairs 36, 38 comprise one pixel 40, 42 of the recorded frame 14 and one pixel 44, 46 of the selected frame 34. In other words pixel pairs 36, 38 are determined between the selected frame 34 and one of the recorded frames 14. Each pixel 40, 42 comprises 2D information 48. Further, based on the 3D model 32, for each pixel 44, 46 of the selected frame 34, 3D element information 50 can be obtained. Thus, each of the pixel pairs 36, 38 comprises 2D information 48 and 3D element information 50. Based on these matching information, extrinsic and intrinsic camera parameters 52 of the stationary monocular camera 12 are calculated by using a pinhole camera model 54. The intrinsic parameters comprise an intrinsic matrix and / or distortion coefficients.
[0070] Fig. 2 shows the general steps of the method according to an embodiment. In the first step 60, a plurality of frames 30 of a scene 10, for example recorded by a mobile camera 22, are retrieved. In step 62, a 3D model 32 of the scene 10 is computed. Further, in step 64, a frame 14 is recorded by a stationary monocular camera 12. In step 66, intrinsic and extrinsic camera parameters 52 are calculated, for example, as explained with respect to fig. 1.
[0071] In step 68, the parameters are normalized, which for example comprises to rectify the recorded frame 14 using estimated distortion coefficients. Further, focal lengths can be normalized to the same value across a plurality of recorded images. Further, the extrinsic parameters can be normalized for example to compensate drift in the pose of the stationary monocular camera 12.
[0072] In step 70, an object 20 is detected and preferably tracked by assigning an identifier to the object. Based on the 2D position in the recorded frame 14, the the depth, the size, and the roll, pitch and yaw angle of the bounding box 18 are estimated or determined. The estimation is based on a neural network and the calculation is based on the camera parameters 52 and the 3D model 32. In step 72, which is a refinement step, a 3D bounding box comprising a position, orientation, length, width and height of the bounding box 18 is calculated based on the camera parameters and the 3D model 32.
[0073] Fig. 3 describes the method for obtaining the calibration parameters 52 of the stationary monocular camera 12 in more detail.
[0074] In step 80, a plurality of frames 30 recorded from different camera poses of a scene 10 are retrieved as retrieved frames 30. In step 82, a 3D model 32 of the scene 10 is computed. In step 84, at least one frame of the scene 10 is recorded as recorded frame 14 using a stationary monocular camera 12.
[0075] In step 86, at least one of the retrieved frames 30 is selected as selected frame 34. In step 88, for each of a number of arbitrary or all pixel in the recorded frame, a pixel in the selected frame 34 is determined, respectively. Consequently, a pair of recorded frame pixel and selected frame pixel are thus obtained, wherein the pixel in a pair are chosen to represent the same feature or point of a feature, namely a salient point. In step 90, 3D element information 50 are assigned to each of the pairs based on the element of the 3D model 32 corresponding to the pixel in the pair associated with the selected frame. Further, 2D information 48 are assigned to the pair based on the position of the pixel in the pair associated to the recorded frame 14.
[0076] In step 92, intrinsic and extrinsic camera parameters 52 are calculated based on the information given in the pairs, preferably based on a perspective projection.
[0077] Fig. 4 describes the steps for providing a 3D bounding box 18 of at least one object 20 in a scene 10. In step 100, an object 20 in a frame 14 of the scene 10 is recorded by a stationary monocular camera 12. In step 102, a reference pixel of the identified object 20 is determined. In step 104, a 3D bounding box 18 surrounding the object 20 is estimated. In step 106, a yaw angle of the 3D bounding box is estimated. In step 108, roll and pitch angle of the bounding box and a depth of the bounding box are estimated and / or calculated based on information of the 3D model 32 and calibration parameters 52. Reference signs
[0078] 10 scene
[0079] 12 stationary monocular camera
[0080] 14 recorded frame(s) 16 computer system
[0081] 18 bounding box(es)
[0082] 20 object(s)
[0083] 22 mobile camera
[0084] 24 vehicle 26 drone
[0085] 28 path
[0086] 30 retrieved frames
[0087] 31 features
[0088] 32 3D model 33 traffic light
[0089] 34 selected frame
[0090] 36 pixel pairs
[0091] 38 pixel pairs
[0092] 40 pixel of recorded frame 42 pixel of recorded frame
[0093] 44 pixel of selected frame
[0094] 46 pixel of selected frame
[0095] 48 2D information 50 3D element information
[0096] 52 camera parameters
[0097] 54 pinhole camera model
[0098] 60 retrieve plurality of frames
[0099] 62 compute 3D model 64 record recorded frame
[0100] 66 calculate intrinsic and extrinsic camera parameters
[0101] 68 normalize parameters
[0102] 70 detect an object
[0103] 72 estimate 3D bounding box and identifier 80 retrieve plurality of frames as retrieved frames
[0104] 82 compute 3D model
[0105] 84 record one frame of the scene as recorded frame
[0106] 86 select at least one of the retrieved frames as selected frame
[0107] 88 determine a pixel in the selected frame 90 assign 3D element information to each of the pairs
[0108] 92 calculate intrinsic and extrinsic camera parameters
[0109] 100 record an object in a frame of the scene
[0110] 102 determine a reference pixel of the identified object 104 estimate a bounding box surrounding the object
[0111] 106 estimate a yaw angle of the 3D bounding box
[0112] 108 determine roll and pitch angle and depth of the 3D bounding box
Claims
Claims1 . Method for a stationary monocular camera (12), in particular a traffic camera, a surveillance camera, a hovering drone, or a stationary mobile camera, recording at least one frame (14) of a scene (10), comprising for obtaining calibration parameters of the camera (12) the steps of: retrieving a plurality of frames (30) recorded from different camera poses of a scene (10) as retrieved frames (30), computing a 3D model (32) of the scene (10), preferably a 3D point cloud or a 3D mesh, based on the retrieved frames (30), wherein the 3D model (32) consists of a plurality of 3D elements and wherein 3D element information (50) are assigned to each 3D element, respectively, recording at least one frame of the scene (10) as a recorded frame (14) using the stationary monocular camera (12), wherein the recorded frame (14) comprises a plurality of pixel, wherein 2D information (48) correspond to each pixel (40, 42) of the recorded frame (14), respectively, selecting at least one of the retrieved frames (30) as a selected frame (34), determining a plurality of pixel pairs (36, 38), wherein each pixel pair (36, 38) comprises one pixel (44, 46) of the selected frame (34) and one pixel (40, 42) of the recorded frame (14) corresponding to each other, assigning to a pixel pair (36, 38), for each of the pixel pairs (36, 38), 3D element information (50) of a 3D element of the 3D model (32) corresponding to the pixel (44, 46) of the selected frame (34) in the pair (36, 38), assigning to a pixel pair (36, 38), for each of the pixel pairs (36, 38), 2D information (48) corresponding to the pixel (40, 42) of the recorded frame (14) in the pair (36, 38) and calculating the calibration parameters, preferably comprising intrinsic and extrinsic parameters of the stationary monocular camera (12), and preferably deriving from the calibration parameters a pose, preferably comprising a position and orientation of the stationary monocular camera (12) and / or an intrinsic matrix and / or distortion coefficients, of the stationary monocular camera (12) based on the 3D element information (50) and 2D information (48) assigned to the pairs (36, 38).
2. Method according to claim 1 , wherein retrieving a plurality of frames (30) comprises recording a plurality of frames (30) using a mobile camera (22), wherein the mobile camera is preferably a camera arranged at a vehicle (24), preferably a drone (26), a car, a bicycle or a plane, or at a human moving through the scene (10) and providing the plurality of frames (30) for retrieving, whereinpreferably the mobile camera (22) comprises a localization device, for example a receiver for GPS, Galileo, Glonass, Beidou or IRNSS System, an inertial measurement unit and / or a detection unit for detecting Ground Control Points, for acquiring positions or poses of the mobile camera (22) and wherein the method preferably comprises to assign a camera pose to some or each of the frames (30) recorded by the mobile camera (22) taken during recording the respective frame (30).
3. Method according to claim 2, wherein recording a plurality of frames (30) comprises positioning the mobile camera (22) to a defined pose, preferably which is substantially similar to the pose of the stationary monocular camera (12) or which is near or nearest to the pose of the stationary monocular camera (12), recording a frame at the defined pose, provide the frame which is recorded for retrieving and select the retrieved frame (30) as the selected frame (34).
4. Method according to any of the previous claims, wherein selecting at least one of the retrieved frames (30) comprises comparing a plurality or each of the retrieved frames (30) with the recorded frame (14) in order to identify at least one of the retrieved frames (30), which perspective of the scene (10) is most similar to the perspective of the recorded frame (14) or which shows the most image features with the recorded frame (14).
5. Method according to any of the previous claims, wherein computing a 3D model (32) based on the retrieved frames (30) comprises to use an automatic algorithm, like a structure from motion (SfM) algorithm or a simultaneous localization and mapping (SLAM) algorithm, preferably for generating a 3D point cloud and more preferably, for generating a 3D mesh, wherein preferably the algorithm uses camera positions or poses assigned to the retrieved frames (30) and / or receives an input for information of Ground Control Points shown in at least one of the retrieved frames (30) to compute the 3D model of the scene.
6. Method according to any of the previous claims, wherein determining a plurality of pixel pairs (36, 38) comprises to determine for each of a number of arbitrary or all pixel (40, 42) in the recorded frame (14) representing a feature of the scene a pixel (44, 46) in the selected frame (34) corresponding to the same feature of the scene, wherein preferably determining a plurality of pixel pairs (36, 38) is executed using an algorithm like SIFT or LoFTR.
7. Method according to previous claim 6, wherein assigning to a pixel pair (36, 38), for each of the pixel pairs (36, 38), a corresponding 3D element information (50) comprises determining for the respective pixel pair (36, 38) the 3D element of the 3D model (32) corresponding to the pixel (44, 46) of the selected frame (34) and assign the 3D element information (50) of the determined 3D element to the pixel pair (36, 38).
8. Method according to previous claim 7, wherein determining for the respective pixel pair (36, 38) the 3D element of the 3D model (32) corresponding to the pixel (44, 46) of the selected frame (34) is executed by using a pinhole camera model and raytracing.
9. Method according to any of the previous claims 1 to 5, wherein determining a plurality of pixel pairs (36, 38) comprises to identify for each of a number of arbitrary or all 3D elements of the 3D model (32) a pixel (44, 46) in the selected frame (34) representing the 3D element and to determine for each of the identified pixel (44, 46) further representing a feature of the scene (10) a pixel (40, 42) in the recorded frame (14) corresponding to the same feature of the scene (10), wherein determining a plurality of pixel pairs (36, 38) is preferably executed using an algorithm like S2DNet.
10. Method according to claim 9, wherein assigning to a pixel pair (36, 38), for each of the pixel pairs (36, 38), corresponding 3D element information (50) comprises assigning the 3D element information (50) to the pixel pair (36, 38), which corresponds to the 3D element of the identified pixel of the pixel pair (36, 38).11 . Method according to any of the previous claims, wherein determining a plurality of pixel pairs (36, 38) and / or assigning to a pixel pair (36, 38), for each of the pixel pairs (36, 38), 3D element information (50) of a 3D element of the 3D model (32) corresponding to the pixel (44, 46) of the selected frame (34) in the pair (36, 38) and / or calculating parameters, and preferably a pose, the intrinsic matrix and / or the distortion coefficients of the stationary monocular camera (12) based on the 3D element information (50) and 2D information (48) assigned to the pairs (36, 38) is executed using a neural network.
12. Method, preferably according to any of the claims 1 to 11 , comprising providing a stationary monocular camera (12) and preferably calibration parameters of the monocular parameter, comprising for providing a 3D bounding box (18) of at least one object (20) in the scene (10) the further steps, wherein preferably some or all further steps, in particular the further estimating steps, are executed with a neural network:identifying an object (20) in a recorded frame (14) of the scene (10) recorded by the stationary monocular camera (12) and preferably assigning an identifier, preferably a unique track identifier, to the identified object, estimate or determine 2D information (48) of a reference pixel, in particular a centre pixel, of the identified object (20), determine a depth of the object (20), in particular a distance between the point of the object (20) represented by the reference pixel and the stationary monocular camera (12), estimate the 3D bounding box (18) surrounding the object (20), wherein preferably estimating the 3D bounding box (18) comprises to estimate a size, in particular, the length, width and height, of the 3D bounding box, estimate a yaw angle of the bounding box (18) and determine a roll and pitch angle of the bounding box (18) wherein determine the roll and pitch angle of the bounding box (18) relative to the camera position is executed by estimating the roll and pitch angle or by calculating the roll and pitch angle , wherein calculating the roll and pitch angle comprises: determine 3D element information (50), in particular 3D polygon information, of a 3D element, in particular a 3D polygon, in the 3D model (32) corresponding to the reference pixel and calculating the pitch and roll angle of the bounding box (18) based on the determined 3D element information (50), the estimated yaw angle and preferably the calibration parameters, and wherein determine the depth of the object is executed by estimating the depth or by calculating the depth, wherein calculating the depth comprises: determine 3D element information (50), in particular 3D polygon information, of a 3D element, in particular a 3D polygon, in the 3D model (32) corresponding to the reference pixel and calculating the depth based on the determined 3D element information (50) and preferably the calibration parameters, in particular the camera position.
13. Method according to claim 12, wherein the 3D element, in particular 3D polygon of the 3D model (32) corresponding to the reference pixel is determined using a pinhole camera model and raytracing and / or wherein the 3D model (32) is a polygon mesh model, in particular a triangle mesh model, comprising a plurality of polygons defined by the 3D elements, in particular 3D points, of the 3D model (32), and / orwherein the method further comprises to provide output data comprising a sequence of frames of the scene (10) each comprising an identified object (20) represented by the 3D bounding box (18) and preferably an identifier assigned to the identified object (20) and / or the 3D bounding box, and / or wherein before identifying an object (20) in the frame a normalization step is executed.
14. Computer program product comprising instructions, wherein the instructions when executed on a processor, cause the processor to execute the method according to any of the claims 1 to 13.
15. Output data comprising a sequence of frames of a scene (10) comprising at least one bounding box (18) representing an object (20) in the scene (10) wherein the data is produced by a method according to any of claims 1 to 13.