Method of calibrating a camera

CN115485727BActive Publication Date: 2026-09-08KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180032395.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-05-01
Filing Date
2021-04-23
Publication Date
2026-09-08
Estimated Expiration
2041-04-23

AI Technical Summary

Technical Problem

然而,这在室外通常是不可能的,并且相机可能安装在机械性能未知的现有基础设施上,相距几米

Benefits of technology

[0015] According to a second embodiment of the invention, the method is characterized as defined in the characterization portion of claim 2. Its advantage is that it can also further distinguish between one-off or low-frequency interference caused by specific routing reasons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115485727B_ABST
    Figure CN115485727B_ABST
Patent Text Reader

Abstract

A method for calibrating at least one of the six degrees of freedom of all or part of the cameras of a formation positioned for scene capture, the method comprising a step of initial calibration prior to the scene capture. The step comprises creating a reference video frame comprising a reference image of a stationary reference object. During the scene capture, the method further comprises a step of further calibration in which the position of the reference image of the stationary reference object within a captured scene video frame is compared to the position of the reference image of the stationary reference object within the reference video frame and a step of adjusting the at least one of the six degrees of freedom of the plurality of cameras of the formation as needed in order to obtain an improved scene capture after the further calibration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for calibrating at least one of the six degrees of freedom of all or part of the cameras in a formation positioned for scene capture, the method comprising an initial calibration step prior to scene capture, wherein the step includes creating a reference video frame comprising a reference image of a still reference object, and

[0002] During scene capture, the method further includes:

[0003] Further calibration steps include comparing the position of the reference image of the still reference object within the captured scene video frame with the position of the reference image of the still reference object within the reference video frame, and

[0004] The steps involve adjusting at least one of the six degrees of freedom of multiple cameras in a formation to obtain improved scene capture after further calibration.

[0005] In this patent application, the term six degrees of freedom refers to three possible translational movements and three possible rotational movements of the camera position. It is important to emphasize that, in this patent application, the camera's degrees of freedom are intended to cover not only physical translational or rotational movements of the camera, but also virtual translational or rotational movements, where the camera does not translate or rotate, but performs a similar action (with similar results) on the image generated by the camera. Therefore, in the case of virtual (electronic, digital) control of the camera, adjusting the six degrees of freedom of the camera is actually adjusting one or more of the three possible camera position parameters and / or one or more of the three possible rotational parameters. An image acquired by a given camera can, for example, be spatially transformed such that it corresponds to a virtual camera slightly rotated compared to the physical camera. This concept is also relevant given virtual view synthesis and maintaining camera calibration. Combinations of physical and virtual movements, such as physical translational movements and virtual rotational movements, are also not excluded.

[0006] The present invention also relates to camera formation and computer programs in connection with the above-described calibration method.

[0007] EP0848884A1 discloses an apparatus for automatic electronic replacement of billboards in video images, including an automatic camera orientation measurement device comprising a motion measurement module operable to measure the field of view of a TV camera relative to a known reference position. The apparatus includes an image processing module for processing video signals generated by the TV camera, wherein the processing module includes a calibration module for periodically and automatically calibrating the motion measurement module.

[0008] US2003210329A1 discloses a multi-camera video system that may include multiple cameras located around a stadium, playing field, or other location. The cameras are remotely controlled in a master-slave configuration. The multi-camera video system also includes a method for calibrating the system. Background Technology

[0009] In the prior art, camera arrays are known, for example, rigs with eight cameras mounted on them. This enables their use in various applications, such as virtual and / or augmented reality recording systems (often referred to as VR / AR applications). Such VR / AR devices can be fixed or portable (handheld). The cameras on such devices are typically pre-calibrated in a factory (or laboratory) with respect to certain parameters (e.g., the focal length of the lens). These parameters are usually fixed thereafter and do not require further calibration by the user of the VR / AR device. However, before actual use, the user must typically perform an initial calibration related to the actual scene setup. Thus, one or more of the camera's six degrees of freedom are adjusted if necessary. After this, the VR / AR device is ready to begin scene recording. Crucially, during recording, any degree of freedom changes only when desired, such as when the camera rotates to another part of the scene. Therefore, during recording, the camera's translational or rotational movements are unlikely to be (significantly) disturbed, as otherwise accurate VR / AR recording could then be suppressed by annoying artifacts. Typically, the equipment is relatively small, with a relatively small number of cameras, and in most cases, recording is done inside houses or buildings, thus minimizing interference over larger areas. However, the chance of interference is higher, especially with portable devices. This problem becomes even more pronounced when the equipment is outdoors and subject to vibrations from passing trucks or wind. It becomes even more challenging in situations requiring large formations of cameras, such as in stadiums like football fields. This is because the distance between cameras is usually greater, and they are no longer physically connected by rigid structures. For example, for small camera setups (small 3D surround-view effects), cameras can be mounted, for example, at 5cm intervals on 30mm x 8mm thick aluminum strips. However, this is often impossible outdoors, where cameras may be mounted on existing infrastructure with unknown mechanical properties, spaced several meters apart. This is because even very small deviations in camera position or orientation can then produce noticeable artifacts. Summary of the Invention

[0010] The purpose of this invention is to overcome or at least reduce the impact of camera interference during scene recording.

[0011] According to a first embodiment of the invention, the method is characterized as defined in the characterization portion of claim 1. One advantage is that if one or more of the cameras become (very) uncalibrated, it is possible to then immediately begin further "online" calibration so that scene recording is not interrupted, thus repeating an initial calibration similar to the one performed during setup before scene recording begins. Another advantage is that further calibration can be performed in time before interference on the cameras becomes large enough to cause substantial and unpleasant artifacts. While it is preferred that all cameras create reference video frames including reference images of a stationary reference object, it is also possible that only a portion of the cameras do this.

[0012] Further calibration has another advantage if it is performed through virtual control of the camera, thus virtually adjusting the camera's degrees of freedom—that is, transforming a reference image of a stationary reference object within a captured scene video frame back to the same position and orientation as in the reference video frame. This is because such calibration can be performed much faster than if it were performed through mechanical control of the camera (electronic control is generally faster than mechanical control). To avoid unwanted external areas in the (adjusted) calibration video frame, the video frame can be rescaled and cropped. Techniques for rescaled and cropped frames are well-known and are described, for example, in patent US 7733405B2 by Van Dyke et al., dated June 8, 2010.

[0013] If further calibration is performed via the camera's mechanical controls, the advantage is that the aforementioned rescaling and cropping techniques are unnecessary. Another advantage of mechanical control is that the calibration range can be greater than in the case of electronic control. In principle, a separate (additional) servo system can be applied to mechanical control, but mechanical control can also be performed by servo systems that are typically already available for use in general camera control.

[0014] Of course, consistent with the previous definition of utilizing camera degrees of freedom, camera adjustments are also intended to cover both physical and / or virtual adjustments. The result of step B can also be that more than one of the six degrees of freedom is adjusted. For example, adjustments to the camera in the lateral direction, as well as adjustments to both yaw and roll. The values ​​in step C can also include signs (positive or negative) to make it clear, for example, in the above example, whether the lateral direction is defined as left or horizontal. The timestamps in step C do not need to reference actual time, such as national local time or Greenwich Mean Time. Only timestamps relative to any (free) determined time are needed. The classification of different root causes of camera interference is a highly valuable feature of this invention. These root causes can be recorded, and (later) this record can be used to understand how to improve the camera's position in the formation or where to better position the entire formation, etc.

[0015] According to a second embodiment of the invention, the method is characterized as defined in the characterization portion of claim 2. Its advantage is that it can also further distinguish between one-off or low-frequency interference caused by specific routing reasons.

[0016] According to the third and fourth embodiments of the invention, the method is characterized as defined in the third characterization portion of claim 3 and the fourth characterization portion of claim 4, respectively. This is advantageous when not all degrees of freedom need to be adjusted for correction. For example, in certain settings, interference can primarily cause deviations in camera orientation. The advantages are faster and easier control of camera adjustment and reduced data volume, which is particularly advantageous when data is transmitted, for example, to a client system.

[0017] According to a fifth embodiment of the invention, the method is characterized as defined in the characterization portion of claim 5. This has the advantage that unnecessary further calibration is not performed when, for example, the orientation deviation of the camera is so small that it does not cause noticeable artifacts in the image.

[0018] According to a sixth embodiment of the invention, the method is characterized as defined in the characterization portion of claim 6. This is particularly advantageous when the formation comprises a large number (e.g., >50) of cameras. Then, for example, it can be decided that, for a second portion of cameras that would require correction by further calibration (translation and / or rotation adjustments), instead of calibrating them using said further calibration, the required correction is performed using information from closely (but not necessarily directly) adjacent cameras. This has the advantage of increasing the (total) calibration speed of the camera formation and reducing the amount of data, which is particularly advantageous, for example, when transmitting data to a client system. For example, consider a camera array comprising 100 cameras distributed along the four sides of a rectangular sports field. After initial calibration of all cameras has been performed, all 100 cameras can be monitored by selecting two endpoint cameras on each side (a total of eight cameras). If, for one side, one of the two monitoring cameras is no longer calibrated, all intermediate cameras can be compared to that camera and recalibrated based on the difference.

[0019] The seventh and eighth embodiments of the invention are characterized as defined in the characterization portions of claims 7 and 8, respectively. The more uniform the distribution, the more accurate the recording. In most cases, interpolation is sufficient, but extrapolation may be advantageous, particularly near the endpoints of the formation.

[0020] However, it should be emphasized that optimal recording quality can be expected if alternative calibration is not applied (and therefore, further calibration is applied instead).

[0021] According to a ninth embodiment of the invention, the method is characterized by autonomously or continuously repeating the initial calibration at specific events. While initial calibration should ideally be avoided as much as possible (preferably performed only once before scene recording), sometimes it becomes insufficient to perform calibration solely through further calibration and / or alternative calibration. This occurs when the camera becomes too uncalibrated. This event can be detected manually or automatically. For example, a natural break during a football match is also a good time to perform initial calibration. However, in an attempt to avoid such occurrences (or events), the initial calibration can also be repeated continuously at a low repetition frequency, preferably autonomously. For example, reference video frames must be updated because stationary reference objects gradually change over time, for example, by changing the lighting environment causing shadows on the stationary objects.

[0022] According to the invention, camera formation is defined in claim 10, which corresponds to claim 1. It should be emphasized that other disclosed methods of the invention, particularly those defined in claims 2-9, can also be advantageously applied to camera formation.

[0023] The present invention also includes a computer program comprising computer program code units that, when run on a computer, are adapted to implement any method of the present invention, particularly the method as defined in claims 1-9. Attached Figure Description

[0024] To better understand the invention and to more clearly illustrate how it can be practiced, reference will now be made to the accompanying drawings by way of example only, in which:

[0025] Figure 1 Two linear arrays of eight cameras are shown, each array positioned along the side of a football field;

[0026] Figure 2 Example images of the leftmost and rightmost cameras in an eight-camera array, as seen from one side of the football field, are shown.

[0027] Figure 3 A multi-camera system for generating 3D data is shown;

[0028] Figure 4 Pairwise correction for cameras used in an 8-camera linear array is shown;

[0029] Figure 5 The set of reference feature points and the set of corresponding feature points in a 2D image plane are schematically shown;

[0030] Figure 6 A schematic diagram of the calibration is shown; and

[0031] Figure 7 Method steps for classifying various patterns identified as indicating different root causes of camera noncalibration are shown. Detailed Implementation

[0032] The invention will be described with reference to the accompanying drawings.

[0033] It should be understood that the detailed description and specific examples are for illustrative purposes only and not for limiting the scope of the invention. These aspects and advantages will become more readily apparent from the following description, claims, and drawings.

[0034] Figure 1 The diagram shows a camera formation with two linear arrays of eight cameras, each array positioned along the side of a football field. The capturing cameras (those that physically exist) are represented by open triangles. Virtual cameras are indicated by filled triangles. Virtual cameras do not actually exist and are created and exist only electronically and in the software. After calibration and depth estimation, a virtual view can be synthesized for the positions between the capturing cameras or even for positions moved forward into the field of play. In many cases, especially on large stadiums like football fields, more cameras can be applied, such as 200 cameras. The cameras do not need to be aligned in a (virtual) straight line. For example, it can be advantageous to place the cameras along the entire side of the football field, where cameras near the edges of the formation (closer to the left / right goals) are positioned closer to the center of the field of play and also more oriented in the direction of the center of the field. Of course, the formation can also completely surround the field of play. The formation can also actually be divided into several formations. For example, Figure 1 The actual display shows two formations of eight cameras.

[0035] Figure 2 Example images are shown captured by the leftmost and rightmost cameras in a long linear array of eight cameras positioned along one side of a football field. As can be seen, the perspective relative to the athletes varies significantly from camera 1 to camera 8. The boundary visible at the far end of the football field can be used to define useful image features for calibration. Therefore, this feature can be used as a stationary object to create one or more reference video frames. For example, the grid panels between seats or stairs might be suitable for this purpose.

[0036] Figure 3The system diagram is shown, where the paired steps of image undistortion and correction, and disparity estimation, depend on the initial calibration and calibration monitoring and control, which also includes further calibration. For simplicity, there is only one connecting arrow from the "Undistortion / Correction Block" above to the "Initial Calibration Block," and also only one connecting arrow from the "Initial Calibration Block" to the "Undistortion / Correction Block" above. However, it will be apparent that these connections also exist with reference to all other "Undistortion / Correction Blocks."

[0037] Figure 3 The input is N camera images I i , where i = 1…N. These camera images are individually undistorted to compensate for lens distortion and focal length of each camera. However, the same spatial remapping process is also performed to correct each camera for each camera pair. For example, cameras 1 and 2 form the first pair in the array. Camera pairs are typically positioned spatially adjacent to each other in the array. The correction step ensures that for each pair, the input camera images are transformed such that they are transformed into a stereo pair, where the optical axes are aligned and the rows in the images have the same position. This is essentially a rotation operation and is often referred to as stereo correction. This process is performed in… Figure 4 The diagram in the middle is shown.

[0038] The calibration process typically relies on a combination of the following known algorithms (see, for example, "Wikipedia"):

[0039] 1. A polygonal method for determining the true range of attitude based on laser measurement;

[0040] 2. Combine N scene points with the perspective N points at known 3D positions for calibration;

[0041] 3. A structure with beam regulation derived from motion.

[0042] After initial calibration during installation, the multi-camera system can be used for various purposes. For example, depth estimation can be performed, followed by the synthesis of virtual views to allow spectators of a competition to experience AR / VR. However, the system can also be used to identify the position of an object (e.g., an athlete's foot or knee) at a given moment and determine its 3D position based on multiple views. Note that for the latter application, two views would theoretically be sufficient, but having multiple viewpoints minimizes the chance of the object of interest being occluded.

[0043] Figure 4 Pairwise correction for an 8-camera linear array is shown.

[0044] This correction transforms each image so that the optical axes, indicated by arrows (R1, R1a; R2, R2a; R3, R3a; R4, R4a), become paired parallel, indicated by arrows (G1, G1a; G2, G2a; G3, G3a; G4, G4a), and orthogonal to the line (dashed line) connecting the two cameras. This correction allows for easier parallax estimation and depth calculation.

[0045] Figure 5 The set of reference feature points and the set of corresponding feature points in the 2D image plane are schematically shown. Correspondences are determined via motion estimation and illustrated by dashed lines. As can be seen, corresponding points in new frames are not always found because occlusion may occur, or motion estimation may fail due to image noise.

[0046] Figure 6 A schematic diagram of the calibration is illustrated using an example showing a changing camera (camera 1 in this case). Original features are detected in a reference frame, and their image positions are represented by hollow circles. The corresponding feature points are then determined via motion estimation in the first and second frames. These corresponding points are represented as closed circles. As shown, for frame 1 of both cameras 1 and 2, and frame 2 of camera 2, the points do not change position. However, for camera 1, points 1, 2, and 3 change position. Therefore, it can be concluded that camera 1 is no longer calibrated. Note that for reference point 4, which is only visible in camera 1, no corresponding feature point was ever found. This point is represented as "invalid" and is therefore not considered when calculating the error metric, on which basis we conclude that the camera is no longer calibrated (see the equation below).

[0047] Therefore, 3D scene points are first projected onto a reference frame (represented by hollow circles). Image-based matching is used to estimate the correspondence between these points and other frames (represented by closed circles). The calibration status of a subsequent frame is determined to be OK when a given portion of the corresponding feature point does not change position. This is the case for both camera 1 and camera 2 in frame 1. However, in frame 2, camera 1 exhibits a rotation error because three points are displaced compared to the reference point.

[0048] Figure 7 The steps for classifying and identifying patterns (in the six degrees of freedom of the camera) to indicate different root causes of camera miscalibration are shown.

[0049] In step A, it is analyzed which cameras in the camera formation have been adjusted due to external disturbances (such as wind).

[0050] In step B, for each adjusted camera, it is analyzed which of the six degrees of freedom were adjusted. For example, in step A, it was determined that for the first camera, only the orientation direction "yaw" was adjusted, while for the second camera, the orientation directions "yaw" and "roll" as well as the translation direction "left and right" were all adjusted.

[0051] In step C, the adjustment values ​​and timestamps are determined. For example, for the first camera, the value for "yaw" is 0.01°, and for the second camera, the values ​​for "yaw" and "roll" are 0.01° and 0.02°, respectively, and the value for "left / right" is +12 mm (e.g., +12 mm is 12 mm to the right, while -12 mm would be 12 mm to the left). The timestamps for the two cameras are, for example, 13:54h + 10 seconds + 17 milliseconds (in principle, the timestamps for the two cameras can also be different). Moreover, for orientation, such as "pitch," a symbol can be used specifically for that value, for example, -0.01°, but alternatively, if the 360° representation is fully utilized, the angle value can also be indicated without a symbol; that is, -0.01° can also be indicated as +359.99°.

[0052] In step D, by analyzing the information obtained from steps A, B, and C, various patterns of camera formation are identified along the camera formation. For example, referring to the first timestamp, all adjustment values ​​related to "yaw" are recorded as the first pattern of camera deviation, all adjustment values ​​related to "roll" are recorded as the second pattern of camera deviation, and so on.

[0053] In step E, the identified patterns are classified, thereby indicating different root causes of camera interference in one or more of the six degrees of freedom of the cameras in the formation. Examples of classification are: "wind," "mechanical stress," "temperature effect," etc. These classifications are determined, for example, by knowledge obtained from earlier measurements. For example, the first pattern is compared with all possible classifications. If the corresponding classifications match perfectly, the first pattern does not necessarily have to be a "hit," a high similarity is sufficient. Various (known) pattern recognition methods can be used. Of course, if no "hit" is found at all, then the pattern can be classified as "unknown cause."

[0054] Optionally, in step F, step E considers multiple analysis sessions from steps A, B, C, and D. This provides additional possibilities for pattern classification. For example, if calibration issues return periodically (e.g., for only two cameras in a large camera formation), the classifier could output: "Frequent group interference".

[0055] A suitable error metric for detecting calibration problems is the sum of the distances between corresponding locations of all valid feature points:

[0056]

[0057] Where, x ref,i Let N be the reference image feature points, xi be the matched image feature points, and N be the reference image feature points. 有效 This represents the number of features successfully matched from the reference video frame to the new video frame. A simple test is now available to evaluate this.

[0058] e < T,

[0059] Where T is the preset threshold [pixels]. This threshold is typically set to a range of 0.1 to 5 pixels. If the test fails, the relevant camera's status should be set to "uncalibrated".

[0060] If the calibration test fails for a single camera but not for other cameras, it can be concluded that the localized disturbance only affects that single camera. Under the (realistic) assumption, only the camera orientation has changed, and the camera can be recalibrated in real time.

[0061] First, observe the distance r from the camera to each 3D scene point i. i The coordinates must remain constant because the scene point and camera position will not change. The camera coordinates of point i can be calculated using the following formula:

[0062]

[0063]

[0064]

[0065] Where f is the focal length [pixels], and u, v are image coordinates [pixels]. This gives the calibration point in both world space and camera space. The well-known Kabsch algorithm (see, for example, "Wikipedia") can then be used to estimate the rotation matrix, which takes a point in world space and transforms it to camera space. When the reprojection error e decreases after projecting the 3D scene points into the image, the new rotation matrix is ​​only accepted as a portion of the updated view matrix. Note that even then, the camera state may still be rated as "uncalibrated." This occurs when e ≥ T.

[0066] Detecting slow trends in camera orientation is also useful. If such a trend exists, then when the reprojection error e exceeds the threshold T, the root cause is likely "mechanical stress or temperature".

[0067] A linear array of five cameras can show the following spatial pattern of error:

[0068] e1<T, e2<T, e3≥T, e4≥T, e5<T

[0069] It can be seen that the problem remains localized (at cameras 3 and 4). It is still possible that "local interference" (near cameras 3 and 4) is causing the problem. Real-time recalibration of both cameras is sufficient to take action. If only these two cameras experience the problem periodically, the classifier will output: "Recurring swarm interference." The appropriate action then is to exclude image data from these two cameras and base view synthesis and / or 3D analysis on the other cameras that are still in calibration. It is possible that the "recurring swarm interference" is related to a specific location within the stadium. By recording the reprojection error throughout the game, it is possible to analyze which spatial locations in the stadium systematically cause the greatest instability. For the next game, the cameras can be placed in different locations.

[0070] Strong winds, cheering crowds (stadium vibrations), or passing trucks can all alter camera orientation. These root causes typically affect all cameras, but may only temporarily affect their orientation. These types of disturbances may temporarily increase the reprojection error for all cameras, but due to the resilience of the system, the reprojection error will decrease again (e.g., when the cheering has stopped). In this situation (where the reprojection error exceeds the threshold by a very small amount for all cameras), we can choose not to take any action; that is, we tolerate the error and pass the information downstream to the view synthesis and analysis modules. Performing the same analysis on the recorded data after the event will again provide valuable insights into whether mechanical changes to the camera configuration should be made (e.g., increasing the quality of each camera or changing the mounting points on the stadium's existing infrastructure). Spatiotemporal analysis (e.g., Fourier transform, spectral analysis) can be used to classify the events that occurred.

[0071] While the spatiotemporal patterns in the reprojection error may be sufficient to provide information for state classification, observing the camera's orientation change signal as a function of space and time may offer even more information. This can reveal specific rotational movements experienced by the camera assembly. For example, in cases where mechanical flexibility is greater in the vertical direction, the motion may be primarily around the horizontal axis. Therefore, such data can provide insights into improving camera mounting for future captures.

[0072] Some applications require each camera in a camera array to be oriented in real-time toward an action occurring on the field. The control algorithm then continuously controls the servo system of each camera to optimally guide each camera toward the action. In this case, the rotation matrix estimate must be updated in real-time for each camera. State classification becomes more complex in this context. To classify states, camera orientation is continuously predicted using servo control signals. The orientation prediction is then used to update the view matrix, and the reprojection error is evaluated for the updated view matrix. External influences such as wind can still be classified via reprojection error. However, a translation-tilt system may also be used to correct errors introduced, for example, by wind. In this case, the control signal is generated to not only point the camera toward the action but also to have a component that corrects for wind in real-time. The estimation and control of wind effects can be accomplished using a Kalman filter (see, for example, Wikipedia).

[0073] It should be emphasized that this invention is not limited to sporting events, but can be applied to many different applications. For example, it can be used in shopping malls, markets, the automotive industry, etc. It is also not limited to outdoor activities, but can be useful for, for example, real estate agents, medical staff in hospitals, etc.

[0074] Referring to the automotive industry, this invention can also be applied to, for example, automobiles. For instance, a camera array can be mounted on the rear bumper. These cameras can replace interior and / or exterior rearview mirrors. For a vehicle with a built-in rear-view camera array, the road is constantly moving in the image. However, if the vehicle's speed is known, the road portion of the image can be made stationary. This situation is then the same as a camera array observing a stationary motion field. Vehicle vibrations, such as those caused by gravel roads, will cause minute changes in camera rotation at high time frequencies. These can be detected and corrected using appropriate methods. Road type identification / classification (flat highway vs. gravel road) will aid in the real-time calibration process. Initial calibration begins when the vehicle turns into another road. Once the vehicle is traveling on a new road, the reference frame is continuously updated (tracked) as the new road becomes visible during travel.

[0075] Referring to the tenth embodiment (claim 10), step F: In the context of the vehicle, the (previously mentioned) low-frequency events are images of the road moving beneath the vehicle, while the (previously mentioned) high-frequency events are vibrations caused by gravel roads or road bumps. Note that gravel roads and road bumps are also expected to cause relative orientation / position changes between cameras mounted on the bumper. When driving on a flat highway, these relative changes are expected to be small. Since the vehicle bumper will be flexible, calibration can be temporarily stopped (due to road bumps) and then automatically resumed (due to the flexible structure of the vehicle bumper).

[0076] It should also be noted that all types of camera formations can be used; for example, drones could also be used, with each drone containing one or more cameras. A major advantage of drones is that the relative positions of the cameras can be adjusted very easily. The term "formation" is considered similar to the terms "arrangement" or "group."

[0077] In many VR / AR applications (using, for example, computers or VR goggles), this invention can also be used for both 2D and 3D applications.

[0078] The fact that certain measures are referenced in mutually different dependent claims does not mean that a combination of these measures cannot be used advantageously.

[0079] Based on a study of the accompanying drawings, the disclosure, and the appended claims, those skilled in the art can understand and implement variations of the disclosed embodiments in practicing the claimed invention. In the claims, the word "comprising" does not exclude other elements or steps, and the word "a" does not exclude a plurality.

Claims

1. A method for calibrating at least one of the six degrees of freedom of all or part of a formation of cameras positioned for scene capture, the method comprising an initial calibration step prior to the scene capture, wherein, The step includes creating a reference video frame, the reference video frame including a reference image of a still reference object, and During scene capture, the method further includes: The further calibration step includes comparing the position of the reference image of the still reference object within the captured scene video frame with the position of the reference image of the still reference object within the reference video frame, and The step of adjusting at least one of the six degrees of freedom of the multiple cameras in the formation to obtain improved scene capture after the further calibration is characterized in that the method further includes: Step A: Analyze which cameras in the formation have been adjusted. Step B: Analyze which of the six degrees of freedom of the camera have been adjusted. Step C: Determine the corresponding value and timestamp for the adjustment. Step D: Identify various formation patterns along the camera by analyzing the information obtained from steps A, B, and C. Step E: Classify the identified patterns, whereby the classification indicates different root causes of camera interference in one or more of the six degrees of freedom of the cameras in the formation.

2. The method according to claim 1, characterized in that, The method further includes: Step F: In step E, consider the multiple analysis sessions of steps A, B, C, and D.

3. The method according to claim 1 or 2, characterized in that, The "at least one" means "three".

4. The method according to claim 3, characterized in that, The three degrees of freedom are yaw, roll, and pitch.

5. The method according to claim 1 or 2, characterized in that... The adjustment step is performed only when at least one 2D image position in the set of 2D feature point image positions has changed, such that the magnitude of the change exceeds a threshold.

6. The method according to claim 1 or 2, characterized in that, Only the first part of the camera is calibrated using the initial calibration and the further calibration, and the remaining second part is calibrated using only one of the following: Only the initial calibration and alternative calibration, or Alternative calibration only The alternative calibration for cameras belonging to the second part is performed by comparing them with one or more neighboring cameras of the first part.

7. The method according to claim 6, characterized in that, The second part of the camera is evenly distributed along all the cameras in the formation.

8. The method according to claim 7, characterized in that, The comparison with the one or more neighboring cameras is performed using interpolation and / or extrapolation techniques.

9. The method according to claim 1 or 2, characterized in that, The initial calibration may be repeated autonomously or continuously at specific events.

10. A camera formation for scene capture positioning, wherein, At least one camera has at least one of six degrees of freedom, wherein the at least one camera is initially calibrated prior to scene capture, characterized in that, during the initial calibration, a reference video frame is created, the reference video frame including a reference image of a still reference object, and characterized in that, during scene capture, the at least one camera is further calibrated, wherein the position of the reference image of the still reference object within the captured scene video frame is compared with the position of the reference image of the still reference object within the reference video frame, and wherein the at least one of the six degrees of freedom of the at least one camera in the formation is adjusted to obtain an improved captured scene.

11. A computer program comprising computer program code units, wherein when the computer program is run on a computer, the computer program code units are adapted to implement the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Method and apparatus for automatic electronic replacement of billboards in a video image

    EP0848884A1

  • Video system and methods for operating a video system

    US20030210329A1

  • Apparatus, method and system

    CN103106654A