Image Ground Levelling Method, Device, Electronic Equipment and Medium for Multi-Camera Calibration
Through detection of key points of human body, especially characteristic points of human eye, the problems of large error and low accuracy in traditional camera calibration methods are solved, and simple and efficient camera calibration and ground leveling are achieved, which are suitable for various working environments.
Patent Information
- Application Number
- CN202210369700.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-08
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-04-08
AI Technical Summary
Traditional camera calibration methods require additional calibration objects, resulting in large errors, low accuracy, and susceptible to environmental and viewing angles, making it difficult to accurately determine the ground plane.
The human body key point detection, especially the human eye as feature points, is used to establish a mapping relationship between image coordinates and world coordinates through multi-camera calibration, calculate the rotation matrix for ground leveling, and avoid the use of additional calibrators.
It realizes simple and efficient camera calibration in various working environments, ensures that the world coordinate system is consistent with the ground, improves calibration accuracy, and simplifies the ground leveling process.
Smart Images

Figure CN114862960B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of visual measurement technology, and more particularly to an image ground leveling method, apparatus, electronic device, and medium for multi-camera calibration. Background Art
[0002] With the development of computer vision technology, its applications are becoming more and more widespread. Generally speaking, in computer vision applications, it is necessary to determine the mutual relationship between the three-dimensional geometric position of a certain point of a spatial object and its corresponding point in the image, that is, to establish a mapping relationship between the image coordinates and the spatial position coordinates. The process of solving this mapping relationship is camera calibration. In computer vision applications, the calibration of camera parameters is a very crucial link, and the accuracy of its calibration results and the stability of the algorithm directly affect the accuracy of the results generated by computer vision. Therefore, doing a good job in camera calibration is a prerequisite for doing subsequent work well. Currently, traditional camera calibration methods include direct linear transformation calibration method, Zhang Zhengyou calibration method, etc. Traditional camera calibration requires the use of a calibration object with known dimensions, and by establishing the correspondence between the known coordinate points on the calibration object and the image points, the internal and external parameters of the camera model are obtained using an algorithm.
[0003] Traditional camera calibration methods can establish a mapping relationship between pixel coordinates and world coordinates, but because the selection methods of the world coordinate system vary, the world coordinates and the spatial position coordinates may not be consistent. A typical problem is that if the xoy plane of the world coordinate system is not parallel to the ground, the z coordinate calculated in computer vision applications is not the height information. In some practical applications, such as judging whether the standing or sitting posture is straight in human pose estimation, it will lead to inaccurate judgment. Specifically, when using a three-dimensional calibration object, such as a three-dimensional calibration frame, generally the plane where the 4 calibration balls at the bottom are located is used as the xoy plane of the world coordinate system. If the manufacturing process is slightly inaccurate, or the frame deforms due to long-term use, it will cause deviation of the xoy plane. When using a planar calibration object, the world coordinate system is based on the calibration board, so it is necessary to calibrate the calibration board parallel or perpendicular to the ground for the first time. It is difficult to ensure perpendicularity, and when the camera is looking straight ahead, if the calibration board is placed on the ground, the tilt angle in the image is too large and it is difficult to detect.
[0004] In addition, to obtain the plane where the ground is located, at least 3 position coordinate points on the ground need to be found. In the process of camera calibration, this requires manually marking or detecting at least 3 feature points in the images of multiple cameras. This first requires additional calibration objects on the ground for manual marking or algorithm detection, and the feature points in the images of multiple cameras are unique for easy matching. This method of manual marking is time-consuming and laborious and has large errors, and algorithm detection is greatly affected by the environment, viewing angle, etc., and requires special detection and matching algorithms. Summary of the Invention
[0005] In view of this, the embodiments of the present disclosure provide an image ground leveling method, device, electronic device and medium for multi-camera calibration, so as to solve the technical problems in the prior art that the image ground leveling method using an additional calibration object for multi-camera calibration has large errors, low accuracy, is easily affected by factors such as environment and perspective, and requires special detection and matching algorithms.
[0006] In view of the above-mentioned defects in the prior art, the inventive concept of the technical solution of this application lies in that, in the absence of an external calibration object, the human body itself can be used as a calibration object. Since ground leveling does not require the known size of feature points, but only requires being parallel to the ground. The human body key points are unique, and the algorithms are mature and easy to detect. Only need to design rules to select the key points parallel to the ground. With the development of deep learning technology, the human body key point detection technology has become more and more mature. For some key points such as face key points, its detection accuracy is even no less than that of humans. The difficulty of detecting different human body key points is different. The detection of key points such as the waist and legs is significantly more difficult than the detection of face key points, and the area of such key points is relatively large. It is difficult to determine whether the same feature point is detected in multiple cameras, and it is difficult to ensure the detection accuracy. Therefore, the inventor chooses the human eyes as the feature points. Because the human eye area is small and the features are obvious, it is easy to detect and has high detection accuracy. When a person stands upright, the heights of both eyes from the ground are basically the same; considering that the head posture of a person may change, therefore, in multiple detections, those with relatively consistent head postures are taken, and both eyes or a single eye are used as feature points. Multiple feature points are used to fit a plane, which can be approximately considered parallel to the ground. Then the rotation matrix from the xoy plane to the fitted plane in the world coordinate system can be obtained, so as to perform the transformation of the world coordinate system.
[0007] To achieve the above object, in the first aspect of the embodiments of the present disclosure, an image ground leveling method for multi-camera calibration is provided, and the method includes: establishing a mapping relationship from image coordinates to world coordinates during multi-camera calibration;
[0008] Obtaining human body image information and performing face key point detection in the human body image information to obtain human eye key point information and head posture auxiliary key point information;
[0009] Calculating 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information and the head posture auxiliary key point information;
[0010] Obtaining the key points of both eyes or a single eye of the human body with consistent postures to construct a fitted plane;
[0011] Calculating the rotation matrix from the xoy plane to the fitted plane in the world coordinate system, and multiplying the rotation matrix by the mapping relationship to obtain the mapping relationship from image coordinates to the leveled world coordinates.
[0012] In a second aspect of the embodiments of the present disclosure, an image ground leveling device for multi-camera calibration is provided, including:
[0013] A building unit configured to establish a mapping relationship from image coordinates to world coordinates during multi-camera calibration;
[0014] A first acquisition unit configured to acquire human body image information and perform face key point detection in the human body image information to obtain human eye key point information and head pose auxiliary key point information;
[0015] A first calculation unit configured to calculate 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information, and the head pose auxiliary key point information;
[0016] A second acquisition unit configured to acquire key points of both eyes or a single eye of a human body with consistent postures to construct a fitting plane;
[0017] A second calculation unit configured to calculate a rotation matrix from the xoy plane in the world coordinate system to the fitting plane, and the rotation matrix multiplied by the mapping relationship is the mapping relationship from image coordinates to the leveled world coordinates.
[0018] In a third aspect of the embodiments of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0019] In a fourth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0020] One of the embodiments of the above various embodiments of the present disclosure has the following beneficial effects: Through the image ground leveling method for multi-camera calibration of the present application, it is applicable to various camera calibration methods. It makes the world coordinate system consistent with the ground, facilitating subsequent various evaluations. Without using additional calibration objects and without manual marking, automatic ground leveling after camera calibration is performed through human key point detection, which is simple and easy to implement and applicable to various working environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.
[0022] Figure 1Schematic flowchart of some embodiments of an image ground leveling method based on multi-camera calibration according to the present disclosure;
[0023] Figure 2 Schematic structural diagram of some embodiments of an apparatus for an image ground leveling method based on multi-camera calibration according to the present disclosure;
[0024] Figure 3 Schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0025] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should clearly understand that the present disclosure can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present disclosure.
[0026] A method for image ground leveling based on multi-camera calibration according to an embodiment of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0027] Figure 1 Schematic flowchart of an image ground leveling method based on multi-camera calibration provided by an embodiment of the present disclosure. As Figure 1 shown, the method for image ground leveling based on multi-camera calibration includes the following steps:
[0028] Step S101, establish a mapping relationship from image coordinates to world coordinates during multi-camera calibration;
[0029] Step S102, obtain human body image information and perform face key point detection in the human body image information to obtain human eye key point information and head pose auxiliary key point information.
[0030] In some embodiments, the obtaining of the human body image information is performed by multiple cameras taking synchronous timed photos at preset interval time periods. The human body image information includes image information of the human body in different positions and multiple different time periods; the number of the multiple different time periods is at least three. The head pose auxiliary key point information includes auxiliary key point information for head pose estimation within all preset time periods, such as nose key points.
[0031] Step S103, calculate 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information, and the head pose auxiliary key point information.
[0032] Step S104, obtain key points of the two eyes or a single eye of a human body with consistent poses to construct a fitting plane;
[0033] In some embodiments, the consistent postures include time points at which the normal vector directions of the planes constructed by the key points of both eyes and the auxiliary key points of the human body are relatively consistent within all preset time periods.
[0034] Step S105: Calculate the rotation matrix from the xoy plane in the world coordinate system to the fitting plane, and the mapping relationship from the image coordinates to the leveled world coordinates is the mapping relationship obtained by multiplying the rotation matrix by the said mapping relationship.
[0035] As an example, in some embodiments, multi-camera calibration (the number of cameras = N) is performed to establish the mapping relationship M from image coordinates to world coordinates: 2D1…2D N →3D. The multi-cameras are set to take synchronized photos at regular intervals, and images I of a person standing straight at multiple different positions are taken t,n , where t is the time and n is the camera number. Face key point detection is performed on the image I t,n to obtain the key points of the human eyes and the auxiliary key points for head pose estimation, such as LEYE_2D t,n , REYE_2D t,n and NOSE_2D t,n . According to the camera calibration mapping relationship M and 2D t,1 ...2D t,N , the 3D key points in the world coordinate system are obtained, such as LEYE_3D t , REYE_3D t and NOSE_3D t . The head pose is estimated based on the facial key points to obtain HEADPOSE t . Among all the time points, k time points with consistent postures are selected, where k >= 3, denoted as t1…tk. The key points of both eyes or a single eye with consistent postures, such as LEYE_3D t1 ...LEYE_3D tk , are used to fit the plane LEYE_SURFACE. The rotation matrix R from the xoy plane in the original world coordinate system to the eye fitting plane LEYE_SURFACE is calculated, and then R⊙M is the mapping relationship from the image coordinates to the leveled world coordinates.
[0036] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated one by one here.
[0037] The following is an embodiment of the apparatus of the present disclosure, which can be used to execute the method embodiment of the present disclosure. For the details not disclosed in the apparatus embodiment of the present disclosure, please refer to the method embodiment of the present disclosure.
[0038] Figure 2 is a schematic diagram of an image ground leveling device for multi-camera calibration provided by an embodiment of the present disclosure. As Figure 2As shown in the figure, the image floor leveling device for multi-camera calibration includes: a building unit 201, a first acquisition unit 202, a first calculation unit 203, a second acquisition unit 204, and a second calculation unit 205. The building unit 201 is configured to establish a mapping relationship from image coordinates to world coordinates during the multi-camera calibration process; the first acquisition unit 202 is configured to acquire human body image information and perform face key point detection in the human body image information to obtain human eye key point information and head pose auxiliary key point information; the first calculation unit 203 is configured to calculate 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information, and the head pose auxiliary key point information; the second acquisition unit 204 is configured to acquire human body binocular or monocular key points with consistent poses to construct a fitting plane; the second calculation unit 205 is configured to calculate a rotation matrix from the xoy plane in the world coordinate system to the fitting plane, and the rotation matrix multiplied by the mapping relationship is the mapping relationship from image coordinates to the leveled world coordinates. It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present disclosure.
[0039] Figure 3 is a schematic diagram of the computer device 3 provided by the embodiment of the present disclosure. As Figure 3 shown, the computer device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps in each of the above method embodiments are implemented. Alternatively, when the processor 301 executes the computer program 303, the functions of each module / unit in each of the above device embodiments are implemented.
[0040] Exemplarily, the computer program 303 can be divided into one or more modules / units. One or more modules / units are stored in the memory 302 and executed by the processor 301 to complete the present disclosure. One or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program 303 in the computer device 3.
[0041] The computer device 3 can be a desktop computer, a notebook, a palm computer, a cloud server, and other computer devices. The computer device 3 can include, but is not limited to, the processor 301 and the memory 302. Those skilled in the art can understand that Figure 3This is only an example of the computer device 3, which does not constitute a limitation on the computer device 3. It may include more or fewer components than those shown in the figure, or combine certain components, or different components. For example, the computer device may also include input / output devices, network access devices, buses, etc.
[0042] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0043] The memory 302 may be an internal storage unit of the computer device 3. For example, the hard disk or memory of the computer device 3. The memory 302 may also be an external storage device of the computer device 3. For example, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 3. Further, the memory 302 may also include both the internal storage unit and the external storage device of the computer device 3. The memory 302 is used to store computer programs and other programs and data required by the computer device. The storage 302 may also be used to temporarily store data that has been output or will be output.
[0044] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used for illustration. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.
[0045] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not described or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0046] Those of ordinary skill in the art will realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this disclosure.
[0047] In the embodiments provided by this disclosure, it should be understood that the disclosed device / computer device and method can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. Multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0048] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0049] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0050] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present disclosure, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice within the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0051] The above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present disclosure, and should all be included in the protection scope of the present disclosure.
Claims
1. An image ground leveling method for multi-camera calibration, characterized in that Including: Establishing a mapping relationship from image coordinates to world coordinates during multi-camera calibration; Obtaining human body image information and performing face key point detection in the human body image information to obtain human eye key point information and head pose auxiliary key point information; The human body image information is an image with the heights of both eyes from the ground being basically the same; Calculating 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information, and the head pose auxiliary key point information; Obtaining the time points when the normal vector directions of the planes constructed by the human body's binocular key points and auxiliary key points are relatively consistent within all preset time periods, and constructing a fitting plane with the human body's binocular or monocular key points; Calculating the rotation matrix from the xoy plane in the world coordinate system to the fitting plane, and multiplying the rotation matrix by the mapping relationship to obtain the mapping relationship from image coordinates to the leveled world coordinates.
2. The method for leveling an image ground by multi-camera calibration according to claim 1, wherein, The obtaining of the human body image information is by multiple cameras taking synchronous timed photos at preset interval time periods.
3. The image ground leveling method for multi-camera calibration according to claim 1, wherein The human body image information includes image information of the human body at different positions and within multiple different time periods.
4. The image ground leveling method for multi-camera calibration according to claim 1, characterized in that The head pose auxiliary key point information includes head pose auxiliary key point information with the same human body pose within all preset time periods.
5. The method for image ground leveling by multi-camera calibration according to claim 3, wherein, The number of the multiple different time periods is at least three.
6. An image ground leveling device for multi-camera calibration, characterized in that, Including: A establishing unit configured to establish a mapping relationship from image coordinates to world coordinates during multi-camera calibration; A first obtaining unit configured to obtain human body image information and perform face key point detection in the human body image information to obtain human eye key point information and head pose auxiliary key point information; The obtaining of the human body image information is an image with the heights of both eyes from the ground being basically the same; A first calculating unit configured to calculate 3D key points in the world coordinate system based on the mapping relationship, the human eye key point information, and the head pose auxiliary key point information; A second obtaining unit configured to obtain the time points when the normal vector directions of the planes constructed by the human body's binocular key points and auxiliary key points are relatively consistent within all preset time periods, and construct a fitting plane with the human body's binocular or monocular key points; A second calculating unit configured to calculate the rotation matrix from the xoy plane in the world coordinate system to the fitting plane, and multiplying the rotation matrix by the mapping relationship to obtain the mapping relationship from image coordinates to the leveled world coordinates.
7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Image calibration method and device and electronic equipment
CN111105467A
Multi-camera system calibration method based on pedestrian head recognition
CN111667540A