Method and apparatus for determining extrinsic parameter of camera inside vehicle, electronic device, and storage medium
By setting a camera on the vehicle's steering column and using the coordinates and perimeter of interior points in the video frames to construct a linear relationship, the camera's extrinsic parameters are automatically calculated. This solves the need for manual operation of calibration boards in existing technologies and realizes the automated determination of camera extrinsic parameters.
Patent Information
- Application Number
- PCT/CN2024/121538
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-15
- Filing Date
- 2024-09-26
- Publication Date
- 2026-02-19
AI Technical Summary
In the existing technology, the re-determination of the extrinsic parameters of the vehicle's internal camera requires manual operation of the calibration board, which is highly demanding and not easy to automate.
By setting a camera on the steering column of the vehicle, the center point coordinates and perimeter of the target interior point are determined using the video frames captured by the camera. A linear relationship is constructed to automatically calculate the camera's extrinsic parameters, including the first rotation matrix and the first translation matrix, reducing operational requirements.
It enables automatic determination of camera extrinsic parameters without human intervention, reducing operational requirements and improving the automation level of camera extrinsic parameter determination.
Smart Images

Figure CN2024121538_19022026_PF_FP_ABST
Abstract
Description
Method and device for determining camera extrinsic parameter in vehicle interior, electronic device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] The present disclosure claims priority to the Chinese patent publication with the application number 202411118585.7 and the title "Method and device for determining camera extrinsic parameter in vehicle interior, electronic device and storage medium" filed on August 15, 2024 with the China National Intellectual Property Office, the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the technical field of computer, in particular to a method and device for determining camera extrinsic parameter in vehicle interior, electronic device and storage medium. BACKGROUND
[0004] For a smart vehicle with gesture control function, the controllable component in the vehicle that the user currently wants to control can be determined through a line-of-sight tracking function, and then the gesture made by the user is recognized through a gesture recognition function, and then the controllable component is controlled. In the process of determining the controllable component by using the line-of-sight tracking function, each controllable component needs to be converted from the world coordinate system to the camera coordinate system, so as to determine the corresponding controllable component according to the gaze coordinates of the user's line-of-sight in the camera coordinate system. In the process of converting each controllable component from the world coordinate system to the camera coordinate system, the camera extrinsic parameter needs to be used for conversion.
[0005] When the position of the camera moves, the camera extrinsic parameter needs to be determined again, otherwise the determined controllable component will be inaccurate or the controllable component cannot be determined. At present, when the camera extrinsic parameter is determined again, a calibration board needs to be used to complete the determination. In this process, the user's operation needs to meet the use standard of the calibration board, which leads to a relatively high operation requirement when the camera extrinsic parameter is determined again using the calibration board.
[0006] SUMMARY
[0007] Therefore, the embodiments of the present disclosure provide a method and device for determining camera extrinsic parameter in vehicle interior, electronic device and storage medium to reduce the operation requirement when the camera extrinsic parameter is determined again.
[0008] In a first aspect, the embodiments of the present disclosure provide a method for determining camera extrinsic parameter in vehicle interior. A steering column of a steering wheel of a vehicle is provided with a camera. The method comprises:
[0009] When a trigger condition for re-determining camera extrinsic parameters of the camera is met, a center point coordinate and a perimeter of a figure formed by at least one target interior point in a first video frame captured by the camera are determined according to the first video frame, wherein when the number of the target interior points is one, the center point coordinate is a pixel coordinate of the target interior point, and the perimeter is an abscissa coordinate of the target interior point.
[0010] According to a first linear relationship between the center point coordinate and an RT matrix, a second linear relationship between the perimeter and the RT matrix, and a third linear relationship between the center point coordinate and the perimeter, a first transformation relationship between a first rotation matrix and a target interior point element for representing a conversion of a reference camera coordinate system to a current camera coordinate system, and a second transformation relationship between a first translation matrix and the target interior point element for representing the conversion of the reference camera coordinate system to the current camera coordinate system are constructed, wherein the target interior point element includes the center point coordinate and the perimeter, and the RT matrix is composed of a rotation matrix and a translation matrix.
[0011] According to an M matrix, an N matrix, the first transformation relationship, and the second transformation relationship, the first rotation matrix and the first translation matrix are determined, and the first rotation matrix and the first translation matrix are taken as camera extrinsic parameters of the camera, wherein the M matrix is a second translation matrix when a specified pixel point in the first video frame captured by the camera at a current position is transferred to the reference camera coordinate system, and the N matrix is a second rotation matrix when the specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system.
[0012] Optionally, the determining of the center point coordinate and the perimeter of the figure formed by the at least one target interior point according to the first video frame captured by the camera comprises:
[0013] The first video frame is input into a trained interior point positioning model to obtain pixel coordinates of the at least one target interior point;
[0014] The center point coordinate and the perimeter of the figure formed by the at least one target interior point are determined according to the pixel coordinates of the at least one target interior point.
[0015] Optionally, the method further comprises:
[0016] A plurality of first videos captured by the camera during movement of the pipe column are obtained.
[0017] The first video is used to train the interior point positioning model to obtain the trained interior point positioning model.
[0018] Optionally, the moving range of the column is less than or equal to a region formed by the maximum stroke of the column in the up-down movement and the telescopic movement.
[0019] Optionally, each of the first videos comprises a plurality of sub-videos taken at a plurality of discrete positions within the moving range.
[0020] Optionally, the column is moved from one vertex of the trapezoid to another vertex on the opposite side of the edge on which the first vertex is located, and the S-shaped movement is performed during the movement, and the distance between the peak and the valley of the S-shaped movement and the maximum stroke of the column is less than a preset distance.
[0021] Optionally, the different first videos are taken under different lighting environments.
[0022] Optionally, the target interior point comprises a position with a pixel difference greater than a preset pixel difference in the second video frame and / or an acute angle point formed by two edges in the second video frame.
[0023] Optionally, the number of target interior points labeled in any of the second video frames is equal to the number of target interior points determined in the first video frame.
[0024] Optionally, the triggering condition comprises at least one of the following:
[0025] When the vehicle is started, when the position of the steering wheel is moved, when the re-determination period of the camera extrinsic parameter is reached, and when the driving speed of the vehicle is greater than a preset speed.
[0026] Optionally, the method further comprises:
[0027] determining a first position of a controllable component in the vehicle in the camera coordinate system according to the camera extrinsic parameter;
[0028] when a face image of a user is captured by the camera, determining a second position of a line of sight of the user in the camera coordinate system by using a line-of-sight positioning method;
[0029] determining a target position overlapping the second position from the first position.
[0030] In a second aspect, the embodiments of the present disclosure provide a device for determining a camera extrinsic parameter of a vehicle interior. The device comprises:
[0031] a first determining unit, configured to determine a center point coordinate and a perimeter of a figure formed by at least one target interior point in a first video frame captured by a camera, when a trigger condition of re-determining camera extrinsic parameters of the camera is met, wherein when the number of the target interior points is one, the center point coordinate is a pixel coordinate of the target interior point, and the perimeter is an abscissa coordinate of the target interior point, and the camera is arranged on a column of a steering wheel of a vehicle;
[0032] a constructing unit, configured to construct a first transformation relationship between a first rotation matrix and a target interior point element for representing a relationship between a reference camera coordinate system and a current camera coordinate system, and a second transformation relationship between a first translation matrix and the target interior point element for representing a relationship between the reference camera coordinate system and the current camera coordinate system, according to a first linear relationship between the center point coordinate and an RT matrix, a second linear relationship between the perimeter and the RT matrix, and a third linear relationship between the center point coordinate and the perimeter, wherein the target interior point element includes the center point coordinate and the perimeter, and the RT matrix is composed of a rotation matrix and a translation matrix;
[0033] a second determining unit, configured to determine the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera according to an M matrix, an N matrix, the first transformation relationship and the second transformation relationship, wherein the M matrix is a second translation matrix when a specified pixel point in the first video frame captured by the camera at a current position is transferred to the reference camera coordinate system, and the N matrix is a second rotation matrix when the specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system.
[0034] Optionally, the first determining unit is configured to determine the center point coordinate and the perimeter of the figure formed by the at least one target interior point in the first video frame captured by the camera, by:
[0035] inputting the first video frame into a trained interior point positioning model to obtain a pixel coordinate of the at least one target interior point;
[0036] determining the center point coordinate and the perimeter of the figure formed by the at least one target interior point according to the pixel coordinate of the at least one target interior point.
[0037] Optionally, the apparatus further includes:
[0038] An acquisition unit is configured to acquire a plurality of first videos captured by the camera during movement of the column;
[0039] A training unit is configured to train a to-be-trained interior point positioning model using each second video frame in the first videos that has been labeled, to obtain the trained interior point positioning model, wherein the labeling of each second video frame is labeling of the target interior point in each second video frame.
[0040] Optionally, the movement range of the column is less than or equal to a region formed by a maximum stroke that the column can reach during up-down movement and telescopic movement.
[0041] Optionally, each first video includes a plurality of sub-videos captured at a plurality of discrete positions within the movement range.
[0042] Optionally, the column is moved from a vertex of a trapezoid formed by the region to another vertex on the opposite side of the side on which the first vertex is located, and the movement is in an S shape, and a distance between a peak and a valley of the S shape and the maximum stroke that the column can reach is less than a preset distance during the S-shaped movement.
[0043] Optionally, different first videos are captured under different lighting environments.
[0044] Optionally, the target interior point includes a position in the second video frame at which a pixel difference is greater than a preset pixel difference and / or an acute angle point formed by two edges in the second video frame.
[0045] Optionally, a number of target interior points labeled in any second video frame is equal to a number of target interior points determined in the first video frame.
[0046] Optionally, the trigger condition includes at least one of the following:
[0047] The vehicle is started, the position of the steering wheel is moved, a re-determination period of the camera extrinsic parameter is reached, and the driving speed of the vehicle is greater than a preset speed.
[0048] Optionally, the apparatus further includes:
[0049] A third determination unit is configured to determine, according to the camera extrinsic parameter, a first position of a controllable component in the vehicle in the camera coordinate system.
[0050] A positioning unit is configured to, when a face image of a user is captured by the camera, determine, by line-of-sight positioning, a second position of a line of sight of the user in the camera coordinate system.
[0051] A fourth determining unit is configured to determine a target position that overlaps with the second position from the first position.
[0052] In a third aspect, an electronic device is provided. The electronic device includes a processor and a memory. The memory stores machine executable instructions that are executable by the processor. The processor executes the machine executable instructions to implement the method for determining camera extrinsic parameters of a vehicle interior according to any one of the first aspect.
[0053] In a fourth aspect, a machine readable storage medium is provided. The machine readable storage medium stores machine executable instructions. When the machine executable instructions are invoked and executed by a processor, the machine executable instructions cause the processor to implement the method for determining camera extrinsic parameters of a vehicle interior according to any one of the first aspect.
[0054] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects.
[0055] In the present disclosure, when it is necessary to determine the camera extrinsic parameters again, a first video frame is captured by a camera arranged on a column of a steering wheel of a vehicle, a center point coordinate of a figure formed by at least one target interior point in the first video frame and a perimeter of the figure are determined, then a first rotation matrix for representing a conversion of a reference camera coordinate system to a current camera coordinate system and a relationship between target interior point elements, a second translation matrix for representing the conversion of the reference camera coordinate system to the current camera coordinate system and the relationship between the target interior point elements are constructed according to a first linear relationship between the center point coordinate and the RT matrix, a second linear relationship between the perimeter and the RT matrix, and a third linear relationship between the center point coordinate and the perimeter, and finally the first rotation matrix and the first translation matrix are determined according to the M matrix, the N matrix, the first transformation relationship and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera. Through the above method, the camera extrinsic parameters can be determined automatically without human intervention, which is conducive to reducing the operation requirements when determining the camera extrinsic parameters.
[0056] In order to make the above objectives, features and advantages of the present disclosure more apparent, more comprehensible, the following will specifically describe preferred embodiments in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0057] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present disclosure, and therefore should not be regarded as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor based on these drawings.
[0058] FIG. 1 is a schematic diagram of a camera setting position provided by an embodiment of the present disclosure;
[0059] FIG. 2 is a flowchart of a method for determining camera extrinsic parameters in a vehicle interior provided by an embodiment of the present disclosure;
[0060] FIG. 3 is a schematic diagram of a conversion relationship of camera extrinsic parameters provided by an embodiment of the present disclosure;
[0061] FIG. 4 is a flowchart of another method for determining camera extrinsic parameters in a vehicle interior provided by an embodiment of the present disclosure;
[0062] FIG. 5 is a flowchart of another method for determining camera extrinsic parameters in a vehicle interior provided by an embodiment of the present disclosure;
[0063] FIG. 6 is a schematic diagram of a column moving range provided by an embodiment of the present disclosure;
[0064] FIG. 7 is a flowchart of another method for determining camera extrinsic parameters in a vehicle interior provided by an embodiment of the present disclosure;
[0065] FIG. 8 is a schematic diagram of a structure of a device for determining camera extrinsic parameters in a vehicle interior provided by an embodiment of the present disclosure;
[0066] FIG. 9 is a schematic diagram of a structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0067] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the following will combine the drawings in the embodiments of the present disclosure to make a clear and complete description of the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The components of the embodiments of the present disclosure described and shown in the drawings can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure provided in the drawings is not intended to limit the scope of the claimed present disclosure, but only represents selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative labor are within the scope of the present disclosure.
[0068] A point p in the world coordinate system is a point in a three-dimensional space, usually represented as p = [Xw ,Y w Z w ], where X w Y w and Z w These are the coordinates of the point in the x, y, and z directions in the world coordinate system, respectively.
[0069] The origin of the world coordinate system (usually denoted as O) w The origin (0, 0, 0) is the reference point of this coordinate system, typically defined as [0, 0, 0]. The location of the origin and the orientation of the world coordinate system are defined by the user or system designer based on the specific application scenario. The definition of the world coordinate system can be arbitrary, but a reference point that is easy to describe and calculate is usually chosen as the origin. For example, in an indoor scene, the origin of the world coordinate system may be defined in a corner or center of the room; in robot navigation, the origin may be defined at the robot's initial position; in photogrammetry, the origin may be defined at the location of a known landmark. In short, the definition of the world coordinate system is relative and depends on the specific application requirements and scenario. Choosing a suitable origin and coordinate system orientation can simplify calculations and descriptions.
[0070] The camera coordinate system is a three-dimensional coordinate system with the camera's optical center (i.e., the center of the camera's lens) as its origin, used to describe the world as seen by the camera. In the camera coordinate system, the horizontal axis of the camera image is usually taken as the X-axis, with the positive direction to the right; the vertical axis of the camera image is taken as the Y-axis, with the positive direction downwards; the direction of the camera's optical axis (from the camera's optical center to the subject) is perpendicular to the camera's image plane, and the Z-axis is the direction from the camera's optical center to the subject, which is the positive direction.
[0071] Camera extrinsic parameters are parameters that describe the camera's position and orientation in the world coordinate system. They define the transformation relationship from the world coordinate system to the camera coordinate system and typically include rotation and translation matrices.
[0072] Rotation Matrix (R): This is a 3×3 matrix used to describe the rotation relationship between the camera coordinate system and the world coordinate system. The rotation matrix can rotate a point in the world coordinate system to the camera coordinate system.
[0073] Translation matrix (T): This is a 3×1 matrix used to describe the translation relationship between the camera coordinate system and the world coordinate system. The translation matrix can translate a point in the world coordinate system to the camera coordinate system.
[0074] The purpose of camera extrinsic parameters is to map point P in the world coordinate system. w =[X w ,Y w Z w Transform point P to the camera coordinate systemc =[X c ,Y c Z c The transformation relationship can be represented as: P c =R·P w For the purpose of calculation, +T is usually constructed as a (4×4) matrix: Since this matrix is an RT matrix, the above transformation relationship can also be expressed as: P c =RT·P w .
[0075] In summary, the RT matrix can be used to transform a point in the world coordinate system to the camera coordinate system. This RT matrix is also called the camera extrinsic parameter.
[0076] The following is a detailed description of this disclosure.
[0077] Figure 1 is a schematic diagram of a camera setting position provided in an embodiment of this disclosure. As shown in Figure 1, a camera is set on the steering column of the vehicle. The image captured by the camera is used to determine the user's position in the camera coordinate system using tracking technology.
[0078] Figure 2 is a flowchart illustrating a method for determining the extrinsic parameters of a camera inside a vehicle according to an embodiment of this disclosure. As shown in Figure 2, the method includes the following steps:
[0079] Step 201: When the trigger condition for re-determining the camera extrinsic parameters of the camera is met, the center point coordinates and perimeter of the graphic formed by at least one target interior decoration point in the first video frame are determined according to the first video frame captured by the camera. When the number of target interior decoration points is one, the center point coordinates are the pixel coordinates of the target interior decoration point, and the perimeter is the abscissa of the target interior decoration point.
[0080] Step 202: Based on the first linear relationship representing the center point coordinates and the RT matrix, the second linear relationship representing the perimeter and the RT matrix, and the third linear relationship representing the center point coordinates and the perimeter, construct a first transformation relationship representing the relationship between a first rotation matrix and the target interior point features when transforming from the reference camera coordinate system to the current camera coordinate system, and a second transformation relationship representing the relationship between a first translation matrix and the target interior point features when transforming from the reference camera coordinate system to the current camera coordinate system. The target interior point features include the center point coordinates and the perimeter, and the RT matrix is composed of the rotation matrix and the translation matrix.
[0081] In step 203, the first rotation matrix and the first translation matrix are determined according to the M matrix, the N matrix, the first transformation relationship and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera, wherein the M matrix is a second translation matrix when a specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system, and the N matrix is a second rotation matrix when the specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system.
[0082] Specifically, when the position of the camera moves or other reasons (such as reaching the determination period of the camera extrinsic parameters) cause the camera extrinsic parameters to be determined again, the camera can capture the cab of the vehicle to obtain a first video frame containing a cab scene, and then the center point coordinates and the perimeter of a figure formed by the target interior points are determined according to the pixel coordinates of the target interior points in the first video frame. For example, when the number of target interior points is one, the center point coordinates are the pixel coordinates of the target interior point, and the perimeter is the horizontal coordinate of the target interior point. When the number of target interior points is two, the figure formed by the target interior points is a straight line, the center point coordinates are the midpoint of the straight line, and the perimeter is the length of the straight line. When the number of target interior points is three and not on a straight line, the figure formed by the target interior points is a triangle, the X value of the center point coordinates is determined by using the average value of the X axis pixel coordinates of the three vertices of the triangle, the Y value of the center point coordinates is determined by using the average value of the Y axis pixel coordinates of the three vertices of the triangle, and thus the center point coordinates are obtained. The perimeter of the triangle can be obtained by using the pixel coordinates of the three vertices of the triangle, and the center point coordinates and the perimeter of the figure formed by other numbers of target interior points are obtained in the same way.
[0083] FIG. 3 is a schematic diagram of a conversion relationship of camera extrinsic parameters provided by an embodiment of the present disclosure. As shown in FIG. 3, the reference camera coordinate system is calibrated when the vehicle is manufactured, and thus the RT matrix of the conversion from the world coordinate system to the reference camera coordinate system is known, which can be referred to as the RT1 matrix. After the position of the camera changes, the current camera coordinate system is generated. When the RT2 matrix of the conversion from the reference camera coordinate system to the current camera coordinate system is known, the camera extrinsic parameters of the conversion from the world coordinate system to the current camera coordinate system can be obtained by using the RT1 matrix and the RT2 matrix.
[0084] When the RT2 matrix is calculated, if there is a calibration board coordinate system of a calibration board, the RT3 matrix of the conversion from the reference camera coordinate system to the calibration board coordinate system and the RT4 matrix of the conversion from the calibration board coordinate system to the current camera coordinate system can be obtained after the reference camera coordinate system and the current camera coordinate system are known. Thus, the RT2 matrix can be obtained according to the RT3 matrix and the RT4 matrix.
[0085] The calibration board cannot exist in the vehicle all the time, and the calibration board actually plays a role of a fixed point, so other fixed things can be used to replace the calibration board. Since the interior points in the vehicle do not change, the interior points can be used to replace the calibration board, so as to obtain the RT2 matrix.
[0086] Since the reference camera coordinate system and the interior point do not change with the camera, when the camera position changes, the relationship between the two does not change, but the relationship between the interior point and the current camera coordinate system changes. Therefore, by the conversion relationship (fixed) between the reference camera coordinate system and the interior point, and the conversion relationship (changed) between the interior point and the current camera coordinate system, the RT2 matrix can be obtained. In this conversion process, it is found that the conversion relationship between the interior point and the current camera coordinate system is a linear relationship. Since the pixel coordinates of the target interior point and the relationship between the interior point and the coordinate system are unchanged, the conversion relationship between the interior point and the current camera coordinate system and the pixel coordinates of the target interior point can also form a linear relationship. The pixel coordinates of the target interior point can be used to obtain the center point and the circumference of the figure formed by the target interior point. Therefore, the center point coordinates and the circumference of the figure formed by the target interior point have a linear relationship with the conversion relationship between the interior point and the current camera coordinate system. According to this linear relationship, the linear relationship between the center point coordinates of the figure formed by the target interior point and the RT2 matrix can be obtained, that is, the first linear relationship. The linear relationship between the circumference of the figure formed by the target interior point and the RT2 matrix can be obtained, that is, the second linear relationship. In addition, the linear relationship between the center point coordinates and the circumference of the figure formed by the target interior point can be obtained according to different numbers of target interior points, that is, the third linear relationship. Therefore, after obtaining the center point coordinates and the circumference of the figure formed by the target interior point corresponding to the first video frame captured by the camera at the current position, according to the above three linear relationships, the first transformation relationship for representing the relationship between the first rotation matrix of the reference camera coordinate system converted to the current camera coordinate system and the target interior point elements, and the second transformation relationship for representing the relationship between the first translation matrix of the reference camera coordinate system converted to the current camera coordinate system and the target interior point elements can be constructed.
[0087] Then, according to the pre-obtained M matrix for representing the transfer of the specified pixel point in the first video frame captured by the camera at the current position to the reference camera coordinate system and the M matrix for representing the transfer of the specified pixel point in the first video frame captured by the camera at the current position to the reference camera coordinate system, the four kinds of data can be used to calculate the accurate first rotation matrix and first translation matrix, which can constitute the camera extrinsic parameters.
[0088] Through the above method, the camera extrinsic parameters can be obtained without using the calibration board, and the camera extrinsic parameters can be automatically determined without human intervention, which is beneficial to reduce the operation requirements when determining the camera extrinsic parameters.
[0089] The derivation process of the M matrix and the N matrix is as follows:
[0090] Through observation, the equation is modeled, and the relationship between RT (herein, RT2 matrix) and y (the center point coordinates of the target interior point formed pattern) and s (the circumference of the target interior point formed pattern) is: (H*y+I)*(J*s+K)=RT, which is simplified as A*y*s+B*y+C*s+D=RT, wherein A, B, C, and D represent the required coefficients.
[0091] The calculation of R is: A*y*s+B*y+C*s+D=R;
[0092] R can be represented by a quaternion, that is: A*y*s+B*y+C*s+D=[w q x q y q z q ];
[0093] After transformation, the following is obtained:
[0094] wherein, is the matrix M.
[0095] The calculation of T is: A*y*s+B*y+C*s+D=T;
[0096] T is a 3x1 matrix, so the following formula can be obtained: A*y*s+B*y+C*s+D=[x T y T z T ];
[0097] Conversion to matrix multiplication can obtain:
[0098] wherein, can be used as the N matrix.
[0099] By collecting different video frames, a plurality of corresponding data of y, s and RT can be obtained. By setting up equations for n groups of y, s and R, T, a group of equations about the M matrix and a group of equations about the N matrix can be obtained. Then, the least square method is used to solve the n groups of data to obtain the M matrix and the N matrix that can be used.
[0100] After obtaining the M matrix and the N matrix that can be used, as long as the center point coordinates and the circumference of the figure formed by the target interior points are obtained, the camera extrinsic parameters can be obtained.
[0101] In a feasible implementation, FIG. 4 is a flowchart of another method for determining camera extrinsic parameters of a vehicle interior provided by an embodiment of the present disclosure. As shown in FIG. 4, when step 201 is performed, the following steps can be implemented:
[0102] Step 401: input the first video frame into the trained interior point positioning model to obtain pixel coordinates of at least one target interior point.
[0103] Step 402: determine the center point coordinates and the circumference of a figure formed by at least one target interior point according to the pixel coordinates of the at least one target interior point.
[0104] Specifically, after the interior point positioning model is trained, the target interior points in the input first video frame can be quickly positioned to determine the pixel coordinates of the target interior points, so as to obtain the corresponding center point coordinates and the circumference, and then the first rotation matrix and the first translation matrix are quickly obtained.
[0105] In a feasible implementation, FIG. 5 is a flowchart of another method for determining camera extrinsic parameters of a vehicle interior provided by an embodiment of the present disclosure. As shown in FIG. 5, when the interior point positioning model is trained, the following steps are implemented:
[0106] Step 501: obtain a plurality of first videos shot by the camera during the movement of the column.
[0107] Step 502: train the interior point positioning model to be trained by using each second video frame in the first video after the labeling to obtain the trained interior point positioning model. When each second video frame is labeled, the labeling is performed on the target interior points in each second video frame.
[0108] Specifically, since the interior point positioning model is trained to identify the pixel coordinates of the interior points in the vehicle, in order to make the trained interior point positioning model have better effect, the video frames in the first video captured by the camera arranged on the column can be used as samples for training. Since the samples used are real images in the vehicle, the recognition accuracy of the interior point positioning model trained in this way is higher. In order to make the samples more comprehensive, the camera can be moved to capture the interior of the vehicle at different angles by moving the column.
[0109] It should be noted that the specific interior point positioning model, the training method of the interior point positioning model, and the moving method of the column can be set according to actual needs, and are not limited here.
[0110] In a feasible implementation, during the movement of the column, the movement range of the column is less than or equal to the area formed by the maximum stroke of the column in the up-down movement and the telescopic movement.
[0111] Specifically, FIG. 6 is a schematic diagram of the movement range of the column provided by an embodiment of the present disclosure. The column of the steering wheel can be adjusted in extension and contraction, and up and down. As shown in FIG. 6, the adjustment range of the column is a rectangular area formed by four points A, B, C, and D in FIG. 6, that is, the rectangular area is the area formed by the maximum stroke of the column. Since the camera is arranged on the column, the column moves with the camera moving. Therefore, the movement range of the column is the movement range of the camera, that is, the camera can be arranged at any position in the above-mentioned rectangular area by adjusting the column. Further, the movement range of the camera is less than or equal to the above-mentioned rectangular area. When the camera moves in the above-mentioned rectangular area, the camera can be as much as possible to capture the interior of the vehicle at different angles, so that the obtained samples are more diversified, which is beneficial to improve the accuracy of the pixel coordinates of the target interior point of the trained interior point positioning model.
[0112] In a feasible implementation, each of the first videos includes a plurality of sub-videos captured at a plurality of discrete positions within the movement range.
[0113] Specifically, in order to ensure the diversity of samples, sub-videos need to be captured at a plurality of positions within the rectangular range shown in FIG. 6.
[0114] In a feasible implementation, when the column is moved, it is started from a vertex of the trapezoid formed by the area and ended at a vertex on the opposite side of the side where the vertex is located. During the movement, it is moved in an S shape. During the S-shaped movement, the distance between the peak and the valley of the S shape and the maximum stroke that the column can reach is less than a preset distance.
[0115] Specifically, in order to guarantee the diversity of the samples, as shown in FIG. 6, any one of the four vertices A, B, C and D can be taken as a starting point, and an S-shaped moving path is used to move within the rectangular range shown in FIG. 6 until the low point corresponding to the starting vertex is reached. During the movement, the distance between the peaks and valleys of the S-shape and the AD side or the BC side is less than a preset distance, or the AD side and the BC side are used as the boundaries of the peaks and valleys of the S-shape, or the AB side and the CD side are used as the boundaries of the peaks and valleys of the S-shape.
[0116] It should be noted that the width of the peak or valley during the movement can be set according to actual needs, which is not specifically limited here.
[0117] In a feasible implementation, the different first videos are taken under different lighting environments.
[0118] Specifically, in order to guarantee the diversity of the samples, a plurality of first videos can be taken under different lighting environments, for example, a plurality of first videos can be taken in the morning, at noon, in the afternoon and at night respectively when taking the first video. In this way, the lighting environment of the samples used for training can be different, thereby facilitating the improvement of the training results of the model.
[0119] In a feasible implementation, the target interior point includes a position with a pixel difference greater than a preset pixel difference in the second video frame and / or an acute corner point composed of two sides in the second video frame.
[0120] Specifically, when labeling the target interior point, the position where the color step is obviously visible from imaging is selected as the target interior point, and the corner point, i.e., the intersection point of two sides, is also selected, and the acute corner point is preferred. The characteristics of such target interior points are relatively obvious, which is conducive to the success rate of identification.
[0121] In a feasible implementation, the number of target interior points labeled in any of the second video frames is equal to the number of target interior points determined in the first video frame.
[0122] Specifically, when labeling the target interior point in the second video frame, the number and position of the labeled interior points in different second video frames are the same according to actual needs. After the interior point positioning model is trained using such second video frames, the interior point positioning model can identify all the interior points, thereby facilitating subsequent processing.
[0123] In a feasible implementation, the trigger condition includes at least one of the following:
[0124] when the vehicle starts, when the position of the steering wheel moves, when the re-determination period of the camera extrinsic parameter is reached, or when the driving speed of the vehicle is greater than a preset speed.
[0125] Specifically, the vehicle starts from 0 to movement, or the vehicle switches from a static state to a moving state, the position of the camera may have moved, at which time the camera parameter 2 needs to be re-determined, or when the position of the steering wheel moves, when the re-determination period of the camera extrinsic parameter is reached, or when the driving speed of the vehicle is greater than a preset speed (i.e., after the vehicle is running) in order to ensure the accuracy of tracking the user's line of sight, the camera extrinsic parameter 2 needs to be re-determined.
[0126] When determining whether the position of the steering wheel moves, since the camera in the present disclosure can be used to track the user's line of sight, the camera will continuously capture video during use of the vehicle. Two video frames at an interval of a preset time length in the continuous video can be compared, for example, the pixel coordinates of a certain interior point are compared to determine whether the steering wheel moves, or when the vehicle is started, the last image frame captured before the vehicle is started and the first image frame captured after the vehicle is started are compared, for example, the pixel coordinates of a certain interior point are compared to determine whether the steering wheel moves.
[0127] In a feasible implementation, FIG. 7 is a flowchart of another method for determining the camera extrinsic parameter of a vehicle interior provided by an embodiment of the present disclosure. As shown in FIG. 7, after the camera extrinsic parameter is determined, the following steps are further included:
[0128] Step 701: determining a first position of a controllable component in the vehicle in the camera coordinate system according to the camera extrinsic parameter.
[0129] Step 702: when a user's face image is captured by the camera, determining a second position of the user's line of sight in the camera coordinate system by line-of-sight positioning.
[0130] Step 703: determining a target position overlapping the second position from the first position.
[0131] Specifically, after the camera extrinsic parameter corresponding to the current camera position is determined, the position of the controllable component in the vehicle in the camera coordinate system can be determined. After the position of the user's line of sight in the current camera coordinate system is determined by line-of-sight positioning, it can be determined that the user is currently gazing at which controllable component, so that the user can control the controllable component by gestures.
[0132] It should be noted that the controllable component includes a vehicle screen, a sunroof, a vehicle-mounted player, a vehicle-mounted camera device, vehicle cabin lighting, and other components in the vehicle that can be controlled by a control.
[0133] FIG. 8 is a structural schematic diagram of a device for determining camera extrinsic parameters of a vehicle interior according to an embodiment of the present disclosure. As shown in FIG. 8, the device includes:
[0134] A first determining unit 81 is configured to, when a trigger condition for re-determining camera extrinsic parameters of a camera is met, determine a center point coordinate of a figure formed by at least one target interior point in a first video frame captured by the camera and a perimeter of the figure, wherein when the number of the target interior points is one, the center point coordinate is a pixel coordinate of the target interior point, and the perimeter is a horizontal coordinate of the target interior point, and the camera is arranged on a column of a steering wheel of a vehicle.
[0135] A constructing unit 82 is configured to construct, according to a first linear relationship between the center point coordinate and an RT matrix, a second linear relationship between the perimeter and the RT matrix, and a third linear relationship between the center point coordinate and the perimeter, a first transformation relationship between a first rotation matrix for representing a conversion of a reference camera coordinate system to a current camera coordinate system and a target interior point element, and a second transformation relationship between a first translation matrix for representing the conversion of the reference camera coordinate system to the current camera coordinate system and the target interior point element, wherein the target interior point element includes the center point coordinate and the perimeter, and the RT matrix is composed of a rotation matrix and a translation matrix.
[0136] A second determining unit 83 is configured to determine the first rotation matrix and the first translation matrix according to an M matrix, an N matrix, the first transformation relationship and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera, wherein the M matrix is a second translation matrix when a specified pixel point in the first video frame captured by the camera at a current position is transferred to the reference camera coordinate system, and the N matrix is a second rotation matrix when the specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system.
[0137] In a feasible implementation, the first determining unit 81 is configured to, when determining the center point coordinate of the figure formed by at least one target interior point in the first video frame captured by the camera and the perimeter of the figure, include:
[0138] inputting the first video frame into a trained interior point positioning model to obtain a pixel coordinate of at least one target interior point;
[0139] determine the center point coordinate and the perimeter of a figure formed by the at least one target interior point according to the pixel coordinate of the at least one target interior point.
[0140] In an implementation, the device further comprises:
[0141] an acquisition unit, configured to acquire a plurality of first videos captured by the camera during movement of the pipe column;
[0142] a training unit, configured to train a to-be-trained interior point positioning model using each second video frame in the first videos after annotation, to obtain the trained interior point positioning model, wherein the annotation on each second video frame is annotation on the target interior point in each second video frame.
[0143] In an implementation, the movement range of the pipe column is less than or equal to a region formed by the maximum stroke of the pipe column in up-down movement and telescopic movement.
[0144] In an implementation, each first video comprises a plurality of sub-videos captured at a plurality of discrete positions in the movement range.
[0145] In an implementation, the pipe column is moved from a vertex of a trapezoid formed by the region to another vertex on the opposite side of the side where the first vertex is located, and the movement is in an S shape, and the distance between the peak and the valley of the S shape and the maximum stroke of the pipe column is less than a preset distance.
[0146] In an implementation, different first videos are captured in different light environments.
[0147] In an implementation, the target interior point comprises a position in the second video frame where the pixel difference is greater than a preset pixel difference and / or an acute angle point formed by two edges in the second video frame.
[0148] In an implementation, the number of target interior points annotated in any second video frame is equal to the number of target interior points determined in the first video frame.
[0149] In an implementation, the trigger condition comprises at least one of the following:
[0150] when the vehicle is started, when the position of the steering wheel is changed, when a re-determination period of the camera extrinsic parameter is reached, and when the driving speed of the vehicle is greater than a preset speed.
[0151] In an implementation, the device further comprises:
[0152] a third determining unit, configured to determine a first position of a controllable component in the vehicle in the camera coordinate system according to the camera extrinsic parameter;
[0153] a positioning unit, configured to determine a second position of a line of sight of a user in the camera coordinate system by line-of-sight positioning when a face image of the user is captured by the camera;
[0154] a fourth determining unit, configured to determine a target position from the first position, the target position being overlapped with the second position.
[0155] The related description of the camera extrinsic parameter determination apparatus inside the vehicle can refer to the related explanation of the camera extrinsic parameter determination method inside the vehicle, and will not be described in detail here.
[0156] FIG. 9 is a structural schematic diagram of an electronic device provided by an embodiment of the present disclosure, including a processor 901, a storage medium 902, and a bus 903. The storage medium 902 stores machine readable instructions executable by the processor 901. When the electronic device runs a camera extrinsic parameter determination method inside a vehicle as in an embodiment, the processor 901 and the storage medium 902 communicate through the bus 903. The processor 901 executes the machine readable instructions to perform steps as in an embodiment.
[0157] In an embodiment, the storage medium 902 can also execute other machine readable instructions to perform other methods as in an embodiment. For specific method steps and principles, refer to the description of the embodiments, which will not be described in detail here.
[0158] Embodiment four of the present disclosure also provides a computer readable storage medium, which stores a computer program. When the computer program is run by a processor, steps shown in the above embodiments are executed.
[0159] In the embodiments provided by the present disclosure, it should be understood that the disclosed apparatus and method can be implemented in other ways. The apparatus embodiments described above are only schematic. For example, the division of the units is only a logical function division. There can be another division during actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or other forms.
[0160] The units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. may be located in one place, or may be distributed on multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0161] In addition, the functional units in the embodiments provided by the present disclosure can be integrated in one processing unit, or each unit can exist alone physically, or two or more units can be integrated in one unit.
[0162] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present disclosure essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0163] It should be noted that: similar reference numerals and letters represent similar items in the following drawings, therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings, in addition, the terms "first", "second", "third" and the like are only used to distinguish description, and cannot be understood as indicating or implying relative importance.
[0164] Finally, it should be noted that: the above-described embodiments are only specific embodiments of the present disclosure, used to illustrate the technical solutions of the present disclosure, and not to limit the same, the protection scope of the present disclosure is not limited thereto, although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any skilled person familiar with the technical field can modify or easily think of changes to the technical solutions recorded in the foregoing embodiments within the technical range disclosed by the present disclosure, or make equivalent replacement to part of the technical features; and these modifications, changes or replacements do not make the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure. All should be covered in the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims. Industrial applicability
[0165] In the present disclosure, when it is necessary to re-determine the camera extrinsic parameters, a first video frame is captured according to a camera arranged on a column of a steering wheel of a vehicle, a center point coordinate of a figure formed by at least one target interior point in the first video frame and a perimeter of the figure are determined, then a first rotation matrix for representing a conversion of a reference camera coordinate system to a current camera coordinate system and a first transformation relationship between target interior point elements are constructed according to a first linear relationship between the center point coordinate and an RT matrix, a second linear relationship between the perimeter and the RT matrix, and a third linear relationship between the center point coordinate and the perimeter, a second transformation relationship between a first translation matrix for representing the conversion of the reference camera coordinate system to the current camera coordinate system and the target interior point elements is constructed, and finally the first rotation matrix and the first translation matrix are determined according to an M matrix, an N matrix, the first transformation relationship and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as camera extrinsic parameters of the camera. Through the above method, the camera extrinsic parameters can be automatically determined without human intervention, which is conducive to reducing the operation requirements when determining the camera extrinsic parameters.
[0166] In addition, it can be understood that the vehicle interior camera extrinsic parameter determination method, device, electronic equipment and storage medium provided by the embodiments of the present disclosure are reproducible and can be used in various industrial applications. For example, the vehicle interior camera extrinsic parameter determination method, device, electronic equipment and storage medium provided by the embodiments of the present disclosure can be used in the field of computer technology.
Claims
1. A method for determining camera extrinsic parameters of a vehicle interior, characterized in that, A camera is arranged on a column of a steering wheel of a vehicle, and the method comprises: When a trigger condition for re-determining camera extrinsic parameters of the camera is met, a center point coordinate and a perimeter of a figure formed by at least one target interior point in a first video frame captured by the camera are determined according to the first video frame, wherein when the number of the target interior points is one, the center point coordinate is a pixel coordinate of the target interior point, and the perimeter is a horizontal coordinate of the target interior point; According to a first linear relationship for representing a relationship between the center point coordinate and an RT matrix, a second linear relationship for representing a relationship between the perimeter and the RT matrix, and a third linear relationship for representing a relationship between the center point coordinate and the perimeter, a first transformation relationship for representing a relationship between a first rotation matrix of a reference camera coordinate system converted to a current camera coordinate system and a target interior point element, and a second transformation relationship for representing a relationship between a first translation matrix of the reference camera coordinate system converted to the current camera coordinate system and the target interior point element are constructed, wherein the target interior point element includes the center point coordinate and the perimeter, and the RT matrix is composed of a rotation matrix and a translation matrix; The first rotation matrix and the first translation matrix are determined according to an M matrix, an N matrix, the first transformation relationship, and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera, wherein the M matrix is a second translation matrix when a specified pixel point in the first video frame captured by the camera at a current position is transferred to the reference camera coordinate system, and the N matrix is a second rotation matrix when the specified pixel point in the first video frame captured by the camera at the current position is transferred to the reference camera coordinate system.
2. The method of claim 1, wherein, The determination of the center point coordinate and the perimeter of the figure formed by the at least one target interior point in the first video frame captured by the camera comprises: The first video frame is input into a trained interior point positioning model to obtain pixel coordinates of the at least one target interior point; The center point coordinate and the perimeter of the figure formed by the at least one target interior point are determined according to the pixel coordinates of the at least one target interior point.
3. The method of claim 2, wherein, The method further comprises: A plurality of first videos captured by the camera during movement of the column are acquired; A trained interior point positioning model is trained using each second video frame in the first videos after labeling, so as to obtain the trained interior point positioning model, wherein the labeling of each second video frame is labeling of the target interior point in each second video frame.
4. The method of claim 3, wherein, A movement range of the column is less than or equal to a region formed by a maximum stroke that can be reached by the column during up-down movement and telescopic movement.
5. The method of claim 4, wherein, Each first video includes a plurality of sub-videos captured at a plurality of discrete positions in the movement range.
6. The method of claim 4, wherein, The pipe column is moved from one vertex of the trapezoid formed by the area to another vertex on the opposite side of the side where the first vertex is located, and is moved in an S shape, wherein the distance between the peak and the valley of the S shape and the maximum stroke that the pipe column can reach is less than a preset distance.
7. The method of claim 3, wherein, The different first videos are captured under different light environments.
8. The method of claim 3, wherein, The target interior points include positions with pixel differences greater than a preset pixel difference in the second video frame and / or acute angle points formed by two edges in the second video frame.
9. The method of claim 3, wherein, The number of target interior points labeled in any second video frame is equal to the number of target interior points determined in the first video frame.
10. The method of claim 1, wherein, The trigger condition includes at least one of the following: When the vehicle is started, when the position of the steering wheel changes, when a re-determination period of the camera extrinsic parameter is reached, and when the driving speed of the vehicle is greater than a preset speed.
11. The method of claim 1, wherein, The method further includes: Determining a first position of a controllable component in the vehicle in the camera coordinate system according to the camera extrinsic parameter; When a face image of a user is captured by the camera, determining a second position of a line of sight of the user in the camera coordinate system using line-of-sight positioning; Determining a target position overlapping the second position from the first position.
12. A device for determining camera extrinsic parameters of a vehicle interior, characterized by The device includes: A first determination unit configured to determine a center point coordinate and a perimeter of a figure formed by at least one target interior point in a first video frame captured by a camera when a trigger condition for re-determining a camera extrinsic parameter of the camera is met, wherein when the number of target interior points is one, the center point coordinate is a pixel coordinate of the target interior point, and the perimeter is a horizontal coordinate of the target interior point, and wherein the camera is arranged on a pipe column of a steering wheel of a vehicle; A construction unit configured to construct a first transformation relationship between a first linear relationship representing a relationship between the center point coordinate and an RT matrix, a second linear relationship representing a relationship between the perimeter and the RT matrix, and a third linear relationship representing a relationship between the center point coordinate and the perimeter, and a second transformation relationship between a first rotation matrix representing a conversion of a reference camera coordinate system to a current camera coordinate system and a target interior point element, and a first translation matrix representing a conversion of the reference camera coordinate system to the current camera coordinate system and the target interior point element, wherein the target interior point element includes the center point coordinate and the perimeter, and the RT matrix is composed of a rotation matrix and a translation matrix. A second determining unit is configured to determine the first rotation matrix and the first translation matrix according to the M matrix, the N matrix, the first transformation relationship and the second transformation relationship, so as to take the first rotation matrix and the first translation matrix as the camera extrinsic parameters of the camera.
13. An electronic device, comprising: A processor and a memory are included, the memory stores machine executable instructions which can be executed by the processor, and the processor executes the machine executable instructions to implement the method for determining the camera extrinsic parameters of the vehicle interior according to any one of claims 1-11.
14. A machine-readable storage medium, characterized in that, The machine readable storage medium stores machine executable instructions, when the machine executable instructions are called and executed by the processor, the machine executable instructions cause the processor to implement the method for determining the camera extrinsic parameters of the vehicle interior according to any one of claims 1-11.
Citation Information
Patent Citations
Camera external parameter calibration method, device, equipment, medium and program product
CN114708339A
Calibration method for external parameters of vehicle-mounted camera and related device
CN114730472A
Multi-camera vision large target positioning method, system and equipment
CN115187658A
Method and device for determining external parameters of camera in vehicle, electronic equipment and storage medium
CN118644560A
Device and method for automatically calibrating camera
JP2010181209A