Image sample annotation method, device, terminal device and image annotation system
By calculating and labeling the three-dimensional coordinates of the target object using the image and position relationship acquired by multiple monocular cameras, the problem that the depth camera cannot obtain the absolute depth of the obstructed part is solved, and efficient and accurate image sample annotation is achieved.
Patent Information
- Application Number
- CN202111135440.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-27
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-09-27
AI Technical Summary
When the existing depth camera has a large degree of freedom of the target object, it may not be able to obtain the absolute depth of the obstructed part, which makes it difficult to label image samples.
By utilizing the images acquired at the same time by at least two monocular cameras and the pose relationship between the cameras, the three-dimensional coordinates of the key points of the target object under the camera coordinate system are calculated and marked using these three-dimensional coordinates.
It realizes that the three-dimensional coordinates of the target object are accurately marked when the depth camera cannot obtain absolute depth, improving the efficiency and accuracy of image sample annotation.
Smart Images

Figure CN113870350B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing technology, and in particular to an image sample annotation method, an apparatus, a terminal device and an image annotation system. Background Art
[0002] At present, most image recognition models based on artificial intelligence algorithms need to be trained with a large number of annotated image data sets to obtain an image recognition model with relatively stable performance. Existing public image data sets are of great help in the training of image recognition models. However, for the training of some specific models, it may be difficult to make the specific models meet the predetermined requirements using existing public image data sets. In this case, specialized personnel are required to collect image data sets and annotate the image samples therein. Although existing depth cameras can detect the depth of target objects in the shooting space, when the target object has a large degree of freedom, the target object may be blocked by itself, resulting in the depth camera being unable to obtain the absolute depth of the blocked part. Summary of the invention
[0003] In view of the above problems, the present application proposes an image sample annotation method, apparatus, terminal device and image annotation system.
[0004] The present application embodiment provides an image sample annotation method, including:
[0005] Calculate a first three-dimensional coordinate of the nth key point in the first camera coordinate system based on a posture relationship between the first camera and the second camera, a first pixel coordinate of the nth key point of the target object in the first image, and a second pixel coordinate of the nth key point in the second image, wherein the first image and the second image are respectively acquired by the first camera and the second camera at the same time, 1≤n≤N, and N is the total number of key points to be labeled;
[0006] The nth key point in the first image is marked using the first three-dimensional coordinate.
[0007] The image sample annotation method described in the embodiment of the present application, when at least three images including the target object are acquired by using at least three cameras at the same time, before calculating the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image, further includes:
[0008] Selecting two images with the least obscured key points to be annotated from a plurality of images including the target object as the first image and the second image;
[0009] A first pixel coordinate of the nth key point in the first image and a second pixel coordinate of the nth key point in the second image are determined respectively.
[0010] The image sample labeling method according to the embodiment of the present application, before labeling the nth key point in the first image by using the first three-dimensional coordinate, further includes:
[0011] Calculate the first three-dimensional coordinate corresponding to the first three-dimensional coordinate in the coordinate system of the i-th camera based on the posture relationship between the first camera and the i-th camera, 3≤i≤I, I is the total number of cameras;
[0012] Determine the pixel coordinates corresponding to the first three-dimensional coordinates in the i-th image, and after marking the corresponding pixel coordinates in the i-th image with a first mark, display the i-th image with the first mark;
[0013] If the first marks of each image correspond to the nth key point, then performing the step of marking the nth key point in the first image by using the first three-dimensional coordinate;
[0014] If there is a first marker in at least one image that deviates from the nth key point, update the first pixel coordinate and / or the second pixel coordinate, and re-execute the calculation of the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image.
[0015] The image sample annotation method described in the embodiment of the present application further includes:
[0016] Calculate a second three-dimensional coordinate of the first three-dimensional coordinate in a second camera coordinate system based on a posture relationship between the first camera and the second camera;
[0017] The nth key point in the second image is marked using the second three-dimensional coordinate.
[0018] The image sample labeling method according to the embodiment of the present application, before labeling the nth key point in the second image by using the second three-dimensional coordinate, further includes:
[0019] Calculate the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the coordinate system of the i-th camera based on the posture relationship between the second camera and the i-th camera, 3≤i≤I, I is the total number of cameras;
[0020] Determine the pixel coordinates corresponding to the second three-dimensional coordinates in the i-th image, and after marking the corresponding pixel coordinates in the i-th image with a second mark, display the i-th image with the second mark;
[0021] If the second marks of each image correspond to the nth key point, then performing the step of marking the nth key point in the second image by using the second three-dimensional coordinate;
[0022] If there is a second marker in at least one image that deviates from the nth key point, update the first pixel coordinate and / or the second pixel coordinate, and re-execute the method of calculating the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image.
[0023] The image sample annotation method described in the embodiment of the present application, the first three-dimensional coordinate of the nth key point in the first camera coordinate system is calculated based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image, including:
[0024] Determine a first transition coordinate using a first pixel coordinate and an intrinsic parameter matrix of the first camera;
[0025] Determine a second transition coordinate using the second pixel coordinate and the intrinsic parameter matrix of the second camera;
[0026] Calculate the four-dimensional coordinates corresponding to the nth key point by using the posture relationship between the first camera and the second camera, the first transition coordinates, and the second transition coordinates;
[0027] The coordinate values of the first three dimensions of the four-dimensional coordinates are divided by the coordinate value of the fourth dimension to obtain the first three-dimensional coordinates.
[0028] The image sample labeling method described in the embodiment of the present application also includes: predetermining a posture relationship between any two cameras, wherein the posture relationship includes a rotation matrix and a translation vector.
[0029] The image sample labeling method described in the embodiment of the present application, wherein the predetermining the posture relationship between any two cameras includes:
[0030] Determine the intrinsic parameter matrix and distortion coefficient of each camera;
[0031] Determine the basic rotation vector and basic translation vector of each camera using the intrinsic parameter matrix and distortion coefficient of each camera;
[0032] The rotation matrix and the translation vector between any two cameras are calculated using the basic rotation vectors and the basic translation vectors of any two cameras.
[0033] The present application also provides an image sample annotation device, including:
[0034] a calculation module, configured to calculate a first three-dimensional coordinate of an nth key point in a first camera coordinate system based on a posture relationship between the first camera and the second camera, a first pixel coordinate of an nth key point of a target object in a first image, and a second pixel coordinate of the nth key point in a second image, wherein the first image and the second image are respectively acquired by the first camera and the second camera at the same time, 1≤n≤N, and N is the total number of key points to be annotated;
[0035] A labeling module is used to label a predetermined position in the first image using the first three-dimensional coordinates.
[0036] The embodiment of the present application further provides a terminal device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is run on the processor, the image sample labeling method described in the embodiment of the present application is executed.
[0037] The embodiment of the present application also provides an image annotation system, comprising multiple cameras and the terminal device described in the embodiment of the present application.
[0038] The embodiment of the present application further provides a readable storage medium storing a computer program, which executes the image sample labeling method described in the embodiment of the present application when running on a processor.
[0039] The image sample annotation method disclosed in the present application utilizes two images of a target object acquired by two monocular cameras at the same time and the posture relationship between the two monocular cameras to determine the three-dimensional coordinates of any position of the target object in the two images, so as to add a three-dimensional annotation to the corresponding position through the three-dimensional coordinates, thereby solving the problem that the depth camera cannot obtain the absolute depth of the obscured part of the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solution of the present invention, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope of protection of the present invention. In each of the drawings, similar components are numbered similarly.
[0041] Figure 1A schematic diagram of a process for determining a camera posture relationship proposed in an embodiment of the present application is shown;
[0042] Figure 2 A schematic diagram of an ARUCO code proposed in an embodiment of the present application is shown;
[0043] Figure 3 A schematic diagram of a process flow of an image sample annotation method proposed in an embodiment of the present application is shown;
[0044] Figure 4 A schematic diagram of the process of determining the first three-dimensional coordinates in an image sample annotation method proposed in an embodiment of the present application is shown;
[0045] Figure 5 A schematic diagram including a first image and a second image of a hand proposed in an embodiment of the present application is shown;
[0046] Figure 6 A schematic diagram showing the relationship between various coordinate systems proposed in an embodiment of the present application;
[0047] Figure 7 A schematic diagram of a process for verifying the accuracy of a first image annotation in an image sample annotation method proposed in an embodiment of the present application is shown;
[0048] Figure 8 A schematic diagram showing a flow chart of another image sample labeling method proposed in an embodiment of the present application;
[0049] Fig. 9 A schematic diagram of a process for verifying the accuracy of second image annotation in an image sample annotation method proposed in an embodiment of the present application is shown;
[0050] Fig.10 A schematic diagram of the structure of an image sample labeling device proposed in an embodiment of the present application is shown;
[0051] Fig.11 A schematic diagram of the structure of a terminal device proposed in an embodiment of the present application is shown;
[0052] Fig.12 A schematic diagram of the structure of an image annotation system proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0053] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0054] The components of the embodiments of the present invention generally described and shown in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0055] Hereinafter, the terms "including", "having" and their cognates, which may be used in various embodiments of the present invention, are intended only to indicate specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be understood as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or adding the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.
[0056] Furthermore, the terms “first”, “second”, “third”, etc. are merely used for distinguishing descriptions and are not to be understood as indicating or implying relative importance.
[0057] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meanings as those generally understood by those skilled in the art to which the various embodiments of the present invention belong. The terms (such as those defined in generally used dictionaries) will be interpreted as having the same meanings as the contextual meanings in the relevant technical field and will not be interpreted as having idealized meanings or overly formal meanings unless clearly defined in the various embodiments of the present invention.
[0058] The image sample annotation method disclosed in the present application utilizes two images of a target object acquired by two monocular cameras at the same time and the posture relationship between the two monocular cameras to determine the three-dimensional coordinates of any position of the target object in the two images, and to add a three-dimensional annotation to the corresponding position through the three-dimensional coordinates, so as to overcome the problem that the depth camera cannot obtain the absolute depth of the obscured part of the target object.
[0059] Furthermore, in order to ensure the accuracy of the added three-dimensional annotations, the present application can also utilize at least three images of the target object acquired by at least three monocular cameras at the same time and the posture relationship between each monocular camera. After determining the three-dimensional coordinates of any position of the target object in any two images, the three-dimensional coordinates corresponding to one of the two images and the posture relationship between the monocular camera corresponding to the image and each other monocular camera can be used to determine the position corresponding to the three-dimensional coordinate in each other image to determine the accuracy of the three-dimensional coordinates, and the three-dimensional coordinates corresponding to the other image of the two images and the posture relationship between the monocular camera corresponding to the image and each other monocular camera can be used to determine the position corresponding to the three-dimensional coordinate in each other image to determine the accuracy of the three-dimensional coordinates.
[0060] It should be noted that the pose relationship between each monocular camera includes the rotation matrix and translation vector, that is, a monocular camera can move from its current position to the position of another monocular camera through the rotation matrix and translation vector between the two monocular cameras, and the pose is the same as that of the other monocular camera. The pose relationship between each monocular camera is also obtained in advance, for example, see Figure 1 , the posture relationship S10-S30 between each monocular camera can be determined in advance through the following posture relationship determination steps:
[0061] S10: Determine the intrinsic parameter matrix and distortion coefficient of each camera.
[0062] The intrinsic matrix and distortion coefficient of each camera can be determined by the chessboard calibration method. For example, a chessboard calibration plate is prepared. To reduce the position and posture relationship errors of each camera, a chessboard with high precision and flat surface needs to be selected as the calibration plate; two cameras with fixed position and posture relationship are used to take pictures of the chessboard calibration plate with changing angles and positions to obtain multiple pictures (which can be 25-30 pictures), and the camera intrinsic matrix and distortion coefficient are calculated using the multiple pictures according to the Zhang Zhengyou calibration method.
[0063] S20: Determine the basic rotation vector and basic translation vector of each camera using the intrinsic parameter matrix and distortion coefficient of each camera.
[0064] For example, the aruco code can be used to calculate the basic rotation vector and basic translation vector of each camera. The calculation steps are as follows:
[0065] Pre-generated Figure 2The Aruco code shown in the figure can be 6*6. A black square or a white square is called a bit. Aruco code is surrounded by a black frame. A 6*6 code mark is covered with a frame with a width of 1. In addition to 6*6, you can also choose 4*4, 5*5, and 7*7 specifications. The side length of the Aruco code is 14cm (other sizes can also be selected). After printing it, paste it on a flat surface.
[0066] Exemplarily, when determining the basic rotation vector and the basic translation vector of the first camera, the Aruco code needs to appear completely in the fields of view of the two cameras at the same time, for example, appear in the fields of view of the first camera and the second camera at the same time.
[0067] Detect the four corner points corners1 and id1 of the aruco code captured by the first camera, and detect the four corner points corners2 and id2 of the aruco code captured by the second camera. The corner points represent the four outermost points of the aruco code, and id represents the number of the code in the dictionary. This step can be implemented using the aruco.detectMarkers function.
[0068] Furthermore, the basic rotation vector r1 and the basic translation vector t1 of the first camera are calculated using the following formula:
[0069] r1,t1,_=aruco.estimatePoseSingleMarkers(corners1,0.14,mtx1,dist1)
[0070] Among them, aruco.estimatePoseSingleMarkers is a function of opencv used to calculate the external parameters of the camera, 0.14 represents the side length of the aruco code (14cm=0.14m), mtx1 represents the intrinsic parameter matrix of the first camera, and dist1 represents the distortion coefficient of the first camera.
[0071] The basic rotation vector r1 of the first camera is converted into a basic rotation matrix R1 by using the Rodriguez formula.
[0072] Furthermore, the basic rotation vector r2 and the basic translation vector t2 of the second camera are calculated using the following formula:
[0073] r2,t2,_=aruco.estimatePoseSingleMarkers(corners1,0.14,mtx2,dist2)
[0074] Among them, mtx2 represents the intrinsic parameter matrix of the second camera, and dist2 represents the distortion coefficient of the second camera.
[0075] It should be noted that the 0.14 in the above function aruco.estimatePoseSingleMarkers() is related to the side length of the aruco code. If the aruco code selects other sizes and the side length of the ruco code changes, the 0.14 in the above function will be replaced by the value corresponding to the side length of the current ruco code.
[0076] The basic rotation vector r2 of the second camera is converted into a rotation matrix R2 by using the Rodriguez formula.
[0077] S30: Calculate the rotation matrix and translation vector between any two cameras using the basic rotation vector and basic translation vector of any two cameras.
[0078] It can be understood that the center of the Aruco code is used as the origin of the world coordinate system, and the camera coordinates (x c ,y c ,z c ) to world coordinates (x w ,y w ,z w ) is as follows:
[0079]
[0080] It can be seen that: Taking the calculation of the rotation matrix and translation vector between the first camera and the second camera as an example, since the first camera and the second camera simultaneously shoot the same Aruco code, the world coordinate systems corresponding to the two cameras are the same, and the following formula can be obtained:
[0081] Where R 1 -1 represents the inverse of the rotation matrix R1, (x c1 ,y c1 ,z c1 ) is the coordinate in the first camera coordinate system, (x c2 ,y c2 ,z c2 ) is the coordinate in the first camera coordinate system, and the formula can be transformed to obtain:
[0082] Since R1, R2, t1 and t2 have been obtained in step S20, the rotation matrix from the second camera to the first camera can be calculated as R 1 R 2 -1 , the translation vector from the second camera to the first camera is t1-R 1 R 2 -1t2. The calculation method of the rotation matrix and translation vector from the first camera to the second camera is similar.
[0083] It should be noted that the target object in the present application can be a face, hand, body, etc. When the target object is a face, the obtained image sample can be used to train a face recognition model, when the target object is a hand, the obtained image sample can be used to train a hand posture recognition model, and when the target object is a body, the obtained image sample can be used to train a body posture recognition model. The following embodiment of the present application is further explained by taking the target object being a hand as an example.
[0084] For an example of this application, see Figure 3 , a method for labeling image samples is proposed, comprising the following steps S100 and S200:
[0085] S100: Calculate a first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, a first pixel coordinate of the nth key point of the target object in the first image, and a second pixel coordinate of the nth key point in the second image.
[0086] Among them, the first image and the second image can be acquired in real time by the first camera and the second camera, and the images acquired by the first camera and the second camera are acquired in real time to be annotated, or they can be acquired in advance by the first camera and the second camera and stored in a predetermined position. When annotating the image, the first image and the second image are acquired from the predetermined position to be annotated.
[0087] It can be understood that the first image and the second image are acquired by the first camera and the second camera at the same time, respectively, and both include the target object. When the target object is a hand, 5, 21 or 26 key points of the hand can be annotated in sequence. Therefore, 1≤n≤N, N is the total number of key points to be annotated. For example, see Figure 2 , a hand schematic diagram including 21 key points is given, where the 21 key points correspond to 21 joints of the hand.
[0088] It should be noted that different values can be set for N according to the requirements of the model to be trained.
[0089] For example, see Figure 4 The above step S100 includes the following steps S110 to S140:
[0090] S110: Determine first transition coordinates using first pixel coordinates and an intrinsic parameter matrix of the first camera.
[0091] See also Figure 5, the left side is a first image with a hand taken by the first camera, and the right side is a second image with a hand taken by the second camera, Figure 5 The triangle in the middle represents the key point to be annotated. The triangles in the first image and the second image correspond to the pixel points at the same position of the actual hand, such as the base joint position of the index finger. Since the position of the triangle is obtained manually, the first pixel coordinates of the triangle in the first image are known to be (u1, v1), and the second pixel coordinates of the triangle in the second image are known to be (u2, v2). It can be understood that Figure 6 , the first pixel coordinate (u1, v1) and the second pixel coordinate (u2, v2) are coordinates on the corresponding image plane.
[0092] When the first pixel coordinates (u1, v1) are known, the first transition coordinates (x1, y1) corresponding to the first pixel coordinates can be calculated using the focal lengths (fx1 and fy1) in the intrinsic parameter matrix of the first camera and the optical center (x01, y01) of the first camera. The formula is as follows:
[0093] x1=(u1-x01) / fx1.
[0094] y1=(v1-y01) / fy1.
[0095] S120: Determine second transition coordinates using second pixel coordinates and an intrinsic parameter matrix of the second camera.
[0096] When the second pixel coordinates (u2, v2) are known, the second transition coordinates (x2, y2) corresponding to the second pixel coordinates can be calculated using the focal lengths (fx2 and fy2) in the intrinsic parameter matrix of the second camera and the optical center (x02, y02) of the second camera. The formula is as follows:
[0097] x2=(u2-x02) / fx2.
[0098] y2=(v2-y02) / fy2.
[0099] S130: Calculate the four-dimensional coordinates corresponding to the nth key point by using the posture relationship between the first camera and the second camera, the first transition coordinates, and the second transition coordinates.
[0100] For example, we can use the opencv function cv2.triangulatePoints(T1,T2,point1,point2,Points_4d) to obtain the four-dimensional coordinates Points_4d corresponding to the nth key point. Where point1 = (x1,y1), point2 = (x2,y2), R 12is the rotation matrix of the first camera relative to the second camera, T 12 is the translation vector of the first camera relative to the second camera.
[0101] S140: Dividing the coordinate values of the first three dimensions of the four-dimensional coordinate by the coordinate value of the fourth dimension respectively to obtain the first three-dimensional coordinate.
[0102] Among them, the four-dimensional coordinate Points_4d is a homogeneous coordinate with four dimensions. The coordinate values of the first three dimensions need to be divided by the coordinate value of the fourth dimension to finally obtain the first three-dimensional coordinate (x c1 ,y c1 ,z c1 ).
[0103] S200: Using the first three-dimensional coordinates to mark the nth key point in the first image.
[0104] The first three-dimensional coordinate can be used as the three-dimensional annotation of the nth key point in the first image. When the nth key point in the first image needs to be two-dimensionally annotated, the first pixel coordinate can be used as the two-dimensional annotation of the nth key point.
[0105] It should be noted that the above steps S100 and S200 need to be performed N times to achieve the labeling of N key points. For example, step S100 may be performed N times, and then step S200 may be performed N times; or step S100 may be performed once, and then step S200 may be performed once, and then steps S100 and S200 may be performed again, and this cycle may be repeated N times to achieve the labeling of N key points.
[0106] The first camera and the second camera used in this embodiment are both monocular cameras, which use two-dimensional images to obtain the absolute depth of each key point, thereby solving the problem that the existing depth camera cannot obtain the absolute depth of the blocked part.
[0107] Furthermore, in order to verify the accuracy of the three-dimensional annotations added to the first image, at least three images including the target object can be acquired at the same time using at least three cameras, two of which are used to determine the three-dimensional coordinates of each key point, and the remaining images are used to verify the accuracy of the three-dimensional coordinates of each key point.
[0108] Exemplarily, when at least three images including the target object are acquired by using at least three cameras at the same time, since the labeled images are to be used as training samples for the training model, two images with the least obscured key points to be labeled can be selected from multiple images including the target object as the first image and the second image to ensure that the key points in each training sample have higher clarity.
[0109] Furthermore, the first pixel coordinates of the nth key point in the first image and the second pixel coordinates in the second image are determined respectively, wherein the first pixel coordinates in the first image and the second pixel coordinates in the second image are manually calibrated in sequence.
[0110] For further information, see Figure 7 The following verification methods S191 to S194 may be used to verify the accuracy of the three-dimensional coordinates corresponding to the first pixel coordinates in the first image.
[0111] S191: Calculate the first three-dimensional coordinate corresponding to the first three-dimensional coordinate in the i-th camera coordinate system based on the posture relationship between the first camera and the i-th camera, 3≤i≤I, I is the total number of cameras.
[0112] Exemplarily, the first three-dimensional coordinate corresponding to the first three-dimensional coordinate in the i-th camera coordinate system may be determined using the following formula:
[0113]
[0114] Among them, (x c1 ,y c1 ,z c1 ) is the first three-dimensional coordinate in the first camera coordinate system, (x ci ,y ci ,z ci ) is the first three-dimensional coordinate corresponding to the first three-dimensional coordinate in the i-th camera coordinate system, R 1i is the rotation matrix from the first camera coordinate system to the i-th camera coordinate system, T 1i is the translation vector from the first camera coordinate system to the i-th camera coordinate system.
[0115] S192: Determine pixel coordinates corresponding to the first three-dimensional coordinates in the i-th image, mark the corresponding pixel coordinates in the i-th image with a first mark, and then display the i-th image with the first mark.
[0116] According to the conversion relationship between the camera coordinate system and the image pixel coordinate system, it can be determined that (x ci ,y ci ,z ci) pixel coordinates in the pixel coordinate system of the ith image, marking the corresponding pixel coordinates in the ith image with the first mark and then displaying the ith image with the first mark, so as to determine whether the first mark of the ith image corresponds to the nth key point.
[0117] S193: Determine whether all first marks of each image correspond to the nth key point.
[0118] If the first marks of each image correspond to the nth key point, it means that the first three-dimensional coordinate accurately marks the nth key point in the first image, and then execute the above step S200. If there is at least one image whose first mark deviates from the nth key point, it means that the first three-dimensional coordinate has an error in marking the nth key point in the first image, and then execute step S194: update the first pixel coordinate and / or the second pixel coordinate, and then re-execute step S100.
[0119] Further, to ensure the accuracy of the annotation, please refer to Figure 8 , after step S200, the method further includes S300 and S400, in which the second image is annotated and the annotated second image is used as a training sample.
[0120] S300: Calculating a second three-dimensional coordinate of the first three-dimensional coordinate in a second camera coordinate system based on a posture relationship between the first camera and the second camera.
[0121] For example, the following formula can be used to determine the second three-dimensional coordinate corresponding to the first three-dimensional coordinate in the second camera coordinate system:
[0122]
[0123] Among them, (xc1, yc1, zc1) is the first three-dimensional coordinate in the first camera coordinate system, (xc2, yc2, zc2) is the second three-dimensional coordinate corresponding to the first three-dimensional coordinate in the second camera coordinate system, R12 is the rotation matrix from the first camera coordinate system to the second camera coordinate system, and T12 is the translation vector from the first camera coordinate system to the second camera coordinate system.
[0124] S400: Marking the nth key point in the second image using the second three-dimensional coordinates.
[0125] The second three-dimensional coordinate can be used as the three-dimensional annotation of the nth key point in the second image. When the nth key point in the second image needs to be two-dimensionally annotated, the second pixel coordinate can be used as the two-dimensional annotation of the nth key point.
[0126] For further information, see Fig. 9 , the accuracy of the second three-dimensional coordinate can also be verified by using the following verification methods S391 to S393 before S400.
[0127] S391: Calculate the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the i-th camera coordinate system based on the posture relationship between the second camera and the i-th camera, 3≤i≤I, I is the total number of cameras.
[0128] Exemplarily, the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the i-th camera coordinate system may be determined using the following formula:
[0129]
[0130] Among them, (x c2 ,y c2 ,z c2 ) is the second three-dimensional coordinate in the second camera coordinate system, (x ci ,y ci ,z ci ) is the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the i-th camera coordinate system, R 2i is the rotation matrix from the second camera coordinate system to the i-th camera coordinate system, T 2i is the translation vector from the second camera coordinate system to the i-th camera coordinate system.
[0131] S392: Determine the pixel coordinates corresponding to the second three-dimensional coordinates in the i-th image, mark the corresponding pixel coordinates in the i-th image with a second mark, and then display the i-th image with the second mark.
[0132] S393: Determine whether the second marks of each image all correspond to the nth key point.
[0133] If the second marks of each image correspond to the nth key point, it means that the second three-dimensional coordinate accurately marks the nth key point in the second image, and then execute the above step S400. If there is at least one image whose second mark deviates from the nth key point, it means that the second three-dimensional coordinate has an error in marking the nth key point in the second image, and then execute step S194: update the first pixel coordinate and / or the second pixel coordinate, and then re-execute step S100.
[0134] For another embodiment of the present application, see Fig.10 , an image sample labeling device 10 is proposed, including: a calculation module 11 and a labeling module 12.
[0135] The calculation module 11 is used to calculate the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image and the second pixel coordinate of the nth key point in the second image, the first image and the second image are respectively acquired by the first camera and the second camera at the same time, 1≤n≤N, N is the total number of key points to be annotated; the annotation module 12 is used to use the first three-dimensional coordinate to annotate the predetermined position in the first image.
[0136] Furthermore, in the case where at least three images including the target object are acquired by using at least three cameras at the same time, before calculating the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image, it also includes: selecting two images with the least obscured key points to be annotated from multiple images including the target object as the first image and the second image; and respectively determining the first pixel coordinate of the nth key point in the first image and the second pixel coordinate of the second image.
[0137] Furthermore, before the nth key point in the first image is marked with the first three-dimensional coordinate, the method further includes: calculating the first three-dimensional coordinate corresponding to the first three-dimensional coordinate in the i-th camera coordinate system based on the posture relationship between the first camera and the i-th camera, 3≤i≤I, I is the total number of cameras; determining the pixel coordinates corresponding to the first three-dimensional coordinate in the i-th image, and displaying the i-th image with the first mark after marking the corresponding pixel coordinates in the i-th image with the first mark; if the first marks of each image correspond to the n-th key point, marking the n-th key point in the first image with the first three-dimensional coordinate is executed; if there is at least one image whose first mark deviates from the n-th key point, updating the first pixel coordinate and / or the second pixel coordinate, and re-executing the calculation of the first three-dimensional coordinate of the n-th key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the n-th key point of the target object in the first image, and the second pixel coordinate of the n-th key point in the second image.
[0138] Furthermore, the calculation module 11 is also used to calculate the second three-dimensional coordinate of the first three-dimensional coordinate in the second camera coordinate system based on the posture relationship between the first camera and the second camera; the labeling module 12 is also used to use the second three-dimensional coordinate to label the nth key point in the second image.
[0139] Furthermore, before the nth key point in the second image is marked with the second three-dimensional coordinate, it also includes: calculating the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the i-th camera coordinate system based on the posture relationship between the second camera and the i-th camera, 3≤i≤I, I is the total number of cameras; determining the pixel coordinates corresponding to the second three-dimensional coordinate in the i-th image, and displaying the i-th image with the second mark after marking the corresponding pixel coordinates in the i-th image with the second mark; if the second marks of each image correspond to the n-th key point, then executing the marking of the n-th key point in the second image with the second three-dimensional coordinate; if there is at least one image whose second mark deviates from the n-th key point, then updating the first pixel coordinate and / or the second pixel coordinate, and re-executing the calculation of the first three-dimensional coordinate of the n-th key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the n-th key point of the target object in the first image, and the second pixel coordinate of the n-th key point in the second image.
[0140] Furthermore, the first three-dimensional coordinate of the nth key point in the first camera coordinate system is calculated based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image, including: determining the first transition coordinate using the first pixel coordinate and the intrinsic parameter matrix of the first camera; determining the second transition coordinate using the second pixel coordinate and the intrinsic parameter matrix of the second camera; calculating the four-dimensional coordinate corresponding to the nth key point using the posture relationship between the first camera and the second camera, the first transition coordinate and the second transition coordinate; and obtaining the first three-dimensional coordinate by dividing the coordinate values of the first three dimensions of the four-dimensional coordinate by the coordinate value of the fourth dimension.
[0141] Furthermore, it also includes: predetermining the posture relationship between any two cameras, wherein the posture relationship includes a rotation matrix and a translation vector.
[0142] Exemplarily, the predetermining the posture relationship between any two cameras includes: determining the intrinsic parameter matrix and distortion coefficient of each camera; determining the basic rotation vector and basic translation vector of each camera using the intrinsic parameter matrix and distortion coefficient of each camera; and calculating the rotation matrix and translation vector between any two cameras using the basic rotation vector and basic translation vector of any two cameras.
[0143] The image sample labeling device 10 disclosed in this embodiment is used to execute the image sample labeling method described in the above embodiment through the cooperation of the calculation module 11 and the labeling module 12. The implementation scheme and beneficial effects involved in the above embodiment are also applicable to this embodiment and will not be repeated here.
[0144] For another embodiment of the present application, see Fig.11 , a terminal device 100 is proposed, including a memory 110 and a processor 120, the memory 110 stores a computer program, and when the computer program runs on the processor 120, the image sample labeling method described in the embodiment of the present application is executed.
[0145] See also Fig.12 The embodiment of the present application also relates to an image annotation system, including multiple cameras and the terminal device described in the embodiment of the present application. The present application can obtain the 3D coordinates of any pixel point without the help of a depth camera, and can flexibly collect gesture images under different backgrounds by moving the position and direction of each camera.
[0146] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and structure diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the structure diagram and / or the flow chart, and the combination of boxes in the structure diagram and / or the flow chart, can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0147] In addition, the functional modules or units in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0148] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a smart phone, a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned readable storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0149] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. An image sample annotation method, characterized in that, it includes: Calculating the first three-dimensional coordinate of the nth key point of the target object in the first camera coordinate system based on the pose relationship between the first camera and the second camera, the first pixel coordinate of the nth key point in the first image, and the second pixel coordinate of the nth key point in the second image. Specifically, it includes: determining a first intermediate coordinate using the first pixel coordinate and the internal parameter matrix of the first camera; determining a second intermediate coordinate using the second pixel coordinate and the internal parameter matrix of the second camera; calculating the four-dimensional coordinate corresponding to the nth key point using the pose relationship between the first camera and the second camera, the first intermediate coordinate, and the second intermediate coordinate; obtaining the first three-dimensional coordinate by dividing the coordinate values of the first three dimensions in the four-dimensional coordinate by the coordinate value of the fourth dimension; wherein, the first image and the second image are respectively acquired by the first camera and the second camera at the same moment, 1≤n≤N, and N is the total number of key points to be annotated; Annotating the nth key point in the first image using the first three-dimensional coordinate.
2. The image sample annotation method according to claim 1, characterized in that, when acquiring at least three images including the target object using at least three cameras at the same moment, before calculating the first three-dimensional coordinate of the nth key point of the target object in the first camera coordinate system based on the pose relationship between the first camera and the second camera, the first pixel coordinate of the nth key point in the first image, and the second pixel coordinate of the nth key point in the second image, it further includes: Selecting two images with the fewest occluded key points to be annotated from the multiple images including the target object as the first image and the second image; Respectively determining the first pixel coordinate of the nth key point in the first image and the second pixel coordinate in the second image.
3. The image sample annotation method according to claim 2, characterized in that, before annotating the nth key point in the first image using the first three-dimensional coordinate, it further includes: Calculating the corresponding first three-dimensional coordinate of the first three-dimensional coordinate in the ith camera coordinate system based on the pose relationship between the first camera and the ith camera, 3≤i≤I, and I is the total number of cameras; Determining the pixel coordinate corresponding to the first three-dimensional coordinate in the ith image, and displaying the ith image with the first mark after marking the corresponding pixel coordinate in the ith image with the first mark; If the first marks of each image all correspond to the nth key point, then execute the step of annotating the nth key point in the first image using the first three-dimensional coordinate; If there is a first marker in at least one image that deviates from the nth key point, update the first pixel coordinate and / or the second pixel coordinate, and re-execute the calculation of the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image.
4. The image sample annotation method according to claim 2, It is characterized in that Also includes: Calculate a second three-dimensional coordinate of the first three-dimensional coordinate in a second camera coordinate system based on a posture relationship between the first camera and the second camera; The nth key point in the second image is marked using the second three-dimensional coordinate.
5. The image sample annotation method according to claim 4, It is characterized in that Before the nth key point in the second image is marked by using the second three-dimensional coordinate, the method further includes: Calculate the second three-dimensional coordinate corresponding to the second three-dimensional coordinate in the coordinate system of the i-th camera based on the posture relationship between the second camera and the i-th camera, 3≤i≤I, I is the total number of cameras; Determine the pixel coordinates corresponding to the second three-dimensional coordinates in the i-th image, and after marking the corresponding pixel coordinates in the i-th image with a second mark, display the i-th image with the second mark; If the second marks of each image correspond to the nth key point, then performing the step of marking the nth key point in the second image by using the second three-dimensional coordinate; If there is a second marker in at least one image that deviates from the nth key point, update the first pixel coordinate and / or the second pixel coordinate, and re-execute the method of calculating the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image, and the second pixel coordinate of the nth key point in the second image.
6. The image sample annotation method according to any one of claims 1 to 5, It is characterized in that Also includes: The position relationship between any two cameras is predetermined, and the position relationship includes a rotation matrix and a translation vector.
7. The image sample annotation method according to claim 6, It is characterized in that The step of predetermining the position relationship between any two cameras includes: Determine the intrinsic parameter matrix and distortion coefficient of each camera; Determine the basic rotation vector and basic translation vector of each camera using the intrinsic parameter matrix and distortion coefficient of each camera; The rotation matrix and the translation vector between any two cameras are calculated using the basic rotation vectors and the basic translation vectors of any two cameras.
8. An image sample annotation device, It is characterized in that include: A calculation module, for calculating the first three-dimensional coordinate of the nth key point in the first camera coordinate system based on the posture relationship between the first camera and the second camera, the first pixel coordinate of the nth key point of the target object in the first image and the second pixel coordinate of the nth key point in the second image, specifically comprising: determining the first transition coordinate using the first pixel coordinate and the intrinsic parameter matrix of the first camera; determining the second transition coordinate using the second pixel coordinate and the intrinsic parameter matrix of the second camera; calculating the four-dimensional coordinate corresponding to the nth key point using the posture relationship between the first camera and the second camera, the first transition coordinate and the second transition coordinate; obtaining the first three-dimensional coordinate by dividing the coordinate values of the first three dimensions of the four-dimensional coordinate by the coordinate value of the fourth dimension; wherein the first image and the second image are respectively acquired by the first camera and the second camera at the same time, 1≤n≤N, and N is the total number of key points to be annotated; A labeling module is used to label a predetermined position in the first image using the first three-dimensional coordinates.
9. A terminal device, It is characterized in that The method comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is run on the processor, the image sample labeling method according to any one of claims 1 to 7 is executed.
10. An image annotation system, It is characterized in that The invention comprises a plurality of cameras and the terminal device as claimed in claim 9.
11. A readable storage medium, It is characterized in that A computer program is stored therein, and when the computer program is run on a processor, the image sample labeling method according to any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Key point automatic labeling method and system, electronic device and storage medium
CN113393563A