Calibration device, three-dimensional coordinate detection system, and computer program
The calibration device enhances the accuracy of transformation matrices between multiple three-dimensional vision sensors by selecting feature points with minimal variation, improving the precision of motion capture systems.
Patent Information
- Application Number
- JP2023215476
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-07-03
AI Technical Summary
Existing technologies for motion capture using multiple three-dimensional vision sensors, such as Kinects, suffer from inaccuracies in the transformation matrix that corrects for the installation position of each sensor, affecting the accuracy of coordinate system integration.
A calibration device that includes a transformation matrix generation unit to integrate the coordinate systems of multiple three-dimensional vision sensors into a reference coordinate system by selecting feature points with less variation, performing coordinate transformation, and calculating a transformation matrix with minimized error.
Improves the accuracy of the transformation matrix, reducing errors in coordinate system conversion and enhancing the precision of motion capture, particularly for human movements, by using reliable joint points and optimizing the transformation matrix.
Smart Images

Figure 2025099090000001_ABST
Abstract
Description
[Technical field]
[0001] The present disclosure relates to a calibration device. [Background technology]
[0002] Conventionally, a technology has been proposed for detecting human movements using a three-dimensional visual sensor. For example, Non-Patent Documents 1 and 2 disclose a technology for detecting human movements using multiple Kinects (registered trademark) having sensors such as a color sensor and a depth sensor. [Prior art documents] [Patent documents]
[0003] [Non-Patent Document 1] Junpei Miyatake and 2 others, "Realizing motion capture using multiple Kinects," [online], HAI Symposium 2016, G-22 (2016), [Retrieved July 23, 2023], Internet<URL:https: / / hai-conference.net / proceedings / HAI2016 / pdf / G-22.pdf> [Non-Patent Document 2] Alexandros Kitsikidis and 3 others, "Dance Analysis using Multiple Kinect Sensors", [online], VISAPP2014, 2014. [Retrieved July 23, 2023], Internet<URL:https: / / ieeexplore.ieee.org / document / <7295020> Summary of the Invention [Problem to be solved by the invention]
[0004] In the technologies described in Non-Patent Documents 1 and 2 above, motion capture is performed using a plurality of Kinects (registered trademark), thereby reducing occlusion (invisible parts) and improving the accuracy of tracking. However, in Non-Patent Documents 1 and 2 above, the accuracy of the transformation matrix for correcting the deviation caused by the installation position of each Kinect (registered trademark) was not sufficient.
[0005] Such a problem is not limited to the case where the detection target is a person, and is a common problem in technologies for detecting the movements of various moving objects such as other animals like dogs and vehicles.
[0006] The present invention has been made to solve the above-described problems, and an object thereof is to provide a technique for improving the accuracy of a transformation matrix that converts the coordinate systems of a plurality of three-dimensional vision sensors into the coordinate system of one three-dimensional vision sensor.
Means for Solving the Problems
[0007] The present invention has been made to solve at least a part of the above-described problems and can be realized in the following forms.
[0008] (1) According to one aspect of the present invention, a calibration device is provided. This calibration device includes a reference 3D vision sensor and a conversion matrix generation unit that obtains a conversion matrix for integrating the coordinate systems of the first to nth (n is a natural number) 3D vision sensors into the reference coordinate system of the reference 3D vision sensor from the 3D information of the object generated by a plurality of 3D vision sensors including the first to nth 3D vision sensors. The plurality of 3D vision sensors generate a plurality of frames having the 3D information synchronously in time. In the calibration device, for each of the first to nth 3D vision sensors, the conversion matrix generation unit uses the initial value of the conversion matrix generated by selecting feature points with less variation from a plurality of feature points extracted from the 3D information of the object, performs coordinate transformation on all of the plurality of frames, integrates them into the reference coordinate system, calculates the variation between the plurality of 3D vision sensors for each frame using the 3D coordinate values integrated into the reference coordinate system for each of the first to nth 3D vision sensors, and sets the initial value of the conversion matrix of the frame with less variation as the conversion matrix.
[0009] According to this configuration, a frame with less variation is selected from all the frames subjected to coordinate transformation using the initial value of the conversion matrix, and the initial value of the conversion matrix used in the selected frame is set as the conversion matrix for integrating the coordinate systems of the first to nth 3D vision sensors into the reference coordinate system of the reference 3D vision sensor. Therefore, a conversion matrix with a small error can be generated. That is, the accuracy of the conversion matrix can be improved.
[0010] (2) In the calibration device of the above aspect, the feature points may be joint points of the human body skeleton. By doing so, the accuracy of motion capture targeting a person can be improved.
[0011] (3) According to another aspect of the present invention, a three-dimensional coordinate detection system for detecting the three-dimensional coordinates of an object is provided. This three-dimensional coordinate detection system includes a reference three-dimensional vision sensor, a plurality of three-dimensional vision sensors including the first to nth (n is a natural number) three-dimensional vision sensors, and the calibration device of the above aspect. According to this configuration, since errors due to coordinate system conversion between a plurality of three-dimensional sensors can be suppressed, the accuracy of motion capture can be improved.
[0012] (4) According to another aspect of the present invention, a computer program for obtaining a transformation matrix between the reference three-dimensional vision sensor and the sensor coordinate systems of each of a plurality of three-dimensional vision sensors including the first to nth (n is a natural number) three-dimensional vision sensors is provided. The plurality of three-dimensional vision sensors generate a plurality of frames each having three-dimensional information of an object in time synchronization. This computer program causes an information processing device to, for each of the first to nth three-dimensional vision sensors, select feature points with little variation from a plurality of feature points extracted from the three-dimensional information of the object to generate an initial transformation matrix value, perform coordinate transformation on all of the plurality of frames, and integrate them into the reference coordinate system of the reference three-dimensional vision sensor; and for each of the first to nth three-dimensional vision sensors, calculate the variation between the plurality of three-dimensional vision sensors for each frame using the three-dimensional coordinate values integrated into the reference coordinate system, and use the initial transformation matrix value of the frame with little variation as the transformation matrix. According to this computer program, the information processing device can generate a transformation matrix with small errors.
[0013] Note that the present disclosure can be realized in various aspects, for example, in the form of a method for generating a transformation matrix, a server device for distributing a computer program for causing a computer to execute the method for generating a transformation matrix, a non-transitory storage medium storing the computer program, and the like.
Brief Description of the Drawings
[0014]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0015] <Embodiment> FIG. 1 is an explanatory diagram conceptually showing an example of a three-dimensional coordinate detection system 1 as an embodiment of the present disclosure. The three-dimensional coordinate detection system includes a reference three-dimensional vision sensor, first to nth (n is a natural number) three-dimensional vision sensors, and a calibration device. In the example shown in FIG. 1, the three-dimensional coordinate detection system 1 includes a reference three-dimensional vision sensor S0, a first three-dimensional vision sensor S1, and a calibration device 100. The reference three-dimensional vision sensor S0 and the first three-dimensional vision sensor S1 are collectively referred to as "a plurality of three-dimensional vision sensors". Also, when the reference three-dimensional vision sensor S0 and the first three-dimensional vision sensor S1 are not distinguished, they are simply referred to as "three-dimensional vision sensor S". In FIG. 1, for simplicity of explanation, an example in which the number of three-dimensional vision sensors is two is shown.
[0016] The three-dimensional vision sensor S includes sensors such as a depth sensor and a stereo camera, and generates three-dimensional information D of the object H. The three-dimensional information includes a plurality of three-dimensional coordinate values. As the three-dimensional vision sensor S, for example, Kinect (registered trademark) can be used. Kinect (registered trademark) has an RGB camera, a depth sensor, a processor for operating dedicated software, and the like. In the example shown in FIG. 1, Kinect (registered trademark) is used as the three-dimensional vision sensor S.
[0017] In the example shown in FIG. 1, the three-dimensional vision sensor S detects a person as the object H, extracts the joint points of the human body skeleton from the detection result (point cloud representing the three-dimensional shape of the person) by estimating the human body skeleton, and outputs the three-dimensional coordinate values of each joint point as three-dimensional information D to the calibration device 100. The three-dimensional vision sensor S detects about thirty points as the joint points of the human body skeleton. When distinguishing the three-dimensional information D generated by each of the reference three-dimensional vision sensor S0 and the first three-dimensional vision sensor S1, the three-dimensional information D output by the reference three-dimensional vision sensor S0 is called "three-dimensional information D0", and the three-dimensional information D output by the first three-dimensional vision sensor S1 is called "three-dimensional information D1".
[0018] A plurality of three-dimensional vision sensors generate a plurality of frames having three-dimensional information in time synchronization. In the example shown in FIG. 1, the reference three-dimensional vision sensor S0 generates k frames 0 (k is a natural number of 2 or more) including the three-dimensional information D0, and the first three-dimensional vision sensor S1 generates k frames 1 (k is a natural number of 2 or more) including the corresponding three-dimensional information D1.
[0019] The calibration device 100 is a computer including a ROM, a RAM, and a CPU, and functionally includes a transformation matrix generation unit 10 and a feature point extraction unit 20. The transformation matrix generation unit 10 generates a transformation matrix for integrating the coordinate systems of the first to nth three-dimensional vision sensors S1 to Sn into the reference coordinate system of the reference three-dimensional vision sensor S0. In the example shown in FIG. 1, a transformation matrix M for integrating the coordinate system of the three-dimensional vision sensor S1 into the reference coordinate system of the reference three-dimensional vision sensor S0 is generated.
[0020] Based on the three-dimensional information D input from the three-dimensional vision sensor S, the feature point extraction unit 20 extracts predetermined joint points as feature points. In the example shown in FIG. 1, among the joint points extracted by the human body skeleton estimation in the three-dimensional vision sensor S, the joint points of the torso part that can be measured relatively stably are extracted as the feature points P. Further, in order to stably obtain the transformation matrix, using the reliability of the joint points input from the three-dimensional vision sensor S, the joint points of the arms and legs with high reliability are selected as the feature points P. For example, the joint points with small standard deviation and average error are selected as the joint points with high reliability. In this example, the feature point extraction unit 20 extracts the 18 points shown in FIG. 1 as the feature points P. In the following description, the feature point P detected by the reference three-dimensional vision sensor S0 is defined as the feature point P0, and the feature point P detected by the three-dimensional vision sensor S1 is defined as the feature point P1.
[0021] As described above, since k frames are generated in the three-dimensional vision sensor S, as shown in the figure, the feature point extraction unit 20 extracts the feature points P for all the frames. Each feature point P extracted in each of the k frames corresponds between the frames. Also, each feature point P corresponds between the three-dimensional vision sensors S.
[0022] The transformation matrix generation unit 10 selects three or more redundant corresponding feature points P0, P1 from the feature points P0, P1 extracted by the feature point extraction unit 20, and uses these feature points to obtain, for each frame, a transformation matrix that minimizes the mean square error between the three-dimensional points after coordinate transformation. Thereby, a robust transformation matrix can be obtained. The transformation matrix generation unit 10 uses the transformation matrix obtained as described above as the initial value of the transformation matrix, and generates a transformation matrix with a smaller error using the coordinate transformation result obtained using the initial value of the transformation matrix (to be described in detail later).
[0023] FIG. 2 is a flowchart showing an example of the flow of the transformation matrix generation process in the calibration device 100. In step S102, the calibration device 100 acquires three-dimensional information D from a plurality of three-dimensional vision sensors S. In the example of FIG. 1, the calibration device 100 acquires three-dimensional information D0 and three-dimensional information D1 from the reference three-dimensional vision sensor S0 and the first three-dimensional vision sensor S1, respectively.
[0024] In step S104, the feature point extraction unit 20 extracts a plurality of feature points P from the input three-dimensional information D. In the example shown in FIG. 1, for each of the k frames 0, the feature point extraction unit 20 extracts 18 feature points P0 each, and for each of the k frames 1, extracts 18 feature points P1 each.
[0025] In step S106, the transformation matrix generation unit 10 selects three or more redundant feature points from the feature points extracted in step S104, and generates, for each frame, an initial value of a transformation matrix for transforming the coordinate system of each of the first to nth three-dimensional vision sensors into the coordinate system of the reference three-dimensional vision sensor S0. In the example shown in FIG. 1, the transformation matrix generation unit 10 generates, for each frame, an initial value of a transformation matrix for transforming the coordinate system of the first three-dimensional vision sensor S1 into the reference coordinate system of the reference three-dimensional vision sensor S0.
[0026] In step S108, the transformation matrix generation unit 10 uses the initial value of the transformation matrix generated for each of the first to nth three-dimensional vision sensors to transform the feature point P detected by the corresponding three-dimensional vision sensor into the reference coordinate system. In the example shown in FIG. 1, for each of the k frames, the transformation matrix generation unit 10 uses the initial value of the transformation matrix to transform the feature point P1 into the reference coordinate system.
[0027] In step S110, the transformation matrix generation unit 10 performs error calculation using the feature points after the transformation in step S108. Taking the case of using the reference three-dimensional vision sensor S0, the first three-dimensional vision sensor S1, and the second three-dimensional vision sensor S2 as an example, the error calculation will be described with reference to FIGS. 3 to 5.
[0028] FIG. 3 is an explanatory diagram showing the arrangements of the reference 3D vision sensor S0, the first 3D vision sensor S1, and the second 3D vision sensor S2. FIG. 3 is a view from above of the room in which the 3D coordinate detection system 1 is arranged. FIG. 4 is an explanatory diagram showing the detection results of the joint points of the human body skeleton by each 3D vision sensor S. FIG. 5 is an explanatory diagram showing a part of the joint points of the human body skeleton transformed using the initial value of the transformation matrix.
[0029] As shown in FIG. 3, the reference 3D vision sensor S0 is arranged in front of the object H (person), the first 3D vision sensor S1 is arranged at the rear right of the object H, and the second 3D vision sensor S2 is arranged at the rear left of the object H.
[0030] FIG. 4 schematically shows the joint points of the human body skeleton detected in a certain frame generated synchronously in time among the reference 3D vision sensor S0, the first 3D vision sensor S1, and the second 3D vision sensor S2. FIG. 4 is created based on the frame image actually generated using Kinect (registered trademark).
[0031] In the calibration device 100, using the initial value of the transformation matrix generated as described above, the coordinate systems (FIG. 4) of the first 3D vision sensor S1 and the second 3D vision sensor S2 are integrated into the reference coordinate system of the reference 3D vision sensor S0, and FIG. 5 shows them superimposed. In FIG. 5, the coordinate system of the reference 3D vision sensor S0 is shown by black circles and solid lines, the transformed coordinate system of the first 3D vision sensor S1 is shown by white circles and dashed lines, the transformed coordinate system of the second 3D vision sensor S2 is shown by black squares and solid lines, and the average of these three is shown by white squares and dashed lines.
[0032] As shown in FIG. 5, the variation is different for each joint point. In FIG. 5, the joint points with low reliability are marked with circles. As shown in FIG. 5, let the coordinates of the joint points by the reference 3D vision sensor S0 be (x j0 , y j0 , z j0 ), and the transformed coordinates of the joint points by the first 3D vision sensor S1 be (x j1 , y j1 , zj1 ) and let the transformed coordinates of the joint points by the second 3D vision sensor S2 be (x j2 , y j2 , z j2 ). Let the average coordinates of these be (x jm , y jm , z jm ). Here, j is the joint point number. Also, let the distance between the joint point and the average by the reference 3D vision sensor S0 be d j0 , the distance between the joint point and the average by the first 3D vision sensor S1 be d j1 , and the distance between the joint point and the average by the second 3D vision sensor S2 be d j2 .
[0033] Using the above average standard deviation and average error of the joint points, calculate the variation for each frame. The average standard deviation is obtained by the following (Equation 1), and the average error is obtained by the following (Equation 2).
Equation
Equation
[0034] Returning to Figure 2, in step S112, the transformation matrix generation unit 10 compares the errors (variations) of the joint points for each frame and selects the frame with the smallest error. Select the initial values of the transformation matrices in the first 3D vision sensor S1 and the second 3D vision sensor S2 used in the selected frame as the transformation matrix M and record it in the calibration device 100. In this way, in step S112, the coordinate transformation error is evaluated by the variation of the joint points (feature points) transformed into the reference coordinate system. The same applies when using the above-mentioned average standard deviation as the variation. The fact that the average standard deviation and the average error are small indicates that the joint points of each 3D vision sensor are gathered at nearby positions.
[0035] Assuming that the points acquired by the reference 3D vision sensor S0 are P0 and the points acquired by the first 3D vision sensor S1 are P1, the transformation matrix M can be expressed as follows.
[0036]
Number
[0037] In step S114, the transformation matrix generation unit 10 determines whether the error (hereinafter referred to as the minimum error) in the frame selected in step S112 is smaller than a predetermined threshold. If the minimum error is greater than or equal to the predetermined threshold, the process returns to step S102 to acquire 3D information at different times. That is, the calibration device 100 repeats steps 112 to S114 until the minimum error (transformation error) becomes smaller than the predetermined threshold. The predetermined threshold can be arbitrarily set. For example, it can be set to 0.02 cm (root mean square error). The threshold is preferably determined according to the variation of the feature points. For example, when the variation is large, the threshold is increased, and when the variation is small, the threshold is decreased.
[0038] If the minimum error is smaller than the predetermined threshold, the process proceeds to step S116. In step S116, the transformation matrix generation unit 10 uses the transformation matrix M selected in step S112 to perform transformation of all coordinate values included in the frame selected in step S112 into the coordinate system of the reference 3D vision sensor S0 for each of the first to nth 3D vision sensors. At this time, transformation is performed not only on the feature points and joint points but also on all coordinate values (point group representing the three-dimensional shape of the object H) acquired by the 3D vision sensor S.
[0039] FIG. 6 is an explanatory diagram showing an example of point cloud data transformed into a reference coordinate system. In FIG. 6, the point cloud acquired by the reference three-dimensional vision sensor S0 is indicated by white circles. Also, using the transformation matrix M selected in step S112, the integrated point cloud obtained by integrating the point cloud acquired by the first three-dimensional vision sensor S1 into the coordinate system of the reference three-dimensional vision sensor S0 is indicated by black triangles. In FIG. 6, each point constituting the point cloud representing the three-dimensional shape of the object H is denoted as SP.
[0040] As shown in the figure, there is an overlapping region between the two, and the transformation matrix M is updated so that the error in the overlapping region is reduced. Here, the overlapping region is a region that is commonly measured by both sensors for each sensor pair (for example, the pair of the reference three-dimensional vision sensor S0 and the first three-dimensional vision sensor S1, the pair of the reference three-dimensional vision sensor S0 and the second three-dimensional vision sensor S2). The transformation matrix generation unit 10 detects this overlapping region and updates the transformation matrix M so that the error in the overlapping region is reduced. In other words, the transformation matrix generation unit 10 searches for a transformation matrix M that minimizes the average distance dp between the re-nearest points of the point cloud in the overlapping region by a known method.
[0041] As described above, according to the calibration apparatus 100 of the present embodiment, frames with little variation are selected from all the frames in which coordinate transformation is performed using the initial transformation matrix value, and the transformation matrix used in the selected frames is used as the transformation matrix M for integrating the coordinate systems of the first to nth three-dimensional vision sensors S1 to Sn into the reference coordinate system of the reference three-dimensional vision sensor S0. Therefore, a transformation matrix with a small error can be generated.
[0042] Further, according to the calibration apparatus 100, when obtaining the initial transformation matrix value or selecting a frame, since the feature point P obtained by selecting a highly reliable joint point from the joint points output by the three-dimensional vision sensor S is used, compared with the case of generating the transformation matrix M using all the joint points output by the three-dimensional vision sensor S, the amount of calculation can be reduced. Therefore, the transformation matrix M can be obtained at high speed.
[0043] Furthermore, in the calibration device 100, when the minimum error that serves as the criterion for frame selection is equal to or greater than a predetermined threshold value, the transformation matrix M is regenerated using a frame group at another time to make the minimum error smaller than the predetermined threshold value, thereby further improving the accuracy of the transformation matrix M.
[0044] In addition, in the calibration device 100, since the transformation matrix M in the overlapping region is updated, the accuracy of the transformation matrix M can be further improved.
[0045] According to the three-dimensional coordinate detection system 1 of the present embodiment, since it includes a plurality of three-dimensional vision sensors S and can reduce occlusion (invisible parts) in the detection of the three-dimensional coordinates of the object H, the detection accuracy of the three-dimensional coordinates can be improved.
[0046] In addition, the three-dimensional coordinate detection system 1 includes a calibration device 100, and since the correction accuracy of the deviation caused by the installation positions of the plurality of three-dimensional vision sensors S is high, the detection accuracy of the three-dimensional coordinates of the object H can be further improved.
[0047] According to the three-dimensional coordinate detection system 1, since calibration is performed by the calibration device 100 using the human body joint points acquired by the three-dimensional vision sensor S, it is not necessary to perform a calibration operation using a specific marker or the like, and coordinate integration between a plurality of three-dimensional vision sensors S can be easily achieved.
[0048] According to the three-dimensional coordinate detection system 1, since a person's movement can be accurately measured, for example, the work load, work time, etc. in factory work can be accurately evaluated.
[0049] <Modification Example of the Present Embodiment> The present disclosure is not limited to the above-described embodiments, and can be implemented in various aspects without departing from the gist thereof. For example, the following modifications are also possible.
[0050] ·In the above-described embodiment, examples in which two 3D vision sensors S are used and examples in which three 3D vision sensors S are used are shown, but four or more 3D vision sensors S may be used.
[0051] ·In the above-described embodiment, a person is exemplified as the object H, and a joint point (human joint point) of the human body skeleton is exemplified as the feature point. However, the feature point is not limited to the human joint point. In addition to the human body skeleton, for example, feature points of the human face or joint points of the hand may be extracted as feature points.
[0052] ·Furthermore, among the human joint points extracted as feature points, the selection of the frame may be performed using the feature points excluding the feature points (for example, the elbow joint) that are known in advance to have a large error.
[0053] ·The detection target is not limited to a person, and various moving objects such as other animals such as dogs and vehicles can be targeted. The feature point may be any point that can define the object. For example, the vertices of a cube, the center of the tire of a vehicle, the headlight, etc., and the points that bend when the object moves can be used.
[0054] ·In the conversion matrix generation process (FIG. 2) of the above-described embodiment, steps S114 and S116 may not be performed. Even in this case, by having steps S108 to S112, the accuracy of the conversion matrix M can be improved.
[0055] ·In the above-described embodiment, an example in which the 3D vision sensor S detects the object H, estimates the human body skeleton from the point cloud representing the detected 3D shape, and extracts the joint points is shown. However, the 3D vision sensor may have at least a sensor (such as an RGB camera or a depth sensor) that detects the 3D shape of the object H. In that case, the calibration device 100 may be configured to have a function of estimating the human body skeleton from the detected point cloud and extracting the joint points. Further, the 3D coordinate detection system 1 may separately include a device that estimates the human body skeleton from the point cloud detected by the 3D vision sensor S and extracts the joint points, in addition to the calibration device 100.
[0056] ·In the above-described embodiment, an example in which the calibration device 100 has the feature point extraction unit 20 has been shown. However, the feature point extraction unit 20 may be provided in the three-dimensional vision sensor S. Further, the three-dimensional coordinate detection system 1 may separately include the feature point extraction unit 20 from the three-dimensional vision sensor S and the calibration device 100.
[0057] ·In each of the above embodiments, part of the configuration realized by hardware may be replaced with software, and conversely, part of the configuration realized by software may be replaced with hardware. Further, when part or all of the functions of the present disclosure are realized by software, the software (computer program) can be provided in a form stored in a computer-readable recording medium. The "computer-readable recording medium" is not limited to portable recording media such as flexible disks and CD-ROMs, but also includes various internal storage devices in a computer such as various RAMs and ROMs, and external storage devices fixed to a computer such as hard disks. That is, the "computer-readable recording medium" has a broad meaning including any recording medium capable of fixedly storing data packets not temporarily.
[0058] As described above, the present disclosure has been described based on the embodiments and modified examples. However, the embodiments of the above-described modes are for facilitating the understanding of the present disclosure and do not limit the present disclosure. The present disclosure can be changed and improved without departing from its spirit and the scope of the claims, and equivalents thereof are included in the present disclosure. Further, if the technical features are not described as essential in this specification, they can be deleted as appropriate.
[0059] The present disclosure can also be realized in the following forms. [Application Example 1] A calibration device, A conversion matrix generation unit that obtains a conversion matrix for integrating the coordinate systems of the first to nth (n is a natural number) 3D vision sensors into the reference coordinate system of the reference 3D vision sensor from the 3D information of the object generated by the reference 3D vision sensor and a plurality of 3D vision sensors including the first to nth 3D vision sensors. The plurality of 3D vision sensors each generate a plurality of frames having the 3D information in time synchronization. In the calibration device, The conversion matrix generation unit For each of the first to nth 3D vision sensors, coordinate conversion is performed on all of the plurality of frames using the initial conversion matrix generated by selecting feature points with little variation from the plurality of feature points extracted from the 3D information of the object, and integrated into the reference coordinate system. For each of the first to nth 3D vision sensors, using the 3D coordinate values integrated into the reference coordinate system, the variation between the plurality of 3D vision sensors is calculated for each frame, and the initial conversion matrix of the frame with little variation is used as the conversion matrix. Calibration device. [Application Example 2] The calibration device according to Application Example 1, wherein the feature points are joint points of a human body skeleton. Calibration device. [Application Example 3] A 3D coordinate detection system for detecting the 3D coordinates of an object, including a reference 3D vision sensor and a plurality of 3D vision sensors including the first to nth (n is a natural number) 3D vision sensors, the calibration device according to Application Example 1 or Application Example 2, and 3D coordinate detection system. [Application Example 4] A computer program for obtaining a conversion matrix between the sensor coordinate systems of a reference 3D vision sensor and a plurality of 3D vision sensors including the first to nth (n is a natural number) 3D vision sensors, The plurality of 3D vision sensors each generate a plurality of frames having 3D information of an object in temporal synchronization. The computer program causes the information processing apparatus to For each of the first to nth 3D vision sensors, using the initial value of the transformation matrix generated by selecting feature points with little variation from a plurality of feature points extracted from the 3D information of the object, perform coordinate transformation on all of the plurality of frames and integrate them into the reference coordinate system of the reference 3D vision sensor; For each of the first to nth 3D vision sensors, using the 3D coordinate values integrated into the reference coordinate system, calculate the variation between the plurality of 3D vision sensors for each frame, and use the initial value of the transformation matrix of the frame with little variation as the transformation matrix; realize computer program.
Explanation of Signs
[0060] D, D0, D1... 3D information H... object P, P0, P1... feature points S... 3D vision sensor S0... reference 3D vision sensor S1... first 3D vision sensor S2... second 3D vision sensor 1... 3D coordinate detection system 10... transformation matrix generation unit 20... feature point extraction unit 100... calibration device M... transformation matrix dp... average distance
Claims
1. A calibration device, comprising a reference three-dimensional vision sensor and a transformation matrix generation unit that obtains a transformation matrix for integrating the coordinate systems of each of the first to n (n is a natural number) three-dimensional vision sensors from three-dimensional information of an object generated by a plurality of three-dimensional vision sensors including the first to n three-dimensional vision sensors, wherein the plurality of three-dimensional vision sensors each generate a plurality of frames having the three-dimensional information in time synchronization, in the calibration device, the transformation matrix generation unit, for each of the first to n three-dimensional vision sensors, performs coordinate transformation on all of the plurality of frames using an initial transformation matrix generated by selecting feature points with little variation from a plurality of feature points extracted from the three-dimensional information of the object, and integrates them into the reference coordinate system, for each of the first to n three-dimensional vision sensors, calculates the variation between the plurality of three-dimensional vision sensors for each frame using the three-dimensional coordinate values integrated into the reference coordinate system, and sets the initial transformation matrix of the frame with little variation as the transformation matrix, a calibration device.
2. The calibration device according to claim 1, wherein the feature points are joint points of a human body skeleton, a calibration device.
3. A three-dimensional coordinate detection system for detecting three-dimensional coordinates of an object, comprising a reference three-dimensional vision sensor, a plurality of three-dimensional vision sensors including the first to n (n is a natural number) three-dimensional vision sensors, and the calibration device according to claim 1 or claim 2, a three-dimensional coordinate detection system.
4. A computer program for obtaining a transformation matrix between the sensor coordinate systems of a reference three-dimensional vision sensor and a plurality of three-dimensional vision sensors including the first to n (n is a natural number) three-dimensional vision sensors, wherein the plurality of three-dimensional vision sensors each generate a plurality of frames having three-dimensional information of an object in time synchronization, the computer program causes an information processing device to, for each of the first to n three-dimensional vision sensors, perform coordinate transformation on all of the plurality of frames using an initial transformation matrix generated by selecting feature points with little variation from a plurality of feature points extracted from the three-dimensional information of the object, and integrate them into the reference coordinate system of the reference three-dimensional vision sensor, For each of the first to nth three-dimensional vision sensors, for each frame, using the three-dimensional coordinate values integrated into the reference coordinate system, calculate the variation among the plurality of three-dimensional vision sensors, and use the initial value of the transformation matrix of the frame with less variation as the transformation matrix, and realize a computer program.