Multi-modal image registration method and system

By using the processing device and the indicator component of the calibration body, and by utilizing the optimized transformation matrix to achieve the transformation between different coordinate systems, the problem of multi-functional measurement in the body temperature detection system is solved, and high-precision multi-mode image alignment and automatic feature point acquisition are realized.

CN115908152BActive Publication Date: 2026-05-29IND TECH RES INST

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
IND TECH RES INST
Filing Date
2022-07-15
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Existing body temperature detection systems are unable to achieve multi-functional measurements, such as body temperature, heart rate, respiration, and behavior detection, and the coordinate systems of different image acquisition devices cannot simultaneously acquire multi-modal images of the same target object.

Method used

Two-dimensional and three-dimensional images are acquired through a processing device. By utilizing the indicator component of the calibration body and the optimized transformation matrix, the conversion between different coordinate systems is achieved, including the coordinate system conversion between two-dimensional image acquisition devices and three-dimensional image acquisition devices.

Benefits of technology

It achieves high-precision multi-modal image alignment, reduces sampling data and time requirements, and can automatically collect feature points in two-dimensional/three-dimensional images, simplifying the machine learning training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908152B_ABST
    Figure CN115908152B_ABST
Patent Text Reader

Abstract

A multi-modal image registration method includes obtaining a plurality of first points corresponding to a central vertex of a correction body and a plurality of second point groups corresponding to a plurality of side vertices from a plurality of two-dimensional images, obtaining a plurality of third points corresponding to the central vertex from a plurality of three-dimensional images, performing a first optimization operation using a first coordinate system associated with the two-dimensional images, the plurality of first points, and the plurality of third points to obtain a first conversion matrix, processing the plurality of three-dimensional images using the first conversion matrix to generate a plurality of first converted images, respectively, performing a second optimization operation using the plurality of first converted images, the plurality of first points, and the plurality of second point groups to obtain a second conversion matrix, and converting a to-be-processed image from a second coordinate system associated with the three-dimensional images to the first coordinate system using the first conversion matrix and the second conversion matrix.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a multimodal image alignment method, and more particularly to a method for aligning a three-dimensional image with a two-dimensional image. Background Technology

[0002] Most of the body temperature detection systems currently on the market only measure body temperature. However, given that body temperature measurement has become a part of daily life, the expansion of system functions should be considered, such as physiological signal measurement and motion detection.

[0003] The basic measurement signals required for a single device to perform multiple functions are mostly body temperature, heart rate, respiration, and behavior, and the necessary non-contact sensing sources are mostly visible light images, thermal images, and 3D images (point clouds). Considering cost, the current mainstream approach is to mix and match image acquisition devices from various affordable brands to try and meet the requirements of a single device performing multiple functions. However, the coordinate systems of each image acquisition device should be different, therefore it is impossible to simultaneously acquire multiple modes of a target object. Summary of the Invention

[0004] In view of the above, the present invention provides a multi-modal image alignment method and system.

[0005] A multi-mode image alignment method according to an embodiment of the present invention includes the following steps performed by a processing device: acquiring multiple two-dimensional images and multiple three-dimensional images associated with a calibration body, the calibration body having a central vertex and multiple side vertices, the two-dimensional images being associated with a first three-dimensional coordinate system, and the three-dimensional images being associated with a second three-dimensional coordinate system; acquiring multiple first points corresponding to the central vertex and multiple second point groups corresponding to the multiple side vertices from the multiple two-dimensional images; acquiring multiple third points corresponding to the central vertex from the multiple three-dimensional images; performing a first optimization operation based on the first three-dimensional coordinate system using a first transformation matrix to be solved, the multiple first points, and the multiple third points to obtain an optimized first transformation matrix; processing the multiple three-dimensional images using the optimized first transformation matrix to generate multiple primary transformation images; performing a second optimization operation based on the first three-dimensional coordinate system using the multiple primary transformation images, the multiple first points, the multiple second point groups, and a default specification parameter group of the calibration body to obtain an optimized second transformation matrix; and performing a transformation operation using the optimized first transformation matrix and the optimized second transformation matrix to transform the image to be processed to the second three-dimensional coordinate system or the first three-dimensional coordinate system.

[0006] A multi-mode image alignment system according to an embodiment of the present invention includes a calibration body, a two-dimensional image acquisition device, a three-dimensional image acquisition device, and a processing device, wherein the processing device is connected to the two-dimensional image acquisition device and the three-dimensional image acquisition device. The calibration body includes a stereo body and multiple indicator components, wherein the stereo body has a central vertex and multiple side vertices, and the indicator components are respectively disposed at the central vertex and the multiple side vertices. The two-dimensional image acquisition device has a first three-dimensional coordinate system and is used to generate multiple two-dimensional images associated with the calibration body. The three-dimensional image acquisition device has a second three-dimensional coordinate system and is used to generate multiple three-dimensional images associated with the calibration body. The processing device is used to obtain a coordinate transformation matrix based on the multiple two-dimensional images and the multiple three-dimensional images, and to transform the image to be processed to the second three-dimensional coordinate system or the first three-dimensional coordinate system using the coordinate transformation matrix.

[0007] Using the above architecture, the multi-modal image alignment method disclosed in this application can obtain the transformation matrix between different coordinate systems through two optimization calculations, achieving high-precision alignment results without the need for complex machine learning training. Furthermore, by using three-dimensional corner features as the basis for obtaining the transformation matrix, compared to the traditional method using a planar chessboard calibration board, the required amount of sampling data is significantly less, meaning less sampling time is required. The multi-modal image alignment system disclosed in this application also achieves the advantages of less required sampling data and less sampling time. Moreover, through a special stereo calibration body design with indicator components, the system can automatically acquire feature points in two-dimensional / three-dimensional images.

[0008] The foregoing description of the contents of this disclosure and the following description of the embodiments are intended to demonstrate and explain the spirit and principles of the present invention, and to provide a further explanation of the scope of the patent application of the present invention. Attached Figure Description

[0009] Figure 1 This is a functional block diagram of a multi-mode image alignment system according to an embodiment of the present invention.

[0010] Figure 2 This is a schematic diagram of the correction position according to an embodiment of the present invention.

[0011] Figure 3 This is a flowchart illustrating a multi-mode image alignment method according to an embodiment of the present invention.

[0012] Figure 4 This is a schematic diagram illustrating the execution of a distance calculation operation in a multi-mode image alignment method according to an embodiment of the present invention.

[0013] Figure 5 This is a flowchart illustrating the second optimization operation in a multi-mode image alignment method according to an embodiment of the present invention.

[0014] Figure 6 This is a schematic diagram illustrating the execution of a projection operation in a multi-mode image alignment method according to an embodiment of the present invention.

[0015] Figure 7 This is a schematic diagram illustrating the execution of specification parameter estimation in a multi-mode image alignment method according to an embodiment of the present invention.

[0016] Figure 8 This is a schematic diagram illustrating the application environment of a multi-mode image alignment method according to an embodiment of the present invention.

[0017] Figure 9 This is a functional block diagram of a multi-mode image alignment system according to another embodiment of the present invention.

[0018] Figure 10 This is a flowchart illustrating a multi-mode image alignment method according to another embodiment of the present invention.

[0019] List of reference numerals

[0020] 1,1': Multimodal image alignment system

[0021] 11: Processing device

[0022] 12,15: Two-dimensional image acquisition device

[0023] 13: Three-dimensional image acquisition device

[0024] 14: Correction body

[0025] 140: Three-dimensional body

[0026] 141a~141: Indicator components

[0027] P1~P6: Correction Position

[0028] v1, v2: Perspective

[0029] CY1: First three-dimensional coordinate system

[0030] CY2: Second three-dimensional coordinate system

[0031] CY2': Coordinate system

[0032] SD: Two-dimensional image

[0033] TD: 3D image

[0034] TD': Secondary image conversion

[0035] D1: First point

[0036] D1': Fifth point

[0037] D2: Third point

[0038] L1: Ray

[0039] d: distance

[0040] D31~D33: Second point

[0041] D31'~D33': Sixth point

[0042] E: Estimated length

[0043] A: Estimate the included angle

[0044] I1~I4: Ideal Position

[0045] S: Side length

[0046] R: included angle

[0047] A1~A6: Image acquisition devices

[0048] S101~S107: Steps

[0049] S601~S604: Steps

[0050] S108~S113: Steps Detailed Implementation

[0051] The following detailed description of the features and advantages of the present invention is sufficient to enable anyone skilled in the art to understand the technical content of the present invention and implement it accordingly. Based on the disclosure, patent claims, and drawings in this specification, anyone skilled in the art can easily understand the relevant objectives and advantages of the present invention. The following embodiments further illustrate the points of the present invention, but are not intended to limit the scope of the present invention in any way.

[0052] Please refer to Figure 1 , Figure 1 This is a functional block diagram of a multi-mode image alignment system 1 according to an embodiment of the present invention. Figure 1 As shown, the multi-mode image alignment system 1 includes a processing device 11, a two-dimensional image acquisition device 12, a three-dimensional image acquisition device 13, and a calibration body 14. The processing device 11 can be connected to the two-dimensional image acquisition device 12 and the three-dimensional image acquisition device 13 in a wired or wireless manner.

[0053] The processing device 11 may include, but is not limited to, a single processor and an integration of multiple microprocessors, such as a central processing unit (CPU) or a graphics processing unit (GPU). The processing device 11 is used to obtain a transformation matrix of the coordinate systems of the two image acquisition devices 12 and 13 based on images generated by the two-dimensional image acquisition device 12 and the three-dimensional image acquisition device 13 capturing the calibration object 14. Detailed execution will be described later. The processing device 11 can use the transformation matrix to map the image generated by the three-dimensional image acquisition device 13 to the image generated by the two-dimensional image acquisition device 12, thereby generating an overlay image.

[0054] The two-dimensional image acquisition device 12 is, for example, a visible light camera, a near-infrared camera, or a thermal imager. The two-dimensional image acquisition device 12 is used to capture images to generate two-dimensional images and has a camera coordinate system and an image plane coordinate system. The three-dimensional image acquisition device 13 is, for example, a three-dimensional point cloud sensor or a depth camera. The three-dimensional image acquisition device 13 is used to capture images to generate three-dimensional images and has a camera coordinate system. The two-dimensional image acquisition device 12 and the three-dimensional image acquisition device 13 can be housed in the same casing or can be housed in different casings.

[0055] The calibration body 14 includes a main body 140 and multiple indicator components 141a to 141d. The three-dimensional main body 140 may have more than four vertices. Figure 1 The solid body 140 is illustrated as a hexahedron, but is not limited thereto. Indicator components 141a to 141d can be respectively disposed at the vertices of the solid body 140, with a minimum number of four. Indicator components 141a to 141d can all be light-emitting components, all be heat-generating components, or all be components that combine both light-emitting and heat-generating functions. Indicator components 141a to 141d can all have the same color and / or temperature, or they can each have different colors and / or temperatures.

[0056] Furthermore, the type of indicator components 141a-141d can depend on the type of two-dimensional image acquisition device 12. In an embodiment where a visible light camera is used as the two-dimensional image acquisition device, indicator components 141a-141d are implemented as light-emitting components. In an embodiment where a near-infrared camera or a thermal imager is used as the two-dimensional image acquisition device 12, indicator components 141a-141d are implemented as heat-generating components. In another embodiment, the multi-mode image alignment system 1 may also include another two-dimensional image acquisition device. If the two two-dimensional image acquisition devices are a visible light camera and a thermal imager, then indicator components 141a-141d are implemented as components that have both light-emitting and heat-generating functions.

[0057] The calibration body 14, the two-dimensional image acquisition device 12, and the three-dimensional image acquisition device 13 of the multi-mode image alignment system 1 can work together to generate data for the processing device 11 to perform multi-mode image alignment. The calibration body 14 can be placed in multiple alignment positions in turn. The two-dimensional image acquisition device 12 can be controlled to capture multiple two-dimensional images by taking pictures of the calibration body 14 placed in different alignment positions in turn. The three-dimensional image acquisition device 13 can be controlled to capture multiple three-dimensional images by taking pictures of the calibration body 14 placed in different alignment positions in turn. The processing device 11 can acquire the multiple two-dimensional images from the two-dimensional image acquisition device 12 and the multiple three-dimensional images from the three-dimensional image acquisition device 13, and perform multi-mode image alignment alignment based on them to obtain the transformation matrix between the camera coordinate system of the two-dimensional image acquisition device 12 and the camera coordinate system of the three-dimensional image acquisition device 13.

[0058] Furthermore, the processing device 11 can obtain the transformation matrix between the image plane coordinate system of the two-dimensional image acquisition device 12 and the camera coordinate system of the three-dimensional image acquisition device 13 based on this transformation matrix and the focal length and projection center of the two-dimensional image acquisition device 12. This transformation matrix can then be used to overlay the images generated by the two devices. The processing device 11 can also obtain multi-mode signals (e.g., containing temperature and spatial information, or containing color and spatial information) corresponding to a specific target object from the overlaid image. The specific target object can be selected by the processing device 11 according to a specific algorithm (e.g., a face recognition algorithm) or by the operator; this invention does not limit this selection. Alternatively, the processing device 11 can present the overlaid image to the operator via a display or output it to other multi-mode signal measurement application devices.

[0059] Please refer to this as well. Figure 1 and Figure 2 ,in Figure 2 This is a schematic diagram of the correction position according to an embodiment of the present invention. Figure 2 As shown, the correction positions P1 to P6 of the calibrator 14 are located within the overlapping range of the viewing angle v1 of the two-dimensional image acquisition device 12 and the viewing angle v2 of the three-dimensional image acquisition device 13, and their number is at least six. Furthermore, at each correction position P1 to P6, the indicator components 141a to 141d of the calibrator 14 are visible to both the two-dimensional image acquisition device 12 and the three-dimensional image acquisition device 13. The number of three-dimensional images can be the same as the number of correction positions P1 to P6, while the number of two-dimensional images depends on the control method of the indicator components 141a to 141d of the calibrator 14.

[0060] In one embodiment, the indicator components 141a-141d of the calibration body 14 have different colors or different temperatures, and are all enabled during the imaging process. The two-dimensional image acquisition device 12 is controlled to capture images of the calibration body 14, which is placed in turn at different calibration positions P1-P6, to generate a number of two-dimensional images equal to the number of calibration positions P1-P6. Each two-dimensional image contains image blocks corresponding to the indicator components 141a-141d of different colors and / or different temperatures.

[0061] In another embodiment, the indicator components 141a to 141d of the calibration body 14 are enabled in a specific order to emit light or heat during a single imaging process. The two-dimensional image acquisition device 12 is controlled to perform multiple imaging processes on the calibration body 14, which is placed in multiple calibration positions P1 to P6 in turn, to generate multiple two-dimensional images. For example, when the calibration body 14 is placed in any of the calibration positions P1 to P6, the indicator component 141a is enabled and the two-dimensional image acquisition device 12 is controlled to generate a two-dimensional image containing an image block corresponding to the indicator component 141a. Then, the indicator component 141b is enabled and the two-dimensional image acquisition device 12 is controlled to generate a two-dimensional image containing an image block corresponding to the indicator component 141b. The generation of the two-dimensional images corresponding to the indicator components 141c and 141d is the same and will not be described in detail. In this embodiment, the number of two-dimensional images is N times the number of calibration positions P1 to P6, where N is the number of indicator components 141a to 141d of the calibration body 14.

[0062] The control of the indicator components 141a to 141b of the aforementioned calibration body 14, the shooting of the two-dimensional image acquisition device 12, and the shooting of the three-dimensional image acquisition device 13 can be controlled by an operator, or can be controlled by a controller storing corresponding control instructions via wired or wireless means. This invention does not limit these controls.

[0063] Please refer to this as well. Figure 1 and Figure 3 ,in Figure 3 This is a flowchart illustrating a multi-mode image alignment method according to an embodiment of the present invention. Figure 3As shown, the multi-mode image alignment method includes step S101: obtaining multiple two-dimensional images and multiple three-dimensional images associated with the calibration body, the calibration body having a central vertex and multiple side vertices, the two-dimensional images corresponding to a first three-dimensional coordinate system, and the three-dimensional images corresponding to a second three-dimensional coordinate system; step S102: obtaining multiple first points corresponding to the central vertex and multiple second point groups corresponding to the multiple side vertices from the multiple two-dimensional images; step S103: obtaining multiple third points corresponding to the central vertex from the multiple three-dimensional images; step S104: using the first transformation matrix to be solved, the multiple first points and the multiple third points, based on the first... The three-dimensional coordinate system performs a first optimization operation to obtain an optimized first transformation matrix; step S105: the multiple three-dimensional images are processed using the optimized first transformation matrix to generate multiple primary transformation images respectively; step S106: using the multiple primary transformation images, the multiple first points, the multiple second point groups, and the default specification parameter group of the correction body, a second optimization operation is performed based on the first three-dimensional coordinate system to obtain an optimized second transformation matrix; and step S107: a transformation operation is performed using the optimized first transformation matrix and the optimized second transformation matrix to transform the image to be processed to the second three-dimensional coordinate system or the first three-dimensional coordinate system. It should be specifically noted that the present invention does not limit the execution order of steps S102 and S103.

[0064] Figure 3 The multimodal image alignment method shown can be applied to Figure 1 The multi-mode image alignment system 1 shown is specifically executed by the processing device 11. Steps S101 to S107 will be further explained below by way of example using the operation of the multi-mode image alignment system 1.

[0065] In step S101, the processing device 11 acquires multiple two-dimensional images and multiple three-dimensional images of the calibration body 14. The methods for acquiring the multiple two-dimensional images and the multiple three-dimensional images are as described above and will not be repeated here. The first three-dimensional coordinate system in step S101 can be the camera coordinate system of the two-dimensional image acquisition device 12, the second three-dimensional coordinate system can be the camera coordinate system of the three-dimensional image acquisition device 13, the central vertex can be the vertex of the calibration body 14 that is provided with the indicator component 141a, and the multiple side vertices can be vertices that are respectively provided with indicator components 141b to 141d.

[0066] In step S102, the processing device 11 obtains multiple first points corresponding to the central vertex and multiple groups of second points corresponding to the multiple side vertices from the multiple two-dimensional images. In embodiments where each two-dimensional image contains image blocks corresponding to indicator components 141a-141d of different colors and / or different temperatures, the processing device 11 may pre-store a lookup table, which records the color and / or temperature corresponding to the central vertex and side vertices respectively. The processing device 11 can use image processing algorithms to find image blocks with different colors and / or temperatures, wherein the image processing algorithms include, but are not limited to, binarization and circle detection. The processing device 11 can determine, according to the lookup table, that these image blocks respectively correspond to the central vertex where the indicator component 141a is provided and the side vertices where the indicator components 141b-141d are provided. The center point of the image block corresponding to the central vertex can be used as the first point, and the center points of the image blocks corresponding to the side vertices can form a group of second points. Furthermore, each point in the first and second point groups has a two-dimensional coordinate system representing its position in the image plane coordinate system of the two-dimensional image acquisition device 12.

[0067] In embodiments where the instruction components 141a-141d are enabled in a specific order during each shooting procedure, the processing device 11 may pre-store the specific order. The processing device 11 may use image processing algorithms to find image blocks with color and / or temperature, wherein the image processing algorithms include, but are not limited to, binarization and circle detection. The processing device 11 may determine, based on the generation time of the two-dimensional image and the pre-stored specific order, whether these image blocks in the two-dimensional image correspond to the central vertex or a side vertex. The center point of the image block corresponding to the central vertex may be designated as a first point, and the center points of the image blocks corresponding to the side vertex may form a second group of points.

[0068] In step S103, the processing device 11 obtains a plurality of third points corresponding to the central vertex from the plurality of three-dimensional images. Further, the plurality of third points have a one-to-one relationship with the plurality of three-dimensional images, and each has a three-dimensional coordinate representing its position in a second three-dimensional coordinate system. The processing device 11 can use each three-dimensional image as a target image and perform the following: obtaining three planes from the target image, the three planes being adjacent to each other and having mutually perpendicular normal vectors; and obtaining the intersection point of the three planes as the corresponding point among the plurality of third points. Furthermore, the processing device 11 can find all planes in the target image, find all combinations consisting of three adjacent planes, calculate the normal vectors of the planes in each combination, filter out combinations where the normal vectors of the planes are mutually perpendicular, and calculate the intersection point of the planes in that combination as the corresponding point among the plurality of third points.

[0069] In step S104, the processing device 11 uses the first transformation matrix to be solved, the plurality of first points, and the plurality of third points to perform a first optimization operation based on a first three-dimensional coordinate system to obtain an optimized first transformation matrix. Further, the plurality of first points and the plurality of third points respectively correspond to the aforementioned plurality of correction positions. The processing device 11 can perform distance calculation operations on the corresponding first point and corresponding third point of each correction position to obtain multiple calculation results, which also respectively correspond to the plurality of correction positions.

[0070] Please refer to this as well. Figure 1 and Figure 4 To further illustrate the distance calculation operation, in which Figure 4 This is a schematic diagram illustrating the execution of a distance calculation operation in a multi-mode image alignment method according to an embodiment of the present invention. Figure 4 As shown, the first point D1 is the point in the two-dimensional image SD that corresponds to the central vertex of the calibration body 14. The two-dimensional image SD corresponds to the first three-dimensional coordinate system CY1 (the camera coordinate system of the two-dimensional image acquisition device 12). The third point D2 is the point in the three-dimensional image TD that corresponds to the central vertex of the calibration body 14. The three-dimensional image TD corresponds to the second three-dimensional coordinate system CY2 (the camera coordinate system of the three-dimensional image acquisition device 13).

[0071] During the distance calculation process, the processing device 11 can obtain the ray L1 connecting the origin of the first three-dimensional coordinate system CY1 and the first point D1, transform the third point D2 using the first transformation matrix to be solved, and calculate the distance d between the transformed third point D2 and the ray L1 as the distance calculation result. The processing device 11 can iteratively adjust the first transformation matrix to be solved using a convergence function (cost function), and use the iteratively adjusted first transformation matrix to be solved as the optimized first transformation matrix, wherein the convergence function indicates that the sum of the distance calculation results corresponding to all correction positions is minimized. The convergence function can be expressed as shown in equation (1):

[0072] Min(Σ{d=Dis(M1*Pt3D,Line)}) (1)

[0073] Where M1 indicates the first transformation matrix to be solved, Pt3D indicates the third point D2, and Line indicates the ray L1.

[0074] Specifically, the first transformation matrix to be solved may include a rotation matrix and a displacement matrix, wherein the rotation matrix is ​​associated with angular parameters along three axes, and the displacement matrix is ​​associated with displacement parameters along three axes. To obtain a solution for the above six parameters to obtain an optimized first transformation matrix, the number of correction positions must be at least six. The detailed parameter composition of the rotation and displacement matrices is something that those skilled in the art can design based on the above six parameters as needed, and this invention does not limit this.

[0075] At Figure 3 In step S105, the processing device 11 transforms the plurality of three-dimensional images using an optimized first transformation matrix. In step S106, the processing device 11 uses the plurality of three-dimensional images transformed by the optimized first transformation matrix (first-transformation images), the plurality of first points, the plurality of second point groups, and the default specification parameter group of the correction body to perform a second optimization operation based on the first three-dimensional coordinate system to obtain an optimized second transformation matrix. Ideally, the optimized first transformation matrix obtained in step S104 should make the sum of the distance calculation results corresponding to all correction positions approach 0. However, in reality, the optimized first transformation matrix is ​​an approximate solution. Therefore, the coordinate system of the three-dimensional image transformed by the optimized first transformation matrix is ​​still somewhat different from the first three-dimensional coordinate system. By using the optimized second transformation matrix obtained in step S106 to further transform the three-dimensional image transformed by the optimized first transformation matrix, the coordinate system corresponding to the three-dimensional image can be made closer to the first three-dimensional coordinate system.

[0076] Please refer to Figure 1 and Figure 5 ,in Figure 5 This is a flowchart illustrating the second optimization operation in a multi-mode image alignment method according to an embodiment of the present invention. The second optimization operation may include step S601: processing the plurality of primary transformation images using the second transformation matrix to be solved to generate a plurality of secondary transformation images respectively; step S602: obtaining a plurality of fourth point groups from the plurality of secondary transformation images respectively based on the plurality of first points and the plurality of second point groups; step S603: obtaining a plurality of estimated specification parameter groups respectively based on the plurality of fourth point groups; and step S604: iteratively adjusting the second transformation matrix to be solved through a convergence function, and using the iteratively adjusted second transformation matrix to be solved as the optimized second transformation matrix, wherein the convergence function indicates that the difference between the plurality of estimated specification parameter groups and the default specification parameter group of the corrector is minimized.

[0077] Furthermore, the plurality of secondary transformation images, the plurality of first points, the plurality of second point groups, the plurality of fourth point groups, and the plurality of estimated specification parameter groups can each correspond to the aforementioned plurality of correction positions. In steps S602 and S603, the processing device 11 can perform a projection operation on the corresponding first point, the corresponding second point group, and the corresponding secondary transformation image for each correction position to obtain the corresponding fourth point group, and then perform a specification parameter estimation operation based on the corresponding fourth point group to obtain the corresponding estimated specification parameter group.

[0078] Please refer to this as well. Figure 1 and Figure 6 To further illustrate the projection operation, in which Figure 6 This is a schematic diagram illustrating the execution of a projection operation in a multi-mode image alignment method according to an embodiment of the present invention. Figure 6 As shown, the first point D1 is the point in the two-dimensional image SD that corresponds to the central vertex of the correction body 14, and the second points D31 to D33 are the points in the two-dimensional image SD that correspond to the side vertices of the correction body 14 and can form a second point group. The two-dimensional image SD corresponds to the first three-dimensional coordinate system CY1, and the three-dimensional image (secondary transformed image TD') transformed by the optimized first transformation matrix and the second transformation matrix to be solved corresponds to the coordinate system CY2'.

[0079] During the projection process, the processing device 11 can project the first point D1 onto the secondary transformed image TD' to obtain the fifth point D1' corresponding to the first point D1, and project the second points D1 to D33 onto the secondary transformed image TD' to obtain multiple sixth points D31' to D33 respectively corresponding to the second points D31 to D33. Alternatively, the processing device 11 can use the point in the secondary transformed image TD' that has the x and y coordinates of the first point D1 as the fifth point D1', and can use the point in the secondary transformed image TD' that has the x and y coordinates of the second points D31 to D33 as the sixth points D31' to D33'. The fifth point D1' and the sixth points D31' to D33' can form a fourth point group. The method for obtaining the fourth point group corresponding to other correction positions is the same as described above and will not be repeated.

[0080] Please refer to this as well. Figure 1 , Figure 5 and Figure 7 To further illustrate the specification parameter estimation process, Figure 7 This is a schematic diagram illustrating the execution of a specification parameter estimation operation in a multi-mode image alignment method according to an embodiment of the present invention. During the execution of the specification parameter estimation operation, the processing device 11 acquires multiple connections between the fifth point D1' and the sixth points D31' to D33', and calculates the estimated length of the multiple connections and the estimated angle between the connections, wherein the estimated length and the estimated angle constitute an estimated specification parameter group. Figure 7The example shows the estimated length E of the connection between the fifth point D1' and the sixth point D31' and the included angle A between the connection between the fifth point D1' and the sixth point D31' and the connection between the fifth point D1' and the sixth point D33'.

[0081] Step S604 is further explained below. The processing device 11 can iteratively adjust the second transformation matrix to be solved using a convergence function, and use the iteratively adjusted second transformation matrix as the optimized second transformation matrix, wherein the convergence function indicates that the difference between the plurality of estimated specification parameter sets and the default specification parameter set of the corrector 14 is minimized. The convergence function used in the second optimization operation is different from the convergence function used in the first optimization operation. The default specification parameter set of the corrector 14 may include a plurality of preset side lengths and a plurality of preset included angles of the corrector 14. Specifically, Figure 7 The ideal positions I1 to I4 of the fifth point D1' and the sixth points D31' to D33' are presented as an example. The multiple preset side lengths can be the side lengths S of multiple connections between the ideal positions I1 to I4, and the multiple preset included angles can be the included angles R of the connections between the ideal positions I1 to I4.

[0082] The difference between the estimated specification parameter set and the default specification parameter set mentioned in step S604 can indicate the weighted sum of the first value and the second value, wherein the first value indicates the sum of the differences between the multiple estimated lengths and the multiple preset side lengths, and the second value indicates the sum of the differences between the multiple estimated angles and the multiple preset angles. The convergence function of step S604 can be expressed as shown in equation (2):

[0083] Min(α*(Σ((S1-E1)+(S2-E2)+(S3-E3))+β*Σ((R1-A1)+(R2-A2)+(R3-A3))) (2)

[0084] S1 to S3 indicate the preset side length, R1 to R3 indicate the preset included angle, E1 to E3 indicate the estimated length, A1 to A3 indicate the estimated included angle, and α and β indicate the weights that can be adjusted as needed.

[0085] Specifically, the unsolved second transformation matrix used to iteratively generate the optimized second transformation matrix may include a rotation matrix and a displacement matrix, or it may only include a rotation matrix. The rotation matrix is ​​associated with angular parameters along three axes. The displacement matrix is ​​associated with displacement parameters along three axes. The detailed parameter composition of the rotation matrix and displacement matrix is ​​something that those skilled in the art can design based on the above six parameters as needed, and this invention does not limit this.

[0086] Please refer to this again. Figure 1 and Figure 3After obtaining the optimized first transformation matrix and the optimized second transformation matrix through the above steps, in step S107, the processing device 11 can use the optimized first transformation matrix and the optimized second transformation matrix to perform transformation, converting the image to be processed from the second three-dimensional coordinate system to the first three-dimensional coordinate system, or from the first three-dimensional coordinate system to the second three-dimensional coordinate system. As mentioned above, the first three-dimensional coordinate system is the camera coordinate system of the two-dimensional image acquisition device 12.

[0087] Furthermore, after converting the image to be processed to the first three-dimensional coordinate system, the processing device 11 can utilize the parameter matrix of the two-dimensional image acquisition device 12 to convert the image to be processed from the first three-dimensional coordinate system to the image plane coordinate system of the two-dimensional image acquisition device 12. The parameter matrix includes the focal length parameter and projection center parameter of the two-dimensional image acquisition device 12. The processing device 11 can optimize the first transformation matrix, optimize the second transformation matrix, and the parameter matrix to form a coordinate transformation matrix for performing the transformation between the image plane coordinate system of the two-dimensional image acquisition device 12 and the camera coordinate system (second three-dimensional coordinate system) of the three-dimensional image acquisition device 13, and store it in internal memory. The processing device 11 can use this coordinate transformation matrix to map the image generated by the three-dimensional image acquisition device 13 to the two-dimensional image generated by the two-dimensional image acquisition device 12 to generate an overlay image.

[0088] For example, the transformation between the image plane coordinate system of the two-dimensional image acquisition device 12 and the camera coordinate system of the three-dimensional image acquisition device 13 can be performed according to equations (3) and (4):

[0089]

[0090]

[0091] Where X, Y, and Z represent coordinates on the 3D image, M1 represents the first optimization transformation matrix, M2 represents the second optimization transformation matrix, and f x and f y c represents the focal length parameter of the two-dimensional image acquisition device 12. x and c y Let x represent the projection center parameter of the two-dimensional image acquisition device 12, and x t and y t Represents coordinates on a two-dimensional image.

[0092] The multi-mode image alignment system and method described in the above embodiments can be applied to environments with multiple image acquisition devices. Please refer to them for further information. Figure 1 and Figure 8 , Figure 8 This is a schematic diagram illustrating the application environment of a multi-mode image alignment method according to an embodiment of the present invention. Figure 8The image acquisition devices A1 to A6 can be configured with two-dimensional and three-dimensional image acquisition devices spaced apart. Any two adjacent image acquisition devices A1 to A6 can serve as the two-dimensional image acquisition device 12 and the three-dimensional image acquisition device 13 in the aforementioned multi-mode image alignment system 1. It should be specifically noted that... Figure 8 The placement of the calibrator 14 is illustrated only as an example, and its placement rules are as described in the foregoing embodiments, and will not be repeated here.

[0093] Using the multi-mode image alignment method described in the aforementioned embodiments, the processing device 11 can obtain the transformation matrix between the coordinate systems of any two adjacent image acquisition devices A1 to A6, and then use a cascaded transformation architecture to transform any one of the image acquisition devices A1 to A6 to the coordinate system of an image acquisition device that is separated from it by one or more image acquisition devices.

[0094] For example, suppose the processing device 11 obtains the transformation matrix M between the coordinate systems of image acquisition device A1 and image acquisition device A2 using a multi-mode image alignment method. 12 And obtain the transformation matrix M between the coordinate systems of image acquisition device A2 and image acquisition device A3. 23 The processing device 11 can use equation (5) to transform the image coordinates generated by the image acquisition device A1 to the coordinate system of the image acquisition device A3:

[0095] P3 = M 23 *M 12 *P1 (5)

[0096] Where P1 represents the image coordinates generated by image acquisition device A1, and P3 represents the coordinates of the image coordinates transformed to the coordinate system of image acquisition device A3. Through the above-described cascading transformation architecture, processing device 11 can transform all image acquisition devices A1 to A6 to the coordinate system of a specific one of image acquisition devices A1 to A6.

[0097] In another embodiment, the multi-mode image alignment system may include three image acquisition devices and can obtain the coordinate transformation matrix between the coordinate systems of the three image acquisition devices. Please refer to... Figure 9 , Figure 9 An exemplary functional block diagram of a multi-mode image alignment system 1' comprising three image acquisition devices is shown. Figure 9 As shown, the multi-mode image alignment system 1' includes a processing device 11, two two-dimensional image acquisition devices 12 and 15, a three-dimensional image acquisition device 13, and a calibration body 14.

[0098] Compared to Figure 1The multi-mode image alignment system 1 shown further includes a two-dimensional image acquisition device 15. The two-dimensional image acquisition device 15, such as a visible light camera, a near-infrared camera, or a thermal imager, can be connected to the processing device 11. It has a camera coordinate system and an image plane coordinate system, and can be controlled to capture multiple two-dimensional images of the calibration body 14 placed in different calibration positions in turn. Specifically, the camera coordinate system and image plane coordinate system of the two-dimensional image acquisition device 15 are different from those of the two-dimensional image acquisition device 12.

[0099] In addition to the operations described in the aforementioned embodiments, the processing device 11 of the multi-mode image alignment system 1' can further obtain a coordinate transformation matrix between the camera coordinate system of the two-dimensional image acquisition device 15 and the camera coordinate system of the three-dimensional image acquisition device 13, based on multiple two-dimensional images generated by the two-dimensional image acquisition device 15 and associated with the calibration body 14, and multiple three-dimensional images generated by the aforementioned three-dimensional image acquisition device 13 and associated with the calibration body 14. Using the coordinate transformation matrix and the optimized first and second transformation matrices described in the aforementioned embodiments, the processing device 11 can perform a transformation between the camera coordinate system of the two-dimensional image acquisition device 12 and the camera coordinate system of the two-dimensional image acquisition device 15. Using the coordinate transformation matrix, the optimized first and second transformation matrices described in the aforementioned embodiments, the parameter matrix of the two-dimensional image acquisition device 12, and the parameter matrix of the two-dimensional image acquisition device 15, the processing device 11 can perform a transformation between the image plane coordinate system of the two-dimensional image acquisition device 12 and the image plane coordinate system of the two-dimensional image acquisition device 15.

[0100] The following further describes the multimodal image alignment method applicable to the multimodal image alignment system 1'. Please refer to the following: Figure 3 , Figure 9 and Figure 10 ,in Figure 10 This is a flowchart illustrating a multimodal image alignment method according to another embodiment of the present invention. The multimodal image alignment method applicable to multimodal image alignment system 1' may include... Figure 3 Steps S101 to S107 shown are Figure 9The steps are as follows: Step S108: Obtain multiple second two-dimensional images associated with the calibration body, the second two-dimensional images corresponding to a third three-dimensional coordinate system; Step S109: Obtain multiple seventh points corresponding to the central vertex and multiple eighth point groups corresponding to the multiple side vertices from the multiple second two-dimensional images; Step S110: Perform a first optimization operation based on the third three-dimensional coordinate system using the third transformation matrix to be solved, the multiple seventh points and the multiple third points, to obtain an optimized third transformation matrix; Step S111: Process the multiple three-dimensional images using the optimized third transformation matrix; Step S112: Perform a second optimization operation based on the third three-dimensional coordinate system using the processed three-dimensional images, the multiple seventh points, the multiple eighth point groups and the default specification parameter group of the calibration body, to obtain an optimized fourth transformation matrix; and Step S113: Perform a transformation between the first three-dimensional coordinate system and the third three-dimensional coordinate system using the optimized first transformation matrix, optimized second transformation matrix, optimized third transformation matrix and optimized fourth transformation matrix. It should be noted that the present invention does not limit the execution order of steps S101 and S108, nor does it limit the execution order of steps S102, S103 and S109, nor does it limit the execution order of the combination of steps S104 to S106 and the combination of steps S110 to S112.

[0101] Steps S101 to S113 can be executed by the processing device 11 of the multi-mode image alignment system 1'. Further implementations of steps S108 to S112 are similar to further implementations of steps S101, S102, and S104 to S106, respectively. In other words, S101, S102, and S104 to S106 obtain the transformation relationship between the first three-dimensional coordinate system of the two-dimensional image acquisition device 12 and the second three-dimensional coordinate system of the three-dimensional image acquisition device 13, while S108 to S112 obtain the transformation relationship between the third three-dimensional coordinate system of the two-dimensional image acquisition device 15 and the second three-dimensional coordinate system based on the same principle; these details will not be elaborated further here. In step S113, the processing device 11 can utilize optimized first transformation matrix, optimized second transformation matrix, optimized third transformation matrix, and optimized fourth transformation matrix to perform the transformation between the first three-dimensional coordinate system and the third three-dimensional coordinate system. Furthermore, optimizing the first transformation matrix and optimizing the second transformation matrix can form a coordinate transformation matrix M used to transform the image to be processed from the second three-dimensional coordinate system to the first three-dimensional coordinate system. a Optimizing the third and fourth transformation matrices can form a coordinate transformation matrix M used to transform the image to be processed from the second three-dimensional coordinate system to the third three-dimensional coordinate system. b The processing device 11 can use the coordinate transformation matrix M a and M bAlgebraic operations are performed to obtain the coordinate transformation matrix M used to transform the image to be processed from the first three-dimensional coordinate system to the third three-dimensional coordinate system. c For example, it can be represented as equation (6):

[0102] M c =M b *M a -1 (6)

[0103] Furthermore, the coordinate transformation matrix M a It may also include the parameter matrix of the two-dimensional image acquisition device 12 and the coordinate transformation matrix M. b It may also include the parameter matrix of the two-dimensional image acquisition device 15. Coordinate transformation matrix M a The relationships between the optimized first transformation matrix, the optimized second transformation matrix, and the parameter matrix are shown in equations (3) and (4) above. The coordinate transformation matrix M b The relationships between the optimized third transformation matrix, the optimized fourth transformation matrix, and the parameter matrix are the same, and will not be repeated here. In this embodiment, the coordinate transformation matrix M obtained through the above algebraic operations... c It can be used to convert between the image plane coordinate system of the two-dimensional image acquisition device 12 and the image plane coordinate system of the two-dimensional image acquisition device 15.

[0104] Using coordinate transformation matrix M a M b and M c The processing device 11 can map the images generated by any two of the two-dimensional image acquisition devices 12 and 15 and the three-dimensional image acquisition device 13 to the images generated by the remaining devices. For example, using the coordinate transformation matrix M b and coordinate transformation matrix M c The processing device 11 can map the images generated by the 3D image acquisition device 13 and the 2D image acquisition device 12 onto the image generated by the 2D image acquisition device 15 to produce an overlay image. In this way, the multi-mode image alignment system 1' can obtain three types of information corresponding to a specific target object from the overlay image at once. For example, if the three image acquisition devices of the multi-mode image alignment system 1' are a thermal imager, a visible light camera, and a 3D point cloud sensor, then the processing device 11 of the multi-mode image alignment system 1' can obtain the temperature, color, and spatial location information of the specific target object from the overlay image at once.

[0105] Using the above architecture, the multi-modal image alignment method disclosed in this application can obtain the transformation matrix between different coordinate systems through two optimization calculations, achieving high-precision alignment results without the need for complex machine learning training. Furthermore, by using three-dimensional corner features as the basis for obtaining the transformation matrix, compared to the traditional method using a planar chessboard calibration board, the required amount of sampling data is significantly less, meaning less sampling time is required. The multi-modal image alignment system disclosed in this application also achieves the advantages of less required sampling data and less sampling time. Moreover, through a special stereo calibration body design with indicator components, the system can automatically acquire feature points in two-dimensional / three-dimensional images.

Claims

1. A multimodal image alignment method, comprising execution by a processing device: Multiple two-dimensional images and multiple three-dimensional images associated with a calibration body are obtained. The calibration body has a central vertex and multiple side vertices. The multiple two-dimensional images are associated with a first three-dimensional coordinate system, and the multiple three-dimensional images are associated with a second three-dimensional coordinate system. Obtain multiple first points corresponding to the central vertex and multiple second point groups corresponding to the multiple side vertices from the multiple two-dimensional images; Multiple third points corresponding to the central vertex are obtained from the multiple three-dimensional images; Using the first transformation matrix to be solved, the plurality of first points, and the plurality of third points, a first optimization operation is performed based on the first three-dimensional coordinate system to obtain an optimized first transformation matrix; The optimized first transformation matrix is ​​used to process the multiple three-dimensional images to generate multiple first-transformation images respectively; Using the multiple first-transformation images, the multiple first points, the multiple second point groups, and the default specification parameter group of the calibration body, based on the first three-dimensional coordinate system, a second optimization operation is performed on the second transformation matrix to be solved, so as to obtain an optimized second transformation matrix; and The optimization first transformation matrix and the optimization second transformation matrix are used to perform transformation operations to transform the image to be processed to the second three-dimensional coordinate system or the first three-dimensional coordinate system.

2. The multi-mode image alignment method as described in claim 1, wherein the plurality of first points and the plurality of third points respectively correspond to a plurality of correction positions, and the first optimization operation is performed based on the first three-dimensional coordinate system using the unsolved first transformation matrix, the plurality of first points, and the plurality of third points to obtain the optimized first transformation matrix comprising: For each of the plurality of correction positions, a distance calculation is performed on the corresponding first point and the corresponding third point to obtain a plurality of calculation results corresponding to the plurality of correction positions, wherein the distance calculation includes: Obtain the ray connecting the origin of the first three-dimensional coordinate system and the corresponding first point; The corresponding third point is transformed using the first transformation matrix to be solved; as well as Calculate the distance between the transformed corresponding third point and the ray; as well as The first transformation matrix to be solved is iteratively adjusted by a convergence function, and the iteratively adjusted first transformation matrix to be solved is used as the optimized first transformation matrix, wherein the convergence function indicates that the sum of the plurality of calculation results is minimized.

3. The multi-mode image alignment method as described in claim 1, wherein, using the plurality of first-transformation images, the plurality of first points, the plurality of second point groups, and the default specification parameter group of the calibration body, based on the first three-dimensional coordinate system, the second optimization operation is performed on the second transformation matrix to be solved to obtain the optimized second transformation matrix comprising: The multiple first-transformed images are processed using the second transformation matrix to generate multiple second-transformed images respectively; Based on the plurality of first points and the plurality of second point groups, a plurality of fourth point groups are obtained from the plurality of secondary transformation images respectively; Based on the aforementioned multiple fourth point groups, multiple estimated specification parameter groups are obtained respectively; as well as The second transformation matrix to be solved is iteratively adjusted by a convergence function, and the iteratively adjusted second transformation matrix to be solved is used as the optimized second transformation matrix, wherein the convergence function indicates that the difference between the plurality of estimated specification parameter sets and the default specification parameter set of the calibration body is minimized.

4. The multi-mode image alignment method as described in claim 3, wherein the plurality of secondary transformation images, the plurality of first points, and the plurality of second point groups respectively correspond to a plurality of correction positions, and the plurality of fourth point groups obtained from the plurality of secondary transformation images based on the plurality of first points and the plurality of second point groups respectively include: For each of the plurality of correction positions, the following is performed: (The first point, the second point group, and the secondary transformed image are then processed.) Project the corresponding first point onto the corresponding quadratic transformed image to obtain the fifth point; and Project multiple points in the second point group onto the corresponding quadratic transformation image to obtain multiple sixth points; The fifth point and the plurality of sixth points form one of the plurality of fourth point groups.

5. The multi-mode image alignment method as described in claim 4, wherein obtaining the plurality of estimated specification parameter groups based on the plurality of fourth point groups includes: For each of the plurality of fourth point groups, specification parameter estimation is performed to obtain one of the plurality of estimated specification parameter groups, wherein the specification parameter estimation includes: Obtain multiple connections between the fifth point and the plurality of sixth points; and Calculate multiple estimated lengths of the multiple connections and multiple estimated included angles between the multiple connections; The default specification parameters include multiple preset side lengths and multiple preset included angles of the calibration body. The difference between the multiple estimated specification parameter groups and the default specification parameter group indicates the weighted sum of the first value and the second value. The first value indicates the sum of multiple differences between the multiple estimated lengths and the multiple preset side lengths, and the second value indicates the sum of multiple differences between the multiple estimated included angles and the multiple preset included angles.

6. The multi-modal image alignment method as described in claim 1, wherein obtaining a plurality of third points corresponding to the central vertex from the plurality of three-dimensional images comprises: Using each of the multiple 3D images as the target image, execute: Three planes are obtained from the target image, the three planes being adjacent to each other and each having three normal vectors that are perpendicular to each other; and Obtain the intersection point of the three planes, which is one of the plurality of third points.

7. The multimodal image alignment method as claimed in claim 1, wherein the plurality of two-dimensional images are a plurality of first two-dimensional images, and the multimodal image alignment method further comprises being executed by the processing device: Acquire multiple second two-dimensional images associated with the calibration body, the multiple second two-dimensional images corresponding to a third three-dimensional coordinate system; From the plurality of second two-dimensional images, obtain a plurality of seventh points corresponding to the central vertex and a plurality of eighth point groups corresponding to the plurality of side vertices; Using the third transformation matrix to be solved, the plurality of seventh points, and the plurality of third points, the first optimization operation is performed based on the third three-dimensional coordinate system to obtain the optimized third transformation matrix; The optimized third transformation matrix is ​​used to process the multiple 3D images; Using the processed multiple 3D images, the multiple seventh points, the multiple eighth point groups, and the default specification parameter group of the calibration body, the second optimization operation is performed based on the third 3D coordinate system to obtain the optimized fourth transformation matrix; and The first optimized transformation matrix, the second optimized transformation matrix, the third optimized transformation matrix, and the fourth optimized transformation matrix are used to perform the transformation between the first three-dimensional coordinate system and the third three-dimensional coordinate system.

8. The multi-modal image alignment method as claimed in claim 1, wherein the first three-dimensional coordinate system is the camera coordinate system of the two-dimensional image acquisition device, and the multi-modal image alignment method further comprises being executed by the processing device: Using the parameter matrix of the two-dimensional image acquisition device, the image to be processed is transformed from the first three-dimensional coordinate system to the image plane coordinate system of the two-dimensional image acquisition device; The parameter matrix includes the focal length and projection center parameters of the two-dimensional image acquisition device.

9. The multimodal image alignment method as described in claim 1, wherein the corrector comprises a plurality of indicator components respectively disposed at the central vertex and the plurality of side vertices, and the multimodal image alignment method further comprises: A two-dimensional image acquisition device is used to perform multiple image capture procedures on the calibration object, which is placed in multiple calibration positions in turn. In each of these multiple image capture procedures, the multiple indicator components are enabled in turn, and the calibration object is captured, thereby generating the multiple two-dimensional images; and The calibration body, which is placed in the multiple calibration positions in turn, is captured by a three-dimensional image acquisition device to generate the multiple three-dimensional images.

10. The multimodal image alignment method as claimed in claim 1, wherein the corrector further comprises a plurality of indicator components, respectively disposed at the central vertex and the plurality of side vertices and respectively having different colors or different temperatures, and the multimodal image alignment method further comprises: The calibration object is photographed by a two-dimensional image acquisition device at multiple calibration positions in turn to generate the multiple two-dimensional images; and The calibration body, which is placed in the multiple calibration positions in turn, is captured by a three-dimensional image acquisition device to generate the multiple three-dimensional images.

11. A multimodal image alignment system, comprising: The corrector includes: A three-dimensional solid having a central vertex and multiple side vertices; and Multiple indicator components are respectively disposed at the central vertex and the multiple side vertices; A two-dimensional image acquisition device has a first three-dimensional coordinate system and is used to generate multiple two-dimensional images associated with the calibration body; A three-dimensional image acquisition device has a second three-dimensional coordinate system and is used to generate multiple three-dimensional images associated with the calibration body; as well as A processing device, connected to the two-dimensional image acquisition device and the three-dimensional image acquisition device, is used to obtain a first coordinate transformation matrix based on the plurality of two-dimensional images and the plurality of three-dimensional images, and to use the first coordinate transformation matrix to transform the image to be processed to the second three-dimensional coordinate system or the first three-dimensional coordinate system. The process of obtaining the first coordinate transformation matrix performed by the processing device includes: Obtain multiple first points corresponding to the central vertex and multiple second point groups corresponding to the multiple side vertices from the multiple two-dimensional images; Multiple third points corresponding to the central vertex are obtained from the multiple three-dimensional images; Using the first transformation matrix to be solved, the plurality of first points, and the plurality of third points, a first optimization operation is performed based on the first three-dimensional coordinate system to obtain an optimized first transformation matrix; The optimized first transformation matrix is ​​used to process the multiple three-dimensional images to generate multiple first-transformation images respectively; as well as Using the multiple first-transformation images, the multiple first points, the multiple second point groups, and the default specification parameter group of the calibration body, a second optimization operation is performed on the second transformation matrix to be solved based on the first three-dimensional coordinate system to obtain the optimized second transformation matrix; The first coordinate transformation matrix includes the optimized first transformation matrix and the optimized second transformation matrix.

12. The multi-modal image alignment system as described in claim 11, further comprising: Another two-dimensional image acquisition device, connected to the processing device, has a third three-dimensional coordinate system and is used to generate multiple second two-dimensional images associated with the calibration body; The processing device is further configured to obtain a second coordinate transformation matrix based on the plurality of second two-dimensional images and the plurality of three-dimensional images, and to perform a transformation between the first three-dimensional coordinate system and the third three-dimensional coordinate system using the first coordinate transformation matrix and the second coordinate transformation matrix.

13. The multi-mode image alignment system of claim 11, wherein the plurality of indicator components have different colors or different temperatures.

14. The multi-mode image alignment system of claim 11, wherein the plurality of first points and the plurality of third points respectively correspond to a plurality of correction positions, and the first optimization operation performed by the processing device includes: For each of the plurality of correction positions, a distance calculation is performed on the corresponding first point and the corresponding third point to obtain a plurality of calculation results corresponding to the plurality of correction positions, wherein the distance calculation includes: Obtain the ray connecting the origin of the first three-dimensional coordinate system and the corresponding first point; The corresponding third point is transformed using the first transformation matrix to be solved; as well as Calculate the distance between the transformed corresponding third point and the ray; as well as The first transformation matrix to be solved is iteratively adjusted by a convergence function, and the iteratively adjusted first transformation matrix to be solved is used as the optimized first transformation matrix, wherein the convergence function indicates that the sum of the plurality of calculation results is minimized.

15. The multimodal image alignment system of claim 11, wherein the second optimization operation performed by the processing device comprises: The multiple first-transformed images are processed using the second transformation matrix to generate multiple second-transformed images respectively; Based on the plurality of first points and the plurality of second point groups, a plurality of fourth point groups are obtained from the plurality of secondary transformation images respectively; Based on the aforementioned multiple fourth point groups, multiple estimated specification parameter groups are obtained respectively; as well as The second transformation matrix to be solved is iteratively adjusted by a convergence function, and the iteratively adjusted second transformation matrix to be solved is used as the optimized second transformation matrix, wherein the convergence function indicates that the difference between the plurality of estimated specification parameter sets and the default specification parameter set of the calibration body is minimized.

16. The multi-mode image alignment system of claim 15, wherein the plurality of secondary transformation images, the plurality of first points, and the plurality of second point groups respectively correspond to a plurality of correction positions, and the acquisition of the plurality of fourth point groups performed by the processing device comprises: For each of the plurality of correction positions, the following is performed: (The first point, the second point group, and the secondary transformed image are then processed.) Project the corresponding first point onto the corresponding quadratic transformed image to obtain the fifth point; and Project multiple points in the second point group onto the corresponding quadratic transformation image to obtain multiple sixth points; The fifth point and the plurality of sixth points form one of the plurality of fourth point groups.

17. The multimodal image alignment system of claim 16, wherein the acquisition of the plurality of estimated specification parameter sets performed by the processing device comprises: For each of the plurality of fourth point groups, specification parameter estimation is performed to obtain one of the plurality of estimated specification parameter groups, wherein the specification parameter estimation includes: Obtain multiple connections between the fifth point and the plurality of sixth points; and Calculate multiple estimated lengths of the multiple connections and multiple estimated included angles between the multiple connections; The default specification parameters include multiple preset side lengths and multiple preset included angles of the calibration body. The difference between the multiple estimated specification parameter groups and the default specification parameter group indicates the weighted sum of the first value and the second value. The first value indicates the sum of multiple differences between the multiple estimated lengths and the multiple preset side lengths, and the second value indicates the sum of multiple differences between the multiple estimated included angles and the multiple preset included angles.

18. The multimodal image alignment system of claim 11, wherein the acquisition of the plurality of third points performed by the processing device comprises: Using each of the multiple 3D images as the target image, execute: Three planes are obtained from the target image, the three planes being adjacent to each other and each having three normal vectors that are perpendicular to each other; and Obtain the intersection point of the three planes, which is one of the plurality of third points.

19. The multi-mode image alignment system as claimed in claim 11, wherein the first three-dimensional coordinate system is the camera coordinate system of the two-dimensional image acquisition device, and the processing device is further configured to use the parameter matrix of the two-dimensional image acquisition device to transform the image to be processed from the first three-dimensional coordinate system to the image plane coordinate system of the two-dimensional image acquisition device, and the parameter matrix includes the focal length parameter and projection center parameter of the two-dimensional image acquisition device.