Image alignment method and apparatus for vnir multimodule hyperspectral camera and medium

By calculating the image sharpness and the correspondence of the identified feature points of the VNIR multi-module hyperspectral camera, and then calculating the homography matrix for image alignment, the problem of low image alignment accuracy in the existing technology is solved, and high-precision image alignment and data fusion are achieved.

CN119693425BActive Publication Date: 2026-05-01RES INST OF TSINGHUA PEARL RIVER DELTA
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
RES INST OF TSINGHUA PEARL RIVER DELTA
Filing Date
2024-11-07
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Because VNIR multi-module hyperspectral cameras do not integrate gyroscopes, existing image alignment methods are limited by the accuracy of image feature point detection results, resulting in low alignment accuracy.

Method used

By acquiring VIS and NIR images containing the target object and the label image, calculating the image sharpness, identifying the label image and matching it, identifying the label feature points, and calculating the homography matrix for image alignment, the system avoids relying on the feature point recognition results.

Benefits of technology

It improves the accuracy of image alignment results, achieves high-precision alignment of VIS and NIR images, and supports subsequent spectral analysis and data fusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693425B_ABST
    Figure CN119693425B_ABST
Patent Text Reader

Abstract

The application discloses an image alignment method and device for a VNIR multimode group hyperspectral camera and a medium, and can be applied to the technical field of image data processing. The application calculates the sharpness of a to-be-processed VIS image and a to-be-processed NIR image, identifies the mark images on the to-be-processed VIS image and the to-be-processed NIR image according to the sharpness, identifies the mark feature points on the mark images after the mark images on the to-be-processed VIS image and the to-be-processed NIR image are identified and corresponded, respectively corresponds the mark feature points with the mark labels and mark positions on the to-be-processed VIS image and the to-be-processed NIR image, calculates a homography matrix according to the mark feature points corresponding to the feature points, and then performs an alignment operation on the to-be-processed VIS image and the to-be-processed NIR image according to the homography matrix, so that the accuracy of the image alignment result can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image alignment method, apparatus, and medium for VNIR multi-module hyperspectral cameras Technical Field

[0001] This application relates to the field of image data processing technology, and in particular to an image alignment method, apparatus and medium for a VNIR multi-module hyperspectral camera. Background Technology

[0002] Among related technologies, the VNIR multi-module hyperspectral camera is an advanced imaging device that combines hyperspectral imaging technology across the visible and near-infrared (VNIR) bands. This camera provides rich information for application fields by capturing spectral data across a wide range of wavelengths. Since the VNIR multi-module hyperspectral camera integrates multiple spectral sensors, image alignment of the spectral images acquired by these modules is necessary. Because the VNIR multi-module hyperspectral camera does not integrate attitude detection devices such as gyroscopes, image alignment is currently achieved through image feature point detection methods. However, this method is limited by the accuracy of the image feature point detection results, leading to low alignment accuracy in practical applications.

[0003] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0004] The main objective of this application is to propose an image alignment method, apparatus, and medium for a VNIR multi-module hyperspectral camera, which can effectively improve the accuracy of image alignment results.

[0005] To achieve the above objectives, one aspect of this application proposes an image alignment method for a VNIR multi-mode hyperspectral camera, the method comprising the following steps:

[0006] Acquire several VIS images and several NIR images to be processed corresponding to the target object captured by a VNIR multi-module hyperspectral camera. Both the VIS images and the NIR images to be processed contain the target object image and the identifier image.

[0007] Calculate the sharpness of the VIS image to be processed and the NIR image to be processed;

[0008] Based on the sharpness, identify the labeled images on the VIS image to be processed and the NIR image to be processed;

[0009] The identification images on the VIS image to be processed are matched with the identification images on the NIR image to be processed;

[0010] Identify the feature points on the identifier image after identifying the identifier;

[0011] The identification feature points are respectively mapped to the identification numbers and positions on the VIS image and the NIR image to be processed;

[0012] The homography matrix is ​​calculated based on the corresponding feature points of the identifier;

[0013] The VIS image and the NIR image to be processed are aligned according to the homography matrix.

[0014] In some embodiments, the identification image includes a QR code icon image or a preset corresponding point image, the QR code icon image corresponding to the QR code icon image is set on a calibration plate, the calibration plate is set on a reference plate, and the target photographed object is set in the hollow area of ​​the calibration plate and in contact with the reference plate.

[0015] In some embodiments, calculating the sharpness of the VIS image to be processed and the NIR image to be processed includes:

[0016] The VIS image and the NIR image to be processed are processed respectively to obtain the corresponding Laplacian images;

[0017] Calculate the Laplacian variance of the Laplacian image corresponding to the VIS image to be processed and the Laplacian image corresponding to the NIR image to be processed. The Laplacian variance is used to measure the sharpness of the VIS image to be processed and the NIR image to be processed.

[0018] In some embodiments, identifying the identifier image on the VIS image to be processed and the NIR image to be processed based on the sharpness includes:

[0019] The Laplacian variances of the VIS image and the NIR image to be processed are sorted respectively.

[0020] Based on the sorting results, the labeled images on the VIS image and the NIR image to be processed are identified respectively.

[0021] In some embodiments, the step of identifying the labeled images on the VIS image to be processed and the NIR image to be processed based on the sorting results includes:

[0022] Based on the sorting results, marker detection is performed on the VIS image to be processed and the NIR image to be processed. The marker detection includes threshold detection and contour detection.

[0023] The identification images of the VIS image and the NIR image to be processed are determined based on the detection results.

[0024] In some embodiments, calculating the homography matrix based on the identified feature points after feature point correspondence includes:

[0025] Iteratively select the feature point to be calculated from the identified feature points corresponding to the feature point;

[0026] The homography matrix is ​​calculated based on the feature points to be calculated.

[0027] In some embodiments, the alignment operation of the VIS image to be processed and the NIR image to be processed according to the homography matrix includes:

[0028] Select one type of image from the VIS image to be processed and the NIR image to be processed as the image to be transferred, and use the remaining type of image as the target image;

[0029] Obtain the pixel coordinates of the image to be transferred;

[0030] The grayscale values ​​of the image to be transferred are transferred to the target image based on the pixel coordinates and the homography matrix.

[0031] To achieve the above objectives, another aspect of this application provides an image alignment device for a VNIR multi-mode hyperspectral camera, the device comprising:

[0032] The first module is used to acquire several VIS images and several NIR images to be processed corresponding to the target object captured by the VNIR multi-module hyperspectral camera. Both the VIS images and the NIR images to be processed contain the target object image and the identifier image.

[0033] The second module is used to calculate the sharpness of the VIS image to be processed and the NIR image to be processed;

[0034] The third module is used to identify the marker images on the VIS image to be processed and the NIR image to be processed based on the sharpness.

[0035] The fourth module is used to identify the corresponding labels on the VIS image to be processed and the labels on the NIR image to be processed.

[0036] The fifth module is used to identify the feature points of the identifier on the identifier image after the identifier is matched;

[0037] The sixth module is used to map the identified feature points to the identifiers and their positions on the VIS image and the NIR image to be processed, respectively.

[0038] The seventh module is used to calculate the homography matrix based on the corresponding feature points of the identifiers;

[0039] The eighth module is used to perform alignment operations on the VIS image to be processed and the NIR image to be processed according to the homography matrix.

[0040] To achieve the above objectives, another aspect of this application provides a computer device, comprising:

[0041] At least one processor;

[0042] At least one memory for storing at least one program;

[0043] When the at least one program is executed by the at least one processor, the at least one processor performs the method described above.

[0044] To achieve the above objectives, another aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method.

[0045] The embodiments of this application include at least the following beneficial effects: This application provides an image alignment method, apparatus, and medium for a VNIR multi-module hyperspectral camera. This scheme acquires several VIS images and several NIR images to be processed, each containing an image of the target object and an identifier image. It then calculates the sharpness of the VIS and NIR images, identifies the identifier images on them based on the sharpness, and matches the identifier images on the VIS and NIR images with the identifier images on the NIR images. It then identifies the identifier feature points on the matched identifier images, matches these feature points with the identifier numbers and positions on the VIS and NIR images, calculates a homography matrix based on the corresponding identifier feature points, and performs alignment operations on the VIS and NIR images based on the homography matrix. This achieves the alignment process of VIS and NIR images without relying on the identification results of the target image feature points, effectively improving the accuracy of the image alignment results. Attached Figure Description

[0046] Figure 1 is a flowchart of an image alignment method for a VNIR multi-module hyperspectral camera provided in an embodiment of this application;

[0047] Figure 2 is a schematic diagram of the calibration plate provided in an embodiment of this application;

[0048] Figure 3 is a schematic diagram of taking pictures using a calibration board according to an embodiment of this application;

[0049] Figure 4 is a VIS image corresponding to a wavelength of 510nm provided in the embodiments of this application;

[0050] Figure 5 is a VIS image corresponding to a wavelength of 590nm provided in the embodiments of this application;

[0051] Figure 6 is a VIS image corresponding to a wavelength of 690nm provided in an embodiment of this application;

[0052] Figure 7 is an NIR image corresponding to a wavelength of 713nm provided in an embodiment of this application;

[0053] Figure 8 is an NIR image corresponding to a wavelength of 828 nm provided in an embodiment of this application;

[0054] Figure 9 is an NIR image corresponding to a wavelength of 920nm provided in an embodiment of this application;

[0055] Figure 10 is a schematic diagram of the ArUco icon recognition results provided in an embodiment of this application;

[0056] Figure 11 is a schematic diagram showing the correspondence between the icon feature points of the VIS image and the NIR image shown in Figure 12 provided in the embodiments of this application;

[0057] Figure 12 is a schematic diagram showing the correspondence between the icon feature points of the NIR image provided in the embodiment of this application and the VIS image shown in Figure 11;

[0058] Figure 13 is a schematic diagram of the alignment result of the VIS image to be processed to the NIR image to be processed shown in Figure 14 provided in the embodiment of this application.

[0059] Figure 14 is a schematic diagram of the alignment result of the VIS image to be processed to the NIR image to be processed shown in Figure 13 provided in the embodiment of this application;

[0060] Figure 15 is a schematic diagram of the image alignment device for a VNIR multi-module hyperspectral camera provided in an embodiment of this application;

[0061] Figure 16 is a schematic diagram of the hardware structure of the computer device provided in an embodiment of this application. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application.

[0063] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0064] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0066] Among related technologies, the VNIR multi-module hyperspectral camera is an advanced imaging device that combines hyperspectral imaging technology in the visible to near-infrared (VNIR) bands. By capturing spectral data across a wide range of wavelengths, it provides a wealth of information for various applications. Specifically, the VNIR multi-module hyperspectral camera can cover the spectral range from visible to near-infrared (wavelength range of approximately 510 nm–920 nm), capturing data at multiple wavelengths per pixel, enabling users to perform detailed spectral analysis.

[0067] Understandably, the VNIR multi-module design allows for the integration of multiple sensors and spectral modules. Currently, mainstream hyperspectral cameras typically integrate two spectral sensing modules, covering the VIS (visible light) and NIR (near-infrared) bands respectively. The system then integrates the information from these two modules to output a unified hyperspectral information of the captured object. Since VNIR multi-module hyperspectral cameras do not integrate attitude detection devices such as gyroscopes, image alignment is currently achieved through image feature point detection. However, this method is limited by the accuracy of the image feature point detection results, leading to low alignment accuracy in practical applications.

[0068] In view of this, this application provides an image alignment method, apparatus, and medium for a VNIR multi-module hyperspectral camera. This application acquires several VIS images and several NIR images to be processed, each containing an image of the target object and an identifier image. It then calculates the sharpness of the VIS and NIR images, identifies the identifier images based on the sharpness, and matches the identifier images on the VIS and NIR images with the identifier images on the NIR images. After matching the identifier images on the VIS and NIR images, it identifies the identifier feature points on the matched identifier images. These feature points are then matched with the identifier numbers and positions on the VIS and NIR images, and a homography matrix is ​​calculated based on the corresponding identifier feature points. Finally, the VIS and NIR images are aligned according to the homography matrix. This achieves the alignment process of the VIS and NIR images without relying on the identification results of the target image feature points, effectively improving the accuracy of the image alignment results.

[0069] The image alignment method for a VNIR multi-module hyperspectral camera provided in this application relates to the field of image data processing technology. This image alignment method for a VNIR multi-module hyperspectral camera can be applied to a terminal, a server, or software running on a terminal or server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, or vehicle terminal, but is not limited to these. The server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The server can also be a node server in a blockchain network. The software can be an application implementing the image alignment method for a VNIR multi-module hyperspectral camera, but is not limited to the above forms.

[0070] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0071] Figure 1 is an optional flowchart of an image alignment method for a VNIR multi-mode hyperspectral camera provided in an embodiment of this application. The method in Figure 1 may include, but is not limited to, steps S110 to S180:

[0072] Step S110: Obtain several VIS images and several NIR images to be processed corresponding to the target object captured by the VNIR multi-module hyperspectral camera. The VIS images and NIR images to be processed each contain the target object image and the identification image.

[0073] Step S120: Calculate the sharpness of the VIS image and the NIR image to be processed;

[0074] Step S130: Identify the marker images on the VIS image and the NIR image to be processed based on their sharpness;

[0075] Step S140: Match the labeled images on the VIS image to be processed with the labeled images on the NIR image to be processed;

[0076] Step S150: Identify the feature points of the identifier on the identifier image after the identifier is identified;

[0077] Step S160: Assign feature points to the label numbers and positions on the VIS image and NIR image to be processed, respectively.

[0078] Step S170: Calculate the homography matrix based on the corresponding feature points;

[0079] Step S180: Align the VIS image and the NIR image to be processed according to the homography matrix.

[0080] It is understood that the VNIR multi-mode hyperspectral camera used in this embodiment has three lens groups, used to capture VIS band (510nm-690nm), NIR band (713nm-920nm), and RGB information respectively. During operation, the camera can simultaneously photograph the same object and save images of the corresponding bands. Due to the different positions of the three lens groups, the image positions will differ when the VNIR multi-mode hyperspectral camera photographs the same object. Therefore, it is necessary to align the images of each different band to provide effective data for subsequent spectral analysis.

[0081] In this embodiment, the identification image can be a QR code-like icon (ArUco) image. The corresponding QR code-like icon is set on a calibration plate, which is placed on a reference plate. When the spectral camera captures an image, the target object is positioned in the cutout area of ​​the calibration plate and in contact with the reference plate. The QR code-like icon is used to record information about the calibration plate. Specifically, an ArUco icon is a two-dimensional matrix-style marking pattern specifically designed for computer vision. It is similar to a QR code, but its function differs. The purpose of ArUco icons is to help computers identify and locate objects. As shown in Figure 2, ArUco icons are designed as simple geometric shapes composed of black and white squares. Compared to QR codes, they are easier to identify and detect, and can be accurately identified even in dark or complex environments. Each ArUco icon contains a unique identifier. This embodiment can detect the icon and identify its identifier and key point positions using computer vision algorithms.

[0082] As can be understood, as shown in Figure 2, this embodiment can generate ArUco icons with different identifiers and place them on the calibration board according to a preset combination (ArUco icons in the four corner positions of Figure 2). The identifier combination corresponding to the ArUco icon can be used as the marking number of the photographed object. Specifically, this embodiment can use the DICT_7X7_1000 method (7×7 means that each icon is a 7×7 matrix; 1000 means that this dictionary contains 1000 different icons, with identifiers from 0 to 999) to construct ArUco icons, balancing information content and icon simplicity. When constructing the calibration board, a certain number of ArUco icons are placed in the corners of the calibration board, and the center of the calibration board is hollowed out to place the photographed object and the shooting reference board. Generally, this invention places four 2×2 ArUco icons in each of the four corners of the calibration board, leaving a certain width of white space around each icon to increase the stability of recognition. The identifiers of the icons can be randomly generated or generated according to a certain meaningful combination.

[0083] After the calibration board design is completed, when using a VNIR multi-module hyperspectral camera for imaging, as shown in Figure 3, place the reference board at the bottom, place the calibration board on top, and then place the subject in the cutout portion of the calibration board. During imaging, ensure that the calibration board, reference board, and subject are captured simultaneously. Specifically, this embodiment can use the VIS lens and NIR lens in the VNIR multi-module hyperspectral camera to simultaneously capture VIS images (VIS images to be processed) and NIR images (NIR images to be processed). In this embodiment, the VIS lens and NIR lens can capture one or more images in the ranges of 510nm-690nm and 713nm-920nm, respectively. Examples include the VIS images shown in Figures 4 (510nm), 5 (590nm), and 6 (690nm), and the NIR images shown in Figures 7 (713nm), 8 (828nm), and 9 (920nm). As can be seen from Figures 4 to 9, this embodiment ensures that the calibration board, the photographed object, and the reference board can be completely captured in both the VIS image and the NIR image during the shooting process.

[0084] In this embodiment, images captured by a VNIR multi-mode hyperspectral camera are selected from the VIS group and NIR group respectively as the VIS image and NIR image to be processed. Then, an alignment operation is performed on the VIS image and NIR image to be processed. During the image selection process, it is necessary to ensure that all selected images can correctly identify the ArUco icon in order to improve the accuracy of the image alignment algorithm.

[0085] Specifically, after obtaining the VIS image and the NIR image to be processed, this embodiment measures the edge strength of the image by calculating its Laplacian transform, thereby determining the image's sharpness. Image sharpness is closely related to the clarity of its edges; the clearer the edges, the sharper the image. Generally, a sharper image indicates a greater distinction between the ArUco icon and the background, which is more conducive to correct ArUco icon recognition. It can be understood that this embodiment can obtain corresponding Laplacian images by processing the VIS image and the NIR image to be processed separately, and then calculate the Laplacian variance of the Laplacian images corresponding to the VIS image and the NIR image. The Laplacian variance in this embodiment is used to measure the sharpness of the VIS image and the NIR image to be processed.

[0086] Specifically, the Laplacian image transformation process is as follows:

[0087]

[0088] In the formula, I(x,y) represents the grayscale value of the VIS image or NIR image to be processed, x and y represent the spatial coordinates of the image, and L(x,y) represents the Laplacian image. In the Laplacian image, the larger the absolute value of the pixel value, the stronger the edge intensity at the corresponding location.

[0089] The Laplace variance is calculated using the following formula:

[0090]

[0091] M and N represent the width and height of the Laplacian image; L(i,j) represents the pixel value of the Laplacian image; μ represents the mean of the Laplacian image; and Variance represents the Laplacian variance. Specifically, the larger the variance, the greater the variation in edge intensity in the image, and the sharper the image; the smaller the variance, the blurrier the image.

[0092] In this embodiment, after obtaining the Laplacian variance representing the sharpness, the Laplacian variances of the VIS image and the NIR image to be processed are sorted respectively. Based on the sorting results, the marker images on the VIS image and the NIR image to be processed are identified respectively. Taking the recognition of the ArUco icon as an example, this embodiment can sort the images of the VIS group and the NIR group from largest to smallest according to the variance calculated in the previous step, and use the ArUco icon recognition algorithm to identify the ArUco icon in the image until the ArUco icon can be correctly identified in both the VIS group and the NIR group. If, in some cases, the ArUco icon cannot be correctly identified due to image shooting quality or other factors, this embodiment can also achieve image alignment by setting key points in the interface. The marker image corresponding to the key point is a preset corresponding point image.

[0093] Specifically, when identifying the label image corresponding to the ArUco icon, this embodiment can perform label detection on the VIS image and NIR image to be processed based on the sorting results, and identify the label image of the VIS image and NIR image to be processed based on the detection results. Label detection includes threshold detection and contour detection. It can be understood that threshold detection can be applied to the grayscale image of the spectrum to perform binarization thresholding to highlight the label area. Contour detection can be performed to identify contours in the binarized image; these contours are potential label candidates. Then, these contours are simplified and filtered according to their shapes to find quadrilaterals and calculate their corner points. After threshold detection and contour detection are completed, the corner points of the quadrilaterals are used to perform perspective transformation on the quadrilaterals, distorting them into squares, so that the algorithm can standardize the size and orientation of the label. Then, the binary code inside the label is extracted. The interior of the ArUco label consists of a grid of black and white cells, each cell encoded with a binary value. The extracted binary code is matched with a predefined ArUco label dictionary to identify the label number, thus obtaining the recognition result of the ArUco icon shown in Figure 10.

[0094] In this embodiment, after correctly identifying the ArUco icon, the corresponding identifier image is matched according to the identifier number in the ArUco icon. Then, the points at the top left, top right, bottom left, and bottom right corners of the ArUco icon contained in the VIS image and the NIR image to be processed are detected as identifier feature points. In this example with 16 ArUco icons, there are a total of 64 feature points to be processed, and the positions of these identifier feature points are recorded. Next, the identifier feature points detected in the VIS image and NIR image are matched one-to-one according to the ArUco icon identifier number and position (as shown in the matching process in Figures 11 and 12).

[0095] In this embodiment, after obtaining the correspondence between feature points in the VIS and NIR images, the homography matrix is ​​estimated using the RANSAC (Random Sample Consensus) algorithm. Specifically, the homography matrix can be calculated by iteratively selecting feature points to be calculated from the identified feature points corresponding to the feature points. For example, this embodiment can use the RANSAC algorithm to iteratively select a subset of corresponding points and use these points to estimate the homography matrix, thereby finding the optimal transformation model. As shown in the following formula, the homography matrix H is a 3×3 matrix representing the transformation relationship between the two images:

[0096]

[0097] In this embodiment, after estimating the homography transformation matrix, the homography matrix can be used to transform one image onto another, thereby achieving image alignment. Specifically, this embodiment obtains a homography matrix aligned from a VIS image to a NIR image.

[0098] It is understood that in this embodiment, one type of image is selected from the VIS image to be processed and the NIR image to be processed as the image to be transferred, and the remaining type of image is selected as the target image. Then, the pixel coordinates of the image to be transferred are obtained, and the gray values ​​of the image to be transferred are transferred to the target image according to the pixel coordinates and the homography matrix. Specifically, this real-time exchange rate can use an image perspective transformation algorithm to apply the homography matrix to the image that needs to be aligned, so as to transform each pixel in the image to be aligned to the corresponding position in the target image. For example, taking the alignment of the VIS image to be processed to the NIR image as an example, this embodiment first multiplies the coordinates of each pixel in the VIS image to be processed by the homography matrix to obtain the corresponding coordinates x' and y', and then transfers the gray values ​​of the corresponding pixels in the VIS image to be processed, to obtain the result of aligning the VIS image to be processed to the NIR image to be processed as shown in Figure 13. In this embodiment, for the case of multiple VIS band images captured, the obtained homography matrix can be applied to each VIS band image to realize the operation of aligning each VIS band image to the NIR band image.

[0099] This application embodiment accurately registers VIS and NIR image information, enabling corresponding information in the two sets of images to correspond. Image data from different spectra can be directly fused, analyzed, and processed, thereby providing richer visual data and providing effective data support for subsequent data analysis and feature extraction.

[0100] Specifically, the practical applications of the VIS and NIR dual-lens shooting device include, but are not limited to:

[0101] Industrial Inspection: In manufacturing, a VIS and NIR dual-lens system can be used to simultaneously inspect surface and internal defects in products. Aligned image data can be overlaid for comprehensive analysis, improving inspection accuracy and efficiency.

[0102] Medical Imaging: In the field of medical imaging, aligning VIS and NIR images allows for better observation of the tissue beneath the skin, aiding in the diagnosis of various conditions such as skin lesions. The combination of information from both spectra provides a more comprehensive diagnostic basis.

[0103] Food Inspection: Utilizing a dual-lens system of VIS and NIR allows for more effective use of spectral imaging information to detect the content of food components, contaminants, and spoilage, ensuring food safety and quality. The addition of VIS and NIR image alignment technology further enhances the detection of small particulate matter such as tea leaves and coffee beans.

[0104] Therefore, this embodiment, through image alignment technology, enables the VIS and NIR dual-lens imaging device to more effectively combine two spectral types of information, achieving more efficient and accurate multispectral imaging applications. This plays a significant role in improving the quality of data analysis and enhancing the application effects in multiple fields.

[0105] Referring to Figure 15, this application embodiment provides an image alignment device for a VNIR multi-module hyperspectral camera, the device comprising:

[0106] The first module 1610 is used to acquire several VIS images and several NIR images to be processed corresponding to the target photographed by the VNIR multi-module hyperspectral camera. Both the VIS images and the NIR images to be processed contain the target object image and the identification image.

[0107] The second module 1620 is used to calculate the sharpness of the VIS image and the NIR image to be processed;

[0108] The third module 1630 is used to identify the marked images on the VIS image and the NIR image to be processed based on sharpness.

[0109] The fourth module 1640 is used to identify and match the labeled images on the VIS image to be processed with the labeled images on the NIR image to be processed.

[0110] The fifth module, 1650, is used to identify the feature points of the corresponding identifier on the identifier image.

[0111] The sixth module 1660 is used to map the marker feature points to the marker numbers and marker positions on the VIS image and the NIR image to be processed, respectively.

[0112] Module 7, 1670, is used to calculate the homography matrix based on the corresponding identifier feature points.

[0113] The eighth module 1680 is used to perform alignment operations on the VIS image and the NIR image to be processed based on the homography matrix.

[0114] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0115] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described image alignment method for a VNIR multi-module hyperspectral camera. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0116] It is understood that the content of the above method embodiments is applicable to this device embodiment. The specific functions implemented by this device embodiment are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0117] Please refer to Figure 16, which illustrates the hardware structure of a computer device according to another embodiment. The computer device includes:

[0118] The processor 1710 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0119] The memory 1720 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1720 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1720 and is called and executed by the processor 1710 to execute the image alignment method for a VNIR multi-module hyperspectral camera according to the embodiments of this application.

[0120] The input / output interface 1730 is used to implement information input and output;

[0121] The communication interface 1740 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0122] Bus 1750 transmits information between various components of the device (e.g., processor 1710, memory 1720, input / output interface 1730, and communication interface 1740);

[0123] The processor 1710, memory 1720, input / output interface 1730 and communication interface 1740 are connected to each other within the device via bus 1750.

[0124] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image alignment method for a VNIR multi-module hyperspectral camera.

[0125] It is understood that the content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.

[0126] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0127] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0128] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0130] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0131] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0132] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0133] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0134] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0135] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0136] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0137] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. An image alignment method for a VNIR multi-module hyperspectral camera, characterized in that, The VNIR multi-module hyperspectral camera is equipped with three lenses, which are used to capture images of the target object in the VIS band, NIR band, and RGB band, respectively. The method includes the following steps: acquiring several VIS images and several NIR images to be processed corresponding to the target object captured by the VNIR multi-module hyperspectral camera, each containing a target object image and a marker image; the marker image includes a QR code-like icon image or a preset corresponding point image, the QR code-like icon image corresponding to the QR code-like icon image is set on a calibration plate, the calibration plate is set on a reference plate, and the target object is set in the hollow area of ​​the calibration plate and is aligned with the reference plate. The VNIR multi-module hyperspectral camera can completely capture the calibration plate, the target object, and the reference plate when capturing several VIS images and several NIR images corresponding to the target object; calculate the sharpness of the VIS images and the NIR images; identify the marker images on the VIS images and the NIR images based on the sharpness; match the marker images on the VIS images with the marker images on the NIR images; identify the marker feature points on the marker images after the matching; and match the marker feature points with the marker numbers and marker positions on the VIS images and the NIR images respectively. Feature point correspondence; a homography matrix is ​​calculated based on the identified feature points after feature point correspondence; alignment is performed on the VIS image and the NIR image to be processed based on the homography matrix; the calculation of the sharpness of the VIS image and the NIR image to be processed includes: processing the VIS image and the NIR image to be processed respectively to obtain corresponding Laplacian images; calculating the Laplacian variance of the Laplacian images corresponding to the VIS image and the NIR image to be processed, the Laplacian variance being used to measure the sharpness of the VIS image and the NIR image to be processed; the sharpness is used to identify the VIS image and the NIR image to be processed. The identification images on the NIR images to be processed include: sorting the Laplacian variances of the VIS images and the NIR images to be processed, respectively; using the ArUco icon recognition algorithm to identify the identification images corresponding to ArUco icons in the images based on the sorting results from largest to smallest, for both the VIS image groups and the NIR images to be processed, until the identification images can be correctly identified in both groups; wherein, if the ArUco icons cannot be correctly identified, the VIS images and the NIR images to be processed are aligned by setting key points in the interface, and the identification images corresponding to the key points set in the interface are preset corresponding point images.

2. The method according to claim 1, characterized in that, The step of identifying the corresponding identifier images of the VIS image group and the NIR image group to be processed based on the sorting results from largest to smallest using the ArUco icon recognition algorithm includes: performing marker detection on the VIS image group and the NIR image group to be processed based on the sorting results from largest to smallest, the marker detection including threshold detection and contour detection; and identifying the identifier images of the VIS image group and the NIR image group to be processed based on the detection results.

3. The method according to claim 1, characterized in that, The step of calculating the homography matrix based on the corresponding feature points includes: iteratively selecting feature points to be calculated from the corresponding feature points; and calculating the homography matrix based on the feature points to be calculated.

4. The method according to claim 1, characterized in that, The alignment operation of the VIS image to be processed and the NIR image to be processed according to the homography matrix includes: selecting one type of image from the VIS image to be processed and the NIR image to be processed as the image to be transferred, and the remaining type of image as the target image; obtaining the pixel coordinates of the image to be transferred; and transferring the grayscale value of the image to be transferred to the target image according to the pixel coordinates and the homography matrix.

5. An image alignment device for a VNIR multi-module hyperspectral camera, characterized in that, The VNIR multi-module hyperspectral camera has three lens groups, respectively used to capture images of the target object in the VIS band, NIR band, and RGB band. The device includes: a first module for acquiring several VIS images and several NIR images to be processed corresponding to the target object captured by the VNIR multi-module hyperspectral camera. Both the VIS and NIR images contain an image of the target object and an identification image. The identification image includes a QR code-like icon image or a preset corresponding point image. The QR code-like icon image is set on a calibration plate, which is set on a reference plate. The target object is set on the calibration plate. The VNIR multi-module hyperspectral camera has a hollowed-out area that contacts the reference plate; wherein, when the VNIR multi-module hyperspectral camera captures several VIS images and several NIR images to be processed corresponding to the target object, it can completely capture the calibration plate, the target object, and the reference plate; the second module is used to calculate the sharpness of the VIS images and the NIR images to be processed; the third module is used to identify the marker images on the VIS images and the NIR images to be processed based on the sharpness; the fourth module is used to match the marker images on the VIS images to be processed with the marker images on the NIR images to be processed; the fifth module is used to identify the marker feature points on the marker images after the marker matching. The sixth module is used to map the identified feature points to the identifier labels and positions on the VIS image and the NIR image to be processed, respectively. The seventh module is used to calculate a homography matrix based on the mapped feature points. The eighth module is used to perform alignment operations on the VIS image and the NIR image to be processed based on the homography matrix. The calculation of the sharpness of the VIS image and the NIR image to be processed includes: processing the VIS image and the NIR image to be processed respectively to obtain corresponding Laplacian images; calculating the Laplacian image corresponding to the VIS image to be processed and the corresponding Laplacian image of the NIR image to be processed. The Laplacian variance of the Laplacian image is used to measure the sharpness of the VIS image and the NIR image to be processed. The step of identifying the marker images on the VIS image and the NIR image to be processed based on the sharpness includes: sorting the Laplacian variances of the VIS image and the NIR image to be processed respectively; and using the ArUco icon recognition algorithm to identify the marker images corresponding to ArUco icons in the images based on the sorting results from largest to smallest, until both the VIS image group and the NIR image group to be processed can correctly identify the marker images.If the ArUco icon cannot be correctly identified, the VIS image and the NIR image to be processed are aligned by setting key points in the interface. The marker image corresponding to the key points set in the interface is a preset corresponding point image.

6. A computer device, characterized in that, include: At least one processor; At least one memory for storing at least one program; when said at least one program is executed by said at least one processor, said at least one processor implements the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Multi-modal image registration method, device and system

    CN112598716A