Image splicing method, device and equipment

Through the target neural network model and energy function optimization method, the image feature points are identified and matched, and the problem of inaccurate feature points recognition in image stitching is solved, and high-quality artifact-free and seamless stitching is achieved, which is suitable for spine X-ray image stitching.

CN120387927AActive Publication Date: 2025-07-29PEKING UNION MEDICAL COLLEGE HOSPITAL

Patent Information

Application Number
CN202510221154.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-29
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify and match feature points in image stitching, resulting in low stitching quality. Especially when spinal X-ray images are spliced, due to field of view and texture repetition, feature points recognition and matching are inaccurate, which affects the stitching effect.

Method used

The target neural network model is used to identify multiple target structures in the image to be stitched, and the homography matrix is determined using the target feature points, and the homography matrix is optimized by constructing the registration energy function and the mixed energy function to realize perspective transformation and stitching of the image.

Benefits of technology

Improve the accuracy and robustness of image stitching, ensure the quality of stitching, achieve artifact-free and seamless stitching, and is suitable for spinal X-ray image stitching in various computer equipment, especially medical equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387927A_ABST
    Figure CN120387927A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and discloses an image stitching method, device and equipment, and the method comprises the steps: recognizing a plurality of target structures in a to-be-stitched image; determining a homography matrix by using the target feature points of the plurality of target structures in different to-be-spliced images; the target feature point is determined based on the geometrical shape of the target structure; performing perspective transformation on one to-be-stitched image in the to-be-stitched images by using the homography matrix to obtain a distorted image; and splicing the distorted image with another to-be-spliced image serving as the reference image. According to the invention, the recognition and registration accuracy of the feature points can be improved, so that the image splicing quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image stitching method, apparatus and device. Background Art

[0002] In image stitching, the homography matrix transformation is a common model for warping images, which includes translation, rotation, scaling, and viewpoint transformation. Traditional image stitching methods use manually defined feature detection methods to calculate the homography matrix. The core idea of these methods is to design optimal features (points, lines, or energy functions) to achieve image alignment. However, since some image features are weak and the textures are repetitive, such as spinal X-ray images taken by a C-arm X-ray machine (the intraoperative spinal images are generally obtained by a small and medium-sized, easily movable C-arm X-ray machine. However, due to its limited field of view (FOV), doctors can only obtain truncated image slices of the spine. Therefore, it is necessary to combine the truncated image slices into a panoramic image through image stitching technology), etc., it becomes particularly difficult to accurately identify and match feature points. The homography matrix for perspective transformation of the image during image stitching is determined based on feature points, that is to say, the accuracy of the homography matrix directly determines the effect of the perspective transformation. If the feature points are not accurately identified and matched, the homography matrix will not be precise enough, which will lead to misalignment between the warped image and the reference image, thereby affecting the stitching quality. Summary of the Invention

[0003] In view of this, the present invention provides an image stitching method, apparatus and device to solve the problem that it is difficult to accurately identify and match feature points in image stitching, resulting in low stitching quality.

[0004] In a first aspect, the present invention provides an image stitching method, which includes:

[0005] Identifying a plurality of target structures in the images to be stitched;

[0006] Determining a homography matrix using the target feature points of the plurality of target structures in different images to be stitched; the target feature points are determined based on the geometric shape of the target structure;

[0007] Performing perspective transformation on one of the images to be stitched using the homography matrix to obtain a warped image;

[0008] Stitching the warped image with another image to be stitched as the reference image.

[0009] In an alternative embodiment, the target structure is a pedicle screw.

[0010] In an alternative embodiment, identifying a plurality of target structures in the images to be stitched includes:

[0011] Using the target neural network model, identify multiple target structures in the images to be stitched respectively;

[0012] Among them, the target neural network model at least includes: a Patch embedding layer, an encoder, a decoder, and a parameter-free attention module;

[0013] The Patch embedding layer is used to divide the images to be stitched into multiple image regions;

[0014] The encoder includes a first VSS block module and a patch merge layer. The multiple image regions divided by the Patch embedding layer are input into the first VSS block module. The first VSS block module is used to extract features from the image regions respectively, and the extracted feature information is input into the patch merge layer for downsampling;

[0015] The decoder includes a Patch expanding layer and a second VSS block module; the Patch expanding layer is used for upsampling, and the second VSS block module is used to extract features from the output of the Patch expanding layer;

[0016] The feature information obtained by downsampling through the patch merge layer in the encoder is transmitted to the decoder through the parameter-free attention module.

[0017] In an alternative embodiment, using the target feature points of multiple target structures in different images to be stitched, determine the homography matrix, including:

[0018] Obtain the initial homography matrix;

[0019] Construct a registration energy function, which is the sum of multiple first distances. A first distance is the distance between a first target feature point in the first image to be stitched and a second target feature point in the warped second image to be stitched; the first image to be stitched and the second image to be stitched are two images to be stitched; the first target feature point and the second target feature point are target feature points within the overlapping part of the first image to be stitched and the warped second image to be stitched, and the second target feature point is the target feature point in the warped second image to be stitched that is closest to the first target feature point; the second image to be stitched is warped by using the initial homography matrix to obtain the warped second image to be stitched;

[0020] Optimize the initial homography matrix by minimizing the value of the registration energy function to obtain the optimal homography matrix.

[0021] In an alternative embodiment, determining a homography matrix using target feature points of multiple target structures in different images to be stitched includes:

[0022] If there are three or more images to be stitched, respectively obtain the minimum registration energy function values between every two images to be stitched;

[0023] For one image to be stitched, use the other image to be stitched with the minimum registration energy function value as the adjacent image to be stitched;

[0024] Determine the homography matrix corresponding to the minimum registration energy function value between two adjacent images to be stitched as the homography matrix between the two adjacent images to be stitched.

[0025] In an alternative embodiment, determining a homography matrix using target feature points of multiple target structures in different images to be stitched includes:

[0026] If there are three or more images to be stitched, respectively obtain the homography matrices between every two adjacent images to be stitched;

[0027] Taking one image to be stitched as the reference image, determine the homography matrix of the third image to be stitched relative to the reference image based on the cumulative multiplication of the homography matrices between two adjacent images to be stitched; the third image to be stitched is the image to be stitched other than the reference image and the adjacent image to be stitched of the reference image.

[0028] In an alternative embodiment, stitching a warped image with another image to be stitched as the reference image includes:

[0029] Construct a hybrid energy function, which includes pixel differences, gradient differences, and depth feature differences between the warped image and the reference image in the gray space;

[0030] Perform energy initialization on the starting boundary of the seam;

[0031] Expand from different starting points to the other end of the image. Each time of expansion, select the pixel with the minimum hybrid energy function value as the next path node until the boundary of the overlapping part between the warped image and the reference image to obtain the optimal path;

[0032] Backtrack along the optimal path to determine the best seam;

[0033] Stitch the warped image and the reference image according to the best seam.

[0034] In an alternative embodiment, stitching a warped image with another image to be stitched as the reference image includes:

[0035] Determine the transition regions on both sides of the seam between the warped image and the reference image;

[0036] Weighted sum is performed on the pixel values of the distorted image and the reference image that belong to the transition region, and used as the pixel value of the stitched image.

[0037] In a second aspect, the present invention provides an image stitching device, the device includes:

[0038] A target structure recognition module, configured to recognize multiple target structures in the images to be stitched;

[0039] A homography matrix determination module, configured to determine a homography matrix by using the target feature points of multiple target structures in different images to be stitched;

[0040] A transformation module, configured to perform perspective transformation on one of the images to be stitched by using the homography matrix to obtain a distorted image;

[0041] A stitching module, configured to stitch the distorted image and another image to be stitched as the reference image.

[0042] In a third aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the image stitching method according to the first aspect or any corresponding embodiment thereof.

[0043] In a fourth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the image stitching method according to the first aspect or any corresponding embodiment thereof.

[0044] In a fifth aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the image stitching method according to the first aspect or any corresponding embodiment thereof.

[0045] The image stitching method, device and equipment provided in this embodiment first recognize multiple target structures in the images to be stitched, then determine the homography matrix based on the target feature points of the recognized target structures, and perform transformation on the images to be stitched based on the determined homography matrix, and finally stitch the transformed distorted image and the reference image. In this process, since the target feature points are determined based on the geometric shape of the target structure, the determination accuracy and matching accuracy can be guaranteed, thereby ensuring the calculation accuracy of the homography matrix and further guaranteeing the stitching quality of the images. In addition, compared with the method of calculating the homography matrix by using the manually defined feature detection method, the manually designed features show low robustness due to their complex feature design, while the method for determining the target feature points in the embodiment of the present invention is simple, so that the image stitching method has high robustness and high versatility. Brief Description of the Drawings

[0046] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the related art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 is a schematic flowchart of an image stitching method according to an embodiment of the present invention;

[0048] Figure 2 is a schematic architecture diagram of a target neural network model according to an embodiment of the present invention;

[0049] Figure 3 is a schematic diagram of the processing process of SimAM on features according to an embodiment of the present invention;

[0050] Figure 4 is a schematic diagram of the registration of target feature points in the overlapping part of two images to be stitched according to an embodiment of the present invention;

[0051] Figure 5 is a schematic diagram of the fusion process of stitching truncated spine X-ray images according to an embodiment of the present invention;

[0052] Figure 6 is a schematic diagram of an overall process when using the image stitching method to stitch truncated spine X-ray images according to an embodiment of the present invention;

[0053] Figure 7 is a structural block diagram of an image stitching device according to an embodiment of the present invention;

[0054] Figure 8 is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Embodiments

[0055] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0056] For the stitching of spine X-ray images taken by a C-arm X-ray machine, the following solutions have also been proposed in the related art:

[0057] 1. Adopt a radiation fluoroscopy X-ray ruler commonly used in orthopedics, and calculate the relative positions between images through the scales on the ruler to achieve image stitching.

[0058] 2. Use a non-fluoroscopic panel placed under the patient, and utilize the absolute coordinates on the panel to complete image stitching.

[0059] 3. By installing a camera and a double-sided mirror on the C-arm, the optical center of the camera is made to coincide with the focus of the X-ray source, thereby realizing the fusion of X-ray images and video images.

[0060] However, the above several methods rely on external markers or structures, which often increase the operation cost. In addition, there are some methods that use deep learning-based stitching strategies. For example, an end-to-end long bone image stitching solution in related technologies uses convolutional neural networks (CNNs) to perform two-dimensional reconstruction on multiple images, and uses the structural similarity index measure (SSIM) and adversarial loss to force the network to generate images that are visually similar to the ground-truth. However, currently, this method is only designed for the femur and has not been extended to other bone types.

[0061] According to an embodiment of the present invention, an embodiment of an image stitching method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of executable computer instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0062] In this embodiment, an image stitching method is provided, which can be used in various computer devices. The computer device can be a medical device. Figure 1 It is a flowchart of the image stitching method according to an embodiment of the present invention. As Figure 1 shown, the process includes the following steps:

[0063] Step S101, identify multiple target structures in the images to be stitched.

[0064] In the embodiment of the present invention, there are no restrictions on the images to be stitched and the target structures. The images to be stitched can be spinal X-ray images (also called spinal X-ray images), or other images. The target structures can be pedicle screws, or other target structures, which are specifically determined according to the specific structures in the images to be stitched.

[0065] There are at least two images to be stitched. Each image to be stitched needs to identify the target structures therein. For example, image segmentation methods such as deep learning models can be used to segment the target structures of the images to be stitched to obtain the target structure masks.

[0066] In some optional specific embodiments, step S101, that is, identifying multiple target structures in the images to be stitched, includes:

[0067] Step S1011, using a target neural network model to respectively identify multiple target structures in the images to be stitched;

[0068] Among them, as Figure 2 shown, the target neural network model at least includes: a block embedding (or patch embedding, i.e., Patchembedding) layer, an encoder, a decoder, and a parameter-free attention module.

[0069] The Patch embedding layer is used to divide the images to be stitched (with a size of H×W×3) into multiple image regions (patches); for example, divided into 4×4 non-overlapping image regions (patches).

[0070] In the embodiments of the present invention, each image region (patches) independently calculates features and maps the number of channels of the image to C.

[0071] The encoder includes a first VSS block module and a block merge (or patch merge) layer. The multiple image regions divided by the Patch embedding layer are input into the first VSS block module. The first VSS block module is used to respectively extract features from the image regions, and the extracted feature information is input into the patch merge layer for downsampling. And expand the feature channels to twice the original.

[0072] The encoder may, for example, include four groups of first VSS block modules and patch merge layers, and each group may include two first VSS block modules.

[0073] The decoder includes a block expansion (also called patch expansion, i.e., Patch expanding) layer and a second VSS block module; the Patch expanding layer is used for upsampling, and the second VSS block module is used to extract features from the output of the Patch expanding layer.

[0074] The decoder may, for example, include four groups of Patch expanding layers and second VSS block modules, and each group may include two second VSS block modules.

[0075] The feature information obtained by downsampling through the patch merge layer in the encoder is transmitted to the decoder through a parameter-free attention module. In the embodiments of the present invention, the feature information obtained by downsampling in the encoder is not directly transmitted to the decoder, but is transmitted to the decoder through a lightweight parameter-free attention module (Simple Attention Module, SimAM) to further focus on the key features of the model.

[0076] After the decoder, there is also a projection layer for restoring the size of the features to the original number of channels.

[0077] In the embodiments of the present invention, in order to achieve the precise segmentation of the target structure (such as a pedicle screw) area, a neural network is constructed to perform semantic segmentation tasks. Specifically, the VM-Unet (Vision Mamba UNet) network architecture is adopted, and a lightweight attention module: SimAM is integrated. The VM-UNet network architecture is a novel UNet architecture based on the state space model, which can not only capture extensive context information but also maintain a linear computational complexity, providing an efficient solution for image segmentation.

[0078] In the embodiments of the present invention, both the VSS block modules (including the first VSS block module and the second VSS block module) include two branches. One branch includes a linear layer, a depthwise separable convolution, an activation function, and a 2D selective scan (SS2D) module for further extracting features from the input. The other branch includes a linear layer and an activation function. The outputs of the two branches are merged by element-wise multiplication, and then a linear layer is used to mix the features, which are added to the residual connection to form the output of the VSS block module.

[0079] In the embodiments of the present invention, SimAM can generate three-dimensional attention weights by calculating the local self-similarity of the feature map. Specifically, SimAM defines a neural energy function to adaptively estimate the importance of each neuron, and the neural energy function is:

[0080]

[0081] where x i and t are the input and output of a single channel of x ∈ R C×H×W ω t is the weight, and M = H × W.

[0082] The model assumes that all pixels in a single channel follow the same distribution. Therefore, calculating the mean and variance of the neurons on a single channel can be applied to all channels, and the final minimum energy function can be expressed as:

[0083]

[0084] Among them,

[0085] each neuron can calculate its attention weight as The values of all channels are E, and finally sigmoid(·) is used for activation, which is expressed by the formula as follows:

[0086]

[0087] where i represents the i-th encoding stage. The activated features are used as the downsampled features, and the features passing through the VSS block in the decoder stage are connected through addition to form a skip connection. The connected features are then input to Patch Expanding for upsampling.

[0088] Such as Figure 3 the schematic diagram of SimAM generating three-dimensional attention weights shown. The cube is the feature map H*W*C. The attention weights are generated in the Fusion stage, and Expansion refers to re-weighting the feature values in the feature channels.

[0089] The embodiment of the present invention proposes a novel Unet architecture - VMS-Unet. This architecture is a U-shaped network based on the Mamba state space model and integrates the SimAM attention mechanism, which can efficiently achieve semantic segmentation of the images to be stitched. This segmentation method shows good robustness in weak texture scenarios and can be applied to images with different resolutions.

[0090] Step S102, determine the homography matrix by using the target feature points of multiple target structures in different images to be stitched. Among them, the target feature points are determined based on the geometric shape of the target structure.

[0091] Specifically, the number of target feature points of a target structure can be one or multiple, which is not limited here. In the case where the number of target feature points of a target structure is one, the target feature point can be the geometric centroid, geometric center, etc. of the target structure, which is not limited here either. In the embodiment of the present invention, the subsequent determination of the homography matrix is based on the target feature points.

[0092] In a specific embodiment, the target feature point of the target structure is the geometric centroid of the target structure. After the mask of the target structure is obtained by the image segmentation method, the geometric centroid of the target structure can be determined based on the mask of the target structure. Specifically, the geometric centroid of the l-th target structure can be calculated, for example, by the following method:

[0093]

[0094] Among them, A l is the area of the mask (polygon) of the target structure, and i is the vertex coordinate number of the mask of the target structure segmented out. respectively represent the coordinates of the i-th vertex of the mask of the l-th target structure.

[0095] In the embodiment of the present invention, in order to ensure that the same screws in a pair of images to be stitched (I1, I2) can be accurately aligned, the non-reference image I2 can be distorted by using the homography matrix H. The homography matrix H is 3*3. Homogenize the pixel coordinates of the non-reference image I2 to P=(x, y, 1), with a shape of 3*1, and the new coordinates of each pixel are P' = H·P. Finally, all pixels are mapped to new positions to form the distorted image I'2. I'2 does not correspond to the original pixels at all positions, and the non-corresponding pixels are generated by bilinear interpolation.

[0096] In some optional specific embodiments, step S102, that is, determining the homography matrix by using the target feature points of multiple target structures in different images to be stitched, includes:

[0097] Step S1021, obtaining an initial homography matrix.

[0098] Step S1022, constructing a registration energy function. The registration energy function is the sum of multiple first distances. A first distance is the distance between a first target feature point in the first image to be stitched I1 (as the reference image) and a second target feature point in the distorted (i.e., homography-transformed) second image to be stitched I2'; as Figure 4 shown, the first image to be stitched I1 and the second image to be stitched I2 are two images to be stitched; the first target feature point and the second target feature point are target feature points within the overlapping part 401 (the part within the red frame) of the first image to be stitched and the distorted second image to be stitched, and the second target feature point is the target feature point in the distorted second image to be stitched that is closest to the first target feature point; the second image to be stitched is transformed by using the initial homography matrix to obtain the distorted second image to be stitched.

[0099] In the embodiment of the present invention, a new registration energy function is introduced to optimize the homography matrix H to minimize the difference between the two images to be stitched during stitching. The new registration energy function is expressed by the formula:

[0100]

[0101] Among them, i and j represent the target feature point numbers in the two images to be stitched, H(·) is the homography matrix, respectively represent the coordinates of the i-th target feature point in the image I1 to be stitched, the coordinates of the j-th target feature point in the image I2 to be stitched, and I1∪H(I2) represents the overlapping part between the image I1 to be stitched and the image H(I2) to be stitched after the homography transformation (i.e., I′2). indicates that the i-th target feature point in the image I1 to be stitched and the j-th target feature point in the image I2 to be stitched belong to the overlapping part. represents the distance between the i-th target feature point in the image I1 to be stitched and the j-th target feature point in the warped image I′2 to be stitched.

[0102] In the embodiments of the present invention, it is necessary to calculate the distance between each target feature point in the overlapping part of the image I1 to be stitched and each target feature point in the overlapping part of the image H(I2) to be stitched after the homography transformation, and then for each target feature point in the overlapping part of the image I1 to be stitched, select a minimum distance (the distance from the target feature point in the overlapping part of the image H(I2) to be stitched after the homography transformation) for summation, that is, to obtain the value of the registration energy function, so as to achieve the highest degree of alignment.

[0103] Step S1023, optimize the initial homography matrix by minimizing the value of the registration energy function to obtain the optimal homography matrix.

[0104] Specifically, in the embodiments of the present invention, the homography matrix is optimized by minimizing the value of the registration energy function, so as to use the optimal homography matrix to achieve the registration between the images to be stitched.

[0105] In some optional specific embodiments, the number of images to be stitched may be more than two and the order between the images is uncertain, that is, it is necessary to stitch multiple unordered images (three or more). At this time, step S102, that is, using the target feature points of multiple target structures in different images to be stitched to determine the homography matrix, includes:

[0106] Step S102a, if there are three or more images to be stitched, respectively obtain the minimum registration energy function values between every two images to be stitched. For the process of obtaining the minimum registration energy function values between every two images to be stitched, please refer to the above.

[0107] Step S102b, for one image to be stitched, use the other image to be stitched with the minimum registration energy function value as the adjacent image to be stitched.

[0108] Step S102c, determine the homography matrix corresponding to the minimum registration energy function value between two adjacent images to be stitched as the homography matrix between the two adjacent images to be stitched.

[0109] When the stitching process is extended to multiple unordered images, the primary task is to determine the stitching order between the images, that is, to solve the pairing problem of multiple images. To achieve this goal, in this embodiment, it is assumed that when all the images are arranged in the correct order, the energy function L of pairwise registration between them Registration will reach the minimum value. Based on this assumption, the cumulative sum of the registration energy function will also be the smallest. Therefore, the optimization goal of this embodiment is to minimize this cumulative sum. In this process, i and j respectively represent the serial numbers of the images. During the optimization process, each image can be regarded as a node in a graph, and the value of the registration energy function between image pairs is used as the distance between nodes. After selecting an image (a node), another image with the smallest registration energy function value (i.e., the closest distance) needs to be selected as the next node.

[0110] In this embodiment, during the process of determining the order of multiple images to be stitched, instead of checking all possible image combinations, the concept of the minimum path algorithm is adopted. Each image is regarded as a node in a graph, and the value of the registration energy function between image pairs is used as the distance between nodes. In this way, the problem of optimizing the total energy function is transformed into the problem of finding the shortest path. Therefore, a strategy based on the greedy algorithm can be adopted: after selecting an image (a node), another image with the smallest registration energy function value (i.e., the closest distance) will be selected as the next node. After determining the image order, the homography matrix between adjacent images is calculated using the registration energy function.

[0111] In some optional specific embodiments, when stitching multiple images, only one of the images to be stitched can be used as the reference image (which can also be called the reference image), and the other images to be stitched are used as non-reference images and need to be subjected to homography transformation. At this time, step S102, that is, using the target feature points of multiple target structures in different images to be stitched to determine the homography matrix, includes:

[0112] If there are three or more images to be stitched, the homography matrix between each adjacent pair of images to be stitched is obtained respectively;

[0113] Taking one of the images to be stitched as the reference image, the homography matrix of the third image to be stitched relative to the reference image is determined based on the cumulative multiplication of the homography matrices between adjacent pairs of images to be stitched; the third image to be stitched is the image to be stitched other than the reference image and the adjacent image to be stitched of the reference image.

[0114] For example, for multiple truncated spinal X-ray images, the uppermost spinal X-ray image can be used as the reference image. The transformation of the third image to be stitched I k can be calculated based on the cumulative multiplication of homographies:

[0115]

[0116] Among them, i and j are the serial numbers of two adjacent images after the order is determined, that is, j = i + 1.

[0117] Step S103: Perform perspective transformation on one of the images to be stitched using a homography matrix to obtain a distorted image.

[0118] Among them, the perspective transformation can specifically be a homography transformation, that is, a transformation using a homography matrix. The image to be stitched that needs to be transformed is a non-reference image.

[0119] Step S104: Stitch the distorted image with another image to be stitched that serves as a reference image.

[0120] In some optional specific implementation manners, step S104, that is, stitching the distorted image with another image to be stitched that serves as a reference image, includes:

[0121] Step S1041: Construct a hybrid energy function, which includes the pixel difference, gradient difference, and depth feature difference between the distorted image and the reference image in the gray space.

[0122] Specifically, based on the pixel difference between the distorted image and the reference image in the gray space, the color difference energy E ρ (u, v) can be obtained:

[0123] E ρ =(u, v)=(I1(u, v)-I2(u, v)) 2

[0124] Among them, u and v are pixel coordinates, I1(u, v) is the pixel value of the reference image I1 at (u, v), and I2(u, v) is the pixel value of the distorted image I2 at (u, v). E ρ is a feature map with the same size as images I1 and I2.

[0125] Based on the gradient difference between the distorted image and the reference image in the gray space, the geometric structure energy E δ (u, v) can be obtained:

[0126] E δ (u, v)=(Δ1(u, v)-Δ2(u, v)) 2

[0127] Among them, Δ1(u, v) is the sum of the squares of the gradients of image I1 in the x and y directions at (u, v) (that is, the sum of the squares of the differences between the pixels in the x and y directions), and Δ2(u, v) is the sum of the squares of the gradients of image I2 in the x and y directions at (u, v).

[0128] Based on the depth feature differences between the distorted image and the reference image in the gray space, the feature energy can be obtained

[0129]

[0130] Among them, Φ1(u, v) is the feature of image I1 at (u, v), and Φ2(u, v) is the feature of image I2 at (u, v). For example, the 24th layer of the pre-trained Resnet50 can be used as the representation of the image semantic content (i.e., the feature).

[0131] The mixed energy function is as follows:

[0132]

[0133] Among them, λ is the weight factor of each energy item.

[0134] Step S1042, perform energy initialization on the starting boundary of the seam. The starting boundary can be the left boundary.

[0135] Step S1043, expand from different starting points (for example, each pixel point of the starting boundary) to the other end of the image. Each time of expansion, select the pixel with the minimum value of the mixed energy function as the next path node until the boundary of the overlapping part between the distorted image and the reference image, and obtain the optimal path.

[0136] Step S1044, perform backtracking along the optimal path to determine the best seam.

[0137] Step S1045, splice the distorted image and the reference image according to the best seam.

[0138] In the embodiment of the present invention, a seam path optimization algorithm based on a mixed energy function is designed for the image fusion process.

[0139] In another alternative specific implementation manner, step S104, that is, splicing the distorted image with another image to be spliced as the reference image, includes:

[0140] Step S104a, determine the transition regions on both sides of the seam between the distorted image and the reference image. For the determination process of the seam, please refer to the above text.

[0141] Step S104b, perform weighted summation on the pixel values belonging to the transition regions on the distorted image and the reference image as the pixel values of the spliced image.

[0142] Traditional stitching-based fusion directly replaces the original images on both sides of the seam, which may result in obvious boundary lines. To eliminate such boundary lines, in the embodiments of the present invention, a fusion stitching region is designed on both sides of the seam, where the image will smoothly transition from I1 to I2 in this region.

[0143] Taking the vertical stitching of images as an example, the formula for the stitching process is expressed as:

[0144] I seam [:, v] = σ(kv) ⊙ I1[:, v] + (1 - σ(kv)) ⊙ I2[:, v], v ∈ [0, ξ]

[0145] where v represents the transition region on both sides of the seam. ξ is the total length of the transition region, σ is the sigmoid function, k is the amplification factor, which can be 10 for example, and I seam is the final stitched image, and I1 and I2 are the reference image and the distorted image.

[0146] In the embodiments of the present invention, in the transition region of stitching, the region closer to the seam has an average weighting of the stitched image, and the region far from the seam is dominated by the respective images. The stitching process of the truncated spinal X-ray image is as Figure 5 shown. Using the image stitching method provided by the embodiments of the present invention, the quality of the stitched image is high, achieving artifact-free and seamless stitching.

[0147] The image stitching method provided in this embodiment first identifies multiple target structures in the images to be stitched, then determines the homography matrix based on the target feature points of the identified target structures, and transforms the images to be stitched based on the determined homography matrix. Finally, the transformed distorted image is stitched with the reference image. In this process, since the target feature points are determined based on the geometric shape of the target structure, the determination accuracy and matching accuracy can be guaranteed, thus ensuring the calculation accuracy of the homography matrix and further guaranteeing the stitching quality of the images. In addition, compared with the method of calculating the homography matrix using manually defined feature detection methods, the manually designed features show lower robustness due to their complex feature design, while the method for determining the target feature points in the embodiments of the present invention is simple, making the image stitching method highly robust and highly versatile.

[0148] The image stitching method provided by the embodiments of the present invention can be adapted to the stitching of truncated spinal X-ray images to achieve end-to-end full-length image stitching. Therefore, this image stitching method can be applied to pedicle screw implantation surgery, and the target structure can be the pedicle screw. The advantage of this method lies in the high quality of the stitched image, achieving artifact-free and seamless stitching, and being able to quickly process the input disordered spinal truncated image slices. Its stability and speed make it very suitable for the professional clinical surgical environment and support intraoperative real-time imaging.

[0149] An embodiment of the present invention also proposes a new Unet architecture - VMS-Unet. This architecture is a U-shaped network based on the Mamba state space model and integrates the SimAM attention mechanism, which can efficiently achieve semantic segmentation of spinal X-ray images. This segmentation method exhibits good robustness in weak texture scenarios and can be applied to images of different resolutions.

[0150] Compared with the current stitching method based on traditional detection operators, the algorithm proposed in the embodiment of the present invention has significant advantages in homography estimation and image stitching in real scenarios and is favored by professional clinical surgeons.

[0151] As Figure 6 shown, an overall process for stitching spinal truncated X-ray images using the image stitching method provided in the embodiment of the present invention is as follows:

[0152] I. Obtain multiple unordered spinal truncated X-ray images.

[0153] II. Input the multiple unordered spinal truncated X-ray images obtained in the previous step into the VM-Unet network including a parameter-free attention module for image segmentation, identify the pedicle screws in each spinal truncated X-ray image, and obtain the geometric centroids of the identified pedicle screws.

[0154] III. Use the principle that the minimum registration energy function value between adjacent two spinal truncated X-ray images is the smallest to determine the correct order between the spinal truncated X-ray images. Specifically, image sorting is performed through pedicle screw registration.

[0155] IV. After determining the correct order between the spinal truncated X-ray images, use one of the spinal truncated X-ray images as the reference image to determine the homography matrix of the other spinal truncated X-ray images relative to the reference image.

[0156] V. Perform a homography transformation on the corresponding spinal truncated X-ray images according to the determined homography matrix, then determine the optimal seam based on the hybrid energy function, and finally stitch each spinal truncated X-ray image in sequence.

[0157] For the detailed steps of the above process, please refer to the above text and will not be elaborated here.

[0158] In this embodiment, an image stitching device is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" may be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0159] This embodiment provides an image stitching device, as Figure 7 shown, including:

[0160] A target structure recognition module 701, configured to recognize multiple target structures in the images to be stitched;

[0161] A homography matrix determination module 702, configured to determine a homography matrix by using the target feature points of multiple target structures in different images to be stitched;

[0162] A transformation module 703, configured to perform a perspective transformation on one of the images to be stitched by using the homography matrix to obtain a distorted image;

[0163] A stitching module 704, configured to stitch the distorted image with another image to be stitched as a reference image.

[0164] In some alternative implementation manners, the target structure is a pedicle screw.

[0165] In some alternative implementation manners, recognizing multiple target structures in the images to be stitched includes:

[0166] Using a target neural network model to respectively recognize multiple target structures in the images to be stitched;

[0167] Wherein, the target neural network model at least includes: a Patch embedding layer, an encoder, a decoder, and a parameter-free attention module;

[0168] The Patch embedding layer is used to divide the image to be stitched into multiple image regions;

[0169] The encoder includes a first VSS block module and a patch merge layer. The multiple image regions divided by the Patch embedding layer are input into the first VSS block module, and the first VSS block module is used to respectively extract features from the image regions, and the extracted feature information is input into the patch merge layer for downsampling;

[0170] The decoder includes a Patch expanding layer and a second VSS block module; the Patch expanding layer is used for upsampling, and the second VSS block module is used for feature extraction of the output of the Patch expanding layer;

[0171] The feature information downsampled by the patch merge layer in the encoder is passed to the decoder through a parameter-free attention module.

[0172] In some alternative embodiments, the homography matrix determination module 702 includes:

[0173] An initial homography matrix acquisition unit for acquiring an initial homography matrix;

[0174] A registration energy function construction unit for constructing a registration energy function, which is the sum of multiple first distances. A first distance is the distance between a first target feature point in a first image to be stitched and a second target feature point in the warped second image to be stitched; the first image to be stitched and the second image to be stitched are two images to be stitched; the first target feature point and the second target feature point are target feature points within the overlapping part of the first image to be stitched and the warped second image to be stitched, and the second target feature point is the target feature point in the warped second image that is closest to the first target feature point; the second image to be stitched is warped using the initial homography matrix to obtain the warped second image to be stitched;

[0175] A homography matrix optimization unit for optimizing the initial homography matrix by minimizing the value of the registration energy function to obtain an optimal homography matrix.

[0176] In some alternative embodiments, the homography matrix determination module 702 includes:

[0177] A minimum registration energy function value acquisition unit for, if there are three or more images to be stitched, respectively acquiring the minimum registration energy function values between every two images to be stitched;

[0178] An adjacent image determination unit for, for an image to be stitched, taking the other image to be stitched with the minimum registration energy function value as the adjacent image to be stitched;

[0179] An adjacent image homography matrix determination unit for determining the homography matrix corresponding to the minimum registration energy function value between two adjacent images to be stitched as the homography matrix between the two adjacent images to be stitched.

[0180] In some alternative embodiments, the homography matrix determination module 702 includes:

[0181] An adjacent image homography matrix acquisition unit, configured to, if there are three or more images to be stitched, acquire the homography matrix between each pair of adjacent images to be stitched respectively;

[0182] A multiplication unit, configured to, in the case of taking one of the images to be stitched as a reference image, determine the homography matrix of the third image to be stitched relative to the reference image in a way of multiplying the homography matrices between two adjacent images to be stitched; the third image to be stitched is an image to be stitched other than the reference image and the adjacent image to be stitched of the reference image.

[0183] In some alternative embodiments, the stitching module 704 includes:

[0184] A hybrid energy function construction unit, configured to construct a hybrid energy function, where the hybrid energy function includes pixel differences, gradient differences, and depth feature differences between the warped image and the reference image in the gray space;

[0185] An initialization unit, configured to perform energy initialization on the starting boundary of the seam;

[0186] A path search unit, configured to expand from different starting points to the other end of the image. Each time of expansion, select the pixel with the minimum hybrid energy function value as the next path node until the boundary of the overlapping part between the warped image and the reference image, so as to obtain an optimal path;

[0187] An optimal seam determination unit, configured to perform backtracking along the optimal path to determine the optimal seam;

[0188] A stitching unit, configured to stitch the warped image and the reference image according to the optimal seam.

[0189] In some alternative embodiments, the stitching module 704 includes:

[0190] A transition region determination unit, configured to determine the transition regions on both sides of the seam between the warped image and the reference image;

[0191] A weighting unit, configured to perform weighted summation on the pixel values belonging to the transition regions on the warped image and the reference image as the pixel values of the stitched image.

[0192] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0193] The image stitching device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0194] An embodiment of the present invention further provides a computer device having the above-mentioned Figure 7 image stitching device shown.

[0195] Please refer to Figure 8 , Figure 8 which is a schematic structural diagram of a computer device provided by an optional embodiment of the present invention. As Figure 8 shown, the computer device includes: one or more processors 10, a memory 20, and an interface for connecting each component, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 8 In

[0196] FIG., a single processor 10 is taken as an example.

[0197] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.

[0198] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.

[0199] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid state drive; the memory 20 may further include a combination of the above types of memory.

[0200] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 8 Taking connection via a bus as an example.

[0201] The input device 30 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the computer device, such as touch screens, keypads, mice, trackpads, touchpads, pointing sticks, one or more mouse buttons, trackballs, joysticks, etc. The output device 40 may include display devices, auxiliary lighting devices (e.g., LEDs), and tactile feedback devices (e.g., vibration motors), etc. The above display devices include, but are not limited to, liquid crystal displays, light emitting diodes, displays, and plasma displays. In some alternative embodiments, the display device may be a touch screen.

[0202] The computer device further includes a communication interface for the computer device to communicate with other devices or communication networks.

[0203] The embodiments of the present invention also provide a computer-readable storage medium. The methods according to the embodiments of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading via a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the methods described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the methods shown in the above embodiments are implemented.

[0204] A part of the present invention can be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operations of the computer. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.

[0205] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. An image stitching method, characterized in that, The method includes: Identifying a plurality of target structures in the images to be stitched; Determining a homography matrix by using target feature points of the plurality of target structures in different images to be stitched; the target feature points are determined based on the geometric shape of the target structures; Performing a perspective transformation on one of the images to be stitched by using the homography matrix to obtain a distorted image; Stitching the distorted image with another image to be stitched that serves as a reference image.

2. The method according to claim 1, characterized in that The target structure is a pedicle screw.

3. The method according to claim 1, characterized in that The identifying a plurality of target structures in the images to be stitched includes: Using a target neural network model to respectively identify the plurality of target structures in the images to be stitched; Wherein, the target neural network model at least includes: a Patch embedding layer, an encoder, a decoder, and a parameter-free attention module; The Patch embedding layer is used to divide the image to be stitched into a plurality of image regions; The encoder includes a first VSS block module and a patch merge layer. The plurality of image regions divided by the Patch embedding layer are input into the first VSS block module, and the first VSS block module is used to respectively extract features from the image regions, and the extracted feature information is input into the patch merge layer for downsampling; The decoder includes a Patch expanding layer and a second VSS block module; the Patch expanding layer is used for upsampling, and the second VSS block module is used to extract features from the output of the Patch expanding layer; The feature information obtained by downsampling through the patch merge layer in the encoder is transmitted to the decoder through the parameter-free attention module.

4. The method according to claim 1, characterized in that The determining a homography matrix by using target feature points of the plurality of target structures in different images to be stitched includes: Obtaining an initial homography matrix; Constructing a registration energy function, which is the sum of a plurality of first distances. One first distance is the distance between a first target feature point in a first image to be stitched and a second target feature point in the distorted second image to be stitched; the first image to be stitched and the second image to be stitched are two images to be stitched; the first target feature point and the second target feature point are the target feature points within the overlapping part of the first image to be stitched and the distorted second image to be stitched, and the second target feature point is the target feature point in the distorted second image to be stitched that is closest to the first target feature point; the second image to be stitched is transformed by using the initial homography matrix to obtain the distorted second image to be stitched; Optimizing the initial homography matrix by minimizing the value of the registration energy function to obtain the optimal homography matrix.

5. The method according to claim 1 or 4, characterized in that, The determining a homography matrix by using target feature points of the plurality of target structures in different images to be stitched includes: If the images to be stitched are three or more, respectively obtain the minimum registration energy function values between every two of the images to be stitched; For one of the images to be stitched, use the other image to be stitched with the minimum minimum registration energy function value as the adjacent image to be stitched; Determine the homography matrix corresponding to the minimum registration energy function value between two adjacent images to be stitched as the homography matrix between the two adjacent images to be stitched.

6. The method according to claim 1 or 4, characterized in that, The determining the homography matrix by using the target feature points of multiple target structures in different images to be stitched includes: If the images to be stitched are three or more, respectively obtain the homography matrices between every two adjacent images to be stitched; Taking one of the images to be stitched as the reference image, determine the homography matrix of the third image to be stitched relative to the reference image based on the cumulative multiplication of the homography matrices between two adjacent images to be stitched; the third image to be stitched is the image to be stitched other than the reference image and the adjacent image to be stitched of the reference image.

7. The method according to claim 1, wherein The stitching the warped image with another image to be stitched as the reference image includes: Construct a hybrid energy function, where the hybrid energy function includes the pixel difference, gradient difference, and depth feature difference between the warped image and the reference image in the gray space; Perform energy initialization on the starting boundary of the seam; Expand from different starting points to the other end of the image. Each time of expansion, select the pixel with the minimum hybrid energy function value as the next path node until the boundary of the overlapping part between the warped image and the reference image to obtain the optimal path; Perform backtracking along the optimal path to determine the best seam; Stitch the warped image and the reference image according to the best seam.

8. The method according to claim 1 or 7, characterized in that, The stitching the warped image with another image to be stitched as the reference image includes: Determine the transition regions on both sides of the seam between the warped image and the reference image; Perform weighted summation on the pixel values belonging to the transition regions on the warped image and the reference image as the pixel values of the stitched image.

9. An image stitching device, characterized in that, The device includes: A target structure recognition module for recognizing multiple target structures in the images to be stitched; A homography matrix determination module for determining the homography matrix by using the target feature points of multiple target structures in different images to be stitched; A transformation module for performing perspective transformation on one of the images to be stitched by using the homography matrix to obtain a warped image; A stitching module for stitching the warped image with another image to be stitched as the reference image.

10. A computer device, characterized in that, It includes: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the image stitching method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Sub-pixel registration and splicing technology based on interpolation and iterative optimization algorithm

    CN109961393A

  • Self-adaptive image splicing fusion method and device, electronic equipment and storage medium

    CN112037130A

  • Panoramic image splicing method under complex background

    CN113689331A

  • Image splicing method and device

    CN116977166A

  • Image splicing method and device, electronic equipment and storage medium

    CN118195892A

Cited By

  • Image splicing method and device and storage medium

    CN120725864A