Multi-modal image data registration method, device, equipment, medium and product

By performing linear detection and segmentation processing on multimodal images, the intersection points are determined for registration, which solves the problem of poor stability in traditional methods and realizes automatic and accurate multimodal image registration.

CN120599007APending Publication Date: 2025-09-05XIAN JIAOTONG LIVERPOOL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510674956.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Traditional multimodal pathological image registration methods rely on manual correction or based on image feature matching, with poor stability and difficulty in accurately extracting matching points.

Method used

By performing linear detection of multimodal images, intersection points are determined, and segmentation processing is performed using the target segmentation model, and registration is performed based on intersection points and segmentation results.

Benefits of technology

Automatic, stable and accurate multimodal image registration is achieved, avoiding the problem of poor stability in traditional methods and improving the accuracy of registration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599007A_ABST
    Figure CN120599007A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal image registration method, device and equipment, a medium and a product, and relates to the technical field of image processing. Performing straight line detection on the multi-modal image to obtain at least one straight line in the multi-modal image; determining at least one target intersection point in the multi-modal image according to the at least one straight line and the horizontal line; performing segmentation processing on the multi-modal image based on a target segmentation model to obtain a segmentation result; and carrying out registration on the multi-modal image based on the segmentation result and the at least one target intersection point. By adopting the technical scheme, the multi-modal image is subjected to straight line detection to obtain the corresponding intersection point, and the multi-modal image is registered according to the intersection point and the segmented image, so that the problem of poor image registration stability caused by the fact that matching points are extracted based on pixel features traditionally is avoided, and registration can be automatically, stably and accurately performed on a common modal image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to a multimodal image data registration method, device, equipment, medium and product. Background Art

[0002] There are multiple modal data in pathological images, and each modal data provides pathological data from different angles. However, due to the different imaging methods of pathological data of different modalities, the same piece of tissue has different positions in different modal images, so the pathological data of different modalities need to be positionally aligned.

[0003] Traditional position registration methods mainly rely on manual correction or image feature matching methods; the manual correction process is cumbersome; image feature matching requires extracting feature values ​​from image pixel values, extracting matching points based on feature value similarity, and calculating the correspondence between different modal images based on matching points for registration. Its stability is poor, and the pixel values ​​of pathological images under different modalities have different physical meanings, making it difficult to accurately extract the correct matching points. Summary of the Invention

[0004] The present invention provides a multimodal image data registration method, apparatus, device, medium and product to solve the registration of pathological images under different modalities and avoid the problem of poor stability caused by image registration based on pixel eigenvalues.

[0005] According to one aspect of the present invention, a multimodal image registration method is provided, comprising:

[0006] Performing straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image;

[0007] Determining at least one target intersection point in the multimodal image based on the at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image;

[0008] Segmenting the multimodal image based on the target segmentation model to obtain a segmentation result; the segmentation result includes a multimodal segmented image and a pixel category corresponding to each pixel in the multimodal segmented image;

[0009] The multimodal images are registered based on the segmentation result and the at least one target intersection point.

[0010] According to another aspect of the present invention, a multimodal image registration apparatus is provided, comprising:

[0011] a detection module, configured to perform straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image comprising a first target image and a second target image;

[0012] an intersection determination module, configured to determine at least one target intersection point in the multimodal image based on the at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image;

[0013] a segmentation module, configured to perform segmentation processing on the multimodal image based on a target segmentation model to obtain a segmentation result; the segmentation result includes a multimodal segmented image and a pixel category corresponding to each pixel in the multimodal segmented image;

[0014] A registration module is configured to register the multimodal image based on the segmentation result and the at least one target intersection point.

[0015] According to another aspect of the present invention, an electronic device is provided, comprising:

[0016] at least one processor; and

[0017] a memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform the multimodal image registration method described in any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal image registration method described in any embodiment of the present invention when executed.

[0020] According to another aspect of the present invention, a computer program product is provided. The computer program product comprises a computer program. When the computer program is executed by a processor, the computer program implements the multimodal image registration method according to any embodiment of the present invention.

[0021] The technical solution of an embodiment of the present invention performs line detection on a multimodal image to obtain at least one line in the multimodal image; determines at least one target intersection point in the multimodal image based on the at least one line and a horizontal line; segments the multimodal image based on a target segmentation model to obtain a segmentation result; and registers the multimodal image based on the segmentation result and the at least one target intersection point. This technical solution, which performs line detection on a multimodal image to obtain corresponding intersection points and registers the multimodal image based on the intersection points and the segmented image, avoids the problem of poor image registration stability caused by traditional pixel feature-based matching point extraction, and can automatically, stably, and accurately register common modality images.

[0022] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0024] Figure 1 is a flowchart of a multimodal image registration method provided according to an embodiment of the present invention;

[0025] Figure 2 is a straight line detection image corresponding to a fluorescent image provided by an embodiment of the present invention;

[0026] Figure 3 is a line detection image corresponding to a spatial transcription image provided by an embodiment of the present invention;

[0027] Figure 4 A fluorescence image intersection recognition diagram provided according to an embodiment of the present invention;

[0028] Figure 5 A spatial transcription image intersection recognition map provided according to an embodiment of the present invention;

[0029] Figure 6 is a flow chart of a method for determining a target intersection point provided according to an embodiment of the present invention;

[0030] Figure 7 is a first segmented image provided according to an embodiment of the present invention;

[0031] Figure 8is a second segmented image provided according to an embodiment of the present invention;

[0032] Figure 9 is a flowchart of a multimodal image registration method provided according to an embodiment of the present invention;

[0033] Figure 10 is an intersection conversion diagram provided according to an embodiment of the present invention;

[0034] Figure 11 is a flow chart of a secondary registration method provided according to an embodiment of the present invention;

[0035] Figure 12 is a structural diagram of a multimodal image registration device provided according to an embodiment of the present invention;

[0036] Figure 13 4 is a schematic structural diagram of an electronic device for implementing the multimodal image registration method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0037] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0038] It should be noted that the terms "first," "second," and the like in the description and claims of the present invention and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the present invention described herein can be practiced in an order other than that illustrated or described herein.

[0039] In addition, it should be noted that the collection, storage, use, processing, transmission, provision and disclosure of the data to be processed involved in the technical solution of the present invention are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0040] Figure 1A flowchart of a multimodal image registration method is provided for an embodiment of the present invention. The embodiment of the present invention is applicable to the case of position registration of common modality images, and is particularly applicable to the case of position registration of pathological image data under different modalities. The method can be executed by a multimodal image registration device provided by an embodiment of the present invention. The multimodal image registration device can be implemented in the form of hardware and / or software, and the multimodal image registration device can be configured in a server. Figure 1 As shown, the method includes:

[0041] S110 , performing straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image.

[0042] Among them, multimodal images are multiple modal data corresponding to pathological images; medical images are formed by integrating different imaging technologies or data sources; the pathological data corresponding to each modal image provides pathological information from different angles, for example: HE images provide clear cell morphology images; immunofluorescence provides the protein expression level of cells; spatial transcription provides the mRAN expression level of cells; the straight lines in the multimodal image are the horizontal lines corresponding to the horizontal coordinates in the multimodal image, and the vertical lines corresponding to the vertical coordinates in the multimodal image; the first target image and the second target image are the target pathological images to be positionally aligned, which can be fluorescence images, spatial transcription images or HE images.

[0043] Specifically, straight line detection is performed on the first target image and the second target image to obtain at least one straight line in the first target image and the second target image.

[0044] In an optional embodiment of the present invention, Hough line detection can be used to perform line detection on multimodal images, or a deep learning method can be used to perform line detection on multimodal images; a preferred embodiment of the present invention is to use DeepLSD (Deep Line Segment Detector, deep learning line detection algorithm) algorithm to perform line detection on multimodal images.

[0045] In a preferred embodiment of the present invention, the first target image is a fluorescence image, and the second target image is a spatial transcription image; the fluorescence image and the spatial transcription image are subjected to line detection by the DeepLSD algorithm to obtain the corresponding heat map and all the lines in the image, as well as the corresponding line coordinates, such as Figure 2 The line detection image corresponding to the fluorescent image shown is Figure 3 The line detection image corresponding to the spatial transcription image shown.

[0046] S120. Determine at least one target intersection point in the multimodal image based on at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image.

[0047] The target intersection points are the coordinates of the intersection points corresponding to the intersecting lines in the multimodal image.

[0048] Specifically, the straight lines are classified according to at least one straight line and a straight line coordinate point, and the horizontal line in the image coordinate system corresponding to the multimodal image. The straight lines in the multimodal image can be divided into two groups, one group close to the horizontal line and the other group close to the vertical line. At least one target intersection coordinate is obtained based on the intersection of each straight line in one group of the multimodal image and each straight line in the other group.

[0049] In a preferred embodiment of the present invention, Figure 2 The line detection image corresponding to the fluorescent image shown is Figure 3 The straight line detection image corresponding to the spatial transcription image shown in FIG is combined with the horizontal lines corresponding to each image to perform intersection recognition, and the following is obtained: Figure 4 The intersection identification diagram of the fluorescence image shown and Figure 5 The intersection identification diagram of the spatial transcription image is shown.

[0050] Optional, such as Figure 6 A target intersection determination method shown includes:

[0051] S121. Determine an angle between the straight line and the horizontal line based on the straight line and the horizontal line.

[0052] Specifically, the angle between the straight line and the horizontal line is calculated based on the straight line in the multimodal image and the horizontal line in the image coordinate system of the corresponding image. In an optional embodiment of the present invention, the specific calculation method is as follows:

[0053]

[0054] Where θ is the angle, is the vector of the line, is the vector of the horizontal line.

[0055] S122. Determine the line type of the line based on the included angle and a preset included angle threshold; the line type includes a horizontal line or a vertical line.

[0056] Specifically, the line type is determined based on a comparison between the absolute value of the angle and a preset angle threshold. An angle with an absolute value less than the preset angle threshold is considered a horizontal line, while an angle with an absolute value greater than the preset angle threshold is considered a vertical line. In a preferred embodiment of the present invention, the preset angle threshold is 45°.

[0057] S123 : Determine an intersection point between different straight lines based on the straight line type and the straight line, and use the intersection point as a target intersection point in the multimodal image.

[0058] Specifically, a horizontal line of the horizontal line type and a vertical line of the vertical line type are determined, and an intersection point of the horizontal line and the vertical line is calculated as a target intersection point in the multimodal image.

[0059] The specific calculation method is as follows:

[0060] Assume that the straight line expressions corresponding to the horizontal line and the vertical line are:

[0061] a1x+b1y+c1=0

[0062] a²x+b²y+c²=0

[0063] The corresponding straight line intersection point can be obtained by solving the following set of equations:

[0064]

[0065] Among them, (x0, y0) is the coordinate of the target intersection point.

[0066] It can be understood that by performing straight line detection on the multimodal image and obtaining at least one target intersection coordinate in the corresponding image coordinate system in the multimodal image, data support is provided for subsequent position alignment, and the intersection coordinates obtained by the straight line facilitate subsequent position calibration.

[0067] S130 , performing segmentation processing on the multimodal image based on the target segmentation model to obtain a segmentation result; the segmentation result includes the multimodal segmented image and the pixel category corresponding to each pixel in the multimodal segmented image.

[0068] Among them, the target segmentation model is a pre-trained segmentation model, the multimodal segmentation image is a grayscale image after the multimodal image is segmented; the pixel point categories include tissue categories and background categories.

[0069] Specifically, a target segmentation model is used to segment the multimodal image. The target segmentation model determines the category of each pixel in the image. The output 1 is the tissue area, and the output 0 is the background area.

[0070] In a preferred embodiment of the present invention, the target segmentation model is used to segment the first target image and the second target image to obtain Figure 7 The first segmented image shown; the second target image is segmented using the target segmentation model to obtain Figure 8The second segmented image is shown, and the corresponding pixel category is obtained. The white area in the figure is the tissue area, and the pixel category can be marked as 1; the black area in the figure is the background area, and the pixel category is marked as 0.

[0071] In an optional embodiment of the present invention, the segmentation method may adopt an adaptive threshold determination method, or segmentation may be performed based on a deep learning U-Net model; in an embodiment of the present invention, the target segmentation model is preferably a U-Net model based on deep learning, and the U-Net backbone model is first initialized using a large pathology model, and then Lora is used to fine-tune the initialized U-Net on a small amount of data.

[0072] Optional object segmentation model training methods include:

[0073] Initialize and train the initial segmentation model based on the initial sample data to obtain the model to be fine-tuned; the initial sample data is unlabeled sample data;

[0074] The labeled sample data is input into the model to be fine-tuned for fine-tuning to obtain the target segmentation model.

[0075] The initial sample data refers to the unprocessed original image sample data; the labeled sample data refers to the manually labeled image sample data.

[0076] Specifically, the initial segmentation model is a neural network model, which can be initialized using a pre-trained model; the initial segmentation model is initialized and trained using the pre-trained model and the initial sample data to obtain a model to be fine-tuned, and the manually labeled sample data is input into the model to be fine-tuned for model fine-tuning to obtain a target segmentation model.

[0077] It can be understood that by training the model with initialized sample data and fine-tuning the model with manually labeled sample data, the performance of the target segmentation model on specific tasks can be optimized, which can adapt to and complete tasks in specific fields more quickly; and compared with the existing neural network initialization which requires a large amount of labeled data training, the initial segmentation model is trained based on the pre-trained model and fine-tuned with labeled sample data, which improves the applicability of the initial sample data acquisition and improves the training efficiency of the segmentation model.

[0078] S140 : Register the multimodal image based on the segmentation result and at least one target intersection point.

[0079] Specifically, the multimodal images are registered according to the pixel categories in the segmentation results and the target intersection coordinates obtained by line detection.

[0080] The technical solution of an embodiment of the present invention performs line detection on a multimodal image to obtain at least one line in the multimodal image; determines at least one target intersection point in the multimodal image based on the at least one line and a horizontal line; segments the multimodal image based on a target segmentation model to obtain a segmentation result; and aligns the multimodal image based on the segmentation result and the at least one target intersection point. This technical solution, which performs line detection on a multimodal image to obtain corresponding intersection points and aligns the multimodal image based on the intersection points and the segmented image, avoids the problem of poor image registration stability caused by traditional pixel feature-based matching point extraction, and automatically and stably and accurately aligns common modality images.

[0081] Optionally, multimodal images are presented with grid lines.

[0082] The grid lines are grid lines attached to the tissue slices corresponding to the multimodal images.

[0083] Specifically, the glass slide used for tissue sectioning is divided into grid lines, the size and spacing specifications of the grid lines are kept consistent, and the tissue is made into tissue sections using the glass slide, and imaging is performed using a digital device to obtain corresponding multimodal images with grid lines.

[0084] It should be noted that the preferred size of the grid lines on the glass slide in the embodiment of the present invention is 20 mm, and the spacing is 10 um-15 um. The embodiment of the present invention does not impose any specific limitation on this, and the size and spacing of the grid lines can be kept consistent.

[0085] It is understandable that by manufacturing the glass slide chip as a chip with grid lines, a basis is provided for subsequent straight line detection, thereby ensuring the stability of subsequent straight line detection of the image.

[0086] Figure 9 This is a flowchart of a multimodal image registration method provided according to an embodiment of the present invention. Based on the above embodiment, the embodiment of the present invention supplements the method of registering multimodal images based on the segmentation results and at least one target intersection. It should be noted that for the part not described in detail in the embodiment of the present invention, please refer to the relevant description of other embodiments. Figure 9 As shown, the method includes:

[0087] S210 , performing straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image.

[0088] S220. Determine at least one target intersection point in the multimodal image based on at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image.

[0089] S230 , performing segmentation processing on the multimodal image based on the target segmentation model to obtain a segmentation result; the segmentation result includes the multimodal segmented image and the pixel category corresponding to each pixel in the multimodal segmented image.

[0090] S240 : Performing initial registration on the second target image based on the segmentation result and at least one target intersection point to obtain initial matching points.

[0091] The initial registration is a coarse registration, and the initial matching points are the coordinates of the matching points converted from the second target image to the coordinate system of the first target image.

[0092] Specifically, according to the pixel point category in the segmentation result, the target intersection point of the pixel point category in the first target image and the second target image under the tissue area is found, and the target intersection point under the corresponding tissue area in the second target image is converted to the coordinate system of the first target image as the initial matching point.

[0093] Optionally, performing a preliminary registration on the second target image based on the segmentation result and the at least one target intersection point to obtain an initial matching point includes:

[0094] Sample each pixel in the multimodal segmentation image to obtain candidate sampling points;

[0095] Determine the target sampling point from the candidate sampling points according to the pixel point categories corresponding to the candidate sampling points;

[0096] The target sampling points are estimated by a sampling point cloud matching algorithm to obtain a rotation matrix and a translation matrix between the first target image and the second target image;

[0097] The target intersection point of the second target image is transformed by a rotation matrix and a translation matrix to obtain an initial matching point of the target intersection point of the second target image in the coordinate system of the first target image.

[0098] Among them, the sampling point cloud matching algorithm is preferably the ICP algorithm.

[0099] Specifically, the pixels in the segmented image are traversed and the pixels are sampled at equal intervals. For example, a point can be sampled every 20 pixels as a candidate sampling point. Based on the pixel category corresponding to the candidate sampling point, if the pixel is a tissue region category, the candidate sampling point is the target sampling point. Based on the target sampling point, the rotation matrix and translation matrix between the first target image and the second target image are estimated through the sampling point cloud matching algorithm. Based on the estimated rotation matrix and translation matrix, the target intersection coordinates of the second target image are matrix transformed to obtain the initial matching point coordinates corresponding to the target intersection of the second target image in the coordinate system of the first target image.

[0100] Specifically, the conversion formula is:

[0101] P′=RP+T

[0102] Wherein, P′ is the initial matching point, R is the rotation matrix, P is the target intersection point of the second target image, and is the translation matrix.

[0103] For example, Figure 5 The spatial transcription image intersection identification map shown is obtained according to the above conversion formula Figure 4 The corresponding initial matching points in the fluorescence image intersection identification diagram are as shown in Figure 10 The intersection conversion diagram shown, the green point is the target intersection coordinates of the spatial transcription image converted to the initial matching point coordinates in the fluorescence image coordinate system, and the blue point is the target intersection coordinates in the fluorescence image.

[0104] It should be noted that the pixel points are sampled at equal intervals, and the interval can be set based on actual needs. The smaller the distance, the more pixel points are obtained, and the higher the rough matching accuracy; the larger the distance, the fewer points, and the lower the rough matching accuracy; the pixel point interval can be 10-20, and the embodiment of the present invention does not impose specific limitations on this.

[0105] It can be understood that the straight lines in the multimodal image are used as auxiliary lines to obtain the coordinates of the intersection of the multimodal image, and the multimodal image is initially aligned based on the intersection coordinates and the segmented pixel point categories, avoiding the situation in which it is difficult to extract the correct matching points by relying solely on pixel features in traditional methods. The intersection coordinates are determined based on the auxiliary lines, and the initial matching points are automatically and accurately determined in combination with the segmentation results.

[0106] S250 , performing secondary registration on the second target image according to the initial matching point and the target intersection point of the first target image.

[0107] Specifically, based on the initial matching point and the target intersection coordinates of the first target image, the second target image is re-registered.

[0108] Optional, such as Figure 11 A secondary registration method shown includes:

[0109] S251 : Determine an intersection distance according to a target intersection point of the first target image and an initial matching point.

[0110] The intersection distance is the distance between the target intersection coordinates of the first target image and the initial matching points of the target intersection coordinates in the second target image in the image coordinate system of the first target image.

[0111] Specifically, based on the existing point-to-point distance formula, the intersection distance between the target intersection point and the initial matching point of the first target image can be directly calculated; the specific calculation formula is as follows:

[0112] K i,j =||q i -p j ||

[0113] Among them, K i,j is the intersection distance, q i is the target intersection point of the first target image, p j is the initial matching point.

[0114] The target intersection point in the first target image is recorded as Q = {q1,q2,…,q n}, the corresponding initial matching point is P′={p′1,p′2,…,p′ m}, then the intersection distance between each pair can be obtained based on the above calculation formula.

[0115] S252: Determine a target matching point corresponding to the target intersection point of the first target image from the initial matching points based on the intersection point distance and a preset distance threshold.

[0116] Specifically, according to the intersection point distance and a preset distance threshold, a target matching point corresponding to the target intersection point in the first target image is found from the initial matching points.

[0117] For example, for each q i Get the nearest but The closest intersection point in the first target image is q i , then q i The matching point is

[0118] S253 : Determine a conversion relationship between the first target image and the second target image based on the target intersection point and the target matching point of the first target image.

[0119] Specifically, a transformation relationship between the first target image and the second target image is calculated according to the target intersection point and the matching point, including a rotation matrix and a translation matrix.

[0120] The correspondence between the target intersection point and the matching point is {(q1, p′1), (q2, p′2), …}, and its calculation formula is as follows:

[0121] q1=R′p′1+T′

[0122] Among them, q1 is the target intersection point; p′1 is the corresponding target matching point, R′ is the rotation matrix, and T′ is the translation matrix.

[0123] S254: Perform secondary registration on the second target image based on the conversion relationship.

[0124] Based on the rotation matrix and transformation matrix obtained above, the corresponding relationship between the target intersection points in the first target image and the second target image can be obtained, and the corresponding relationship is as follows:

[0125] Q=R′(RP+T)+T′

[0126] It can be understood that by finding the matching point closest to the target intersection in the first target image, the corresponding point in the image coordinate system of the first target image is accurately extracted based on the distance threshold. The secondary registration effectively reduces the cumulative error of the first registration through iterative optimization and improves the accuracy of the overall registration.

[0127] The embodiment of the present invention determines the intersection points in the multimodal image based on the detected straight lines, further performs a primary registration of the multimodal image to determine the matching points based on the intersection point coordinates and the pixel point categories, and performs a secondary registration of the multimodal image based on the transformation relationship determined based on the matching points. Compared with image registration relying on image pixel values, the determination of matching points based on straight lines as auxiliary lines is not easily affected by the image, and can automatically and accurately determine the matching points to complete the registration of the multimodal image.

[0128] Figure 12 A schematic diagram of the structure of a multimodal image registration device is provided for an embodiment of the present invention. This embodiment of the present invention is applicable to the positional registration of images of common modalities, and is particularly applicable to the positional registration of pathological images of different modalities. The multimodal image registration device can be implemented in hardware and / or software and can be configured in a server. The multimodal image registration device 300 includes a detection module 310, an intersection determination module 320, a segmentation module 330, and a registration module 340.

[0129] A detection module 310 is configured to perform straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image;

[0130] An intersection determination module 320 is configured to determine at least one target intersection point in the multimodal image based on at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in the image coordinate system corresponding to the multimodal image;

[0131] The segmentation module 330 is used to perform segmentation processing on the multimodal image based on the target segmentation model to obtain a segmentation result; the segmentation result includes the multimodal segmented image and the pixel category corresponding to each pixel in the multimodal segmented image;

[0132] The registration module 340 is configured to register the multimodal images based on the segmentation result and at least one target intersection point.

[0133] The technical solution of an embodiment of the present invention performs line detection on a multimodal image to obtain at least one line in the multimodal image; determines at least one target intersection point in the multimodal image based on the at least one line and a horizontal line; segments the multimodal image based on a target segmentation model to obtain a segmentation result; and aligns the multimodal image based on the segmentation result and the at least one target intersection point. This technical solution, which performs line detection on a multimodal image to obtain corresponding intersection points and aligns the multimodal image based on the intersection points and the segmented image, avoids the problem of poor image registration stability caused by traditional pixel feature-based matching point extraction, and automatically and stably and accurately aligns common modality images.

[0134] Optionally, the registration module 340 includes a primary registration unit and a secondary registration unit;

[0135] a primary registration unit, configured to perform primary registration on the second target image based on the segmentation result and at least one target intersection point to obtain an initial matching point;

[0136] The secondary registration unit is used to perform secondary registration on the second target image according to the initial matching point and the target intersection point of the first target image.

[0137] Optionally, the initial registration unit is specifically used to sample each pixel point in the multimodal segmented image to obtain candidate sampling points; determine the target sampling point from the candidate sampling points according to the pixel point category corresponding to the candidate sampling point; estimate the target sampling point through the sampling point cloud matching algorithm to obtain the rotation matrix and translation matrix between the first target image and the second target image; perform matrix transformation on the target intersection point of the second target image through the rotation matrix and translation matrix to obtain the initial matching point of the target intersection point of the second target image in the coordinate system of the first target image.

[0138] Optionally, a secondary registration unit is specifically used to determine the intersection distance based on the target intersection point of the first target image and the initial matching point; based on the intersection distance and a preset distance threshold, determine the target matching point corresponding to the target intersection point of the first target image from the initial matching point; determine the conversion relationship between the first target image and the second target image based on the target intersection point and the target matching point of the first target image; and perform secondary registration on the second target image based on the conversion relationship.

[0139] Optionally, the intersection determination module 320 is specifically used to determine the angle between a straight line and a horizontal line based on the straight line and the horizontal line; determine the line type of the straight line based on the angle and a preset angle threshold; the line type includes a horizontal straight line or a vertical straight line; determine the intersection between different straight lines based on the line type and the straight line, and use the intersection as the target intersection in the multimodal image.

[0140] Optional object segmentation model training methods include:

[0141] Initialize and train the initial segmentation model based on the initial sample data to obtain the model to be fine-tuned; the initial sample data is unlabeled sample data;

[0142] The labeled sample data is input into the model to be fine-tuned for fine-tuning to obtain the target segmentation model.

[0143] Optionally, multimodal images are presented with grid lines.

[0144] The multimodal image registration device provided in the embodiment of the present invention can execute the multimodal image registration method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0145] According to an embodiment of the present invention, the present invention further provides an electronic device, a readable storage medium and a computer program product.

[0146] Figure 13 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0147] like Figure 13 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0148] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0149] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the multimodal image registration method.

[0150] In some embodiments, the multimodal image registration method can be implemented as a computer program that is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the multimodal image registration method described above can be performed. Alternatively, in other embodiments, the processor 11 can be configured to perform the multimodal image registration method in any other appropriate manner (e.g., by means of firmware).

[0151] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0152] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0153] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0154] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0155] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0156] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0157] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0158] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A multimodal image registration method, characterized in that: include: Performing straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image; Determining at least one target intersection point in the multimodal image based on the at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image; Segmenting the multimodal image based on the target segmentation model to obtain a segmentation result; the segmentation result includes a multimodal segmented image and a pixel category corresponding to each pixel in the multimodal segmented image; The multimodal images are registered based on the segmentation result and the at least one target intersection point.

2. The method according to claim 1, characterized in that The registering the multimodal image based on the segmentation result and the at least one target intersection point includes: Performing a primary registration on the second target image based on the segmentation result and the at least one target intersection point to obtain an initial matching point; The second target image is re-registered according to the initial matching point and the target intersection point of the first target image.

3. The method according to claim 2, characterized in that The performing initial registration on the second target image based on the segmentation result and the at least one target intersection point to obtain an initial matching point includes: Sampling each pixel in the multimodal segmented image to obtain candidate sampling points; Determining a target sampling point from the candidate sampling points according to the pixel point categories corresponding to the candidate sampling points; Estimating the target sampling points by a sampling point cloud matching algorithm to obtain a rotation matrix and a translation matrix between the first target image and the second target image; Performing matrix transformation on the target intersection point of the second target image using the rotation matrix and the translation matrix to obtain an initial matching point of the target intersection point of the second target image in the coordinate system of the first target image.

4. The method according to claim 2, characterized in that The performing secondary registration on the second target image according to the initial matching point and the target intersection point of the first target image includes: determining an intersection distance based on the target intersection point of the first target image and the initial matching point; Determining, from the initial matching points, a target matching point corresponding to the target intersection point of the first target image based on the intersection distance and a preset distance threshold; determining a conversion relationship between the first target image and the second target image based on a target intersection point of the first target image and the target matching point; A secondary registration is performed on the second target image based on the conversion relationship.

5. The method according to claim 1, characterized in that The determining at least one target intersection point in the multimodal image according to the at least one straight line and the horizontal line includes: determining an angle between the straight line and the horizontal line based on the straight line and the horizontal line; Determining the line type of the line based on the angle and a preset angle threshold; the line type includes a horizontal line or a vertical line; An intersection point between different straight lines is determined based on the straight line type and the straight line, and the intersection point is used as a target intersection point in the multimodal image.

6. The method according to claim 1, characterized in that The training method of the target segmentation model includes: Initializing and training the initial segmentation model based on the initial sample data to obtain a model to be fine-tuned; the initial sample data is unlabeled sample data; The labeled sample data is input into the model to be fine-tuned for fine-tuning to obtain the target segmentation model.

7. The method according to any one of claims 1 to 6, characterized in that The multimodal image is provided with grid lines.

8. A multimodal image registration device, characterized in that: include: a detection module, configured to perform straight line detection on a multimodal image to obtain at least one straight line in the multimodal image; the multimodal image includes a first target image and a second target image; an intersection determination module, configured to determine at least one target intersection point in the multimodal image based on the at least one straight line and a horizontal line; wherein the horizontal line is a horizontal line in an image coordinate system corresponding to the multimodal image; a segmentation module, configured to perform segmentation processing on the multimodal image based on a target segmentation model to obtain a segmentation result; the segmentation result includes a multimodal segmented image and a pixel category corresponding to each pixel in the multimodal segmented image; A registration module is configured to register the multimodal image based on the segmentation result and the at least one target intersection point.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the multimodal image registration method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multimodal image registration method according to any one of claims 1 to 7 when executed.

11. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the multimodal image registration method according to any one of claims 1 to 7.