Image Processing Method, Apparatus, Computer Device, and Storage Medium
By determining and matching the target object attribute information in different images in medical images, high-precision fusion of multi-source images is achieved, and the problem of large errors in the prior art is solved.
Patent Information
- Application Number
- CN202111129299.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-09-26
AI Technical Summary
The prior art has a problem of large errors in multi-source fusion of medical images, and it is difficult to fuse images from different sources with high precision.
By determining the attribute information of the target object in different images, matching and fusion of object instances is performed, and precise fusion of image object instances is performed using the degree of matching.
High-precision fusion of image object instances is realized, and the accuracy and accuracy of image fusion are improved.
Smart Images

Figure CN114240809B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular, to an image processing method, apparatus, computer device, and storage medium. Background Art
[0002] Target detection-related algorithms can detect regions of interest from input images and have wide applications in fields such as medical image analysis, such as lesion detection. However, the process of determining the lesion course in existing medical images often involves multiple image sources, such as multiple temporal images for follow-up needs, multiple images at different contrast stages for contrast imaging, multiple images of different modalities for multi-modal fusion diagnosis, etc., which involves fusing images; the current fusion method has a problem of large errors. Summary of the Invention
[0003] The embodiments of the present disclosure at least provide an image processing method, apparatus, computer device, and storage medium.
[0004] In a first aspect, an embodiment of the present disclosure provides an image processing method, including:
[0005] Determining first attribute information of a first object instance of a target object in a first image and determining second attribute information of a second object instance of the target object in a second image;
[0006] Based on the first attribute information and the second attribute information, matching the first object instance and the second object instance to obtain a matching degree between the first object instance and the second object instance;
[0007] Based on the matching degree, fusing the first object instance and the second object instance.
[0008] In this way, the first object instance and the second object instance can be fused with higher precision.
[0009] In an optional implementation manner, it further includes: obtaining a first original image and a second original image;
[0010] Determining transformation relationship information between the first original image and the second original image;
[0011] Performing transformation processing on the first original image based on the transformation relationship information to obtain a first image, and using the second original image as the second image;
[0012] Or,
[0013] Performing transformation processing on the second original image based on the transformation relationship information to obtain a second image, and using the first original image as the first image.
[0014] In this way, the first image and the second image can be registered to facilitate subsequent processing.
[0015] In an alternative embodiment, the first attribute information or the second attribute information includes at least one of the following:
[0016] The position information of the object instance, the size information, the probability that the object instance belongs to the target object, the feature data, the gray-scale data of the image region corresponding to the object instance, the radiomics information;
[0017] The object instances include: the first object instance and the second object instance.
[0018] In this way, based on multiple attribute information, the matching accuracy between object instances can be improved.
[0019] In an alternative embodiment, the matching of the first object instance and the second object instance based on the first attribute information and the second attribute information includes:
[0020] Based on the first attribute information and the second attribute information, determine at least one of the following matching information between the first object instance and the second object instance: similarity, matching priority, distance, equivalent radius;
[0021] Based on the matching information, obtain the matching degree between the first object instance and the second object instance.
[0022] In an alternative embodiment, the obtaining of the matching degree between the first object instance and the second object instance based on the matching information includes:
[0023] Use the matching priority as a weight to perform weighted processing on the similarity between the first object instance and the second object instance to obtain the weighted similarity between the first object instance and the second object instance; and
[0024] Based on the distance and the equivalent radius, determine the adjacency relationship information between the first object instance and the second object instance;
[0025] Based on the weighted similarity and the adjacency relationship information, determine the matching degree between the first object instance and the second object instance.
[0026] In this way, the matching degree between the first object instance and the second object instance can be obtained more accurately.
[0027] In an alternative embodiment, there are multiple first object instances and multiple second object instances;
[0028] Based on the first attribute information and the second attribute information, matching the first object instance and the second object instance to obtain the matching degree between the first object instance and the second object instance, including:
[0029] Based on multiple first object instances and multiple second object instances, a plurality of object instance pairs are formed;
[0030] For each object instance pair, according to the first attribute information corresponding to the first object instance included in the object instance pair and the second attribute information corresponding to the second object instance included in the object instance pair, matching the first object instance and the second object instance included in the object instance pair to obtain the matching degree of the object instance pair.
[0031] In this way, object instance pairs with a matching degree meeting certain conditions can be screened out, and the accuracy of fusion can be improved.
[0032] In an optional implementation manner, based on the matching degree, fusing the first object instance and the second object instance, including:
[0033] Based on the matching degrees respectively corresponding to multiple object instance pairs and a preset matching degree threshold, grouping multiple first object instances and multiple second object instances to obtain multiple object instance groups;
[0034] For each object instance group, when the object instance group includes at least one first object instance and at least one second object instance, fusing the first object instance and the second object instance included in the object instance group.
[0035] In an optional implementation manner, the fusing of the first object instance and the second object instance includes:
[0036] Fusing the first attribute information of the first object instance and the second attribute information of the second object instance.
[0037] In a second aspect, an embodiment of the present disclosure provides an image processing apparatus, including:
[0038] A first determination module, configured to determine the first attribute information of the first object instance of the target object in the first image and determine the second attribute information of the second object instance of the target object in the second image;
[0039] A matching module, configured to match the first object instance and the second object instance based on the first attribute information and the second attribute information to obtain the matching degree between the first object instance and the second object instance;
[0040] A fusion module, configured to fuse the first object instance and the second object instance based on the matching degree.
[0041] In an alternative embodiment, it further includes:
[0042] An acquisition module, configured to acquire a first original image and a second original image;
[0043] A second determination module, configured to determine transformation relation information between the first original image and the second original image;
[0044] A transformation module, configured to perform transformation processing on the first original image based on the transformation relation information to obtain a first image, and use the second original image as the second image; or, perform transformation processing on the second original image based on the transformation relation information to obtain a second image, and use the first original image as the first image.
[0045] In an alternative embodiment, the attribute information includes at least one of the following:
[0046] Position information of the object instance, size information, probability that the object instance belongs to the target object, feature data, gray-scale data of the image region corresponding to the object instance, radiomics information;
[0047] The object instance includes: the first object instance and the second object instance.
[0048] In an alternative embodiment, the matching module includes:
[0049] A determination unit, configured to determine at least one of the following matching information between the first object instance and the second object instance based on the first attribute information and the second attribute information: similarity, matching priority, distance, equivalent radius;
[0050] A matching unit, configured to obtain a matching degree between the first object instance and the second object instance based on the matching information.
[0051] In an alternative embodiment, the matching unit is specifically configured to:
[0052] Use the matching priority as a weight to perform weighted processing on the similarity between the first object instance and the second object instance to obtain a weighted similarity between the first object instance and the second object instance; and
[0053] Determine adjacency relation information between the first object instance and the second object instance based on the distance and the equivalent radius;
[0054] Based on the weighted similarity and the adjacency relationship information, determine the matching degree between the first object instance and the second object instance.
[0055] In an alternative implementation, there are multiple first object instances and multiple second object instances;
[0056] The matching module is specifically configured to:
[0057] Based on multiple first object instances and multiple second object instances, form multiple pairs of object instances;
[0058] For each pair of object instances, according to the first attribute information corresponding to the first object instance included in the pair of object instances and the second attribute information corresponding to the second object instance included in the pair of object instances, match the first object instance and the second object instance included in the pair of object instances to obtain the matching degree of the pair of object instances.
[0059] In an alternative implementation, the fusion module is specifically configured to:
[0060] Based on the matching degrees corresponding to multiple pairs of object instances and a preset matching degree threshold, group multiple first object instances and multiple second object instances to obtain multiple groups of object instances;
[0061] For each group of object instances, when the group of object instances includes at least one first object instance and at least one second object instance, fuse the first object instance and the second object instance included in the group of object instances.
[0062] In an alternative implementation, the fusion module is specifically configured to:
[0063] Fuse the first attribute information of the first object instance and the second attribute information of the second object instance.
[0064] In a third aspect, an embodiment of the present disclosure further provides a computer device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps in the first aspect, or any possible implementation manner in the first aspect, are executed.
[0065] In a fourth aspect, an embodiment of the present disclosure further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the steps in the first aspect, or any possible implementation manner in the first aspect, are executed.
[0066] The image processing method, apparatus, computer device, and storage medium provided by the embodiments of the present disclosure match the attribute information of the first object instance and the second object instance to obtain the matching degree between the first object instance and the second object instance, and fuse the first object instance and the second object instance based on the matching degree. In this way, by matching the attribute information of the first object instance and the second object instance, the matching degree between the first object instance and the object instance is obtained, so that the first object instance and the second object instance can be more precisely fused according to the matching degree.
[0067] To make the above objects, features, and advantages of the present disclosure more obvious and understandable, the following specifically enumerates preferred embodiments and, in conjunction with the accompanying drawings, makes a detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for the embodiments. The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0069] Figure 1 Shows a flowchart of an image processing method provided by an embodiment of the present disclosure;
[0070] Figure 2 Shows a schematic diagram of image fusion proposed by an embodiment of the present disclosure;
[0071] Figure 3 Shows a flowchart of a specific method for determining the transformation relationship information between a first original image and a second original image provided by an embodiment of the present disclosure;
[0072] Figure 4 Shows a schematic diagram of multi-object instance matching proposed by an embodiment of the present disclosure;
[0073] Figure 5 Shows a schematic diagram of an image processing apparatus provided by an embodiment of the present disclosure;
[0074] Figure 6 Shows a schematic diagram of another image processing apparatus provided by an embodiment of the present disclosure;
[0075] Figure 7 Shows a specific schematic diagram of the matching module in the image processing apparatus provided by an embodiment of the present disclosure;
[0076] Figure 8 FIG. 1 shows a schematic diagram of a computer device provided by an embodiment of the present disclosure. DETAILED IMPLEMENTATION MANNER
[0077] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are only some, but not all, of the embodiments of the present disclosure. Components of the embodiments of the present disclosure generally described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations. Therefore, the detailed description of the embodiments of the present disclosure provided herein is not intended to limit the scope of the claimed present disclosure, but merely represents selected embodiments of the present disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of the present disclosure without creative efforts fall within the scope of the present disclosure.
[0078] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not require further definition and explanation in subsequent drawings.
[0079] As used herein, the term "and / or" merely describes an association relationship and indicates that three relationships may exist. For example, A and / or B may represent: A alone, both A and B present simultaneously, or B alone. Additionally, the term "at least one" as used herein represents any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C may represent including any one or more elements selected from the set consisting of A, B, and C.
[0080] It has been found through research that for images from different sources, the detection performance and results of the target detection algorithm may vary. For example, lesions may disappear or new ones may appear during follow-up, and the manifestations of lesions are different in different contrast phases; in addition, when the same patient takes different images, there are spatial displacements and changes in the background situation of the patient himself; these two factors make it impossible to simply combine the detection results from different sources. For the detection of multi-source images, there are currently two main solutions: one is to perform detection on each source image separately, and the reader establishes the connection by himself; the other is to use a registration algorithm to establish the spatial relationship between the two images and then map from one source to another. The first method does not solve the problem of multi-source image detection, and the second method partially solves the problem, but the current multi-source image fusion method generally directly superimposes the registered images; this image fusion method results in poor fusion accuracy.
[0081] Based on the above research, the present disclosure provides an image processing method. By matching the attribute information of the first object instance and the second object instance, the matching degree between the first object instance and the object instance can be obtained, so that the first object instance and the second object instance can be more precisely fused according to the matching degree.
[0082] To facilitate the understanding of this embodiment, first, a detailed introduction to an image processing method disclosed in the embodiments of the present disclosure is provided. The execution subject of the image processing method provided in the embodiments of the present disclosure is generally a computer device with certain computing capabilities. In some possible implementation manners, the image processing method can be implemented by a processor calling computer-readable instructions stored in a memory.
[0083] See Figure 1 as shown in Figure 1 is a flowchart of an image processing method provided in an embodiment of the present disclosure. The method includes steps S101 to S103, where:
[0084] S101: Determine the first attribute information of the first object instance of the target object in the first image and determine the second attribute information of the second object instance of the target object in the second image.
[0085] Among them, the first image and the second image are images obtained by photographing the same target object. Exemplarily, the first image and the second image are images obtained by photographing the target object at different times; or, they are images obtained by photographing the target object from different angles. Exemplarily, in the medical field, the target object can refer to the same diseased organ of the same patient, or the same body part, etc.; the first image and the second image can be multiple images with different shooting parameters such as angles and distances taken at the same time for the same organ or the same body part, or can be multiple images taken of the same organ or the same body part at different times. For example, during a single examination of the same patient, multiple images taken based on different shooting angles to obtain a comprehensive view of the lesion; or, during multiple examinations of the same patient, multiple images taken of the same lesion to observe the development process of the lesion. Correspondingly, the first object instance and the second object instance refer to a possible lesion on the first image and the second image respectively.
[0086] The attribute information includes at least one of the following:
[0087] The position information of the object instance, the size information, the probability that the object instance belongs to the target object, the feature data, the gray data of the image area corresponding to the object instance, the radiomics information;
[0088] The object instance includes: the first object instance in the first image and the second object instance in the second image.
[0089] In one embodiment of the present disclosure, the object instance is the result detected by a target detection algorithm, and certain post-processing and analysis may be performed on the detection result to obtain the object instance and its corresponding attribute information. Among them, the target detection algorithm can be any method for detecting specific targets from images, which is not limited herein. When detecting two images, the corresponding target detection algorithms can be different, but the types of detected targets need to be the same to obtain the same type of attribute information.
[0090] Exemplarily, the target detection algorithm can be Mask RCNN, Retina Net, etc.
[0091] In an embodiment of the present disclosure, the above target detection algorithm can be used to detect the image, so as to obtain the attribute information corresponding to the object instance.
[0092] Exemplarily, among various possible attribute information, the position information can be the coordinate information of the object instance in the image, that is, (x, y), which can be the coordinate information corresponding to the center point of the bounding box; the size information can be parameters such as the size and radius of the object instance that characterize the size of the object instance; the probability that the object instance belongs to the target object can be the probability or degree that the object instance has the target attribute determined based on a certain measurement standard, such as the malignancy degree corresponding to a lesion, etc.; the feature data can be the hidden layer encoding, that is, the feature vector corresponding to the object instance extracted by a certain network in the detection network; the gray-scale data of the image region corresponding to the object instance refers to the gray-scale values corresponding to each pixel point in the image region; the radiomics information is a descriptive quantity obtained by using radiomics methods to describe the features according to the image and range of the object instance. In addition, there may be other attribute information, which will not be elaborated herein.
[0093] In another embodiment of the present disclosure, target detection processing can also be performed on the image. Among them, the target detection method can be a method using semantic segmentation, that is, determining which category each pixel in the image belongs to, and based on the result of the target detection processing, obtaining the position information of the object instance in the image and the probability that the object instance belongs to the target object, and based on the position of the object instance in the image, determining the gray-scale data and / or radiomics information of the image region corresponding to the object instance.
[0094] In another embodiment of the present disclosure, before determining the attribute information of the object instance, it further includes:
[0095] Obtain a first original image and a second original image;
[0096] Determine the transformation relationship information between the first original image and the second original image;
[0097] Perform transformation processing on the first original image based on the transformation relationship information to obtain a first image, and use the second original image as the second image;
[0098] Or,
[0099] Perform transformation processing on the second original image based on the transformation relationship information to obtain a second image, and use the first original image as the first image.
[0100] Wherein, the first original image and the second original image are respectively original images taken for a target object. Since there is a certain positional deviation between the first original image and the second original image, for example, the second original image is translated 10 mm as a whole relative to the first original image, etc., it is necessary to register the first original image and the second original image so that the first original image and the second original image are in the same position in space, facilitating the extraction and matching of subsequent attribute information.
[0101] Specifically, the second original image can be registered based on the position information of the first original image to obtain a first image and the second image after transformation processing, or the first original image can be registered based on the position information of the second original image to obtain a second image and the first image after transformation processing.
[0102] Refer to Figure 2 shown, Figure 2 is a schematic diagram of image fusion proposed in an embodiment of the present disclosure. In Figure 2 , it can be obtained that after registering object instance a based on object instance A, the global spatial relationship 2 corresponding to object instance a is adjusted to the same global spatial relationship 1 as object instance A, and its corresponding coordinates and dimensions change, while other attributes independent of space remain unchanged.
[0103] See Figure 3 shown, which is a flowchart of a specific method for determining the transformation relationship information between the first original image and the second original image provided by an embodiment of the present disclosure. The method includes:
[0104] S301: Perform multi-level feature extraction on the first image to obtain first target feature maps respectively corresponding to the multi-level feature extraction; and perform multi-level feature extraction on the second image to obtain second target feature maps respectively corresponding to the multi-level feature extraction.
[0105] Among them, the first original image and the second original image may include different images captured at different times or from different angles for the same object. For example, when applying the image registration method provided in the examples of the present disclosure to the medical field, the first original image and the second original image may be medical images obtained by taking multiple shots during a single scan of the same lesion of the same patient, or may be different medical images obtained during scans of the same lesion of the patient at different times, which is not limited herein.
[0106] In the embodiments of the present disclosure, after the first original image and the second original image are determined, multi-level feature extraction can be performed on the two images respectively to obtain a first target feature map and a second target feature map corresponding to the multi-level feature extraction respectively. Among them, multi-level feature extraction can, for example, adopt the method of neural network self-learning to realize the gradient backpropagation of multi-level features of the image and extract the high-level semantic features of the image.
[0107] Specifically, performing multi-level feature extraction on an image to obtain a feature map corresponding to the multi-level feature extraction respectively includes:
[0108] For each level of feature extraction in the multi-level feature extraction, determine the first input data and the second input data of this level of feature extraction; the first input data includes: the image, or the encoded feature map output by the next level of feature extraction; the second input data includes: the decoded feature map output by the previous level of feature extraction;
[0109] Perform encoding processing corresponding to this level of feature extraction on the first input data to obtain an encoded feature map corresponding to this level of feature extraction;
[0110] Perform fusion processing on the encoded feature map corresponding to this level of feature extraction and the second input data to obtain a fused feature map;
[0111] Perform decoding processing corresponding to this level of feature extraction on the fused feature map to obtain a decoded feature map corresponding to this level of feature extraction;
[0112] Determine the decoded feature map corresponding to this level of feature extraction as the target feature map corresponding to this level of feature extraction.
[0113] In this example, the feature extractor can separately extract features from the first original image and the second original image to obtain multi-level features corresponding to the two images respectively. The feature extractor is a feature pyramid network, which is divided into two stages: encoding and decoding, and connects the low-level and high-level networks by means of skip connections. The feature extractor receives an image as input data, downsamples (also known as "subsampling") and extracts features layer by layer in the encoding module, and upsamples (also known as "upsampling") and extracts features layer by layer in the decoding module. The output of each decoding module will be fed into the iterative registration network to form multi-level, coarse-to-fine pyramid features.
[0114] For the encoding feature extraction process, for the highest level, that is, the coarsest level in the "coarse-to-fine" structure of the feature "pyramid" network, the corresponding input data is the original image, that is, the first original image or the second original image. For other levels, the corresponding input data is the encoded feature map output by the previous level. Thus, the iterative process of encoding feature extraction can be realized. For the decoding feature extraction process, its lowest level, that is, the finest level in the "coarse-to-fine" structure of the feature "pyramid" network, is the decoded feature map obtained from the corresponding lowest-level encoded feature map. For other levels, the corresponding input data is the decoded feature map output by the next level, and finally the feature extraction result for the original image is output. Thus, the iterative process of decoding feature extraction can be realized.
[0115] The multi-level feature extraction described in the embodiments of the present disclosure indicates a multi-level feature extraction process from "fine" to "coarse". Each level of feature extraction includes an encoding process and a decoding process corresponding to that level of feature extraction, that is, from the bottom to the top of the "pyramid" network.
[0116] In this example, each level of feature extraction includes an encoding network corresponding to that level of feature extraction and a decoding network corresponding to that level of feature extraction;
[0117] Among them, each encoding network includes: an encoding module and a parallel domain adaptation module.
[0118] The decoding network corresponding to the first-level encoding network includes: a decoding module.
[0119] The decoding networks corresponding to other encoding networks except the first-level encoding network include: a decoding module and a gated fusion module. The encoding process is a process from "coarse" to "fine", while the decoding process decodes sequentially from "fine" to "coarse". The multi-level feature extraction described in the embodiments of the present disclosure corresponds to the decoding process.
[0120] Among them, for the encoding network in each level of feature extraction, it is used to perform encoding processing corresponding to that level of feature extraction on the first input data to obtain the encoded feature map corresponding to that level of feature extraction.
[0121] For the decoding network in each - level feature extraction, it is used to perform a fusion process on the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map, and perform a decoding process corresponding to the feature extraction at this level on the fused feature map to obtain the decoded feature map corresponding to the feature extraction at this level.
[0122] After determining the first input data corresponding to each level during the encoding process, encoding processing can be performed based on the first input data to obtain the encoded feature map corresponding to each level.
[0123] In an embodiment of the present disclosure, the performing the encoding process corresponding to the feature extraction at this level on the first input data to obtain the encoded feature map corresponding to the feature extraction at this level includes:
[0124] Performing downsampling processing on the first input data to obtain a downsampled feature map;
[0125] Performing channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map;
[0126] Based on the downsampled feature map and the attention weights, obtaining the encoded feature map.
[0127] When the first input data is input into the encoding module in each - level feature extraction, the encoding module performs downsampling processing on the first input data to obtain the corresponding downsampled feature map, and inputs the output downsampled feature map into the parallel domain adaptation module corresponding to each - level encoding module. Among them, this parallel domain adaptation module can enhance the expression ability of the encoding network for certain specific features during the encoding stage; for example, if the image includes human organs; then the feature extractor can be trained to enhance the expression ability of the texture features of the organs. After inputting the downsampled feature map into the parallel domain adaptation module, the parallel domain adaptation module performs channel attention processing on the downsampled feature map to obtain the attention weights corresponding to each data channel in the downsampled feature map. Specifically, the above - mentioned process includes:
[0128] Performing global average pooling processing on the downsampled feature map to obtain a first feature sub - map;
[0129] Based on the first feature sub - map, determining the candidate attention weights corresponding to each data channel respectively; and based on the first feature sub - map, determining the feature - domain weights of the candidate attention corresponding to each data channel respectively;
[0130] Based on the candidate attention weights corresponding to each data channel respectively and the feature - domain weights of the candidate attention corresponding to each data channel respectively, obtaining the attention weights corresponding to each data channel in the downsampled feature map.
[0131] In the embodiment of the present disclosure, the parallel domain adaptation module includes two mechanisms: a channel attention mechanism and a domain perception mechanism. In the parallel domain adaptation module, the channel attention mechanism determines the channel attention weights of multiple channels in the form of grouped convolution, thereby generating multiple candidate channel weights. The domain perception mechanism can combine multiple candidate channel weights according to the properties of the feature map to obtain the final result, so that when the performance differences between the first original image and the second original image in the image domain are large, similar features can still be obtained through the feature extractor, facilitating the subsequent registration process.
[0132] In this embodiment, the parallel domain adaptation module is applied to the encoding stage of multi-level feature extraction. Specifically, in the encoding process from the (i + 1)-th level to the i-th level, the corresponding parallel domain adaptation module is applied. For example, in the encoding process from the 3rd level to the 2nd level, the parallel domain adaptation module 3 is applied. As Figure 4As shown, it is a specific example of a parallel domain adaptation module provided by an embodiment of the present disclosure. Here, it is assumed that the downsampled feature map corresponding to the encoding module of the i-th layer is represented as (N, C, D, H, W), where N represents the number of feature maps included in one feature extraction (Number of instances in batch), C represents the number of channels of the feature map (Channel), and D, H, W respectively represent the depth, height, and width of the feature map (Depth, Height, Width), and B represents the number of branches of channel attention (Branches of attention). After inputting the encoded feature map into this parallel domain adaptation module, after obtaining the downsampled feature map through the encoding module, global average pooling processing (Global Average Pooling) can be performed on the downsampled feature map, that is, the feature map is dimensionally reduced by taking the average of all pixels in the spatial dimensions of length, width, and height, so as to reduce the dimension of the data, making the feature map change from (N, C, D, H, W) to (N, C). Thus, the overall information in the channel dimension can be extracted to obtain the first feature sub-map. Then, the first feature sub-map is respectively subjected to repeat (Repeat) and flatten (Flatten) processing through two branches. In the first branch, after repeating the first feature sub-map, the first feature sub-map is transformed into the form of (N, BC, 1), where repeating means repeating the dimensionally reduced data B times, corresponding to B channel attentions. Here, the repeating operation is to change the data format for convenient calculation. Then, the feature map in this form is subjected to B groups of convolutions, activation, and then B groups of convolutions. Among them, the activation function can adopt the rectified linear unit (Rectified Linear Unit, ReLU). Finally, the channel of the convolved feature map is reorganized to obtain a feature map of (N, C, B), and thus the candidate attention weights corresponding to each data channel can be obtained. In the second branch, after flattening the first feature sub-map, the first feature sub-map is transformed into the form of (N, C), where flattening means rearranging the order of the data. Here, the flattening operation is to change the data format for convenient calculation. Then, the feature map in this form is subjected to fully connected and activation processing. Among them, the ReLU activation function and the excitation function (Softmax) of the deep learning output layer can be respectively adopted to obtain a feature map of the form (N, B), and thus the feature domain weights of the candidate attentions corresponding to each data channel can be obtained. After obtaining the candidate attention weights and feature domain weights corresponding to each data channel, the above two data are subjected to inner product processing to obtain feature data of (N, C). Then, this data is subjected to activation processing, and the activation function can be the sigmoid growth curve (Sigmoid), so that the encoded feature map of the i-th level can be obtained.
[0133] After determining the feature domain weights of the candidate attentions corresponding to each channel, the attention weights corresponding to each data channel in the downsampled feature map can be determined based on the feature domain weights of the candidate attentions corresponding to each data channel. Taking multiple images of the human liver as an example, for the same lesion, due to different scanning and shooting times or different shooting angles, although it is the same position, the presented images are different. Therefore, based on the feature domain weights determined by the domain perception module, the feature domain weights of the candidate attentions corresponding to each data channel are proportioned to highlight the weight occupied by the target position, that is, the weight occupied by the liver position, and weaken the weights occupied by other positions, such as muscle and blood positions. In this way, the target position can be prominently displayed. Thus, even when the display effects of the first original image and the second original image are quite different, similar features can be obtained through the feature extractor for registration.
[0134] Here, the channel attention mechanism can adopt a method for determining the channel attention weights in multiple parallel groups. For example, assuming there are 12 channels, usually, the above 12 channels can be convolved separately to obtain the candidate attention weights corresponding to the 12 channels. To improve the processing speed, the 12 channels can be divided into three branch channel attentions, with 4 channels in each branch. In this way, the three branch channel attentions can be convolved simultaneously, improving the processing speed.
[0135] After determining the attention weights corresponding to each data channel in the downsampled feature map, the encoded feature map can be obtained based on the downsampled feature map and the attention weights, that is, the features of each channel in the downsampled feature map are recombined according to the corresponding attention weights to obtain the encoded feature maps corresponding to each level.
[0136] In the decoding process corresponding to each level of feature extraction, in addition to the second input data input to each decoding module, the corresponding encoded feature map also participates in the generation process of the encoded feature map.
[0137] Here, since the encoded feature map has a high spatial resolution but a low degree of semantic information expression, and the upsampled feature map of the previous decoding module has a low spatial resolution but a high degree of semantic information expression. Therefore, in order to combine the respective advantages of the encoded feature map and the upsampled feature map of the previous decoding module, in the embodiments of the present disclosure, a gated fusion module for improving the fusion effect of high-level low-resolution features and low-level high-resolution features is adopted in the decoding stage.
[0138] Specifically, the fusion process of the encoded feature map corresponding to the feature extraction at this level and the second input data to obtain a fused feature map includes:
[0139] Based on the encoded feature map corresponding to this level of feature extraction and the second input data, obtain the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction;
[0140] Based on the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction and the encoded feature map, obtain a second feature sub-map;
[0141] Concatenate the second feature sub-map and the second input data to obtain the fused feature map.
[0142] Exemplarily, through a gated fusion module, fuse the encoded feature map of the encoded module data corresponding to this level of decoding module with the second feature data input to this decoding module to obtain the weights occupied by the encoded feature map corresponding to this level of feature extraction. Multiply the encoded feature map by this weight to obtain a second feature sub-map, and concatenate it with the second feature data input to this decoding module, then the fused feature map can be obtained. In this way, after multiplying the encoded feature map from the encoding module by the weight, it is concatenated with the decoded feature map of the decoding module in the channel dimension and sent to the next decoding module, enabling the feature maps from the two sources to be fused more efficiently.
[0143] In the disclosed embodiments, the main function of the gated fusion module is to determine the weights corresponding to each feature point in the encoded feature map. Specifically, the step of obtaining the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction based on the encoded feature map corresponding to this level of feature extraction and the second input data includes:
[0144] Concatenate the encoded feature map corresponding to this level of feature extraction and the second input data to obtain a third feature sub-map;
[0145] Perform convolution processing on the third feature sub-map to obtain a fourth feature sub-map, and based on the fourth feature sub-map, obtain the local autocorrelation coefficient between the encoded feature map and the second input data;
[0146] Based on the local autocorrelation coefficient, obtain the weights corresponding to each feature point in the encoded feature map corresponding to this level of feature extraction.
[0147] Wherein, the feature value of any feature point in the third feature sub-map represents the autocorrelation coefficient of the image region corresponding to this feature point.
[0148] After obtaining the third feature sub - graph obtained by splicing the encoded feature map corresponding to the feature extraction at this level and the second input data, perform convolution, normalization, and activation processing based on the Rectified Linear Unit (ReLU) on the third feature sub - graph to obtain the fourth feature sub - graph corresponding to the third feature sub - graph. Here, convolution, normalization, and activation processing are an operation pattern in a Convolutional Neural Network (CNN). The number of times the above operations occur represents the depth of the network. The deeper the network, the stronger the expression ability, and at the same time, the larger the number of parameters. Here, in order to improve the expression ability of the network, two or more layers of convolution, normalization, and activation processing can be selected.
[0149] After determining the fourth feature sub - graph, based on the fourth feature sub - graph, the encoded feature map and the local autocorrelation coefficient of the second input data can be obtained, including:
[0150] Perform maximum processing on the channel dimension of the fourth feature sub - graph to obtain the maximum value of the channel dimension of the fourth feature sub - graph; and perform average processing on the channel dimension of the fourth feature sub - graph to obtain the average value of the channel dimension of the fourth feature sub - graph;
[0151] Perform splicing processing on the maximum value of the channel dimension and the average value of the channel dimension to obtain the splicing result in the channel dimension, and perform convolution and normalization processing on the splicing result in the channel dimension to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0152] Among them, both the maximum in the channel dimension and the average in the channel dimension are operations for dimensionality reduction of data. They operate based on the channel dimension, aiming to reduce the number of parameters and retain spatial features, so that the data changes from (N, C, D, H, W) to (N, 1, C, H, W). Here, both the maximum value of the channel dimension and the average value of the channel dimension can be expressed as (N, 1, C, H, W). After determining the maximum value of the channel dimension and the average value of the channel dimension, the above two values can be spliced to obtain the spliced result (N, 2, C, H, W). Then, convolution and normalization processing can be performed on this spliced result to obtain the encoded feature map and the local autocorrelation coefficient of the second input data.
[0153] After obtaining the local autocorrelation coefficient, the activation function (sigmoid) can be used to activate the autocorrelation coefficient to obtain the gating activation values corresponding to the respective feature points in the encoded feature map; the gating activation values are used to represent the weights corresponding to the respective feature points in the encoded feature map. In this way, after multiplying the encoded feature map by the gating activation values obtained through the activation function and then concatenating with the second input data, the fused feature map can be obtained.
[0154] The embodiment of the present disclosure also provides a specific example of a gating fusion module. In this example, the encoded feature map output by the i-th level encoding module and the decoded feature map of the data after downsampling by the i-th level decoding module are concatenated. Here, concatenation means concatenation in the channel dimension, and it is required that the sizes of other channel dimensions are the same. For example, if the two data are respectively (N, C1, D, H, W) and (N, C2, D, H, W), the size after concatenation is (N, C1 + C2, D, H, W). Then, convolution, normalization, and activation processing are performed on the concatenated feature map, and channel dimension maximum and channel dimension average processing are also performed on it. After obtaining the channel dimension maximum value and the channel dimension average value, the above two data are concatenated, and convolution, normalization, and activation processing are performed again. After performing multiplication processing, the decoded feature map of the i + 1-th level is obtained.
[0155] After obtaining the fused feature map in the decoding stage, decoding processing corresponding to the feature extraction at this level is performed on the fused feature map to obtain the decoded feature map corresponding to the feature extraction at this level, and the decoded feature map corresponding to the feature extraction at this level is determined as the target feature map corresponding to the feature extraction at this level, which is convenient for subsequently inputting the target feature map corresponding to each level into the registration network.
[0156] Thus, in the embodiment of the present disclosure, the method of multi-level feature extraction, that is, the pyramid feature method, has a high degree of neural network reuse and fast iterative registration speed. At the same time, parallel domain adaptation modules and gating fusion modules are used in the encoding and decoding stages to enhance the generalization ability and robustness of the neural network and improve the abstraction ability in the feature extraction process.
[0157] Continuing from the above S301, the specific method for determining the transformation relationship information between the first original image and the second original image further includes:
[0158] S302: For each level of feature extraction, based on the first target feature map corresponding to the feature extraction at this level, the second target feature map, and the first transformation relationship information corresponding to the feature extraction at this level, determine the second transformation relationship information corresponding to the feature extraction at this level; wherein, the first transformation relationship information corresponding to the feature extraction at this level includes: the second transformation relationship information corresponding to the previous level of feature extraction, or the original transformation relationship information between the first image and the second image.
[0159] Among them, the first transformation relation information corresponding to this level of feature extraction includes: the second transformation relation information corresponding to the previous level of feature extraction, or the original transformation relation information between the first original image and the second original image.
[0160] Exemplarily, the original transformation relation information is also called the initial deformation relation. When the global deformation is small, the identity transformation can be used as the initial deformation relation, that is, the first original image and the initial image of the second original image are used as input data; when there is a large global deformation, the initialization methods commonly used in traditional methods can be used, such as: initializing the displacement based on the center of the image or the center of gravity of the gray level, and using the moment estimation of the gray level to initialize the rotation, etc. In this way, the problems of poor accuracy and even non-matching of the Field of View (FoV) when the previous linear registration network encounters too large global deformation can be solved.
[0161] Specifically, for each level of feature extraction, based on the first target feature map, the second target feature map corresponding to this level of feature extraction, and the first transformation relation information corresponding to this level of feature extraction, determining the second transformation relation information corresponding to this level of feature extraction includes:
[0162] Based on the first target feature map, the second target feature map, and the first transformation relation information, determining the transformation residual between this level of feature extraction and the previous level of feature extraction;
[0163] Based on the transformation residual and the first transformation relation information, determining the second transformation relation information corresponding to this level of feature extraction.
[0164] Among them, the iterative registration framework accepts an initial transformation relation in the form of a deformation field and several pairs of feature maps as input (the registration of each stage corresponds to the feature maps of each stage). The iterative registration framework includes multiple stages, and each stage includes a registration module and a combination module. The registration module accepts the feature map pair of this stage and the transformation relation of the previous stage (for the first stage, it is the initial transformation relation) as input, and outputs the residual with respect to the previous transformation relation, that is:
[0165]
[0166] where f(·) is the registration module, represents the feature map output by the target image feature extractor in the i-th stage, represents the feature map output by the source image feature extractor in the i-th stage, represents the application of the transformation relation, Φ i represents the cumulative transformation relation in the i-th stage, φ i represents Φ i with respect to Φi-1 The residuals. The superscripts of φ and Φ indicate the direction of the transformation relationship, that is represents the transformation relationship from the source image to the target image, represents the transformation relationship from the target image to the source image.
[0167] Here, the deformation field refers to in image registration, adding the predicted deformation field to the regular spatial grid to obtain the sampling grid, and using the sampling grid containing deformation information for the floating image to obtain the deformed image. The size of the deformation field corresponding to a two-dimensional image of size [W, H] is [W, H, 2], where the size of the third dimension is 2, representing the displacements in the x-axis and y-axis directions respectively. Similarly, the size of the deformation field corresponding to a three-dimensional image of size [D, W, H] is [D, W, H, 3], where the size of the third dimension is 3, representing the displacements in the x-axis, y-axis, and z-axis directions respectively.
[0168] The registration module is divided into two types: linear transformation and deformation transformation. They are the same in terms of input and output. Therefore, in application, the registration modules of linear and deformation transformations can be freely combined according to requirements, such as connecting a linear registration module and three deformation transformation modules in series. To ensure that the transformation process is diffeomorphic, the two modules do not directly output φ, but obtain φ according to the transformation, that is, the output is the matrix corresponding to φ. Here, by ensuring that the transformation process is diffeomorphic, the number of operations can be reduced, and at the same time, since the output is a matrix, the process is reversible, providing a role for verification.
[0169] The specific process of the registration process is as follows:
[0170] The original output of the linear registration module is the forward rotation, scaling, and shearing matrix A fw and the forward translation vector b fw . A fw can be directly predicted by the network, or the parameters of rotation, scaling, and shearing can be predicted separately and then combined. The reverse linear transformation relationship is obtained by inverting the linear transformation, that is:
[0171] A bw =(A fw ) -1 (2)
[0172] b bw =-(A fw ) -1 b fw (3)
[0173] After obtaining A and b, applying them to the mesh grid and then subtracting the mesh grid can obtain the deformation relationship φ in the form of the deformation field.
[0174] The original output of the deformation transformation is the forward deformation relationship V in the form of a velocity field fw , and on this basis, the deformation relationship φ in the form of a deformation field is obtained through integration, where the deformation relationship φ satisfies the following formula (4) and formula (5):
[0175]
[0176]
[0177] where represents the application of the deformation relationship.
[0178] The combination module accumulates φ of each previous step to obtain the cumulative deformation relationship Φ, that is (note that for forward and reverse deformations, the combination order is opposite):
[0179]
[0180]
[0181] where and are the initial deformation relationships, and the final network output is also obtained by accumulating through the combination module.
[0182] Thus, in the embodiments of the present disclosure, after deformation initialization, a registration neural network is used to perform multi-level fitting residuals, that is, it has the advantage of fast registration with neural network learning, and can also avoid the disadvantage of low accuracy when the global deformation is too large in neural network learning, making the registration fast and accurate.
[0183] Continuing from the above S302, the specific method for determining the transformation relationship information between the first original image and the second original image further includes:
[0184] S303: Based on the second transformation relationship information corresponding to the last-level feature extraction, register the first image and the second image.
[0185] Specifically, after multi-level iterative registration, based on the second transformation relationship information output after the last-level registration, that is, the final transformation information, the first original image and the second original image can be registered.
[0186] When registering the first original image and the second original image, for example, the second transformation relationship information corresponding to the last-level feature extraction can be used to perform transformation processing on the first original image to obtain the transformed image of the first original image, and then the transformed image and the second original image are matched one by one in position to obtain the registration result.
[0187] In addition, the second transformation relationship information corresponding to the last-level feature can also be utilized to perform transformation processing on the second original image to obtain the image after transformation of the second original image, and then the images after transformation and the first original image are matched one by one in position to obtain the registration result.
[0188] In the embodiment of the present disclosure, the registration method involved is applied to a pre-trained registration neural network, and the registration neural network includes two branch networks, namely, a feature extraction neural network and a multi-level registration neural network; the feature extraction neural network is used to perform multi-level feature extraction processing on the first original image and the second original image respectively; the multi-level registration neural network is used to determine the target transformation relationship information between the first original image and the second original image based on the multi-level features extracted from the first original image and the second original image, wherein the target transformation relationship information is used to register the first original image and the second original image.
[0189] Following the above S101, the image processing method proposed in the embodiment of the present disclosure further includes:
[0190] S102: Based on the first attribute information and the second attribute information, match the first object instance and the second object instance to obtain the matching degree between the first object instance and the second object instance.
[0191] Specifically, after obtaining the first attribute information and the second attribute information, the first object instance and the second object instance can be matched based on the obtained attribute information. That is, based on the first attribute information and the second attribute information, at least one of the following matching information between the first object instance and the second object instance is determined: similarity, matching priority, distance, equivalent radius, and based on the matching information between the first object instance and the second object instance, the matching degree between the first object instance and the second object instance is obtained.
[0192] Exemplarily, the embodiment of the present disclosure provides a specific manner to obtain the matching degree between the first object instance and the second object instance. In this embodiment, the matching information includes: similarity, matching priority, distance, and equivalent radius;
[0193] The obtaining of the matching degree between the first object instance and the second object instance based on the matching information between the first object instance and the second object instance includes:
[0194] Using the matching priority as a weight to perform weighted processing on the similarity between the first object instance and the second object instance to obtain the weighted similarity between the first object instance and the second object instance; and
[0195] Determine the adjacency relationship information between the first object instance and the second object instance based on the distance and the equivalent radius;
[0196] Determine the matching degree between the first object instance and the second object instance based on the weighted similarity and the adjacency relationship information.
[0197] Exemplarily, the similarity, matching priority, distance, and equivalent radius between every two object instances can be calculated by traversing all possible object instances in two images for pairing, forming four matrices. In this example, let A represent the first object instance and a represent the second object instance. Specifically:
[0198] (1) Similarity is a metric value. When the attributes of two object instances are closer, the value of the similarity is larger. For numerical controllability, it can be scaled or truncated to the range of [0, 1] or [-1, 1]. Here, taking the radius of the object instance as an example, for the first object instance A and the second object instance a, their similarity can be:
[0199]
[0200] In this case, the similarity of their radii is between [0, 1], and the closer the two object instances are in terms of radius, the greater the similarity.
[0201] Similarly, the similarity between other attribute information can also be calculated. For example, when the attribute information includes: in the case where the attribute information of the object instance includes the bounding box, the intersection over union (IoU) of the bounding boxes corresponding to the first object instance and the second object instance can be calculated, and this IoU can be used as the similarity between the two; in the case where the attribute information of the object instance includes the contour, the Dice coefficient of the contours corresponding to the first object instance and the second object instance can be calculated, and this Dice coefficient can be used as the similarity between the two; in the case where the attribute information of the object instance includes feature data, such as in the case of an encoded vector, the cosine similarity of the encoded vectors corresponding to the first object instance and the second object instance can be calculated, and the value of this cosine similarity can be used as the similarity between the two, etc.
[0202] In addition, when calculating the similarities corresponding to multiple attributes, the similarities corresponding to multiple attributes can be weighted and summed, and the result of the weighted sum can be used as the similarity between the first object instance and the second object instance.
[0203] (2) The matching priority is also a measure. The higher the importance of the object instance, the larger its value. For example, for a lesion, the matching priority can represent the malignant probability of the lesion. The higher the malignant probability and the larger the lesion, the higher the corresponding matching priority. Similarly, the matching priority can also be scaled and truncated to ensure numerical controllability. Taking the malignant probability of a lesion as an example, the matching priority can be expressed as:
[0204]
[0205] Since the malignant probability is a value between [0, 1], from the above formula, it can be seen that the numerical distribution of the matching priority is between [0.5, 1]. Similarly, different matching priority measures can be set for different attributes, or matching priority measures can be set for multiple attributes, or multiple matching priority measures can be weighted, such as the deterioration speed and spread speed of the lesion, which will not be elaborated here.
[0206] (3) The distance is the distance between all possible pairs of object instances. Exemplarily, the Euclidean distance between the centers of the object instances can be used to represent the distance between two object instances:
[0207] Distance(A, a) = ||Center coordinate A - Center coordinate a|| 2 ;
[0208] where, ||·|| k represents the k-th norm, and the 2-norm of a vector is the square root of the sum of the squares of all components of the vector.
[0209] (4) The equivalent radius is usually the average value of the radii or diagonal lengths of two object instances in a certain form. For example, when the manifestation form of the object instance is center + radius, its equivalent radius is the geometric mean of the radii of the two object instances; when the manifestation form of the object instance is a bounding box, its equivalent radius can be the diagonal length, etc.: Among them, the equivalent radius being the geometric mean of the radii of two object instances can be expressed as:
[0210]
[0211] According to the above method, four matrices can be obtained, namely the similarity matrix, the matching priority matrix, the distance matrix, and the equivalent radius matrix. For each matrix, the element in the i-th row and j-th column represents the corresponding value of the i-th object instance of Image 1 for the j-th object instance of Image 2. These four matrices can be used for further combination in subsequent calculations.
[0212] Based on the four obtained matrices, the similarity matrix and the matching priority matrix can be combined into a weighted similarity matrix. The weighted similarity is the similarity after considering the matching priority, and thus the weighted similarity can be expressed as:
[0213] Weighted similarity = Similarity * Matching priority.
[0214] The adjacency relationship matrix is the relative position relationship between two object instances calculated based on the distance and the equivalent radius. The adjacency relationship is used to measure whether two object instances are close or coincident in position, is negatively correlated with the distance, and is positively correlated with the equivalent radius. For example, it can be expressed as follows:
[0215]
[0216] In addition, for the convenience of subsequent calculations, a threshold can also be set to correct the adjacency relationship. For example, it can be expressed as follows:
[0217]
[0218] Based on the weighted similarity and the adjacency relationship, the final matching degree can be obtained. The matching degree determines the quality of the match between two object instances. The matching degree can be expressed as:
[0219] Matching degree = Weighted similarity * Adjacency relationship.
[0220] Thus, the matching degree between two object instances can be obtained.
[0221] In another embodiment of the present disclosure, when there are multiple first object instances and multiple second object instances, the matching the first object instances and the second object instances based on the first attribute information and the second attribute information to obtain the matching degree between the first object instances and the second object instances includes:
[0222] Based on the multiple first object instances and the multiple second object instances, multiple object instance pairs are formed;
[0223] For each object instance pair, based on the first attribute information corresponding to the first object instance included in the object instance pair and the second attribute information corresponding to the second object instance included in the object instance pair, the first object instance and the second object instance included in the object instance pair are matched to obtain the matching degree of the object instance pair.
[0224] Specifically, based on the matching degrees respectively corresponding to multiple object instances and a preset matching degree threshold, the multiple first object instances and the multiple second object instances can be grouped to obtain multiple object instance groups; for each object instance group, when the object instance group includes at least one first object instance and at least one second object instance, the first object instance and the second object instance included in the object instance group are fused.
[0225] Specifically, after obtaining the matching degree matrix, the final matching result can be obtained from the matrix. That is, based on a threshold, the pairs with a matching degree greater than the threshold are selected as a feasible pair, and the matrix is transformed into an undirected bipartite graph, where each node on the graph is an object instance and the weight of the edge is the matching degree; then, according to the weight of the edge from high to low, sampling with replacement or sampling without replacement is performed for pairing. Sampling with replacement will fuse all possible connected object instances into a new object instance; sampling without replacement preferentially pairs the two object instances with higher matching degrees. When an object instance is successfully paired, it can no longer be paired with other object instances. Here, the method of sampling with or without replacement is a configurable item and is determined according to requirements. In sampling without replacement, an object instance can be paired with at most one other object instance, which is applicable to tracking the changes of pulmonary nodules during follow-up; sampling with replacement allows many-to-many matching and is applicable to cross-modal situations, such as the capsule of a tumor not being visible in a specific modality, and there is no limitation on the application direction.
[0226] Exemplarily, after a pair of object instances is successfully matched, the successfully matched object instances will not continue to be matched and can be removed from the object instance pool allowed for matching, but the object instances that have not been successfully matched can still remain in the object instance pool and continue to wait for matching until all the object instances remaining in the object instance pool no longer meet the matching conditions.
[0227] Exemplarily, refer to Figure 4 , Figure 4 which is a schematic diagram of multi-object instance matching proposed in the embodiment of the present disclosure. After obtaining the matching degree matrix, based on a preset threshold of 0.5, multiple object instance pairs that meet the threshold are screened out, including 0.7 between Aa, 0.6 between Ba, and 0.6 between Cc, and the final matching result is obtained based on the methods of sampling with replacement and sampling without replacement.
[0228] In the embodiment of the present disclosure, since the results of multiple attributes are incorporated during matching, considering similarity, matching priority, position, and size relationship, etc., the matching accuracy is higher.
[0229] Following the above S102, the image processing method proposed in the embodiment of the present disclosure further includes:
[0230] S103: Based on the matching degree, fuse the first object instance and the second object instance.
[0231] Specifically, after determining the matching degree, the first attribute information of the first object instance and the second attribute information of the second object instance can be fused.
[0232] Exemplarily, after the matching is completed, attribute fusion can be performed. By default, Image 1 is used as the reference space. For attributes related to space, such as bounding boxes, centers, contours, etc., coordinate transformation will be performed according to the registration relationship before fusion. When performing attribute fusion, different fusion methods are adopted according to the type of the attribute, which is not limited herein. For example, for bounding boxes and contours, the union of all object instances is obtained; for dimensions, they are either recalculated based on the fused bounding box or contour, or the maximum value of all object instances is obtained; for malignancy degree, etc., the maximum value of all object instances is obtained; for the hidden layer encoding of the detection module, the mean value of multiple object instances is obtained. Refer to Figure 2 As shown, for the fused object instance, its coordinates and dimensions are transformed into the coordinates and dimensions after registration according to the registration relationship, its corresponding malignancy degree selects the maximum value between object instance A and object instance a, and the feature vectors are also averaged, thus completing the fusion process between the first object instance and the second object instance.
[0233] In another embodiment of the present disclosure, when there are multiple images, the merging method can be used for fusion to obtain the final result. For example, pairwise object instances are fused, and then the results of pairwise fusion are fused; or, any two object instances are randomly selected for fusion, and then the fusion result is fused with another object instance until all object instances are fused. These two merging methods depend on whether multiple sources are in a parallel relationship or a hierarchical relationship, which will not be elaborated herein. Thus, multiple images, such as multi-modal, multi-phase contrast imaging, or the results of multiple temporal follow-ups, can be associated, and the matching method is more flexible, which can better assist users in joint image reading and diagnosis.
[0234] The embodiments of the present disclosure are applied, for example, to a liver imaging diagnosis platform. Patients will perform detections on multi-phase contrast results respectively. On the one hand, it solves the problem that single-phase imaging may miss diagnoses. For example, focal nodular hyperplasia is not obvious in the portal phase, and metastatic tumors are often not obvious in the arterial phase, and the relationship of lesions in multiple phases is established for subsequent multi-phase joint diagnosis. It can also be applied to a pulmonary imaging diagnosis platform. According to multiple temporal pulmonary images, nodule follow-up is performed to analyze the volume and sign changes of nodules, which is of great significance for judging the patient's condition.
[0235] In the embodiments of the present disclosure, for multiple images, first, a target detection algorithm is used to obtain the detection result corresponding to each image, that is, the attribute information of the object instance, and then a registration algorithm is used to obtain the spatial transformation relationship between any two images. According to the attributes corresponding to the detection results and the relative spatial relationship, the similarity, matching priority, distance, and equivalent radius of all possible detection target pairings in the two images are calculated in sequence, and the matching degree is calculated by combination to obtain the detection target matching result of the two images, and then attribute fusion is performed to obtain the fused object instance. In this way, by matching the attribute information of the first object instance and the second object instance, the matching degree between the first object instance and the object instance is obtained, so that the first object instance and the second object instance can be more precisely fused according to the matching degree.
[0236] Those skilled in the art can understand that in the above method of the specific embodiment, the writing order of each step does not mean a strict execution order and does not constitute any limitation to the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0237] Based on the same inventive concept, an image processing device corresponding to the image processing method is also provided in the embodiments of the present disclosure. Since the principle of solving problems by the device in the embodiments of the present disclosure is similar to the above image processing method in the embodiments of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0238] Refer to Figure 5 、 Figure 6 、 Figure 7 As shown in Figure 5 is a schematic diagram of an image processing device provided by an embodiment of the present disclosure; Figure 6 is a schematic diagram of another image processing device provided by an embodiment of the present disclosure; Figure 7 In the image processing device provided by the embodiments of the present disclosure, it is a specific schematic diagram of the matching module. The device includes: a first determination module 510, a matching module 520, and a fusion module 530; wherein,
[0239] The first determination module 510 is configured to determine the first attribute information of the first object instance of the target object in the first image and determine the second attribute information of the second object instance of the target object in the second image;
[0240] The matching module 520 is configured to match the first object instance and the second object instance based on the first attribute information and the second attribute information to obtain the matching degree between the first object instance and the second object instance;
[0241] The fusion module 530 is configured to fuse the first object instance and the second object instance based on the matching degree.
[0242] In an alternative embodiment, as Figure 6 shown, it further includes:
[0243] An acquisition module 540, configured to acquire a first original image and a second original image;
[0244] A second determination module 550, configured to determine transformation relation information between the first original image and the second original image;
[0245] A transformation module 560, configured to perform transformation processing on the first original image based on the transformation relation information to obtain a first image, and use the second original image as the second image; or, perform transformation processing on the second original image based on the transformation relation information to obtain a second image, and use the first original image as the first image.
[0246] In an alternative embodiment, the attribute information includes at least one of the following:
[0247] Position information of the object instance, size information, probability that the object instance belongs to the target object, feature data, gray-scale data of the image region corresponding to the object instance, radiomics information;
[0248] The object instance includes: the first object instance and the second object instance.
[0249] In an alternative embodiment, as Figure 7 shown, the matching module 520 includes:
[0250] A determination unit 521, configured to determine at least one of the following matching information between the first object instance and the second object instance based on the first attribute information and the second attribute information: similarity, matching priority, distance, equivalent radius;
[0251] A matching unit 522, configured to obtain a matching degree between the first object instance and the second object instance based on the matching information between the first object instance and the second object instance.
[0252] In an alternative embodiment, the matching unit 522 is specifically configured to:
[0253] Use the matching priority as a weight to perform weighted processing on the similarity between the first object instance and the second object instance to obtain a weighted similarity between the first object instance and the second object instance; and
[0254] Determine adjacency relation information between the first object instance and the second object instance based on the distance and the equivalent radius;
[0255] Based on the weighted similarity and the adjacency relationship information, determine the matching degree between the first object instance and the second object instance.
[0256] In an optional implementation, there are multiple first object instances and multiple second object instances;
[0257] The matching module 520 is specifically configured to:
[0258] Based on multiple first object instances and multiple second object instances, form multiple pairs of object instances;
[0259] For each pair of object instances, according to the first attribute information corresponding to the first object instance included in the pair of object instances and the second attribute information corresponding to the second object instance included in the pair of object instances, match the first object instance and the second object instance included in the pair of object instances to obtain the matching degree of the pair of object instances.
[0260] In an optional implementation, the fusion module 530 is specifically configured to:
[0261] Based on the matching degrees corresponding to multiple pairs of object instances and a preset matching degree threshold, group multiple first object instances and multiple second object instances to obtain multiple groups of object instances;
[0262] For each group of object instances, in the case where the group of object instances includes at least one first object instance and at least one second object instance, fuse the first object instance and the second object instance included in the group of object instances.
[0263] In an optional implementation, the fusion module 530 is specifically configured to:
[0264] Fuse the first attribute information of the first object instance and the second attribute information of the second object instance.
[0265] In the embodiments of the present disclosure, by matching the attribute information of the first object instance and the second object instance, the matching degree between the first object instance and the object instance is obtained, so that the first object instance and the second object instance can be fused more precisely according to the matching degree.
[0266] For the description of the processing flow of each module in the device and the interaction flow between each module, reference may be made to the relevant descriptions in the above method embodiments, which will not be elaborated here.
[0267] Corresponding to Figure 1 in the image processing method, the embodiments of the present disclosure further provide a computer device, such as Figure 8As shown in the figure, it is a schematic structural diagram of a computer device provided by an embodiment of the present disclosure, including:
[0268] A processor 801, a memory 802, and a bus 803; the memory 802 is used to store execution instructions, including an internal memory 8021 and an external memory 8022; here, the internal memory 8021 is also called the main memory, which is used to temporarily store the operation data in the processor 801 and the data exchanged with the external memory 8022 such as a hard disk. The processor 801 exchanges data with the external memory 8022 through the internal memory 8021. When the computer device runs, communication between the processor 801 and the memory 802 is carried out through the bus 803, so that the processor 801 executes the following instructions:
[0269] Determine the first attribute information of the first object instance of the target object in the first image and determine the second attribute information of the second object instance of the target object in the second image;
[0270] Based on the first attribute information and the second attribute information, match the first object instance and the second object instance to obtain the matching degree between the first object instance and the second object instance;
[0271] Based on the matching degree, fuse the first object instance and the second object instance.
[0272] An embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the steps of the image processing method described in the above method embodiment. Among them, the storage medium can be a volatile or non-volatile computer-readable storage medium.
[0273] An embodiment of the present disclosure also provides a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the image processing method described in the above method embodiment. For details, refer to the above method embodiment, which will not be repeated here.
[0274] Among them, the above computer program product can be specifically implemented in a way of hardware, software, or a combination thereof. In an optional embodiment, the computer program product is specifically embodied as a computer storage medium. In another optional embodiment, the computer program product is specifically embodied as a software product, such as a Software Development Kit (SDK), etc.
[0275] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein. In several embodiments provided by the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some communication interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0276] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0277] In addition, in each embodiment of the present disclosure, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0278] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art or part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present disclosure. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0279] Finally, it should be noted that the above-described embodiments are only specific implementation manners of the present disclosure, used to illustrate the technical solutions of the present disclosure, rather than limiting it. The protection scope of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the technical field familiar with the present disclosure can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should all be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure should be subject to the protection scope of the claims.
Claims
1. An image processing method, characterized in that, comprising: determining first attribute information of multiple first object instances of a target object in a first image and second attribute information of multiple second object instances of the target object in a second image; wherein the first object instances or the second object instances are obtained by performing target detection on the first image or the second image; matching the first object instances and the second object instances based on the first attribute information and the second attribute information to obtain a matching degree between the first object instances and the second object instances; grouping the multiple first object instances and the multiple second object instances based on the matching degrees between multiple pairs of the first object instances and the second object instances and a preset matching degree threshold to obtain multiple object instance groups; for each of the object instance groups, when the object instance group includes at least one of the first object instances and at least one of the second object instances, fusing the first object instances and the second object instances included in the object instance group.
2. The image processing method according to claim 1, characterized in that, further comprising: acquiring a first original image and a second original image; determining transformation relation information between the first original image and the second original image; performing transformation processing on the first original image based on the transformation relation information to obtain the first image, and using the second original image as the second image; or, performing transformation processing on the second original image based on the transformation relation information to obtain the second image, and using the first original image as the first image.
3. The image processing method according to claim 1 or 2, characterized in that, the first attribute information or the second attribute information includes at least one of the following: position information of the object instance, size information, probability that the object instance belongs to the target object, feature data, gray scale data of an image region corresponding to the object instance, radiomics information; the object instances include: the first object instances and the second object instances.
4. The image processing method according to claim 1, characterized in that, the matching of the first object instances and the second object instances based on the first attribute information and the second attribute information includes: determining at least one of the following matching information between the first object instances and the second object instances based on the first attribute information and the second attribute information: similarity, matching priority, distance, equivalent radius; obtaining the matching degree between the first object instances and the second object instances based on the matching information.
5. The image processing method according to claim 4, characterized in that, the obtaining the matching degree between the first object instances and the second object instances based on the matching information includes: using the matching priority as a weight to perform weighted processing on the similarity between the first object instances and the second object instances to obtain a weighted similarity between the first object instances and the second object instances; and Determine the adjacency relationship information between the first object instance and the second object instance based on the distance and the equivalent radius; Determine the matching degree between the first object instance and the second object instance based on the weighted similarity and the adjacency relationship information.
6. The image processing method according to claim 1, wherein, the matching the first object instance and the second object instance based on the first attribute information and the second attribute information to obtain the matching degree between the first object instance and the second object instance includes: Construct a plurality of object instance pairs based on a plurality of the first object instances and a plurality of the second object instances; For each object instance pair, match the first object instance and the second object instance included in the object instance pair according to the first attribute information corresponding to the first object instance included in the object instance pair and the second attribute information corresponding to the second object instance included in the object instance pair, to obtain the matching degree of the object instance pair.
7. The image processing method according to claim 1, wherein, fusing the first object instance and the second object instance included in the object instance group includes: Fusing the first attribute information of the first object instance and the second attribute information of the second object instance.
8. An image processing apparatus, wherein, comprises: A first determination module, configured to determine the first attribute information of a plurality of first object instances of a target object in a first image and determine the second attribute information of a plurality of second object instances of the target object in a second image; The first object instance or the second object instance is obtained by performing target detection on the first image or the second image; A matching module, configured to match the first object instance and the second object instance based on the first attribute information and the second attribute information, to obtain the matching degree between the first object instance and the second object instance; A fusion module, configured to group the plurality of first object instances and the plurality of second object instances based on the matching degree between the plurality of first object instances and the second object instances and a preset matching degree threshold, to obtain a plurality of object instance groups; For each of the object instance groups, when the object instance group includes at least one first object instance and at least one second object instance, fuse the first object instance and the second object instance included in the object instance group.
9. A computer device, wherein, comprises: A processor, a memory and a bus, the memory stores machine-readable instructions executable by the processor, when the computer device runs, the processor communicates with the memory through the bus, and when the machine-readable instructions are executed by the processor, the steps of the image processing method according to any one of claims 1 to 7 are executed.
10. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is run by a processor, the steps of the image processing method according to any one of claims 1 to 7 are executed.
Citation Information
Patent Citations
Multi-modal medical image registration method and device
CN110533641A