Geometric Registration Method and Model for Natural Images Based on Multi-Intelligent Reinforcement Learning
Through multi-intelligent reinforcement learning methods, image structure information is extracted, attention-related maps are calculated, and registration parameters are generated through policy networks and value fusion networks, which solves the problem of poor results in traditional image registration methods and achieves more efficient and accurate image registration.
Patent Information
- Application Number
- CN202210910738.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-07-29
AI Technical Summary
The traditional feature-based image registration method lacks attention to key points, resulting in poor image registration effect.
The natural image geometric registration method based on multi-intelligent reinforcement learning is adopted to generate accurate registration parameters through structural information extraction, attention-related graph calculation, policy network calculation and value fusion network calculation.
It improves the accuracy and efficiency of image registration, can better reflect the subject content and secondary parts in the image, and is suitable for geometric registration tasks of natural images.
Smart Images

Figure CN115482262B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image recognition, and particularly to a geometric registration method and model for natural images based on multi-intelligent reinforcement learning. Background Art
[0002] Image registration generally refers to the process of matching two image sets to the same coordinate system according to their image contents. Common image registration tasks include multi-view registration, multi-temporal registration, multi-modal registration, and scene-model registration. According to the data source, it can also be divided into medical image registration, remote sensing image registration, and natural image registration. In terms of principle, it can be divided into region-based registration and feature-based image registration. The region-based method focuses more on the matching of similar regions in the image rather than extracting features in the image, and is more suitable for the registration of some images lacking local structure and shape information.
[0003] In the traditional method, the feature-based image registration method generally has four steps: feature extraction, feature matching, transformation model estimation, and image resampling and transformation. Feature extraction refers to detecting key structural information in the image, such as features like points, lines, and regions. The invariance and overlap criteria of these features provide the preconditions for subsequent feature matching. Feature matching is to match the detected features through a measure of the degree of correlation described by the feature space distribution in their neighboring regions, that is, to find the pairwise correspondence between them by using the spatial relationship or other feature descriptions of the two sets of features detected in the source image and the target image. In addition, there are also some methods based on invariant descriptors to match some stable and unique features in the image. However, due to the lack of attention to key points, the image registration effects of these methods are not good. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide a geometric registration method and model for natural images based on multi-intelligent reinforcement learning, and the geometric registration method and model can accurately complete the image registration operation.
[0005] To achieve the above purpose, the embodiments of the present invention provide a geometric registration method for natural images based on multi-intelligent reinforcement learning, including:
[0006] Obtain a source image and a registration image;
[0007] Perform structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features;
[0008] Perform an attention-related graph calculation operation according to the first features and the second features to obtain an attention-related graph;
[0009] Perform multiple policy network calculation operations respectively according to the attention-related map to obtain a corresponding plurality of first vectors;
[0010] Perform a value fusion network calculation operation according to the first vector and the attention-related map to obtain a registration parameter.
[0011] Optionally, performing a structure information extraction operation on the source image and the registration image respectively to obtain corresponding first features and second features includes:
[0012] Input the source image and the registration image into a U-net network respectively to obtain the first feature and the second feature, wherein the last layer of the U-net network is replaced by two convolutional layers and a pooling layer arranged alternately.
[0013] Optionally, performing an attention-related map calculation operation according to the first feature and the second feature to obtain an attention-related map includes:
[0014] Perform a convolution operation and a scale transformation operation on the first feature to obtain a first standard feature;
[0015] Perform a convolution operation, a scale transformation operation and a transpose operation on the second feature to obtain a second standard feature;
[0016] Perform a matrix multiplication operation and a softmax function activation operation on the first standard feature and the second standard feature to obtain an attention map;
[0017] Stack the first feature and the second feature to obtain a correlation map;
[0018] Perform a matrix multiplication operation on the attention map and the correlation map to obtain a first attention-related map;
[0019] Perform a matrix multiplication operation on the attention map and the first correlation map to obtain the attention-related map.
[0020] Optionally, performing multiple policy network calculation operations respectively according to the attention-related map to obtain a corresponding plurality of first vectors includes:
[0021] Perform a convolution operation, a multi-layer perceptron operation, a recurrent neural network operation and a fully connected operation on the attention-related map in sequence to obtain the first vector.
[0022] Optionally, performing a value fusion network calculation operation according to the first vector and the attention-related map to obtain a registration parameter includes:
[0023] Perform a multi-layer perceptron operation on the attention-related map to obtain a first weight;
[0024] Perform a weighted operation according to the first weight and the first vector to obtain a second vector;
[0025] Perform a multi-layer perceptron operation on the attention-related map to obtain a first bias parameter;
[0026] Perform an addition operation and a first function activation operation according to the first bias parameter and the second vector to obtain a third vector;
[0027] Perform a multi-layer perceptron operation on the attention-related map to obtain a second weight;
[0028] Perform a weighted operation according to the second weight and the second vector to obtain a third vector;
[0029] Perform a multi-layer perceptron operation, a second function activation operation, and a multi-layer perceptron operation on the attention-related map to obtain a second bias parameter;
[0030] Perform a third function activation operation and an addition operation according to the second bias parameter and the third vector to obtain the registration parameter.
[0031] Optionally, the geometric registration method further includes:
[0032] Obtain a training set and a test set;
[0033] Train according to the training set and the test set using the loss function of formula (1) to update the parameters of the structure information extraction operation, the attention-related map calculation operation, the policy network calculation operation, and the value fusion network calculation operation,
[0034]
[0035]
[0036] where L(θ) is the loss function with respect to the input quantity θ, b is the size of the sampling batch in the experience replay strategy, r is the reward function, γ is the discount factor, Q tot is the value function, τ is the observation-action history, u is the current decision-making action, s is the current observation value, u′ is the maximum value of the current decision-making action, τ′ is the maximum value in the observation-action history, s′ is the maximum value of the current observation value, and θ - is the maximum value of the input quantity.
[0037] On the other hand, the present invention also provides a geometric registration model for natural images based on multi-agent reinforcement learning, as Figure 8 shown, the geometric registration model includes:
[0038] An input layer for obtaining a source image and a registration image;
[0039] A structure information extraction layer for performing structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features;
[0040] An attention-related map calculation layer for performing attention-related map calculation operations based on the first features and second features to obtain an attention-related map;
[0041] Multiple policy network layers for performing multiple policy network calculation operations based on the attention-related map respectively to obtain corresponding multiple first vectors;
[0042] A value fusion network for performing value fusion network calculation operations based on the first vectors and the attention-related map to obtain registration parameters.
[0043] Optionally, the structure information extraction layer includes a U-net network, and the last layer of the U-net network is replaced by two convolutional layers and a pooling layer arranged alternately;
[0044] The attention-related map calculation layer is used for:
[0045] Performing a convolution operation and a scale transformation operation on the first features to obtain first standard features;
[0046] Performing a convolution operation, a scale transformation operation, and a transpose operation on the second features to obtain second standard features;
[0047] Performing a matrix multiplication operation and a softmax function activation operation on the first standard features and the second standard features to obtain an attention map;
[0048] Stacking the first features and the second features to obtain a correlation map;
[0049] Performing a matrix multiplication operation on the attention map and the correlation map to obtain a first attention-related map;
[0050] Performing a matrix multiplication operation on the attention map and the first correlation map to obtain the attention-related map.
[0051] Optionally, the policy network layer is used for sequentially performing a convolution operation, a multi-layer perceptron operation, a recurrent neural network operation, and a fully connected operation on the attention-related map to obtain the first vector.
[0052] Optionally, the value fusion network is used for:
[0053] Performing a multi-layer perceptron operation on the attention-related map to obtain a first weight;
[0054] Performing a weighted operation based on the first weight and the first vector to obtain a second vector;
[0055] Perform multi-layer perceptron operations on the attention-related map to obtain a first bias parameter;
[0056] Perform an addition operation and a first function activation operation based on the first bias parameter and the second vector to obtain a third vector;
[0057] Perform multi-layer perceptron operations on the attention-related map to obtain a second weight;
[0058] Perform a weighted operation based on the second weight and the second vector to obtain a third vector;
[0059] Perform multi-layer perceptron operations, a second function activation operation, and multi-layer perceptron operations on the attention-related map to obtain a second bias parameter;
[0060] Perform a third function activation operation and an addition operation based on the second bias parameter and the third vector to obtain the registration parameter.
[0061] Through the above technical solutions, the geometric registration method and model for natural images based on multi-agent reinforcement learning provided by the present invention first extract the structural information of the source image and the registration image, so that the registration structure can better reflect the main content in the two images; then through the calculation of the attention-related map, the calculated attention-related map can reflect the main and secondary parts in the image; finally, through the calculation of the policy network and the value fusion network, the output registration parameter can communicate and combine the action values of each agent to fuse into the global action value, and by the method of fusion weight and bias, the accuracy of the registration parameter is improved.
[0062] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The drawings are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the specification, and are used to explain the embodiments of the present invention together with the following specific implementation, but do not constitute a limitation to the embodiments of the present invention. In the drawings:
[0064] Figure 1 is a flowchart of a geometric registration method for natural images based on multi-agent reinforcement learning according to an embodiment of the present invention;
[0065] Figure 2 is a schematic structural diagram of a U-net network according to an embodiment of the present invention;
[0066] Figure 3 is a flowchart of a method for obtaining an attention-related map according to an embodiment of the present invention;
[0067] Figure 4 Schematic diagram of the attention-related map calculation layer according to an embodiment of the present invention;
[0068] Figure 5 Schematic diagram of the policy network layer according to an embodiment of the present invention;
[0069] Figure 6 Flowchart of a method for obtaining registration parameters according to an embodiment of the present invention;
[0070] Figure 7 Schematic diagram of the value fusion network according to an embodiment of the present invention;
[0071] Figure 8 Block diagram of the geometric registration model of natural images based on multi-agent reinforcement learning according to an embodiment of the present invention. Detailed implementation manners
[0072] The following will describe in detail the specific implementation manners of the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.
[0073] As Figure 1 shown is the flowchart of the geometric registration method of natural images based on multi-agent reinforcement learning according to an embodiment of the present invention. In this Figure 1 , the geometric registration method may include:
[0074] In step S10, obtain a source image and a registration image;
[0075] In step S11, perform structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features;
[0076] In step S12, perform an attention-related map calculation operation according to the first feature and the second feature to obtain an attention-related map;
[0077] In step S13, perform multiple policy network calculation operations according to the attention-related map respectively to obtain corresponding multiple first vectors;
[0078] In step S14, perform a value fusion network calculation operation according to the first vector and the attention-related map to obtain registration parameters.
[0079] In this as Figure 1In the method shown, step S10 is used to obtain a source image and a registration image, facilitating subsequent image registration operations through registration parameters. Among them, considering that the main content in the registration image needs to be highlighted in subsequent registration operations, it is necessary to first extract the structural information of the source image and the registration image. When extracting this structural feature in step S11, in order to ensure as much as possible that the extracted structural feature can contain more structural information about the main body of the image, so as to improve the performance of subsequent feature matching and registration. In an example of the present invention, a U-Net (semantic segmentation) network can be used to extract this structural feature. Additionally, since the U-net network itself outputs a semantic segmentation map, and only the middle structural feature of the U-net is required in this example, the last layer of the U-net network can be replaced by two convolutional layers (CONV) and pooling layers (maxpool) arranged alternately, as Figure 2 shown. The U-Net network can include three parts: a contracting path, an expanding path, and skip connections. The contracting path consists of three convolutional blocks, each convolutional block containing a convolutional layer and a pooling layer. Each time the feature undergoes downsampling, the number of channels becomes twice the original. The expanding path is also three times of upsampling. The upsampling method selects deconvolution with a size of 2×2. Each time the feature undergoes upsampling, the number of channels becomes half of the original, and it needs to be concatenated with the feature of the same-level contracting path, that is, skip connection. After connection, the number of channels of the feature remains the same as before upsampling. Then, the feature will pass through two convolutional layers with a size of 3×3 and a ReLU activation layer. In the last layer, the feature passes through a convolutional layer with a size of 1×1 to make the feature into a single-channel output.
[0080] Step S12 can be used to calculate an attention-related map, thereby obtaining a correlation map that can express the relationship between all features, that is, this attention-related map, for subsequent registration operations. Although the specific method for obtaining the attention-related map can be various forms known to those skilled in the art. But in a preferred example of the present invention, step S12 can further include as Figure 3 shown in. In this Figure 3 step S12 can include:
[0081] In step S20, a convolution operation and a scale transformation operation are performed on the first feature to obtain a first standard feature;
[0082] In step S21, a convolution operation, a scale transformation operation, and a transpose operation are performed on the second feature to obtain a second standard feature;
[0083] In step S22, a matrix multiplication operation and a softmax function activation operation are performed on the first standard feature and the second standard feature to obtain an attention map;
[0084] In step S23, the first feature and the second feature are superimposed to obtain a correlation map;
[0085] In step S24, matrix multiplication is performed on the attention map and the correlation map to obtain a first attention-correlation map;
[0086] In step S25, matrix multiplication is performed on the attention map and the first attention-correlation map to obtain an attention-correlation map.
[0087] And the attention-correlation map calculation layer corresponding to the method shown as Figure 3 can have a structure as shown as Figure 4 . Specifically, the first feature f A obtains a first standard feature through a convolution operation (1*1, CONV) and a scale transformation operation (Reshape); the second feature f B obtains a second standard feature through a convolution operation (1*1, CONV), a scale transformation operation (Reshape), and a transpose operation (Transpose). The first standard feature and the second standard feature go through matrix multiplication and a Softmax function activation operation, thereby obtaining an attention map (Attention Map) that measures the importance of the representation features. Among them, the Softmax function of the Softmax function activation operation can be as shown in formula (1),
[0088]
[0089] where m ji is the influence of the feature at the i-th position on the feature at the j-th position , and N is the dimension of the feature. Further, in order to balance the attention map and the correlation map, an adaptive weight α can be set in front of the attention map.
[0090] The first feature f A and the second feature f B are superimposed, thereby obtaining a correlation map (Correlation Map) representing the relevant part of the first feature and the second feature; then matrix multiplication is performed on the attention map and the correlation map to obtain a first attention-correlation map; finally, matrix multiplication is performed on the attention map and the first attention-correlation map, thereby obtaining the attention-correlation map (Attention-Correlation Map).
[0091] Step S13 is used to perform multiple policy network calculation operations on the attention-related graph to obtain the corresponding first vector. Specifically, it can be to sequentially perform a convolutional operation (CNN Module), a multi-layer perceptron operation (MLP), a recurrent neural network operation (GRU), and a fully connected operation (MLP) on the attention-related graph to obtain the first vector. The network model structure corresponding to this step S13 can be as shown in Figure 5 as shown.
[0092] Step S14 can be used to perform a value fusion network calculation operation based on the first vector and the attention-related graph to obtain the registration parameter. Specifically, this step S14 can include the steps as shown in Figure 6 as shown in.
[0093] In this Figure 6 this step S14 can include:
[0094] In step S30, perform a multi-layer perceptron operation on the attention-related graph to obtain the first weight;
[0095] In step S31, perform a weighted operation based on the first weight and the first vector to obtain the second vector;
[0096] In step S32, perform a multi-layer perceptron operation on the attention-related graph to obtain the first bias parameter;
[0097] In step S33, perform an addition operation and a first function activation operation based on the first bias parameter and the second vector to obtain the third vector;
[0098] In step S34, perform a multi-layer perceptron operation on the attention-related graph to obtain the second weight;
[0099] In step S35, perform a weighted operation based on the second weight and the second vector to obtain the third vector;
[0100] In step S36, perform a multi-layer perceptron operation, a second function activation operation, and a multi-layer perceptron operation on the attention-related graph to obtain the second bias parameter;
[0101] In step S37, perform a third function activation operation and an addition operation based on the second bias parameter and the third vector to obtain the registration parameter.
[0102] Corresponding to the method as shown in Figure 6 the network structure of the value fusion network can be as shown in Figure 7As shown. Since it is considered that the non-linearity of the fusion equation in the fusion stage is not strong and the fusion relationship can be traced, the two-stage fusion process is not directly completed using the fusion network. Additionally, if the global state and the multi-objective action values output by the policy network are directly handed over to a neural network for processing, it will undoubtedly increase the learning difficulty and even result in underfitting. Therefore, a hypernetwork (such as Figure 7 the network structure shown therein) is selected to generate the fusion parameters. The hypernetwork can use a neural network to process the state information and output the corresponding fusion parameters, and can selectively perform some additional processing on the fusion parameters, such as achieving monotonicity constraints through absolute value activation or adding a linear component to the fusion. The state information is processed and the corresponding fusion parameters are output, and some additional processing can be selectively performed on the fusion parameters, such as achieving monotonicity constraints through absolute value activation or adding a linear component to the fusion.
[0103] In this embodiment, for the method of training a model composed of a structure information extraction operation, an attention-related graph calculation operation, a policy network calculation operation, and a value fusion network calculation operation, although there are various methods known to those skilled in the art, in this embodiment, in order to improve the training efficiency of the model. The method of training the model can be to first obtain a test set and a training set, and then use the loss function of formula (1) to train according to the training set and the test set to update the parameters of the structure information extraction operation, the attention-related graph calculation operation, the policy network calculation operation, and the value fusion network calculation operation,
[0104]
[0105]
[0106] where L(θ) is the loss function with respect to the input quantity θ, b is the size of the sampling batch in the experience replay policy, r is the reward function, γ is the discount factor, Q tot is the value function, τ is the observation-action history, u is the current decision-making action, s is the current observation value, u′ is the maximum value of the current decision-making action, τ′ is the maximum value in the observation-action history, s′ is the maximum value of the current observation value, and θ - is the maximum value of the input quantity.
[0107] On the other hand, the present invention also provides a geometric registration model for natural images based on multi-intelligent reinforcement learning. The geometric registration model may include an input layer 01, a structure information extraction layer 02, an attention-related graph calculation layer 03, a plurality of policy network layers 04, and a value fusion network 05. Among them, the input layer 01 may be used to obtain a source image and a registration image. The structure information extraction layer 02 may be used to perform structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features. The attention-related graph calculation layer 03 may be used to perform attention-related graph calculation operations according to the first features and the second features to obtain an attention-related graph. The plurality of policy network layers 04 may be used to perform multiple policy network calculation operations according to the attention-related graph to obtain corresponding first vectors. The value fusion network 05 may perform value fusion network calculation operations according to the first vectors and the attention-related graph to obtain registration parameters.
[0108] In this embodiment, the structure information extraction layer 02 may include a U-net network. The last layer of the U-net network may be replaced by two convolutional layers and a pooling layer arranged alternately, that is, as Figure 2 shown in the structure. The structure of the attention-related graph calculation layer 03 may be as Figure 4 shown, and may be used for the method shown in Figure 3 . Specifically, the attention-related graph calculation layer 03 may be used for:
[0109] In step S20, perform a convolution operation and a scale transformation operation on the first feature to obtain a first standard feature;
[0110] In step S21, perform a convolution operation, a scale transformation operation, and a transpose operation on the second feature to obtain a second standard feature;
[0111] In step S22, perform a matrix multiplication operation and a softmax function activation operation on the first standard feature and the second standard feature to obtain an attention map;
[0112] In step S23, stack the first feature and the second feature to obtain a correlation graph;
[0113] In step S24, perform a matrix multiplication on the attention map and the correlation graph to obtain a first attention-related graph;
[0114] In step S25, perform a matrix multiplication operation on the attention map and the first attention-related graph to obtain an attention-related graph.
[0115] In this embodiment, the plurality of policy network layers 04 may be used to perform a convolution operation, a multi-layer perceptron operation, a recurrent neural network operation, and a fully connected operation on the attention-related graph in sequence to obtain a first vector, and its network structure may be as Figure 5as shown
[0116] In this embodiment, the structure of the value fusion network 05 can be as Figure 7 shown, for performing the method as shown Figure 6 in. In this Figure 6 , the value fusion network 05 can be used for:
[0117] In step S30, perform a multi-layer perceptron operation on the attention-related map to obtain a first weight;
[0118] In step S31, perform a weighted operation according to the first weight and the first vector to obtain a second vector;
[0119] In step S32, perform a multi-layer perceptron operation on the attention-related map to obtain a first bias parameter;
[0120] In step S33, perform an addition operation and a first function activation operation according to the first bias parameter and the second vector to obtain a third vector;
[0121] In step S34, perform a multi-layer perceptron operation on the attention-related map to obtain a second weight;
[0122] In step S35, perform a weighted operation according to the second weight and the second vector to obtain a third vector;
[0123] In step S36, perform a multi-layer perceptron operation, a second function activation operation, and a multi-layer perceptron operation on the attention-related map to obtain a second bias parameter;
[0124] In step S37, perform a third function activation operation and an addition operation according to the second bias parameter and the third vector to obtain a registration parameter.
[0125] Through the above technical solution, the geometric registration method and model of natural images based on multi-intelligent reinforcement learning provided by the present invention first extract the structural information of the source image and the registration image, so that the registration structure can better reflect the main content in the two images; then through the calculation of the attention-related map, the calculated attention-related map can reflect the main and secondary parts in the image; finally, through the calculation of the policy network and the value fusion network, the output registration parameter can communicate and combine the action values of each agent to fuse into the global action value, and by the method of fusing weights and biases, the accuracy of the registration parameter is improved.
[0126] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0127] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0128] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means realizes the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0130] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.
[0131] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0132] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0133] In addition, in order to further verify the technical effects of the geometric registration method and model of natural images based on multi-intelligent reinforcement learning provided by the present invention, the effectiveness of the geometric registration method and model can also be verified by setting the methods of formulas (2) and (3).
[0134]
[0135]
[0136] Among them, is the index of the i-th key point (feature point) and the k-th key point when the threshold is T, p is the serial number of the object, and d pi is the Euclidean distance between the predicted value and the marked point of the key point with id i in the p-th object. is the normalized reference of p objects, and δ is the perturbation factor. is the overall index.
[0137] In this embodiment, the PCK (Percentage of Correct Keypoints) method can be used as an evaluation index. The PCK method refers to calculating the proportion of the total number of points whose normalized distance between the labeled points in the registered image and the corresponding labeled points in the target image is less than the threshold value under a given threshold. It is proved by the above method that when the input image is a photovoltaic panel dataset, the PCK index score of the geometric registration method and model provided by the present invention is 74.25%. In contrast, the score of this geometric registration method and model has increased by 13.25% compared with the CNNgeometric model in the prior art, and has increased by 7.75% compared with the Unet-MRL model in the prior art. The post-model with the attention mechanism suppresses some irrelevant regions, such as the surrounding turf and trees, etc. The model mainly focuses on the local regions such as the photovoltaic panels in the image, which is beneficial to predicting all deformation parameters. Compared with the VGG-Attention-MRL model in the prior art, the score of this geometric registration method and model has increased by 6%. This geometric registration method and model use U-net to pre-train a semantic segmentation network as part of the feature extraction network, effectively extracting the common feature information between multi-modal data, and obtaining more refined features with the help of the implicit structural information after segmentation, so that the agent can learn the feature matching differences from it to complete better registration. The model in this chapter applies semantic segmentation and structural features in feature extraction and environmental reward feedback, and uses the attention mechanism to improve the accuracy of feature correlation calculation. For partially overlapping image pairs, or images with very little main content, the model can also successfully align the overlapping photovoltaic panels. The model in this chapter can achieve large translations, rotations, scalings, etc. between image pairs with different environments, and the model is not limited to registering between exactly the same image pairs, but can achieve registration between partially overlapping image pairs.
[0138] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the element.
[0139] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A geometric registration method for natural images based on multi-intelligent reinforcement learning, characterized in that, The geometric registration method includes: Obtaining a source image and a registration image; Performing structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features; Performing an attention-related graph calculation operation according to the first features and the second features to obtain an attention-related graph; Performing multiple policy network calculation operations according to the attention-related graph respectively to obtain corresponding multiple first vectors; Performing a value fusion network calculation operation according to the first vectors and the attention-related graph to obtain registration parameters; Performing an attention-related graph calculation operation according to the first features and the second features to obtain an attention-related graph includes: Performing a convolution operation and a scale transformation operation on the first features to obtain first standard features; Performing a convolution operation, a scale transformation operation, and a transpose operation on the second features to obtain second standard features; Performing a matrix multiplication operation and a softmax function activation operation on the first standard features and the second standard features to obtain an attention map; Stacking the first features and the second features to obtain a correlation graph; Performing a matrix multiplication operation on the attention map and the correlation graph to obtain a first attention-related graph; Performing a matrix multiplication operation on the attention map and the first attention-related graph to obtain the attention-related graph; Performing a value fusion network calculation operation according to the first vectors and the attention-related graph to obtain registration parameters includes: Performing a multi-layer perceptron operation on the attention-related graph to obtain a first weight; Performing a weighted operation according to the first weight and the first vectors to obtain second vectors; Performing a multi-layer perceptron operation on the attention-related graph to obtain a first bias parameter; Performing an addition operation and a first function activation operation according to the first bias parameter and the second vectors to obtain third vectors; Performing a multi-layer perceptron operation on the attention-related graph to obtain a second weight; Performing a weighted operation according to the second weight and the second vectors to obtain third vectors; Performing a multi-layer perceptron operation, a second function activation operation, and a multi-layer perceptron operation on the attention-related graph to obtain a second bias parameter; Performing a third function activation operation and an addition operation according to the second bias parameter and the third vectors to obtain the registration parameters.
2. The geometric registration method according to claim 1, wherein Performing structure information extraction operations on the source image and the registration image respectively to obtain corresponding first features and second features includes: Respectively inputting the source image and the registration image into a U-net network to obtain the first features and the second features, wherein the last layer of the U-net network is replaced by two convolutional layers and pooling layers arranged alternately.
3. The geometric registration method according to claim 1, wherein Performing multiple policy network calculation operations according to the attention-related graph respectively to obtain corresponding multiple first vectors includes: Sequentially performing a convolution operation, a multi-layer perceptron operation, a recurrent neural network operation, and a fully connected operation on the attention-related graph to obtain the first vectors.
4. The geometric registration method according to claim 1, characterized in that, The geometric registration method further includes: Obtaining a training set and a test set; The loss function using formula (1) is trained according to the training set and the test set to update the parameters of the structure information extraction operation, the attention-related map calculation operation, the policy network calculation operation, and the value fusion network calculation operation. Among them, \(L(\theta)\) is the loss function with respect to the input quantity \(\theta\), \(b\) is the size of the sampling batch in the experience replay strategy, \(r\) is the reward function, \(\gamma\) is the discount factor, \(Q\) tot is the value function, \(\tau\) is the observation-action history, \(u\) is the current decision-making action, \(s\) is the current observation value, \(u'\) is the maximum value of the current decision-making action, \(\tau'\) is the maximum value in the observation-action history, \(s'\) is the maximum value of the current observation value, and \(\theta\) - is the maximum value of the input quantity.
5. A geometric registration model for natural images based on multi-intelligent reinforcement learning, characterized in that, The geometric registration model includes: An input layer for obtaining a source image and a registration image; A structure information extraction layer for respectively performing structure information extraction operations on the source image and the registration image to obtain corresponding first features and second features; An attention-related map calculation layer for performing attention-related map calculation operations according to the first features and the second features to obtain an attention-related map; Multiple policy network layers for respectively performing multiple policy network calculation operations according to the attention-related map to obtain corresponding multiple first vectors; A value fusion network for performing value fusion network calculation operations according to the first vectors and the attention-related map to obtain registration parameters; The structure information extraction layer includes a U-net network, and the last layer of the U-net network is replaced by two convolutional layers and a pooling layer arranged alternately; The attention-related map calculation layer is used for: Performing a convolution operation and a scale transformation operation on the first feature to obtain a first standard feature; Performing a convolution operation, a scale transformation operation, and a transpose operation on the second feature to obtain a second standard feature; Performing a matrix multiplication operation and a softmax function activation operation on the first standard feature and the second standard feature to obtain an attention map; Stacking the first feature and the second feature to obtain a correlation map; Performing a matrix multiplication operation on the attention map and the correlation map to obtain a first attention-related map; Performing a matrix multiplication operation on the attention map and the first attention-related map to obtain the attention-related map; The value fusion network is used for: Performing a multi-layer perceptron operation on the attention-related map to obtain a first weight; Performing a weighted operation according to the first weight and the first vector to obtain a second vector; Performing a multi-layer perceptron operation on the attention-related map to obtain a first bias parameter; Performing an addition operation and a first function activation operation according to the first bias parameter and the second vector to obtain a third vector; Performing a multi-layer perceptron operation on the attention-related map to obtain a second weight; Performing a weighted operation according to the second weight and the second vector to obtain a third vector; Performing a multi-layer perceptron operation, a second function activation operation, and a multi-layer perceptron operation on the attention-related map to obtain a second bias parameter; Performing a third function activation operation and an addition operation according to the second bias parameter and the third vector to obtain the registration parameters.
6. The geometric registration model according to claim 5, wherein The policy network layer is used for sequentially performing a convolution operation, a multi-layer perceptron operation, a recurrent neural network operation, and a fully connected operation on the attention-related map to obtain the first vector.
Citation Information
Patent Citations
Structural information guided cross-domain image geometric registration method
CN113592927A
Unsupervised learning medical image registration method and system
CN113763441A