Similar image transformation parameter calculation method and device, equipment, medium and product

By constructing a deep learning model to automatically extract high-dimensional feature differences between images, the problem of low efficiency and poor accuracy in calculating transformation parameters of similar images is solved, and high-precision transformation parameter estimation is achieved in complex backgrounds and non-rigid deformation scenarios.

CN120823249APending Publication Date: 2025-10-21CHENGDOU HUAQIYUN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510982509.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing technologies suffer from low computational efficiency and poor accuracy in calculating transformation parameters for similar images, especially when dealing with complex image structures or non-rigid changes, making it difficult to meet the needs of practical applications.

Method used

A deep learning model is constructed that includes a feature extraction network, a feature error calculation layer, and a transformation parameter regression network. This model automatically extracts high-dimensional feature differences between images and accurately regresses the scaling, translation, and rotation parameters between images.

Benefits of technology

It improves the intelligence and computational efficiency of image registration and structure alignment, enhances the accuracy of transformation parameter estimation in complex backgrounds and non-rigid deformation scenarios, and has strong robustness and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823249A_ABST
    Figure CN120823249A_ABST
Patent Text Reader

Abstract

The invention discloses a similar image transformation parameter calculation method and device, equipment, a medium and a product, and relates to the technical field of image processing, and the method comprises the steps: obtaining a to-be-calculated image and a reference contrast image, inputting the to-be-calculated image and the reference contrast image to a transformation parameter calculation model, and carrying out the transformation parameter calculation; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer and a transformation parameter regression network, and a transformation parameter calculation result is output. The parameter calculation result comprises a scaling calculation result, a displacement calculation result and a rotation calculation result. According to the method, the high-dimensional feature difference between the images can be automatically extracted, the scaling, displacement and rotation parameters between the images can be accurately regressed, and the method has relatively high robustness and adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, apparatus, device, storage medium, and computer product for calculating similar image transformation parameters. Background Art

[0002] With the rapid development of image processing and computer vision technology, image registration and transformation parameter estimation based on deep learning have been widely used in virtual human generation, medical image alignment, AR / VR modeling and other fields.

[0003] In related technologies, feature point detection is usually performed using algorithms such as SIFT and SURF to obtain feature descriptors of the feature points, and then algorithms such as RANSAC are used to obtain transformation parameters such as scaling and translation. However, when processing complex image structures or non-rigid changes, the computational efficiency is significantly limited.

[0004] Furthermore, some existing neural network-based methods attempt to directly regress transformation parameters from images. However, most models are simple, using only a single image as input or failing to consider inter-image similarity, making it difficult to accurately model the difference between two images. Existing methods lack sufficient understanding and utilization of the details of image geometric transformations, resulting in large errors in transformation parameter estimation and failing to meet the dual requirements of accuracy and generalization for practical applications. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, equipment, storage medium and computer product for calculating similar image transformation parameters, aiming to solve the technical problems of low computational efficiency and poor accuracy in the calculation of similar image transformation parameters in related technologies.

[0006] To achieve the above objectives, this application proposes a method for calculating similar image transformation parameters, which includes: Acquire an image to be calculated and a reference control image; Inputting the image to be calculated and the reference image into a transformation parameter calculation model to calculate the transformation parameters; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer and a transformation parameter regression network; Output transformation parameter calculation results; parameter calculation results include scaling calculation results, displacement calculation results, and rotation calculation results.

[0007] In one embodiment, the feature extraction network sequentially includes a three-layer convolution layer-activation function layer-maximum pooling layer structure and a one-layer convolution layer-activation function layer-average pooling layer structure; The feature extraction network is used to extract features from the image to be calculated and the reference image to obtain a 256-dimensional feature vector of the image to be calculated and a 256-dimensional feature vector of the reference image.

[0008] In one embodiment, the feature error calculation layer is used to determine a 256-dimensional feature vector based on the difference between the feature vector of the image to be calculated and the feature vector of the reference control image.

[0009] In one embodiment, the transformation parameter regression network includes a three-layer linear layer-activation function layer structure; The transformation parameter regression network is used to learn the mapping relationship of image transformation and output a 5-dimensional transformation parameter feature vector.

[0010] In one embodiment, the training step of the transformation parameter calculation model includes: Input the training sample into the feature extraction network to calculate the image features; the training sample includes an image pair consisting of an original image and a transformed image; each image in the image pair shares the network parameters of the feature extraction network when calculating the features; For each image pair, calculate the image feature difference between the image pairs; The image feature difference is input into the parameter regression network to obtain the initial transformation parameter calculation result; Based on a preset loss function, the error value between the initial transformation parameter calculation result and the actual transformation parameter is determined; Continue to iterate the model parameters until the error value is less than the preset value or the number of iterations is greater than the preset number of iterations.

[0011] In one embodiment, the loss function consists of a scaling error function, a translation error function, a rotation error function, and a total error function.

[0012] In addition, to achieve the above-mentioned purpose, the present application further provides a similar image transformation parameter calculation device, the device comprising: An image acquisition module, used to acquire the image to be calculated and the reference control image; A parameter calculation module is used to input the image to be calculated and the reference control image into a transformation parameter calculation model to calculate the transformation parameters; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer and a transformation parameter regression network; The result output module is used to obtain the transformation parameter calculation results; the parameter calculation results include scaling calculation results, displacement calculation results and rotation calculation results.

[0013] In addition, to achieve the above-mentioned purpose, the present application continues to provide a similar image transformation parameter calculation device, the device including: a memory, a processor, and a computer program stored in the memory and runnable on the processor, the computer program being configured to implement the steps of the similar image transformation parameter calculation method as claimed in any one of claims 1 to 6.

[0014] In addition, to achieve the above-mentioned purpose, the present application continues to provide a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the similar image transformation parameter calculation method as claimed in any one of claims 1 to 6 are implemented.

[0015] In addition, to achieve the above-mentioned purpose, the present application continues to provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the similar image transformation parameter calculation method as claimed in any one of claims 1 to 6.

[0016] One or more technical solutions proposed in this application have at least the following technical effects: The method for calculating similar image transformation parameters provided in this application builds a deep learning model consisting of a feature extraction network, a feature error calculation layer, and a transformation parameter regression network. This model automatically extracts high-dimensional feature differences between images and accurately regresses the scale, displacement, and rotation parameters between them. Compared to traditional methods, this method is more robust and adaptable, maintaining high transformation parameter estimation accuracy even in complex backgrounds and non-rigid deformation scenarios, thereby improving the intelligence and computational efficiency of image registration and structural alignment. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 Schematic diagram of the flow of a method for calculating similar image transformation parameters in one embodiment of the present application.

[0020] Figure 2 Schematic diagram of the model structure of the similar image transformation parameter calculation method in one embodiment of the present application.

[0021] Figure 3 This is a structural diagram of the similar image transformation parameter calculation device of this application.

[0022] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0023] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.

[0024] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0025] The present invention provides a method for calculating similar image transformation parameters. Figure 1 , Figure 1 This is a flowchart of the first embodiment of the method for calculating similar image transformation parameters of the present application.

[0026] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device capable of implementing the above functions, a similar image transformation parameter calculation device, etc. The following uses the similar image transformation parameter calculation device as an example to illustrate this embodiment and the following embodiments.

[0027] In this embodiment, the similar image transformation parameter calculation method includes steps S10 to S30: Step S10: Acquire the image to be calculated and the reference image.

[0028] Step S20: input the image to be calculated and the reference image into a transformation parameter calculation model to perform transformation parameter calculation; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer and a transformation parameter regression network.

[0029] Step S30, outputting transformation parameter calculation results; the parameter calculation results include scaling calculation results, displacement calculation results, and rotation calculation results.

[0030] It should be noted that the image to be calculated refers to the image for which transformation parameters are analyzed. Its geometry may differ from that of the reference image in terms of scale, displacement, or rotation. The reference image is the reference image against which the image to be calculated is compared. It is typically the original image or an image in a standard pose, serving as a geometric reference. The transformation parameters (e.g., scale, translation, and rotation) of the image to be calculated are calculated relative to the reference image.

[0031] The transformation parameter calculation model is a neural network model built based on deep learning, which is used to analyze the geometric change relationship between two images. Figure 2As shown, the network consists of a feature extraction network, a feature error calculation layer, and a transformation parameter regression network. The feature extraction network is used to extract high-level semantic features from images and consists of multiple convolutional layers, activation function layers, and pooling layers. Its output is a fixed-dimensional feature vector that reflects the overall structural information of the image. The feature error calculation layer is used to calculate the difference between the feature vectors of two images. This difference represents changes in the image structure, shape, or spatial position and serves as the basis for subsequent transformation parameter prediction. The transformation parameter regression network is used to regress the geometric transformation parameters between the images based on the feature difference vector, obtaining the transformation parameter calculation results, which include scaling, displacement, and rotation calculations.

[0032] In a feasible embodiment, the feature extraction network includes three layers of convolution layer-activation function layer-maximum pooling layer structure and one layer of convolution layer-activation function layer-average pooling layer structure in sequence; the feature extraction network is used to perform feature extraction on the image to be calculated and the reference control image to obtain a 256-dimensional feature vector of the image to be calculated and a 256-dimensional feature vector of the reference control image.

[0033] The feature error calculation layer is used to determine a 256-dimensional feature vector based on the difference between the feature vector of the image to be calculated and the feature vector of the reference control image.

[0034] The transformation parameter regression network consists of a three-layer linear layer-activation function layer structure.

[0035] The transformation parameter regression network is used to learn the mapping relationship of image transformation and output a 5-dimensional transformation parameter feature vector.

[0036] Specifically, the input layer of the feature extraction network is a convolutional layer Conv2d, which contains 32 convolution kernels with a convolution kernel size of 3. The input layer is followed by a ReLU layer, followed by a 2*2 maximum pooling layer MaxPool2d, followed by a convolutional layer Conv2d with 32 input channels, 64 output channels, and a convolution kernel size of 3, followed by a ReLU layer, followed by a 2*2 maximum pooling layer MaxPool2d, followed by a convolutional layer Conv2d with 64 input channels, 128 output channels, and a convolution kernel size of 3, followed by another ReLU layer, followed by a 2*2 maximum pooling layer MaxPool2d, followed by a convolutional layer Conv2d with 128 input channels, 256 output channels, and a convolution kernel size of 3, followed by another ReLU layer, and finally a ReLU layer, followed by an adaptive average pooling layer AdaptiveAvgPool2d, which averages the output size of each channel to a 1-dimensional scalar value.

[0037] The deformation parameter regression network starts with a linear layer with an input dimension of 256 and an output dimension of 128, followed by a ReLU layer, and then another linear layer with an input dimension of 128 and an output dimension of 64; then another ReLU layer, and finally another linear layer with an input dimension of 64 and an output dimension of 5.

[0038] Specifically, two pairs of similar images are first fed into the feature extraction network to generate two 256-component feature vectors. These two feature vectors are then subtracted to generate a 256-dimensional error vector. This 256-dimensional error vector is then fed into the regression network to generate a 5-dimensional feature vector. The five parameters of this 5-dimensional feature vector correspond to the five values ​​of the deformation parameters (sx, sy, tx, ty, r), where sx is the X-axis scaling factor, sy is the Y-axis scaling factor, tx is the number of pixels translated in the X direction, ty is the number of pixels translated in the Y direction, and r is the angle of rotation along the Z axis.

[0039] In a feasible embodiment, the training step of the transformation parameter calculation model includes steps T10-T50: Step T10: input the training sample into the feature extraction network to calculate the image features.

[0040] The training samples include image pairs consisting of an original image and a transformed image. Each image in the image pair shares the network parameters of the feature extraction network during feature calculation.

[0041] Step T20: Calculate the image feature difference between the image pairs.

[0042] Step T30: input the image feature difference into the parameter regression network to obtain the initial transformation parameter calculation result.

[0043] Step T40: determining the error value between the initial transformation parameter calculation result and the actual transformation parameter based on a preset loss function.

[0044] Step T50, continuously iterating the model parameters until the error value is less than a preset value or the number of iterations is greater than a preset number of iterations.

[0045] Specifically, each training sample consists of an image pair consisting of an original image OImg and its corresponding transformed image GImg. In one example, the transformed image GImg is generated from OImg using specified affine transformation parameters. For each frame of the original image OImg, the scaling, rotation, and translation ranges are first set. Then, uniformly distributed sampling is performed within these ranges to determine the scaling factor, the pixel positions for translation, and the rotation angle. The original image OImg is then scaled, rotated, and translated using these parameters to obtain its corresponding transformed image GImg. For each image pair, the transformation parameter information is recorded: (sx0, sy0, tx0, ty0, r0), where sx0 is the scaling factor in the X direction, sy0 is the scaling factor in the Y direction, tx0 is the number of pixels translated in the X direction, ty0 is the number of pixels translated in the Y direction, and r0 is the rotation angle along the Z axis.

[0046] The model is then trained as follows: A training sample is input, and the feature computation network computes the features of the original image and the transformed image. The computation of the features of the original and transformed image shares the same network parameters. The difference between the original and transformed image features is calculated. The interpolated value calculated in the previous step is input into the parameter regression network to output the transformation parameters sx, sy, tx, ty, and r. The error is calculated using the loss function, and the model is iterated until the error is less than a given value or the number of iterations exceeds a preset value.

[0047] Among them, the loss function consists of a scaling error function, a translation error function, a rotation error function, and a total error function. Specifically, the error function is as follows: The design loss function is calculated as follows: Error in scaling parameters (scaling error function): Ls=(log(sx)-log(sx0))*(log(sx)-log(sx0))+(log(sy)-log(sy0))*(log(sy)-log(sy0)) Error in translation parameters (translation error function): Lt=(tx-tx0)*(tx-tx0)+(ty-ty0)*(ty-ty0) Error in rotation parameters (rotation error function): Lr=(r-r0)*(r-r0) Total error (total error function): L=Ls+Lt+Lr It can be understood that the similar image transformation parameter calculation method provided in this embodiment can automatically extract the structural feature differences between images based on deep neural networks, and accurately regress the geometric transformation parameters of the images (including scaling, translation and rotation). Compared with the traditional method that relies on manually designed features and alignment algorithms, this solution has stronger adaptability and generalization capabilities, can effectively improve the accuracy and robustness of image alignment and geometric analysis, and is suitable for rapid comparison and analysis tasks in various image change scenarios.

[0048] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the method for calculating similar image transformation parameters of the present application. More forms of simple transformations based on this technical concept are all within the scope of protection of the present application.

[0049] This application also provides a similar image transformation parameter calculation device, please refer to Figure 3 , the similar image transformation parameter calculation device includes: The image acquisition module is used to acquire the image to be calculated and the reference control image.

[0050] The parameter calculation module is used to input the image to be calculated and the reference control image into the transformation parameter calculation model to perform transformation parameter calculation; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer and a transformation parameter regression network.

[0051] The result acquisition module is used to obtain the transformation parameter calculation results. The parameter calculation results include scaling calculation results, displacement calculation results, and rotation calculation results.

[0052] The similar image transformation parameter calculation device provided in this application utilizes the similar image transformation parameter calculation method described in the aforementioned embodiments, thereby resolving the technical issues of low computational efficiency and poor accuracy in similar image transformation parameter calculation in related technologies. Compared to related technologies, the similar image transformation parameter calculation device provided in this application achieves the same beneficial effects as the similar image transformation parameter calculation method described in the aforementioned embodiments. Other technical features of the similar image transformation parameter calculation device are the same as those disclosed in the aforementioned embodiments and are not further elaborated upon here.

[0053] The present application provides a similar image transformation parameter calculation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the similar image transformation parameter calculation method in the above-mentioned embodiment.

[0054] The similar image transformation parameter calculation device shows a schematic structural diagram of a similar image transformation parameter calculation device suitable for implementing the embodiments of the present application. The similar image transformation parameter calculation device in the embodiments of the present application can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (such as in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers.

[0055] The similar image transformation parameter calculation device provided in this application utilizes the similar image transformation parameter calculation method described in the aforementioned embodiment, resolving the technical issues of low computational efficiency and poor accuracy in similar image transformation parameter calculation in related technologies. Compared to related technologies, the similar image transformation parameter calculation device provided in this application achieves the same beneficial effects as the similar image transformation parameter calculation method described in the aforementioned embodiment. Other technical features of this similar image transformation parameter calculation device are the same as those disclosed in the aforementioned embodiment and are not further elaborated upon here.

[0056] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0057] The above are only some embodiments of the present application and are not intended to limit the patent scope of the present application. All equivalent structural transformations made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A method for calculating similar image transformation parameters, characterized in that: The method includes: Acquire an image to be calculated and a reference control image; Inputting the image to be calculated and the reference control image into a transformation parameter calculation model to perform transformation parameter calculation; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer, and a transformation parameter regression network; Output transformation parameter calculation results; the parameter calculation results include scaling calculation results, displacement calculation results and rotation calculation results.

2. The method according to claim 1, wherein The feature extraction network sequentially includes a three-layer convolution layer-activation function layer-maximum pooling layer structure and a one-layer convolution layer-activation function layer-average pooling layer structure; The feature extraction network is used to extract features from the image to be calculated and the reference control image to obtain a 256-dimensional feature vector of the image to be calculated and a 256-dimensional feature vector of the reference control image.

3. The method according to claim 2, wherein The feature error calculation layer is used to determine a 256-dimensional feature vector based on the difference between the feature vector of the image to be calculated and the feature vector of the reference control image.

4. The method according to claim 3, wherein The transformation parameter regression network includes a three-layer linear layer-activation function layer structure; The transformation parameter regression network is used to learn the mapping relationship of image transformation and output a 5-dimensional transformation parameter feature vector.

5. The method according to claim 1, wherein The training step of the transformation parameter calculation model includes: Inputting a training sample into a feature extraction network to calculate image features; the training sample includes an image pair consisting of an original image and a transformed image; each image in the image pair shares the network parameters of the feature extraction network during feature calculation; For each of the image pairs, calculating the image feature difference between the image pairs; Inputting the image feature difference into the parameter regression network to obtain the initial transformation parameter calculation result; Based on a preset loss function, the error value between the initial transformation parameter calculation result and the actual transformation parameter is determined; The model parameters are continuously iterated until the error value is less than a preset value or the number of iterations is greater than a preset number of iterations.

6. The method according to claim 5, wherein The loss function consists of a scaling error function, a translation error function, a rotation error function and a total error function.

7. A similar image transformation parameter calculation device, characterized in that: The device includes: An image acquisition module, used to acquire the image to be calculated and the reference control image; a parameter calculation module, configured to input the image to be calculated and the reference control image into a transformation parameter calculation model to perform transformation parameter calculation; wherein the transformation parameter calculation model is composed of a feature extraction network, a feature error calculation layer, and a transformation parameter regression network; The result output module is used to obtain the transformation parameter calculation results; the parameter calculation results include scaling calculation results, displacement calculation results and rotation calculation results.

8. A device for calculating similar image transformation parameters, characterized in that: Equipment includes: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for calculating similar image transformation parameters according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the similar image transformation parameter calculation method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the method for calculating similar image transformation parameters according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Convolutional neural network-based medical image registration method and system, and electronic device

    CN107798697A

  • Sonar image registration method and system based on regression correction network

    CN111311652A

  • Image registration method improvement

    CN1799068A