Image Similarity Measurement Method, System and Device Robust to Geometric Transformations

The method enhances geometric transformation robustness in image similarity metrics by using a modified ResNet framework with Hartley pooling and convolutional neural networks, addressing vulnerabilities in existing methods and aligning better with human visual perception.

CN119919686BActive Publication Date: 2025-07-15ZHEJIANG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510409976.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-15
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing image similarity measurement methods are not robust enough for geometric transformations, and it is difficult to maintain accuracy under image rotation, flip, translation, etc., and cannot meet the needs of network security and privacy services.

Method used

A graph classification neural network framework based on Hartley pooling is constructed, and the convolution layer with a step size of 2 in ResNet is changed to 1, and a Hartley pooling layer is followed by each group of convolutional layers. The graph classification model is trained and migrated to the image similarity measurement task. The geometric information of the image is retained through the Hartley pooling layer to construct an image similarity calculation neural network.

Benefits of technology

The geometric transformation robustness of image similarity metrics is improved, making it closer to human visual judgments, and is significantly better than the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919686B_ABST
    Figure CN119919686B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and device for measuring image similarity that is robust to geometric transformations. The present invention designs a graph classification neural network framework based on Hartley pooling that is robust to geometric transformations, changes the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connects a Hartley pooling layer with equivalent dimensionality reduction effect after each group of convolutional layers. The present invention uses the designed new framework to train a graph classification model, takes the output of the feature layer of the trained graph classification model as the input feature when calculating the image similarity measure, and trains a shallow image similarity calculation neural network. The image similarity measure scheme provided by the present invention is closer to human visual judgment and is significantly more robust to geometric transformations than the current leading technical level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and more specifically, to an image similarity measurement method, system and device that are robust to geometric transformations. Background Art

[0002] Image similarity measurement is the basis for the implementation of many network security and privacy services. A typical application scenario is Client-side Scanning (CSS) adopted by Apple and Facebook to monitor the spread of Child Sexual Abuse Material (CSAM) in chat software. The specific operation mechanism of CSS is as follows: the operator establishes a database of illegal images; whenever a user sends an image in the chat software, the system uses this image as input and conducts a search based on perceptual hashing on the database; if a match with a similarity above the threshold is found during this process, then it will trigger further review work by the operator. Here, the "perceptual hash proximity" is a commonly used type of image similarity measurement.

[0003] The performance of mechanical traditional image similarity measurements such as PNSR, SSIM, etc. is insufficient to support such applications - what we desire is the "perceptual distance" close to the determination of image similarity by human vision. Currently, the most widely recognized image similarity measurement that meets this requirement is Learned Perceptual Image Patch Similarity (LPIPS). As a relatively new image similarity measurement based on neural networks, its approximation effect on human visual perception is significantly better than traditional similarity measurements. Its value is relatively robust under per-pixel editing that does not change the pixel position, such as color adjustment and blurring. However, it is significantly vulnerable to geometric transformations such as rotation, flipping, and translation. The existence of this shortcoming is not ideal for many of its potential application scenarios because humans can easily obtain the information contained in the original image from the image that has undergone geometric transformation. Summary of the Invention

[0004] The purpose of the present invention is to provide an image similarity measurement method, system and device that are robust to geometric transformations in view of the deficiencies of the prior art.

[0005] According to the first aspect of this specification, there is provided an image similarity measurement method that is robust to geometric transformations, including:

[0006] Construct a graph classification neural network framework that is robust to geometric transformations, change the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connect a Hartley pooling layer with an equivalent dimensionality reduction effect after each group of convolutional layers;

[0007] Based on an image classification dataset, a graph classification model is trained using the constructed graph classification neural network framework;

[0008] The training result of the feature layer of the graph classification model is migrated to the image similarity measurement calculation task.

[0009] Further, the calculation process of the Hadamard pooling layer includes: performing a discrete Hadamard transform on each channel of the input value, making the low-frequency data centered through a translation operation, cropping off the high-frequency data according to the dimensionality reduction requirement, then performing an inverse translation to remove the centering, and performing an inverse discrete Hadamard transform to complete the dimensionality reduction.

[0010] Further, migrating the training result of the feature layer of the graph classification model to the image similarity measurement calculation task is specifically as follows:

[0011] Extract several feature layers from the trained graph classification model. Take two images for which the similarity is to be calculated as inputs. For the output value of each layer, first perform normalization along the number of channels, then calculate the distance between the features of the two inputs at each pixel point, use the Hadamard pooling layer to unify the height and width of all features, and then perform concatenation along the dimension of the number of channels;

[0012] Construct a computer vision convolutional neural network with the shape of the concatenated image as the input requirement as the image similarity calculation neural network.

[0013] Further, extract the first rectified linear unit layer, layer1, layer2, layer3, and layer4 from the trained graph classification model, and merge them as the initial layer.

[0014] Further, the image similarity calculation neural network includes a convolutional layer, a Hadamard pooling layer, and a fully connected layer connected in sequence.

[0015] Further, denote the image similarity calculation neural network as , where are the parameters to be trained. Let be the parameters obtained through training; for two images for which the similarity is to be calculated, define the intermediate function and the similarity measurement function , where is the concatenated image obtained according to ;

[0016] Prepare training samples in the form of , where is the reference image, is between for The relative similarity weight; during the training process, make the prediction function form , and optimize the following binary cross-entropy loss function: .

[0017] According to the second aspect of this specification, there is provided an image similarity measurement system that is robust to geometric transformations, which is used to implement the image similarity measurement method described in the first aspect. The system includes:

[0018] A graph classification neural network framework construction module, which is used to construct a graph classification neural network framework that is robust to geometric transformations, change the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connect a Hadamard pooling layer with an equivalent dimensionality reduction effect after each group of convolutional layers;

[0019] A graph classification model training module, which is used to train a graph classification model based on an image classification data set by using the constructed graph classification neural network framework that is robust to geometric transformations;

[0020] An image similarity measurement module, which is used to transfer the training result of the feature layer of the graph classification model to the image similarity measurement calculation task.

[0021] According to the third aspect of this specification, there is provided an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the image similarity measurement method described in the first aspect.

[0022] According to the fourth aspect of this specification, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image similarity measurement method described in the first aspect.

[0023] According to the fifth aspect of this specification, there is provided a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, it implements the image similarity measurement method described in the first aspect.

[0024] The beneficial effects of the present invention are as follows: The present invention designs a graph classification neural network framework based on Hadamard pooling that is robust to geometric transformations, changes the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connects a Hadamard pooling layer with an equivalent dimensionality reduction effect after each group of convolutional layers. The present invention uses the designed new framework to train a graph classification model, uses the output of the feature layer of the trained graph classification model as the input feature when calculating the image similarity measurement, and trains a shallow image similarity calculation neural network. The image similarity measurement scheme provided by the present invention is closer to human visual judgment, and the robustness to geometric transformations is significantly better than the current leading technology level. Description of the Drawings

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0026] Figure 1 It is the overall flowchart of the image similarity measurement method that is robust to geometric transformations shown in an exemplary embodiment;

[0027] Figure 2 It is the schematic implementation principle diagram of the image similarity measurement method that is robust to geometric transformations shown in an exemplary embodiment;

[0028] Figure 3 It is the schematic calculation principle diagram of the Hartley pooling layer in the graph classification neural network framework shown in an exemplary embodiment;

[0029] Figure 4 It is the schematic structural diagram of the image similarity measurement system that is robust to geometric transformations shown in an exemplary embodiment;

[0030] Figure 5 It is the schematic structural diagram of an electronic device shown in an exemplary embodiment. Detailed implementation manners

[0031] To better understand the technical solutions of this application, the following will describe the embodiments of this application in detail with reference to the drawings.

[0032] It should be clear that the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.

[0033] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments, and are not intended to limit this application. The singular forms of "a", "the", and "said" used in the embodiments of this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0034] The present invention proposes an image similarity measurement method, system and device that are robust to geometric transformations. The implementation solution of the present invention is as Figure 1 shown, and can be summarized as: designing a graph classification neural network framework that is robust to geometric transformations based on Hartley pooling; training a graph classification model using the designed new framework based on an image classification dataset; and migrating the training results of the feature layer of the obtained graph classification model to the image similarity measurement calculation task.

[0035] The principle behind the solution of the present invention is a phenomenon widely discovered by scholars in the field: the feature layers of neural network models trained for a specific computer vision task often acquire feature quantities that can be widely transferred to other computer vision tasks. However, in order to achieve the purpose of the present invention being robust to geometric transformations, a framework that is sufficiently robust to geometric transformations needs to be used when training the graph classification model. To this end, the present invention introduces Hartley pooling based on the Discrete Hartley Transform in the neural network framework to enhance the robustness of the neural network framework to geometric transformations. On the one hand, the dimensionality reduction mechanism for the input is an essential part of the neural network framework for computer vision applications, but at the same time it is a step that causes loss of geometric information of the input. On the other hand, the Hartley transform is a frequency domain transform, and for naturally occurring images, usually the low-frequency part contains most of the meaningful information while the high-frequency part contains mostly noise. Therefore, the present invention designs a dimensionality reduction mechanism that is sufficiently robust to geometric transformations based on this rule. As Figure 2 shown, the following details the specific implementation process of the method for measuring image similarity that is robust to geometric transformations proposed by the present invention.

[0036] 101. Construct a graph classification neural network framework that is robust to geometric transformations.

[0037] Optimize the ResNet neural network framework, which is most commonly used in the field of computer vision, according to the task characteristics of the present invention, so as to obtain a new framework that has better robustness to geometric perturbations of the input values. ResNet originally achieved step-by-step dimensionality reduction of the input values through a series of convolutional layers with a stride of 2. The present invention changes the stride of the convolutional layers with an original stride of 2 to 1, and connects a Hartley pooling layer with an equivalent dimensionality reduction effect after each group of convolutional layers. The intuitive principle of Hartley pooling is as Figure 3 shown.

[0038] Specifically, the Hartley pooling layer performs the following calculations. Let be the tensor representation of the input value of the Hartley pooling layer, where are the number of channels, height, and width of the image respectively. Perform a discrete Hartley transform on each channel l:

[0039]

[0040]

[0041] Secondly, according to the usual practice of frequency domain analysis, obtain a centered representation with low frequencies in the center and high frequencies at the edges by translation:

[0042]

[0043] Among them denotes rounding downwards, and mod denotes taking the remainder.

[0044] Then, according to the dimensionality reduction requirement (here it is assumed to be reduced to ), the data corresponding to the high frequencies is trimmed:

[0045]

[0046] After that, inverse translation is performed on this to remove the centered representation:

[0047]

[0048] Finally, the inverse discrete Hartley transform is performed (here, the inverse transform corresponding to the previous Hartley transform is performed, so the denominator in the trigonometric function is instead of ):

[0049]

[0050]

[0051]

[0052] 102. Train a graph classification model using the new framework constructed in step 101.

[0053] Denote the new framework constructed in step 101 as ResNetR2GT (ResNet Robust to Geometric Transformation). In this step, based on an image classification dataset, use this new framework to train a graph classification model and determine its parameter values. Specifically, the optimizer adopted in this embodiment is the Stochastic Gradient Descent algorithm, the hyperparameters are the learning rate that decreases by a factor of 1 / 10 every 30 iterations (epochs) from an initial value of 0.1 and a momentum of 0.9, and the loss function uses the Cross Entropy Loss function.

[0054] 103. Extract and merge the feature layers from the graph classification model trained in step 102 as the initial layer, and build a shallow image similarity calculation neural network on top of this.

[0055] After obtaining the specific parameter values of ResNetR2GT in step 102, the outputs of several of its feature layers are extracted as the input features when calculating the image similarity metric. Specifically, in this embodiment, the outputs of a total of five layers, namely the first linear rectification function (ReLU) layer of ResNetR2GT and the subsequent layer1, layer2, layer3, and layer4, are extracted (here the names of the layers follow the conventional nomenclature of the ResNet framework). The following uses to represent the output values of these five layers when the input is x. The output value is a 3D tensor (i.e., number of channels * height * width). Denote the number of channels, height, and width of the output of as 、 respectively.

[0056] Let and

[0057]

[0058] be two input images for which the similarity is to be calculated. For each layer, first perform normalization along the number of channels:

[0059]

[0060] Here note that the output of is still a 3D tensor with the number of channels, height, and width being 、 respectively. Use Hartley pooling to unify the height and width of all features for subsequent concatenation operations. The features after size unification are denoted as . Specifically, unify the height and width of all features to a size not exceeding the minimum height and minimum width of all features. In this embodiment, use Hartley pooling to unify the (height, width) of the features into :

[0061]

[0062] where refers to the Hartley pooling layer with the output (height, width) being . Therefore, the output of is a 3D tensor with the number of channels, height, and width being 、 3D tensor. Since the Hadamard pooling can better preserve the geometric information of the data, the robustness of this step to the geometric transformation of the input is maximally guaranteed. Stack these and splice them along the dimension of the number of channels to form a 3D tensor with the number of channels, height, and width being , , 3D tensor .

[0063] Since has the shape of conventional image data (i.e., the number of channels * height * width), the present invention constructs a computer vision convolutional neural network with its shape as the input requirement as the image similarity calculation neural network. Specifically, construct the following image similarity calculation neural network , where are the parameters to be trained:

[0064] The first layer: convolutional layer. In this embodiment, a convolutional layer of is adopted, and the stride is taken as 1;

[0065] The second layer: Hadamard pooling layer. In this embodiment, a Hadamard pooling layer with an output size of is adopted;

[0066] The third layer: fully connected layer with an output dimension of 1.

[0067] Let be the parameters obtained by training. Since the similarity metric requires a function with two images as the input, the present invention defines the following intermediate function and the final similarity metric function :

[0068]

[0069]

[0070] The reason for additionally defining the intermediate function outside the similarity metric function is that the present invention uses the intermediate function to train the parameter .

[0071] 104. Train the image similarity calculation neural network using the image similarity judgment data set.

[0072] The entire calculation formula of the image similarity metric of the present invention has been elaborated in step 103. The only remaining technical detail is the training method of the parameter .

[0073] First, the data format of the image similarity evaluation dataset for training is described. A single training sample needs to have the form of , where are three images, r is the reference image, is the relative similarity weight between and with respect to r: close to 0 indicates that is more similar to r than ; close to 1 indicates that is more similar to r than ; around 0.5 indicates that is equally similar to r as and . When collecting data, the present invention takes a group of and asks multiple human subjects to judge which of and is more similar to r, and then takes the relative frequency of different judgments as the value.

[0074] During the training process, a logistic function layer is stacked on the image similarity calculation neural network given in step 103 to obtain class probability values with a value range of . Specifically, let be a training sample, as defined in step 103. Let the prediction function be in the following form:

[0075]

[0076] Then, the following binary cross entropy loss function is optimized:

[0077]

[0078] The optimizer used for training is ADAM (Adaptive Moment Estimation), and the hyperparameters are a learning rate with an initial value of 0.0001, an exponential decay rate of the first moment estimate of 0.5 and an exponential decay rate of the second moment estimate of 0.999 .

[0079] On the other hand, an embodiment of the present invention provides an image similarity measurement system 2000 that is robust to geometric transformations, which is used to implement the above-mentioned method for measuring image similarity that is robust to geometric transformations, as shown in Figure 4 . The system includes:

[0080] The graph classification neural network framework construction module 2001 is used to construct a graph classification neural network framework that is robust to geometric transformations. The stride of the convolutional layer with a stride of 2 in ResNet is changed to 1, and a Hadamard pooling layer with an equivalent dimensionality reduction effect is connected after each group of convolutional layers;

[0081] The graph classification model training module 2002 is used to train a graph classification model based on an image classification dataset using the graph classification neural network framework constructed by the graph classification neural network framework construction module 2001;

[0082] The image similarity measurement module 2003 is used to transfer the training results of the feature layer of the graph classification model trained by the graph classification model training module 2002 to the image similarity measurement calculation task.

[0083] Regarding the system in the above embodiments, the specific ways in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0084] For the system embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the descriptions of the method embodiments. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0085] Correspondingly, the present application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the method for measuring image similarity that is robust to geometric transformations as described above. As Figure 5 shown, it is a hardware structure diagram of any device with data processing capabilities where the method for measuring image similarity that is robust to geometric transformations provided by the embodiments of the present invention is located. Except for Figure 5 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0086] Correspondingly, the present application further provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the method for measuring image similarity robust to geometric transformation as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.

[0087] After considering the specification and practicing the content disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary.

[0088] It should be understood that the present application is not limited to the exact structure already described and shown in the drawings, and various modifications and changes can be made without departing from its scope.

[0089] The above is only the preferred embodiment of the present invention. Although the present invention has been disclosed above with preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into an equivalent embodiment with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the protection of the technical solution of the present invention.

Claims

1. An image similarity measurement method robust to geometric transformation, characterized in that, The method includes: Construct a graph classification neural network framework that is robust to geometric transformations. Change the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connect a Hadamard pooling layer with equivalent dimensionality reduction effect after each group of convolutional layers. The calculation process of the Hadamard pooling layer includes: performing a discrete Hadamard transform on each channel of the input value, making the low-frequency data centered through a translation operation, cropping off the high-frequency data according to the dimensionality reduction requirement, then performing an inverse translation to remove the centering, and performing an inverse discrete Hadamard transform to complete the dimensionality reduction; Based on an image classification dataset, use the constructed graph classification neural network framework to train a graph classification model; Transfer the training results of the feature layer of the graph classification model to the image similarity metric calculation task.

2. The method according to claim 1, wherein Transfer the training results of the feature layer of the graph classification model to the image similarity metric calculation task, specifically: Extract several feature layers from the trained graph classification model. Take two images for which the similarity is to be calculated as inputs. For the output value of each layer, first perform normalization along the number of channels, then calculate the distance between the features of the two inputs at each pixel point. Use the Hadamard pooling layer to unify the height and width of all features, and then concatenate them along the dimension of the number of channels; Construct a computer vision convolutional neural network with the shape of the concatenated image as the input requirement as the image similarity calculation neural network.

3. The method according to claim 2, wherein Extract the first rectified linear unit layer, layer1, layer2, layer3, and layer4 from the trained graph classification model and merge them as the initial layer.

4. The method according to claim 2, wherein The image similarity calculation neural network includes a convolutional layer, a Hadamard pooling layer, and a fully connected layer connected in sequence.

5. The method according to claim 2, wherein Denote the neural network for calculating the image similarity as , where are parameters to be trained, and let be the parameters obtained through training; for two images for which the similarity is to be calculated, define the intermediate function and the similarity metric function , where is the spliced image obtained according to ; Prepare training samples in the form of where is the reference image, is the relative similarity weight between for ; during training, let the form of the prediction function be , and optimize the following binary cross-entropy loss function: .

6. An image similarity measurement system robust to geometric transformation, characterized in that, For implementing the image similarity metric method as described in any one of claims 1-5, the system includes: A graph classification neural network framework construction module for constructing a graph classification neural network framework that is robust to geometric transformations, changing the stride of the convolutional layer with a stride of 2 in ResNet to 1, and connecting a Hadamard pooling layer with equivalent dimensionality reduction effect after each group of convolutional layers; A graph classification model training module for training a graph classification model based on an image classification dataset using the constructed graph classification neural network framework that is robust to geometric transformations; An image similarity metric module for transferring the training results of the feature layer of the graph classification model to the image similarity metric calculation task.

7. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the image similarity metric method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image similarity metric method as described in any one of claims 1-5.

9. A computer program product comprising computer programs / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the image similarity metric method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Remote sensing image scene classification method based on multi-similarity measurement deep learning

    CN111723675A