An image similarity calculation method based on multi-view feature fusion

By employing a multi-view feature fusion-based image similarity calculation method, which combines target recognition, cosine similarity, and Siamese neural networks, the problem of insufficient data augmentation and adaptability in confined space operations in oil and gas fields has been solved, achieving higher accuracy and more flexible image similarity calculation.

CN120726353BActive Publication Date: 2025-11-18SOUTHWEST PETROLEUM UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511177501.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-21
Publication Date
2025-11-18
Estimated Expiration
2045-08-21

AI Technical Summary

Technical Problem

Existing image similarity calculation methods suffer from inappropriate data augmentation and insufficient adaptability in confined spaces in oil and gas fields, affecting model performance and applicability.

Method used

A deep learning model is constructed by using a multi-view feature fusion method, combining target recognition technology, cosine similarity calculation, and Siamese neural network. The model integrates the output features of target detection, cosine similarity, and Siamese neural network through a multilayer perceptron to perform image similarity calculation.

Benefits of technology

It improves the accuracy and adaptability of image similarity calculation, and can effectively cope with problems such as lighting changes, device occlusion and viewpoint differences in complex environments, ensuring accurate image similarity calculation in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726353B_ABST
    Figure CN120726353B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-view feature fusion's image similarity calculation method, it is related to image similarity calculation technical field, steps are as follows: S1, collect the scene data applied to image similarity calculation;S2, processing and sampling are carried out to data, obtain optimal positive-negative sample set;S3, respectively build and train target detection module, twin neural network module and cosine similarity module;S4, construct multi-view feature fusion model, integrate the output result of above three modules;S5, the image to be carried out image similarity calculation is input to the multi-view feature fusion model that has been trained to obtain accurate image similarity calculation prediction result.The application adopts the above steps a kind of based on multi-view feature fusion's image similarity calculation method, combines target recognition technology, cosine similarity calculation and the output of twin neural network constructs multi-view feature fusion model, to improve the precision of image similarity calculation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image similarity calculation technology, and in particular to an image similarity calculation method based on multi-view feature fusion. Background Technology

[0002] In the development of oil and gas fields in the modern petroleum industry, the identification of confined space operations is crucial for ensuring industrial control safety and compliance with operational procedures. With the development of computer vision technology, image similarity calculation methods have been widely applied to tasks such as image classification, object detection, and image retrieval, and are gradually being introduced into safety management and operational supervision at oil and gas field sites. Despite these significant advancements, considerable challenges remain when applying these methods to the identification of complex operations in confined spaces.

[0003] (1) Inappropriate data augmentation: Existing data augmentation methods may not be suitable for professional scenarios. For example, in the analysis of equipment images in oil and gas fields, overly aggressive operations (such as large-scale cropping or excessive color transformation) may destroy key features, affect the contrast learning effect and reduce model performance.

[0004] (2) Insufficient adaptability: Existing methods are difficult to comprehensively handle different types of professional image data. The complex and ever-changing environment of oil and gas fields requires models to be able to flexibly cope with various situations. Summary of the Invention

[0005] The purpose of this invention is to provide an image similarity calculation method based on multi-view feature fusion. By combining target recognition technology, cosine similarity calculation, and the output of Siamese neural network to construct a deep learning model, the accuracy of image similarity calculation is improved, and stronger adaptability and scalability are provided for different application scenarios.

[0006] To achieve the above objectives, this invention provides an image similarity calculation method based on multi-view feature fusion, the steps of which are as follows:

[0007] S1. Collect scene data for image similarity calculation;

[0008] S2. Process and sample the data, filter and retain key feature images to obtain the optimized set of positive and negative samples;

[0009] S3. Build and train the object detection module, the Siamese neural network module, and the cosine similarity module respectively. The object detection module outputs the object detection confidence and bounding box coordinates, the Siamese neural network module outputs the Manhattan distance similarity score, and the cosine similarity module outputs the cosine similarity value.

[0010] S4. Construct a multi-view feature fusion model, integrate the output results of the object detection module, the Siamese neural network module, and the cosine similarity module, and generate the final similarity score through a multilayer perceptron;

[0011] S5. Using the optimized sample set, the constructed multi-view feature fusion model is trained by optimizing the binary cross-entropy loss function to obtain the final model.

[0012] S6. Input the images to be compared into the pre-trained final model, and obtain accurate image similarity calculation and prediction results through comprehensive analysis.

[0013] Preferably, S2 specifically includes the following steps:

[0014] S201. Sampling: Screening images for the application scenario and retaining images with significant differences in image similarity calculation;

[0015] S202, Processing: The collected images are processed into an input format corresponding to the object detection module, the Siamese neural network module, and the cosine similarity module.

[0016] Preferably, in S301, the process of building the target detection module is as follows: create a deep separable convolutional neural network for target recognition and a data loader, establish the bounding box regression loss function, the classification cross-entropy loss function, the confidence loss function and the optimizer, and set and optimize the hyperparameters;

[0017] Depthwise separable convolutional networks can be expressed by the following formula:

[0018]

[0019] in, Indicates the position in the input tensor ( w,h ) in the c Values ​​on the channel; Indicates the first n The convolutional kernel at the _th ... c Weights on each input channel; It is the bias term, corresponding to the nth input channel, and the output... It is an element in the output feature map after convolution and biasing;

[0020] The bounding box regression loss function is used to optimize the difference between the predicted bounding box and the true bounding box, and is expressed by the following formula:

[0021]

[0022] in, S Indicates the size of the grid. BThis indicates the number of bounding boxes predicted for each grid cell. Indicates the first i In the _ grid cell, the _ ... j Does each bounding box have the responsibility of predicting the target? x, y These are the coordinates of the center point of the bounding box. w,h These represent the width and height of the bounding box, respectively. These are weighting coefficients used to balance the losses of different parts;

[0023] The classification cross-entropy loss function is expressed by the following formula:

[0024]

[0025] in, S Indicates the grid size; Indicates the first i Does each grid cell contain a target? It is the model's prediction of the first i The target in each grid cell belongs to category c The probability of; It is a real label, indicating the first i Does the target in each grid cell belong to a category? c ;

[0026] The confidence loss function is expressed by the following formula:

[0027]

[0028] N is the number of samples; C is the number of categories; It is the true label of the i-th sample; The predicted probability that the i-th sample belongs to category c; It is a balancing factor used to adjust the weights between positive and negative samples; It is an aggregation parameter used to control the degree of attention given to difficult samples.

[0029] Preferably, in S302, the process of building the cosine similarity module is as follows:

[0030] (1) Define the image and output paths, including the sample image path, the image path to be compared, and the output path;

[0031] (2) Preprocess the sample images by using OpenCV to read the images and then using skimage to convert the image data into an unsigned 8-bit integer format;

[0032] (3) Read all images in the target folder and preprocess them. Use OpenCV to read the images, and then use skimage to read and save them as arrays.

[0033] (4) Traverse the images in the target folder and calculate the cosine similarity value. Use cosine_similarity in sklearn.metrics.pairwise to get the similarity score of each image and save the images with similarity scores greater than the specified threshold.

[0034] Preferably, in S303, the process of building the Siamese neural network module is as follows:

[0035] (1) The input end of the Siamese neural network model is connected to two sister convolutional neural network models with the same architecture, hyperparameters and weights, and the output end uses a fully connected layer to output feature vectors;

[0036] (2) Construct a dataset loader. The data in the dataset exists in the form of data pairs. Each data pair contains two images and image category labels. The convolutional neural network model uses the Manhattan distance formula to calculate the similarity between two feature vectors and obtain the Manhattan distance similarity score.

[0037] (3) Select the contrastive loss function as the loss function. The formula for the contrastive loss function is as follows:

[0038]

[0039] in, N It is the number of sample pairs. y The label indicates whether two input samples belong to the same class: y = 1 indicates the same category, y = 0 indicates different categories, d It is the distance between two input samples in the feature space. margin It is a predefined threshold used to control the minimum distance between samples of different categories;

[0040] (4) Train the model. Use the prepared paired dataset and the selected contrastive loss function to perform initial model training. During the training process, the SGD optimization algorithm is used to update the network weights to minimize the loss function.

[0041] (5) Optimization: After completing the initial training, evaluate the model performance on an independent validation set and adjust the hyperparameters or improve the network architecture based on the evaluation results.

[0042] Preferably, in S4, the process of constructing the multi-view feature fusion model is as follows:

[0043] The multi-view feature fusion model consists of three fully connected layers. The outputs of the object detection module, the Siamese neural network module, and the cosine similarity calculation module are used as the original features input to the first fully connected layer. The first fully connected layer maps the original features with a dimension of 3 to a 128-dimensional space through a linear transformation. The feature dimension of the second fully connected layer is 64-dimensional, and the feature dimension of the third fully connected layer is 32-dimensional. After each fully connected layer, a batch normalization layer, a ReLU activation function, and a Dropout layer are connected in sequence. Finally, the Sigmoid function is used to restrict the output of the output layer to the range of [0, 1].

[0044] Preferably, in S5, a binary cross-entropy loss function is used to optimize the training of the multi-view feature fusion model, and the optimizer, learning rate, number of learning epochs, and early stopping mechanism are adjusted according to the training results.

[0045] The formula for the binary cross-entropy loss function is as follows:

[0046]

[0047] in, It is aimed at the first i The loss value for each sample. is the true label of the i-th sample, with a value of 0 or 1, representing the negative class and the positive class, respectively; It is the model's predicted probability that the i-th sample belongs to the positive class, which comes from the output of the multi-view feature fusion model;

[0048] for N The average loss for each sample is calculated using the following formula:

[0049]

[0050] in, L This represents the average loss value for the entire dataset or batch of samples, used to measure the difference between the model's predictions and the true labels; N Indicates the number of samples.

[0051] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0052] (1) By integrating target detection, cosine similarity and the output features of Siamese neural network through multilayer perceptron (MLP), the three different dimensions of features are fused together, so that the model can fully capture the feature information of the image in different aspects, thereby achieving more accurate image similarity calculation in complex scenes.

[0053] (2) It has strong scene adaptability. Through multi-view feature fusion and adaptive weight learning, the model can effectively cope with changes in lighting (such as different lighting conditions such as nighttime and strong light).

[0054] Problems such as device occlusion (part of the device is blocked by other objects) and differences in perspective (different shooting angles) are addressed to ensure accurate image similarity calculation in various complex environments.

[0055] (3) The entire model adopts a modular design, with the three modules of object detection, cosine similarity and Siamese neural network being independent yet collaborative. This design makes the model highly flexible, and users can replace one of the components as needed according to specific application scenarios and requirements.

[0056] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a flowchart illustrating an embodiment of an image similarity calculation method based on multi-view feature fusion according to the present invention.

[0059] Figure 2 This is a schematic diagram of the system architecture according to an embodiment of the present invention. Detailed Implementation

[0060] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0062] Example 1: As Figure 1 As shown, an image similarity calculation method based on multi-view feature fusion is described, with the following steps:

[0063] S1. Collect scene data for image similarity calculation, including but not limited to images in JPG or PNG format. In the specific scenario of confined space operations in the oil industry, the main goal of collecting image data is to train a model capable of accurately calculating image similarity to identify compliance status, equipment condition, etc., during the operation. Confined space operation scenarios include, but are not limited to, internal inspection of oil pipelines, oil tank maintenance, and underground oil depot operations. These scenarios are characterized by complex lighting conditions, diverse equipment layouts, and the presence of occlusion; identifying these characteristics helps in subsequent targeted data collection.

[0064] S2. Process and sample the data, filter and retain key feature images to obtain the optimized set of positive and negative samples;

[0065] S201. Sampling: Screening images for application scenarios, identifying images with significant differences based on image similarity calculations;

[0066] S202, Processing: The collected images are processed into input formats corresponding to the object detection module, the Siamese neural network module, and the cosine similarity module. The data formats of the three corresponding modules are processed into, but are not limited to, images and corresponding label files, image cropping, and corresponding high-dimensional feature vectors. This invention processes the images into 640×640 images and corresponding label files (txt) for object detection training, 200×160 images for calculating cosine similarity, and 640×640 images for training the Siamese neural network.

[0067] S3. Build and train the object detection module, the Siamese neural network module, and the cosine similarity module respectively. The input data for the three modules should be consistent, but the input format should be different. The object detection module outputs the object detection confidence score and bounding box coordinates, the Siamese neural network module outputs the Manhattan distance similarity score, and the cosine similarity module outputs the cosine similarity value.

[0068] Object recognition: First, key objects or regions in the image are extracted using object recognition technology (deep separable convolutional neural network). This helps to focus on the most important parts of the image rather than processing the entire image, thereby improving the relevance and accuracy of features.

[0069] Cosine similarity: Used to measure the angular difference between two vectors, it is an effective metric, particularly suitable for comparing data in high-dimensional spaces. By calculating the cosine similarity between feature vectors, the semantic similarity between different images can be captured.

[0070] Siamese neural networks process paired inputs by sharing weights and mapping them to a common feature space for direct comparison. Siamese networks are effective at capturing subtle differences between images.

[0071] S301, the process of building the object detection module is as follows: Create a deep separable convolutional neural network for object recognition and a dataloader, establish the bounding box regression loss function, the classification cross-entropy loss function, the confidence loss function and the optimizer, and set the hyperparameters including the learning rate (lr) of 0.0001, the number of learning epochs of 500, and the early stopping mechanism (patience) of 30.

[0072] Depthwise separable convolutional networks can be expressed by the following formula:

[0073]

[0074] in, Indicates the position in the input tensor ( w,h ) in the c Values ​​on the channel; Indicates the first n The convolutional kernel at the _th ... c Weights on each input channel; It is the bias term, corresponding to the nth input channel, and the output... It is an element in the output feature map after convolution and biasing;

[0075] The bounding box regression loss function is used to optimize the difference between the predicted bounding box and the true bounding box, and is expressed by the following formula:

[0076]

[0077] in, S Indicates the size of the grid. B This indicates the number of bounding boxes predicted for each grid cell. Indicates the first i In the _ grid cell, the _ ... j Does each bounding box have the responsibility of predicting the target? x, y These are the coordinates of the center point of the bounding box. w,h These represent the width and height of the bounding box, respectively. These are weighting coefficients used to balance the losses of different parts;

[0078] The classification cross-entropy loss function is expressed by the following formula:

[0079]

[0080] in, SIndicates the grid size; Indicates the first i Does each grid cell contain a target? It is the model's prediction of the first i The target in each grid cell belongs to category c The probability of; It is a real label, indicating the first i Does the target in each grid cell belong to a category? c ;

[0081] The confidence loss function is expressed by the following formula:

[0082]

[0083] N It is the sample size; C It is the number of categories; It is the first i The true label of each sample; It is the first i Each sample belongs to category c The predicted probability; It is a balancing factor used to adjust the weights between positive and negative samples; It is an aggregation parameter used to control the degree of attention given to difficult samples.

[0084] S302, The process of building the cosine similarity module is as follows:

[0085] (1) Define the image and output paths, including the sample image path, the image path to be compared, and the output path;

[0086] (2) Preprocess the sample images by using OpenCV to read the images and then using skimage to convert the image data into an unsigned 8-bit integer format;

[0087] (3) Read all images in the target folder and preprocess them. Use OpenCV to read the images, and then use skimage to read and save them as arrays.

[0088] (4) Traverse the images in the target folder and calculate the cosine similarity value. Use cosine_similarity in sklearn.metrics.pairwise to get the similarity score of each image and save the images with similarity scores greater than the specified threshold.

[0089] S303, the process of building the Siamese neural network module is as follows:

[0090] (1) Design a convolutional neural network model, which will serve as the shared weight part of the Siamese network. Both inputs will be processed through the exact same network parameters. The input of the Siamese neural network model is connected to two sister convolutional neural network models with identical architecture, hyperparameters, and weights, and the output is a fully connected layer that outputs feature vectors.

[0091] (2) Construct a dataset loader. The data in the dataset exists in the form of data pairs. Each data pair contains two images and image category labels. The data required for Siamese network training exists in pairs. Each pair contains two images and a label indicating whether they belong to the same category.

[0092] The convolutional neural network model uses the Manhattan distance formula to calculate the similarity between two feature vectors and obtain the Manhattan distance similarity score;

[0093] (3) Select the contrastive loss function as the loss function. The formula for the contrastive loss function is as follows:

[0094]

[0095] in, N It is the number of sample pairs. y The label indicates whether two input samples belong to the same class: y = 1 indicates the same category, y = 0 indicates different categories, d It is the distance between two input samples in the feature space. margin It is a predefined threshold used to control the minimum distance between samples of different categories;

[0096] (4) Train the model. Use the prepared paired dataset and the selected contrastive loss function to perform initial model training. During the training process, the SGD optimization algorithm is used to update the network weights to minimize the loss function.

[0097] (5) Optimization: After completing the initial training, evaluate the model performance on an independent validation set and adjust the hyperparameters or improve the network architecture based on the evaluation results.

[0098] S4. Construct a multi-view feature fusion model, integrate the output results of the object detection module, the Siamese neural network module, and the cosine similarity module, and generate the final similarity score through a multilayer perceptron;

[0099] The process of constructing a multi-view feature fusion model is as follows:

[0100] like Figure 2 As shown, Is as , , Input data, It is the target detection module; It is the cosine similarity module; It is a Siamese neural network module. Each of the three modules receives the same input data conforming to its own data format, and after processing by its respective algorithm and model, produces three different outputs. The deep learning multilayer perceptron module is responsible for processing the outputs of the first three modules, and will... Target recognition results cosine similarity score and The similarity scores of the twin networks are combined. log It is the natural logarithm function. The logarithm is used to ensure that the loss is small when the predicted probability is close to the actual label, but increases rapidly when the predicted probability deviates significantly from the actual label. This is because the logarithmic function's properties within its domain amplify low-probability events.

[0101] The multi-view feature fusion model consists of three fully connected layers. The outputs of the object detection module, the Siamese neural network module, and the cosine similarity calculation module are used as raw features input to the first fully connected layer. The first fully connected layer maps the raw features with a dimension of 3 to a 128-dimensional space through a linear transformation. The second fully connected layer has a feature dimension of 64, and the third fully connected layer has a feature dimension of 32. After each fully connected layer, a batch normalization layer, a ReLU activation function, and a Dropout layer are sequentially applied. Batch normalization is applied after each layer to accelerate training and stabilize gradients, and the ReLU activation function is used to support non-linear modeling. The Dropout mechanism prevents overfitting and improves the model's generalization ability by randomly discarding neuron connections. Finally, the Sigmoid function is used to restrict the output of the output layer to the range [0, 1], which is the similarity score, representing the degree of similarity between objects.

[0102] S5. Using the optimized sample set, the constructed multi-view feature fusion model is trained with the binary cross-entropy loss function to obtain the final model.

[0103] The binary cross-entropy loss function is used to optimize the training of the multi-view feature fusion model. Its function is to calculate the loss value by measuring the difference between the probability distribution output by the model and the actual label when predicting the probability of a sample belonging to one of two possible categories.

[0104] The formula for the binary cross-entropy loss function is as follows:

[0105]

[0106] in, This is the loss value for the i-th sample. is the true label of the i-th sample, with a value of 0 or 1, representing the negative class and the positive class, respectively; It is the model's predicted probability that the i-th sample belongs to the positive class, which comes from the output of the multi-view feature fusion model;

[0107] for N The average loss for each sample is calculated using the following formula:

[0108]

[0109] in, L This represents the average loss value for the entire dataset or batch of samples. This is the true label of the i-th sample, with a value of 0 or 1, representing the negative class and the positive class, respectively. It is the model's predicted probability that the i-th sample belongs to the positive class. N Indicates the number of samples.

[0110] S6. Input the images to be compared into the pre-trained final model, and obtain accurate image similarity calculation and prediction results through comprehensive analysis.

[0111] In summary, this invention proposes a novel multi-view feature fusion model, which has the following beneficial effects:

[0112] Multi-source information integration: This model integrates multiple feature information such as target detection, cosine similarity calculation and Siamese network output, and analyzes and understands image content from multiple perspectives. It is particularly suitable for recognition of confined space operations in complex environments such as oil and gas fields.

[0113] Optimized architecture design: A three-layer linear transformation structure is adopted, with batch normalization and ReLU activation function applied after each layer. Combined with the Dropout mechanism to prevent overfitting, the expressive power and generalization ability of the model are enhanced.

[0114] Accurate evaluation: Finally, the Sigmoid function maps the features to similarity scores in the (0,1) interval, providing a quantitative evaluation standard to facilitate quick judgment of the similarity between objects.

[0115] This approach not only solves the problem that traditional data augmentation strategies may destroy key features, but also improves the model's adaptability and recognition accuracy to different types of image data, thereby better ensuring the safety and efficiency of confined space operations in oil and gas field development.

[0116] The remaining technical features in the above embodiments can be flexibly selected by those skilled in the art to meet different specific practical needs. However, it is obvious to those skilled in the art that these specific details are not necessary to implement the present invention. In other instances, to avoid obscuring the present invention, well-known components, structures, or parts are not specifically described, and all are within the scope of technical protection defined by the claims of the present invention.

[0117] Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of this invention should be within the protection scope of the appended claims. In the above description, numerous specific details have been set forth to provide a thorough understanding of the invention. However, it will be apparent to those skilled in the art that these specific details are not necessary to practice the invention. In other instances, to avoid obscuring the invention, well-known techniques, such as specific construction details, operating conditions, and other technical conditions, have not been specifically described.

[0118] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. An image similarity calculation method based on multi-view feature fusion, characterized in that, The steps are as follows: S1. Collect scene data for image similarity calculation; S2. Process and sample the data, filter and retain key feature images to obtain the optimized set of positive and negative samples; S3. Build and train the object detection module, the Siamese neural network module, and the cosine similarity module respectively. The object detection module outputs the object detection confidence and bounding box coordinates, the Siamese neural network module outputs the Manhattan distance similarity score, and the cosine similarity module outputs the cosine similarity value. S4. Construct a multi-view feature fusion model, integrating the outputs of the object detection module, the Siamese neural network module, and the cosine similarity module. The final image similarity score is generated using a multilayer perceptron. The process of constructing the multi-view feature fusion model is as follows: The multi-view feature fusion model consists of three fully connected layers. The outputs of the object detection module, the Siamese neural network module, and the cosine similarity calculation module are used as raw features input to the first fully connected layer. The first fully connected layer maps the raw features with a input dimension of 3 to a 128-dimensional space through a linear transformation. The feature dimension of the second fully connected layer is 64-dimensional, and the feature dimension of the third fully connected layer is 32-dimensional. After each fully connected layer, a batch normalization layer, a ReLU activation function, and a Dropout layer are connected in sequence. Finally, the Sigmoid function is used to restrict the output of the output layer to the range of [0, 1]. S5. Using the optimized sample set, the constructed multi-view feature fusion model is trained by optimizing the binary cross-entropy loss function to obtain the final model. S6. Input the image to be calculated for image similarity into the pre-trained final multi-view feature fusion model, and obtain accurate image similarity calculation prediction results through comprehensive analysis.

2. The image similarity calculation method based on multi-view feature fusion according to claim 1, characterized in that: S2 specifically includes the following steps: S201. Sampling: Filtering images for application scenarios; S202, Processing: The collected images are processed into an input format corresponding to the object detection module, the Siamese neural network module, and the cosine similarity module.

3. The image similarity calculation method based on multi-view feature fusion according to claim 1, characterized in that: S301, the process of building the object detection module is as follows: create a deep separable convolutional neural network for object recognition and a data loader, establish bounding box regression loss function, classification cross-entropy loss function, confidence loss function and optimizer, and set and optimize hyperparameters; Depthwise separable convolutional networks can be expressed by the following formula: (1) in, Indicates the position in the input tensor ( w,h ) in the c Values ​​on the channel; Indicates the first n The convolutional kernel at the _th ... c Weights on each input channel; It is a bias term, corresponding to the first... n One input channel, one output channel It is an element in the output feature map after convolution and biasing; The bounding box regression loss function is used to optimize the difference between the predicted bounding box and the true bounding box, and is expressed by the following formula: (2) in, S Indicates the size of the grid. B This indicates the number of bounding boxes predicted for each grid cell. Indicates the first i In the grid cell, the th j Does each bounding box have the responsibility of predicting the target? x, y These are the coordinates of the center point of the bounding box. w,h These represent the width and height of the bounding box, respectively. These are weighting coefficients used to balance the losses of different parts; The classification cross-entropy loss function is expressed by the following formula: (3) in, S Indicates the grid size; Indicates the first i Does each grid cell contain a target? It is the model's prediction of the first i The target in each grid cell belongs to category c The probability of; is the true label, representing the first... i Does the target in each grid cell belong to a category? c ; The confidence loss function is expressed by the following formula: (4) N It is the sample size; C It is the number of categories; It is the first i The true label of each sample; It is the first i Each sample belongs to category c The predicted probability; It is a balancing factor used to adjust the weights between positive and negative samples; It is an aggregation parameter used to control the degree of attention given to difficult samples.

4. The image similarity calculation method based on multi-view feature fusion according to claim 1, characterized in that: S302, The process of building the cosine similarity module is as follows: (1) Define the image and output paths, including the sample image path, the image path for which image similarity calculation is to be performed, and the output path; (2) The sample images were preprocessed by using the OpenCV library in Python to read the images and then using skimage to convert the image data into an unsigned 8-bit integer format. (3) Read all images in the target folder and preprocess them. Use Python's OpenCV library to read the images, and then use skimage to read and save them as arrays. (4) Traverse the images in the target folder and calculate the cosine similarity value. Use cosine_similarity in sklearn.metrics.pairwise to get the similarity score of each image and save the images with similarity scores greater than the specified threshold.

5. The image similarity calculation method based on multi-view feature fusion according to claim 1, characterized in that: S303, the process of building the Siamese neural network module is as follows: (1) The input end of the Siamese neural network model is connected to two sister convolutional neural network models with the same architecture, hyperparameters and weights, and the output end uses a fully connected layer to output feature vectors; (2) Construct a dataset loader. The data in the dataset exists in the form of data pairs. Each data pair contains two images and image category labels. The convolutional neural network model uses the Manhattan distance formula to calculate the similarity between two feature vectors and obtain the Manhattan distance similarity score. (3) Select the contrastive loss function as the loss function. The formula for the contrastive loss function is as follows: (5) in, N It is the number of sample pairs. y The label indicates whether two input samples belong to the same class: y = 1 indicates the same category, y = 0 indicates different categories, d It is the distance between two input samples in the feature space. margin It is a predefined threshold used to control the minimum distance between samples of different categories; (4) Train the model. Use the prepared paired dataset and the selected contrastive loss function to perform initial model training. During the training process, the SGD optimization algorithm is used to update the network weights to minimize the loss function. (5) Optimization: After completing the initial training, evaluate the model performance on an independent validation set and adjust the hyperparameters or improve the network architecture based on the evaluation results.

6. The image similarity calculation method based on multi-view feature fusion according to claim 1, characterized in that: S5. The multi-view feature fusion model is trained using a binary cross-entropy loss function. The optimizer, learning rate, number of learning epochs, and early stopping mechanism are adjusted based on the training results. The formula for the binary cross-entropy loss function is as follows: (6) in, It is aimed at the first i The loss value for each sample. is the true label of the i-th sample, with a value of 0 or 1, representing the negative class and the positive class, respectively; It is the model's predicted probability that the i-th sample belongs to the positive class, which comes from the output of the multi-view feature fusion model; for N The average loss for each sample is calculated using the following formula: (7) in, L This represents the average loss value for the entire dataset or batch of samples, used to measure the difference between the model's predictions and the true labels; N Indicates the number of samples.

Citation Information

Patent Citations

  • Long tail distribution visual classification method based on sample perception distillation

    CN115995018A

  • Fast aggregation classification method based on feature similarity calculation

    CN119380113A